Owner turns beat what audits found elsewhere

Every published audit of a general news aggregator finds the same thing: a few large legacy outlets take a third to over half of the slots, the top slot is more concentrated than the list, and fresh copy wins. The cause is ranking by recency and authority, not personalisation, so a same-for-everyone page does not escape it unless it balances on purpose. Huduku's round 23 rules do that. They count owners rather than feeds, give each owner a turn before any gets a second, cap cards per owner, and weight stories by how many owners carried them. Together they answer the mechanism the audits name, and in one respect they go further than any aggregator has disclosed: Google's public cap folds subdomains into a site, but nobody folds sister brands into an owner. The research also shows where the design is weaker or incomplete. Equal turns per owner is one side of an unresolved debate (equal against merit-proportional exposure). A fixed order within each round leaves a lasting position advantage that the fairness literature says to rotate away. The guarantee is only as good as the ownership data, and the Kannada Prabha gap shows that already. Freshness still leaks in through fetch latency and newest-first search. India has no published aggregator audit and no binding concentration threshold, so Huduku's own counts would be among the few public numbers on the question.

Audits find two to five outlets taking most of the room

The pattern holds across countries, years and products. In the 2016 US primaries, CNN and the New York Times held 44.2% of Google "In the News" links and 64.4% of the first slot (Diakopoulos et al.). In Trielli and Diakopoulos's 2017 Top Stories audit, 20 sources took more than half of impressions, the top 20% of 678 sources took 86%, and 83.5% of articles were under 24 hours old. The authors concluded that "organizations that can generate fresh copy may be more apt to have that material selected" (CJR Tow Center; ACM DL). A 2024 audit of the Google News tab in Brazil, the UK and the US found a Gini of 0.822, with 2.1% of outlets supplying half of results. Position-weighted CR10 was about 0.77–0.81, against an unweighted 0.22–0.36. Top-slot articles averaged 13.3 hours old, against 17.2 hours lower down (Hernandes & Corsi, arXiv). That gap between weighted and unweighted shares is the single most useful fact for Huduku: whoever holds the top positions holds most of the attention.

Personalisation is not the driver. In Nechushtai and Lewis's US study, liberals and conservatives saw the same Google News stories 99.9% of the time, and five organisations made up 69% of recommendations. Only 3 of the 14 dominant organisations were born-digital (JournalismAI summary). A UK replication found nine London legacy outlets supplying 75% of links, with little personalisation (Evans, Jackson & Murphy). Robertson and colleagues likewise found that exposure to partisan news on Google Search is driven "not primarily by algorithmic curation but by users' own choices" (Nature 2023). Concentration is therefore a property of the shared ranking. That is the part Huduku controls.

Design choices move the numbers. In Apple News, the algorithmic Trending section gave CNN alone 16.1% and the top three 45.2%. The editor-run Top Stories section gave its top source 9.8% and its top three 23.7%. Shannon equitability rose from 0.689 to 0.780 (p < 0.001) (Bandy & Diakopoulos, ICWSM 2020). Local news is a disputed case. Fischer, Jaidka and Lelkes found national outlets ranked above local ones even for local topics (Nature Human Behaviour). A Google author replied that the inequality measure was built wrongly, and that correcting it cut the effect by 56% on average (Magnusson 2022). The lesson is methodological: concentration figures depend on how the metric is built, so Huduku should publish its method alongside its numbers.

Production practice is thinner than the research. The only quantified anti-flooding rule a major aggregator has disclosed is Google's of June 2019: "usually" no more than two listings from one site in top results, with subdomains folded into the root domain, and exceptions when "especially relevant" (Search Engine Journal). It was reported to cover core web results, not Top Stories (Search Engine Roundtable). Google has also boosted "significant original reporting" (Google), and it now tells syndication partners to use noindex because it could not reliably find the original (SEJ). Ground News and SmartNews tackle balance with outlet ratings and left–right sliders, which Huduku's rules forbid (Ground News; TechCrunch). Ground News's hand-coded ownership for 2,200+ outlets is the nearest precedent for Huduku's owner facts.

The research names three biases, and Huduku answers each differently

The computer-science literature frames the problem as unfair exposure for providers. Burke separates fairness to consumers from fairness to providers, and says provider fairness can require going "beyond a purely personalization-oriented approach" (arXiv:1707.00093). Singh and Joachims model exposure with a position discount of 1/log(1+j), the share of readers who look at position j. Under that model small score differences become large exposure differences (arXiv:1802.07281). Morik and colleagues use a news example: a 2% relevance gap gave one group "vastly less exposure", and click feedback then compounds it (SIGIR 2020).

Popularity bias comes from feedback loops: items with more clicks get recommended more (Abdollahpouri et al.). Huduku has no click signal, so it has no such loop. It has a supply-side analogue instead. An owner with many titles looks like many voices unless counts go by owner, and ranking Top stories by owners carrying a story closes that. Volume flooding works through freshness and count: an owner that publishes continuously fills any recency-sorted list. Round-robin by owner plus a ceiling removes the reward for volume, and story grouping stops one agency copy filling several slots. That matches the standard defences against near-duplicate flooding, such as SimHash clustering and per-site caps (Manku et al., WWW 2007). Freshness bias is what the audits measured most directly. Huduku's decay of owners ÷ (1 + quiet hours ÷ 8) still rewards recency, but it multiplies it by breadth of coverage rather than letting the newest item win.

Each round 23 rule mapped to the evidence

Huduku ruleClosest research or standardAssessment
Count owners, not feedsGoogle folds subdomains into the root domain (SEJ); antitrust HHI aggregates firms (DOJ); TRAI's 2014 HHI is per entity, per language market (TRAI)Goes beyond any disclosed aggregator practice. It fails silently when ownership data is wrong (Kannada Prabha / Asianet).
One turn per owner before any gets a secondPM-2's election-style seat allocation, with every owner given equal votes (Dang & Croft); "open" diversity (Loecherbach et al.); the deliberative model's "equal representation" (Vrijenhoek et al.)Fair at every prefix of the list, like FA*IR (Zehlike et al.). It takes the equal-exposure side of a live debate.
Order within a round by weightPosition bias of 1/log(1+j) (Singh & Joachims); amortised attention (Biega et al.)Ties and near-ties keep favouring the same owners. The literature says to rotate or amortise.
At most 2 per owner per section (3 in the lead)Google's "usually no more than two per site"Close to the only public industry rule.
Soft ceiling of 10 cards per owner (~10%)Hard cap that relaxes, like Google's "unless especially relevant"Far tighter than legal ownership caps: Germany 30% (MStV § 60), UK 20% (Ofcom), TRAI about 32%. These are design norms, not legal analogues.
Weight = owners ÷ (1 + quiet hours ÷ 8)Breadth of coverage is a reflective agenda signal; the audits show freshness favours high-volume publishers (CJR)Uses merit that isn't volume, a middle path the literature supports. The 8-hour constant is a choice, not something the research derived.
Top stories by owners carrying a story; fronted by the reporter whose owner has fewest cardsReflective diversity; Google's original-reporting boost (Google)Sound. Fronting by room rather than by first reporter trades credit to the originator for balance. Google's own trouble with syndication shows that "first" is hard to know anyway (SEJ).
Audit: effective number of owners (1/HHI); fail past 15% for one ownerHHI, CR_k, Gini and Shannon equitability in audits (arXiv; ICWSM)Comparable with the literature, but unweighted. The audits show the top slots are where concentration bites.
Hourly fetch; future timestamps clamped to fetch timeGoogle: "Don't specify future dates" (Search Central)Clamping matches the guidance. A ~2.5-hourly reach for some feeds leaves a latency bias.
No ratings, no boosts, no personal signalsCoE: criteria "clear, non-discriminatory, viewpoint neutral, transparent, and objectively justifiable" (CoE 2021)The closest soft-law description of Huduku's approach.

Equal turns is a defensible choice, not a settled one

Singh and Joachims call equal average exposure across groups "demographic parity". They warn it "can cause substantial utility losses, particularly when groups differ significantly in relevance", and they favour exposure in proportion to merit (arXiv:1802.07281). Biega and colleagues take attention proportional to relevance as their target (arXiv:1805.01788). Saito and Joachims then show that a merit-proportional policy can leave some items worse off than a random ranking, and argue that no principled reason makes exposure linear in merit (KDD 2022). The ranking-fairness survey concludes that the choice is "contingent on the chosen value framework" (Zehlike et al.). News-diversity scholars add a democratic argument: "one actor should not be able to dominate public discourse" (Dagstuhl manifesto).

Huduku has a good answer to the merit objection. Output volume is not merit, and the audits show the advantage equal turns removes is production capacity. Huduku also keeps merit ordering inside each turn and uses owner breadth, which isn't volume, as its merit signal. The cost is real, though. When one owner breaks several big stories, its second-best item sits below another owner's marginal one, and a small owner with one weak item still gets a full turn. Huduku should state this trade openly in the guide rather than present balance as costless. The theory also limits what the rule can claim. One turn per owner is supply diversity at source level. It is not viewpoint diversity, since two owners can say the same thing, and it is not exposure diversity, which depends on what readers click (Loecherbach et al.). Being the same for everyone does remove the "fragmentation" risk by construction (Vrijenhoek et al.).

Position bias inside rounds is the clearest gap

Equal turns fixes how many cards each owner gets, not where they sit. Under the 1/log(1+j) model, the first card of a round gets more attention than the last (Singh & Joachims). Because Huduku orders owners within a round by weight, and breaks ties the same way each time, the same owners keep the better seats. Biega's answer is to amortise: no single ranking can be fair, but attention summed over many rankings can be (arXiv:1805.01788). FairCo does it online by adding λ × the accumulated exposure deficit to each item's score (Morik et al.). Diaz and colleagues make expected exposure the measure for rankings that vary (CIKM 2020). Huduku can do this without click data. Break ties by the owner's position-weighted exposure over the past day (the lowest goes first), or rotate the tie order with the edition stamp. Both stay deterministic for cached pages, so readers can still check them.

Ownership data and freshness are where guarantees leak

Grouping by owner is a group-fairness definition, and its unit is economic control. Re-ranking research groups providers by protected attributes or geography, and none of it groups by corporate owner (Gómez, Boratto et al.). That leaves Huduku without a tested method, and dependent on its owner table. The Kannada Prabha / Asianet gap is a direct failure of the guarantee: one group gets two turns and two ceilings. Ownership in India is hard to trace. The Media Ownership Monitor found that "cross-shareholdings obscure beneficial owners" and that there were no measurement standards (RSF), and the CMPF notes that concentration figures often rest on estimates (CMPF). TRAI's proposed control test of 20% equity or de facto control (TRAI 2014) and the Council of Europe's 5% disclosure threshold (CM/Rec(2018)1) give checkable rules for what counts as "one owner".

Freshness bias returns by three routes Huduku has not closed. First, feeds reached only every 2.5 hours lose to hourly ones under a recency-decayed weight, which is the latency advantage the audits found (arXiv). Second, search and topic pages run newest-first with no per-owner limit, so they work exactly like the recency lists the audits criticise. Third, outlets stuck in one section compete for a single section's two slots, so a structural feed artefact becomes an exposure penalty. None of these is a rule fault, but each undoes a rule's intent on some pages.

What professionals recommend next, in priority order

First, fix the inputs the rules depend on. Group Kannada Prabha with Asianet News Network, and record each owner link with its source and date, using a stated control test. Give non-news feeds their own fetch budget so every public feed is reached hourly. The audits show latency becomes exposure, and the ownership literature shows grouping errors fail silently.

Second, measure what the audits measure. Add position-weighted shares to check:balance, using 1/log(1+j) so the method is the standard one. Report CR3/CR5, the top-slot share, Shannon equitability over owners, and 1/HHI, both weighted and unweighted, for English and Kannada separately. Use the DOJ bands of 1,000 and 1,800 as a familiar reference, not a target (DOJ). RADio's rank-aware "representation" and "alternative voices" metrics fit a same-for-everyone page; its authors say they are "not to serve as thresholds" (arXiv:2209.13520). Run the check on a schedule, as the audits do: hourly scraping over weeks (Diakopoulos et al.), and Jaccard or rank-biased overlap for stability (arXiv). No standard threshold for "acceptable" concentration exists, so Huduku's 15% is its own choice and should say so.

Third, amortise position. Order owners within a round by accumulated exposure, or rotate the order, so the same owners don't keep the first seat.

Fourth, extend balance to the remaining recency lists. Apply owner turns, or at least a per-owner cap, to search and topic pages. Count "N reports" by distinct owners. Google's two-per-site rule shows a light cap is acceptable even in search.

Fifth, disclose the rules in plain language. Name each rule, its constants and its reason, and publish the per-owner counts. This meets the DSA's "main parameters" benchmark (DSA Observatory), the CoE's "how the criteria … are used ... and by whom or what" (CoE 2021), and Diakopoulos and Koliska's data, model, inference and interface layers (Digital Journalism 2017). Publishing counts, not judgements, fits "Huduku counts; readers judge".

The literature also holds lower-priority ideas: a coarse-and-fine fairness layer that balances outlets within an owner as well as owners (UMAP 2024), and checking a change against "dominance over uniform ranking", so no owner fares worse than under random order (Saito & Joachims).

India has the concentration but none of the audits

No published algorithmic audit of source concentration exists for Google News India, Dailyhunt, Inshorts or Indian search results. These reach many readers: Google News 53%, Dailyhunt 25% and Inshorts 19% among India's mainly English-speaking online sample (Reuters Institute DNR 2022). The supply they draw on is highly concentrated. Four Hindi dailies hold 76.45% of Hindi readership, the top two papers in each regional language hold over 50%, and Indian law has no concentration thresholds (RSF Media Ownership Monitor). An aggregator that mirrors that supply reproduces it, which makes owner-balancing more important in Kannada than in English. Platform power adds to this. The CCI found that publishers "have no choice but to accept the terms and conditions imposed by Google" (TechCrunch). Google News Initiative programmes favoured legacy publishers (Agarwal, IJoC 2025).

The nearest Indian standard is TRAI's 2014 recommendation. It measures news concentration by HHI in 12 language markets, Kannada among them, with HHI above 1,800 counting as concentrated and a 1,000-point limit on any one entity's contribution. It was never enacted, and TRAI reopened the question in 2022 (TRAI 2014; TRAI 2022). Its unit, per owner per language market, is the one Huduku already uses, so Huduku's numbers can be read against it. The legal position is unsettled. The IT Rules 2021 name "news aggregators" as publishers bound by a Code of Ethics, and leave "curation" undefined (MeitY). The Bombay and Madras High Courts stayed Rules 9(1) and 9(3) in 2021 (IFF; Tribune), and a draft amendment of March 2026 would widen Part III (IFEX). The current status of the stays, the Supreme Court petitions and the draft could not be confirmed. Whether Huduku counts as a "news aggregator" under the Rules is a question for its pending legal review. No Indian rule requires or forbids owner-balancing.

Several things remain unknown. No study tests whether readers see equal-per-owner pages as lower quality, and no CS paper evaluates round-robin by owner on news. The EU's binding rules (DSA Arts. 27, 34–35 and 38; EMFA from 8 August 2025) reach large platforms, not Huduku (EUR-Lex). They serve as benchmarks, not obligations.

Conclusion

The research turns Huduku's balance question from "is it fair?" into "fair by which norm, where on the page, and on what data?" On the norm, Huduku has chosen open, equal-per-owner diversity, with breadth of coverage as its merit signal. That is a coherent, citable position, provided the guide states the trade-off. On the page, the open risk has moved from how many cards an owner gets to where its cards sit and to the pages balance doesn't yet reach. On the data, the guarantee is only as true as the owner table and only as even as the fetch schedule.

Huduku's own counts could also do something rare. India has no published aggregator audit, and Kannada has the kind of supply concentration TRAI wanted to measure. A regular, position-weighted, owner-level count of Huduku's pages, published with its method, would be a small public data series on both. It would also give the open questions (equal against proportional exposure, rotation, the 15% line) evidence that is now missing.

Corrected since

  • Diakopoulos and colleagues, “I Vote For”: The data ran from 31 May to 8 July 2016, mostly after the primaries had ended, not through them.
  • Trielli and Diakopoulos: The 83.5% is a share of impressions, not of articles.
  • Hernandes and Corsi: The unweighted top-ten share runs from 21% to 36%, not 22% to 36%. The 17.2 hours is the average from sixth place down; places two to five averaged 14.0 hours.
  • Nechushtai and Lewis: The 69% is an average of the top five organisations in each search, not a share of all recommendations.
  • Evans, Jackson and Murphy: A similar study in the UK, not a replication.
  • Robertson and colleagues, 2023: The Springer link in the notes goes to a login page; the paper is open at nature.com.
  • Fischer, Jaidka and Lelkes: National outlets dominated results unless people searched specifically for topics of local interest; the notes say local outlets ranked lower even for local topics.
  • Magnusson: Correcting the inequality measure lowered the measure by 56% on average; the report says it cut the effect.
  • Search Engine Roundtable: Google’s site-diversity change was reported to cover only the main web listings, not other search features; the article doesn’t mention Top Stories.
  • Search Engine Journal: A site’s subdomains are generally counted as the site, not always.
  • Ground News; SmartNews: Ground News’s left-to-right ratings come from third parties; SmartNews’s are its editors’.
  • Abdollahpouri and colleagues: The study is of ratings, not clicks, and describes no feedback loop: items that already have more ratings get recommended more.
  • Singh and Joachims: They offer exposure in proportion to merit as an alternative to equal exposure.
  • Morik and colleagues: The 2% relevance gap and the click feedback are separate results. FairCo adds each group’s accumulated shortfall in exposure, measured against its merit.
  • Manku and colleagues: The paper is about detecting near-duplicates, not clustering them.
  • Zehlike and colleagues, FA*IR: The anthology link is dead; the paper is on arXiv. FA*IR keeps a minimum share for a protected group at every point in the list.
  • Gómez, Boratto and colleagues: Fairness research groups providers by traits such as gender, age or where they come from.
  • CMPF: It says the data needed to measure concentration are often missing or unreliable; it doesn’t say the figures rest on estimates.
  • TRAI, 2014; RSF: The 1,000-point limit applies where both the TV and the newspaper markets are concentrated, across both. That the recommendation was “never enacted” couldn’t be confirmed; what is sourced is that Indian law still sets no such thresholds.
  • Medienstaatsvertrag, § 60: The gesetze-bayern.de link failed (503); the official text is the media authorities’ copy. The law presumes dominant power over opinion at a 30% audience share; it isn’t a cap.
  • Ofcom, legal framework: The page refused every fetch (403), so the UK’s 20% couldn’t be checked and was left out of the research page.
  • Council of Europe, 2021: The five criteria are for deciding what counts as public-interest content to give prominence to, not for ranking in general. The site refused the fetch; the wording was read in two published copies.
  • Council of Europe, 2018: The 5% disclosure threshold couldn’t be checked (403) and was left out.
  • India’s competition regulator (TechCrunch): The CCI said it appeared that publishers had no choice, in ordering an investigation; it was a first view, not a finding.
  • IFF; The Tribune: IFF reports the Bombay High Court’s stay; the Madras High Court’s is in The Tribune.
  • IFEX; SFLC.in: IFEX refused the fetch; SFLC.in publishes the same statement.
  • European Media Freedom Act; DSA Observatory: The Act has no rules on recommender systems. Its Article 18 covers how very large platforms treat media content; explaining ranking is the Digital Services Act’s Article 27.

Every correction