Aggregator audits and practice: source concentration and publisher flooding
Notes compiled 4 October 2026. Every claim has a link. Items I could not verify from a fetched source are listed under Gaps rather than stated as fact.
1. What have audits found about source concentration in aggregators?
Takeaway
Every audit of a general news aggregator or search news box that I could check found strong head-concentration. A handful of large legacy outlets take a third to over half of slots. Top positions are more concentrated than the list as a whole. Fresh content is strongly favoured. Personalisation adds little. The concentration comes from ranking and recency, not from tailoring to users. Human-curated sections (Apple News Top Stories) were measurably less concentrated than algorithmic ones (Apple News Trending).
Cited Findings
Google "In the News" / Top Stories (US)
- 2016 US primaries (Diakopoulos, Trielli, Stark, Mussenden, "I Vote For"): hourly collection, 31 May to 8 July 2016, 5,604 links (972 unique articles). CNN and the New York Times held 2,476 links (44.2%) and 30% of unique articles. In the first, most prominent slot they held 1,211 of 1,868 (64.4%). 60 sources appeared nine times or fewer. 30.5% of links were marked under three hours old — Academia.edu
- Trielli & Diakopoulos, "Search as News Curator", CHI 2019: 200+ news queries from Google Trends, scraped once a minute for 24 hours, November 2017. "Just 20 news sources account for more than half of article impressions." The top 20% of sources (136 of 678) took 86% of impressions. CNN, NYT and Washington Post together took 23%. On single queries it was sharper: for "rex tillerson", two sources took 75%. 83.5% of articles were under 24 hours old and 13.1% under an hour. The authors concluded that "organizations that can generate fresh copy may be more apt to have that material selected" — CJR Tow Center; paper: ACM DL, NSF PAR
Google News (US, personalisation)
- Nechushtai & Lewis, Computers in Human Behavior 90 (2019): 168 participants searched Google News on their own accounts for Clinton and Trump in 2016 and reported the top five stories. The five most-recommended organisations made up 69% of recommendations ("Five news organizations alone accounted for 49%" in another cut). Liberals and conservatives saw the same stories 99.9% of the time. Only 3 of the 14 dominant organisations were born-digital, so the algorithm reproduced the legacy structure — JournalismAI summary; ResearchGate
- Haim, Graefe & Brosius, "Burst of the Filter Bubble? Effects of personalization on the diversity of Google News", Digital Journalism 6(3), 330–343 (2018), DOI 10.1080/21670811.2017.1338145 — citation. I could not fetch the abstract (403); see Gaps.
Google News (UK)
- Evans, Jackson & Murphy, Digital Journalism 11(9) (2023): 78 participants searched four terms in a 9-hour window on 20 March 2019. The top five outlets made up 52% of 775 recommendations. Nine outlets, all London-based legacy media, made up 75% of links. Evidence of personalisation was limited — Bournemouth eprints PDF; T&F
Google News tab (Brazil, UK, US)
- Hernandes & Corsi (Cambridge LCFI), arXiv 2024: Selenium scraping of the News tab for Google Trends keywords, 11 March to 1 April 2024. 143,976 results, 2,298 queries, 4,296 outlets, plus a validation set (29 April to 11 May 2024, 77,887 results). Unweighted CR4 / CR10 were Brazil 0.163 / 0.287, UK 0.209 / 0.358 and US 0.118 / 0.215. Position-weighted CR10 was about 0.77–0.81. The overall Gini was 0.822, with "2.1% of outlets account for 50% of results". Top-position articles averaged 13.3 hours old, against 17.2 hours at position 6 and below. There was a slight leftward skew (mean time-averaged output bias −0.05 on politics) — arXiv 2410.23842
Google News local vs national (US)
- Fischer, Jaidka & Lelkes, Nature Human Behaviour (2020): searches on 32 topics, without location and then with each of the 3,141 US counties. Local outlets were systematically ranked below national ones, even for local topics. Lelkes said: "I don't think Google News intentionally engineered an algorithm to help the New York Times over the South Jersey Times" — Nature; Annenberg
- Contested. Magnusson (Google LLC), "Local news in Google News", NHB, 11 August 2022, argues the inequality measure was built wrongly. Correcting it "reduces their measure on average 56% and up to 86% for some topics". He also says "the vast majority of even the first few results for local topics are from local outlets" — Nature. The authors replied — Nature reply (not fetched).
Google Search partisanship (audit vs user choice)
- Robertson et al., "Auditing Partisan Audience Bias within Google Search", PACM HCI / CSCW 2018: 187 participants, a browser extension, 21 root queries expanded to 549 by autocomplete, January–February 2017, 15,337 SERP pairs. It found "little evidence for the 'filter bubble' hypothesis". Results leaned slightly left overall, and ranking shifted them slightly right (0.02). Bias varied by component; Twitter cards leaned right for 17 of 21 root queries — PDF
- Robertson et al., Nature (2023), 2018 and 2020 waves: "Exposure to and engagement with partisan or unreliable news on Google Search are driven not primarily by algorithmic curation but by users' own choices" — Nature
Apple News (editors vs algorithm)
- Bandy & Diakopoulos, "Auditing News Curation Systems: A Case Study Examining Algorithmic and Editorial Logic in Apple News", ICWSM 2020. They collected data by automation (Appium), 9 March to 9 May 2019. Trending (algorithmic) had 83 sources: CNN alone held 16.1% (505 stories) and the top three held 45.2%. Top Stories (editors) had 87 sources: the top source held 9.8% and the top three 23.7%. Shannon equitability was 0.689 for Trending and 0.780 for Top Stories (Hutcheson t = 11.17, p < 0.001). Trending carried 50.7 stories a day with hourly churn; Top Stories carried 20.4, updated at set times. Trending leaned to celebrity stories — ICWSM PDF; scraper code
- The brief's "Kawakami et al." Apple News audit appears to be this Bandy & Diakopoulos paper. I found no separate Kawakami Apple News audit.
Germany
- Unkel & Haim, "Googling Politics" (Social Science Computer Review 39(5), 2021), on the 2017 Bundestag campaign: parties, sources and issue ownership on Google — SAGE (not fetched in detail).
India-specific evidence (separate)
- I found no published algorithmic audit of source concentration in Google News India, Dailyhunt, Inshorts or Indian search results.
- Aggregator reach, from the Reuters Institute Digital News Report 2022, India (mainly English-speaking online users, not nationally representative): "Google News (53%), Daily Hunt (25%), InShorts (19%), and NewsPoint (17%)". 72% get news on mobile — RISJ 2022 India
- DNR 2025, India: YouTube 55%, WhatsApp 46%, Instagram 37%, Facebook 36% for news. Trust is 43%. 53% name WhatsApp as the biggest misinformation threat. The country page I fetched did not give aggregator figures — RISJ 2025 India
- Platform power over Indian publishers: the CCI opened an investigation in January 2022 on a complaint from the Digital News Publishers Association. It found that "news publishers have no choice but to accept the terms and conditions imposed by Google". Over half of members' traffic came from search — TechCrunch
- Agarwal, "Rooting Platform Dependencies in the Digital News Economy: Google News Initiative in India", IJoC 19 (2025): 16 GNI India programmes (2018–21), 36 documents and 20 interviews. Legacy publishers got News Showcase (India, May 2021 per the HT report below; the paper summary says 2020). Digital natives competed for short grants. The paper argues GNI reinforces incumbents — IJoC PDF; HT on Showcase India
Inferences
- The pattern repeats across 2016–2024, the US, UK and Brazil, and Google and Apple: CR3 of about 23–45% for list-level shares and Gini of about 0.8. The top slot is more concentrated still (64% for two outlets in 2016). This is a structural tendency of recency- and authority-ranked lists, not a one-off.
- Personalisation is not the main driver. The concentration is the same for everyone, so a non-personalised design (like Huduku's) does not by itself avoid it. Explicit balancing is needed.
- Apple's editors vs algorithm comparison is the cleanest evidence that a design choice moves concentration measurably (equitability 0.78 vs 0.69).
Gaps
- Haim, Graefe & Brosius (2018) and Puschmann (2019, Google search during the German election, data donation): I could not fetch the abstracts or numbers. From background knowledge both report little personalisation, but that is not verified here.
- No Yahoo News or Microsoft Start audit of source concentration found.
- No academic Google Discover concentration audit found (only industry traffic reports).
- No Indian aggregator audit found. Media Ownership Monitor India (RSF, 2019) concentration indicators failed to load — MoM India.
2. Audit methods and metrics: what is standard, and are there thresholds?
Takeaway
The methods follow Sandvig et al.'s 2014 taxonomy. Scraping audits run at high frequency (minute-level) for news boxes. Crowdsourced or browser-extension audits are used for personalisation. Concentration is most often reported as top-k share (CR_k), then Gini, HHI and Shannon equitability. The only established numeric thresholds are antitrust HHI bands. The diversity-metric literature (RADio) explicitly declines to set thresholds.
Cited Findings
- Sandvig, Hamilton, Karahalios & Langbort (2014) set out five audit designs:
- code audit;
- non-invasive user audit (surveys or consented data);
- scraping audit;
- sock-puppet audit (programmed fake users);
- crowdsourced or collaborative audit (paid or volunteer testers).
- Designs used in the news audits above:
- minute-level scraping over 24 hours from a fixed location, to cancel noise (Trielli & Diakopoulos) — CJR;
- hourly scraping for 5 weeks (2016 primaries) — Academia;
- app automation with Appium (Apple News) — GitHub;
- crowd participants on real accounts (Nechushtai & Lewis; Evans et al.);
- a browser-extension crowd audit with incognito pairs (Robertson 2018) — PDF;
- location-spoofed searches across all counties (Fischer et al.) — Annenberg.
- Metrics in use:
- top-k share and "X% of sources take Y% of impressions" (Trielli & Diakopoulos);
- Shannon equitability, tested with Hutcheson's t (Bandy & Diakopoulos) — ICWSM;
- HHI, Gini, CR4/CR8/CR10, position-weighted shares, and Jaccard and rank-biased overlap for stability over time (Hernandes & Corsi) — arXiv;
- inequality measures by topic (Fischer et al.). Their construction was disputed, which shows these metrics are sensitive to how they are built — Magnusson 2022.
- HHI thresholds (US DOJ/FTC, 2023 Merger Guidelines): 1,000–1,800 is "moderately concentrated" and above 1,800 is "highly concentrated". An increase of more than 100 in a highly concentrated market is presumed to add market power — DOJ
- RADio (Vrijenhoek et al., RecSys 2022) gives five normative metrics: calibration, fragmentation, activation, representation and alternative voices. Each uses Jensen–Shannon divergence, bounded 0–1. The authors critique intra-list diversity as "the opposite of similarity". They state the metrics are "not to serve as thresholds or strict guidelines" — arXiv 2209.13520; ACM
Inferences
- For comparability with the literature, a project measuring its own pages could report CR3/CR5, the share of the top slot, Gini or Shannon equitability over owners, and HHI with the DOJ bands as a reference point. Each should be computed both unweighted and position-weighted, because the top slot is where concentration bites.
- An HHI computed over owners (not domains) maps directly onto the antitrust convention, since antitrust aggregates firms.
Gaps
- No standard or recommended threshold for "acceptable" source concentration in a news feed was found. Beyond antitrust HHI, any threshold would be the designer's own choice.
- I did not find a published effective-number-of-sources (1/HHI) convention for news feeds, though it follows directly from HHI.
3. What have aggregators changed and disclosed? (Including freshness, date manipulation and flooding)
Takeaway
Google has publicly capped same-site results (typically two, with subdomains folded into the root domain, since June 2019). It has boosted original reporting (September 2019), pushed syndication partners to noindex (2023), added spam policies against scaled content (March 2024), and changed Discover to cut clickbait and favour local, topic-expert sites (February 2026). Apple uses human editors explicitly to avoid algorithmic failure. Ground News, AllSides and SmartNews address viewpoint balance with outlet ratings, not volume caps. None of them disclosed a per-owner cap like Huduku's. Google's grouping by root domain (not owner) is the nearest public precedent.
Cited Findings
Google: site diversity (June 2019; the brief's "September" date is wrong)
- Announced and live by about 3–6 June 2019: "you usually won't see more than two listings from the same site in our top results". The exception is "when our systems determine it's especially relevant". Google "will generally treat subdomains as part of a root domain", but subdomains count as separate "when deemed relevant". Danny Sullivan said it was "not really about ranking" — Search Engine Journal; Search Engine Land; Engadget
- Coverage I found did not say whether this applied to Top Stories. Grouping is by domain, not by corporate owner — SEJ
Google: original reporting (12 September 2019)
- Richard Gingras: ranking updates to "highlight significant original reporting", which "may stay in a highly visible position longer". The rater guidelines were updated: §5.1 gives the "very high quality" rating to original news, and §2.6.1 says Pulitzers or "a history of high quality original reporting" are evidence of reputation. Google said "there is no absolute definition of original reporting" — Google blog
- A week later, syndicated copies (Yahoo News reprints of Variety, Footwear News and The Blast) were still outranking the originals in Top Stories. Sullivan: "If people deliberately chose to syndicate their content, it makes it difficult to identify the originating source. That's why we recommend the use of canonical or blocking" — SEJ, 18 September 2019
- From July 2023 Google recommended that syndication partners use noindex rather than canonical, because syndicated pages "can differ from the original content with lots of other material on the page" — SEJ, 11 July 2023; SER
Google: dates and freshness
- Google says its systems favour fresh content: 83.5% of Top Stories articles were under 24 hours old (2017), and top positions averaged 13.3 hours old against 17.2 lower down (2024) — CJR; arXiv
- Google's byline-date documentation: "Don't specify future dates, or the date of the action described on the page. The dates must describe the publication or update date of the page" — Google Search Central. The March 2019 post "Help Google Search know the best date for your web page" exists — Search Central blog — but my fetch returned only a generic summary, so I don't quote it.
- John Mueller (Google) said changing dates on pages won't improve rankings — SEJ (headline only, not fetched).
Google: flooding and scaled content
- In March 2024 Google announced a core update and new spam policies (scaled content abuse, site reputation abuse, expired domain abuse) — Search Central blog. My fetch didn't return the text; secondary coverage reports a stated aim of cutting unhelpful content by 40% — SEJ (headline).
Google: Discover (5 February 2026)
- Google is "reducing sensational content and clickbait", prioritising "locally relevant content from websites based in the user's country", and judging expertise "on a topic-by-topic basis" within a site — PPC Land
Google News Initiative / Showcase
- GNI India and Showcase favoured legacy publishers, according to Agarwal (2025) — IJoC
Apple News
- Human curation is the stated policy. Editor-in-chief Lauren Kern: "We're so much more subtly following the news cycle and what's important". Roger Rosner: "We're not just going to let it be a total crazy land" (October 2018) — 9to5Mac. The audit confirmed lower concentration in the editor-run section — ICWSM
Ground News
- Bias ratings come from AllSides, Ad Fontes and Media Bias Fact Check, plus a "factuality" score. Ownership is hand-coded for 2,200+ outlets. "Blindspot" marks stories covered mostly by one side. It groups ~60,000 articles a day from 50,000+ sources into stories — Ground News about
- Huduku's rules forbid bias or reliability ratings, so only Ground News's ownership display and story grouping are relevant precedents.
SmartNews
- "News From All Sides" (September 2019, US): a slider over five left-to-right buckets. Editors classify outlets by hand, and an algorithm picks the headlines. No mechanism against single-publisher domination was described — TechCrunch
AllSides
- Presents left, centre and right headlines side by side using its own outlet ratings; it rates SmartNews "Lean Left" — AllSides (not fetched in detail).
Inferences
- The only public, quantified anti-flooding rule from a major aggregator is Google's "usually no more than two per site, subdomains folded into root". Grouping by owner (sister brands on different domains) goes beyond anything Google has disclosed.
- Every audit found that recency-driven ranking advantages high-volume, fast publishers, which the audit authors state outright. A design that decays weight by quiet hours, or caps per owner, directly answers the mechanism the audits identified.
- Syndication and agency copy undermine "original first". Google's own fix moved from canonical to noindex. That shows aggregators can't reliably identify the originator without publisher cooperation, so near-verbatim detection (as Huduku's
sameWordingdoes) is a reasonable independent signal.
Gaps
- Flipboard, Microsoft Start, Yahoo News and SmartNews: I found no public disclosures of per-publisher caps or source-diversity rules.
- No peer-reviewed study was found that quantifies timestamp manipulation (re-dating) by news publishers, or that measures how feed latency affects aggregator share. The evidence is Google guidance and audit observations on recency.
- No academic study was found on content farms or churnalism flooding aggregators specifically (as opposed to search generally), beyond the Yahoo syndication case.
- Dailyhunt and Inshorts curation and publisher policies were not found; the Hoot article on them didn't load — The Hoot.
Corrected since
- Diakopoulos and colleagues, “I Vote For”: The data ran from 31 May to 8 July 2016, mostly after the primaries had ended, not through them.
- Trielli and Diakopoulos: The 83.5% is a share of impressions, not of articles.
- Nechushtai and Lewis: The 69% is an average of the top five organisations in each search, not a share of all recommendations.
- Robertson and colleagues, 2023: The Springer link in the notes goes to a login page; the paper is open at nature.com.
- Fischer, Jaidka and Lelkes: National outlets dominated results unless people searched specifically for topics of local interest; the notes say local outlets ranked lower even for local topics.
- Search Engine Journal: A site’s subdomains are generally counted as the site, not always.
- India’s competition regulator (TechCrunch): The CCI said it appeared that publishers had no choice, in ordering an investigation; it was a first view, not a finding.