Provider-side fairness and source concentration in ranking and recommender systems, as applied to news aggregation
Scope note: these notes cover the CS literature on popularity/exposure bias and provider fairness, the main mitigation techniques, and how they relate to a non-personalised news front page (same list for every reader) that uses "one turn per owner before any gets a second, with a soft per-owner ceiling". Verification level is marked where a primary text could not be fetched (several ACM DL pages returned 403).
1. What popularity bias, exposure bias and provider-side unfairness are; the key papers
Takeaway
Ranked lists concentrate attention because users examine top positions far more than lower ones, so small differences in score become large differences in exposure; and feedback loops (rich-get-richer) amplify this over time. The literature frames the harm as unfairness to providers (P-fairness, Burke 2017) and formalises it as exposure allocation (Singh & Joachims 2018; Biega et al. 2018).
Cited Findings
- Multisided fairness (Burke, FATML 2017). Distinguishes C-fairness (consumers), P-fairness (providers) and CP-fairness (both). P-fairness applies when "fairness needs to be preserved for the providers only", with Kiva.org's microloan field partners as the example; it notes the need to "go beyond a purely personalization-oriented approach" when fairness demands it, and sketches an auction/budget mechanism giving protected provider groups equivalent "purchasing power". — Burke, arXiv:1707.00093 (HTML)
- Later framing of the multi-stakeholder problem: Sonboli, Burke, Ekstrand, Mehrotra, "The multisided complexity of fairness in recommender systems", AI Magazine 2022. — Wiley; PDF
- Popularity bias (Abdollahpouri, Burke, Mobasher). Defined as "popular items are recommended frequently while less popular, niche products, are recommended rarely or not at all"; arises because collaborative filtering favours items with more ratings over long-tail items. Published at AAAI/ACM AIES, 2019 ("Managing Popularity Bias in Recommender Systems with Personalized Re-ranking" / "Popularity Bias in Ranking and Recommendation"). — arXiv:1901.07555; ACM DL AIES 2019. Earlier regularisation approach: "Controlling Popularity Bias in Learning-to-Rank Recommendation" (RecSys 2017). — Semantic Scholar
- Fairness of Exposure in Rankings (Singh & Joachims, KDD 2018). Frames fairness as constraints on how exposure is allocated, maximising user utility subject to the constraint. Exposure uses a position-bias model v(j) = 1/log(1+j), consistent with DCG, i.e. "the fraction of users who examine the document shown at a particular position". The authors state fairness requires "trade-offs between the utility of the users and the rights of the items" and quantify a "Cost of Fairness". — arXiv:1802.07281
- Equity of Attention (Biega, Gummadi, Weikum, SIGIR 2018). Individual (per-subject) fairness: attention should be proportional to relevance. Since no single ranking can be individually fair (positions are discrete and attention falls steeply), they amortise: "attention accumulated across a series of rankings is proportional to accumulated relevance", solved online with integer linear programming against a ranking-quality constraint; experiments show real rankings are substantially unfair and the method improves fairness while keeping competitive quality. — arXiv:1805.01788; ACM DL; PDF
- Rich-get-richer dynamics (Morik, Singh, Hong, Joachims, SIGIR 2020, best paper). "Items ranked highly are more likely to collect additional feedback, which in turn can influence future rankings". Uses a news example: with left- and right-leaning article groups, a modest ~2% relevance difference under pure relevance ranking gives one group "vastly less exposure". Proposes FairCo, a proportional controller ranking by relevance + λ·(accumulated exposure error), amortised over time; reported to cut unfairness while keeping quality and to be robust to initial conditions and group-size asymmetry. — PDF (Cornell); SIGIR 2020 awards
- Expected exposure (Diaz, Mitra, Ekstrand, Biega, Carterette, CIKM 2020). Evaluates stochastic rankings by expected exposure, the principle that equally relevant items should get equal expected exposure. — ACM DL; Boise State
- Surveys. Zehlike, Yang & Stoyanovich, "Fairness in Ranking: A Survey", ACM Computing Surveys 55(5–6), 2022 (published in two parts: Part I score-based ranking, Part II learning-to-rank and recommender systems). It systematises interventions by the "value frameworks" that motivate them, stressing there is no single correct definition. — arXiv:2103.14000; Part I, ACM DL. Ekstrand, Das, Burke & Diaz, "Fairness in Information Access Systems", Foundations and Trends in Information Retrieval (2022): a taxonomy of fairness dimensions for search and recommendation, emphasising multi-stakeholder settings, ranking, personalisation and interaction. — arXiv:2105.05779; now publishers
- Real-world source concentration in news aggregation. Trielli & Diakopoulos's audit of Google Top Stories (queries scraped November 2017; CHI 2019 "Search as News Curator"): "just 20 news sources account for more than half of article impressions"; the top 20% of sources (136 of 678) got 86% of impressions; the top three (CNN, NYT, Washington Post) got 23%; 83.5% of articles were under 24 hours old, so freshness is a big factor, which favours outlets able to publish continuously. — CJR / Tow Center, 2019; ACM DL CHI 2019 (403 on fetch; numbers from the CJR write-up)
- A 2024 audit of Google search news results in Brazil, the UK and the US (221,863 results) used HHI and Gini and found "significant concentration trends" with "a preference for popular, often national outlets" that may "reinforce existing media inequalities"; it argues concentration has to be measured horizontally across many queries, since a single results page hides it. — arXiv:2410.23842
Inferences
- In a non-personalised edition, the volume effect works mainly through freshness and count: an owner with many feeds that publish continuously fills recency-sorted lists, much like the freshness effect behind the Google Top Stories audit. It needs no collaborative-filtering popularity loop.
- The position-bias model (1/log(1+j)) is a useful way to state, as a count, how much more exposure the top card gets than card 10. It applies to any ranked list, personalised or not.
- Rich-get-richer feedback loops (Morik et al.) need click feedback. A system that ranks by "how many outlets carried a story" with no click signal avoids this loop, but it has a supply-side analogue: owners with more outlets are counted more unless counts are by owner.
Gaps
- Could not fetch the full CHI 2019 paper (403); impression numbers are taken from the CJR summary by the lab's authors.
- The authors and venue of arXiv:2410.23842 were not shown in the fetched abstract page.
2. Mitigation techniques, trade-offs, failure modes; comparison with owner round-robin plus a soft ceiling
Takeaway
The main families are (a) hard per-source caps, (b) diversification re-ranking (MMR, xQuAD, PM-2 proportionality), (c) constrained fair-exposure re-ranking (FA*IR, Singh & Joachims LP), (d) amortised or controller-based fairness over time (Biega; FairCo), and (e) calibration (Steck). An owner round-robin with a soft ceiling sits closest to a hard cap combined with a proportional seat-allocation scheme in which every owner's "vote share" is equal. It is a strong, transparent equal-exposure intervention, and it gives up merit proportionality.
Cited Findings
- Per-source caps in production. In June 2019 Google said it "generally won't show more than two listings from the same site in our top results", with exceptions "where our systems determine it's especially relevant"; subdomains generally count with the root domain; applied to core web results only, not Top Stories, featured snippets and other features. — Search Engine Roundtable; Search Engine Journal
- MMR (Carbonell & Goldstein, SIGIR 1998). Greedy re-ranking: each next item maximises λ·relevance − (1−λ)·max similarity to items already chosen; the standard redundancy-penalty approach. — ACM DL; dblp
- xQuAD (Santos, Macdonald & Ounis, WWW 2010). Explicit aspect-coverage diversification: greedily picks items that cover aspects not yet covered. — ACM DL. Abdollahpouri et al. adapted xQuAD with "aspects" = short-head vs long-tail: score P(v|u) + λP(v,S'|u), with λ tuned "to achieve the desired trade-off between accuracy and better coverage of long-tail, less popular items"; it outperformed a regularisation baseline while "maintaining acceptable recommendation accuracy". — arXiv:1901.07555
- PM-2, "Diversity by proportionality: an election-based approach to search result diversification" (Dang & Croft, SIGIR 2012). Treats result positions as seats and aspects as parties, filling them with the Sainte-Laguë highest-quotient method so each aspect's share of the list is proportional to its weight. — ACM DL 10.1145/2348283.2348296 (403 on fetch; the description follows the paper's title and its standard account; the Sainte-Laguë method itself: Wikipedia)
- FA*IR (Zehlike et al., CIKM 2017). A post-processing top-k algorithm enforcing "ranked group fairness": at every prefix of the list, the share of the protected group must not fall statistically below a target proportion p (a binomial test with multiple-testing adjustment), while otherwise keeping score order. — IR Anthology; code
- Fair exposure via optimisation. Singh & Joachims solve a linear programme over doubly stochastic ranking matrices (sampled as a distribution over rankings) to maximise utility subject to exposure constraints. — arXiv:1802.07281
- Amortised and online control. Biega et al. amortise over a sequence of rankings with an ILP — arXiv:1805.01788; FairCo adds λ·exposure-deficit to relevance at each step — Morik et al. 2020.
- Calibration (Steck, RecSys 2018, Netflix). Greedy re-ranking to make the genre distribution of a user's list match the distribution in their history, minimising KL divergence with a λ trade-off against relevance; it is personalised by design. — ACM DL (403 on fetch; description from the paper's well-known formulation)
- Provider-group re-ranking with personalised tolerance. OFAiR (Sonboli, Eskandanian, Burke, Liu, Mobasher, UMAP 2020) weights protected provider groups per user according to their tolerance for diversity; reported 16–41% more protected-item exposure than baselines (FAR/PFAR) at equal accuracy loss, on loan and movie data. — arXiv:2005.12974. Tooling: librec-auto (RecSys 2020 demo). — ACM DL
- A known failure mode of exposure fairness. Saito & Joachims (KDD 2022) show that an exposure-proportional policy can leave some items worse off than a uniformly random ranking; they argue there is no principled reason exposure should be linear in merit, and propose envy-freeness and "dominance over uniform ranking" as axioms, satisfied by maximising Nash social welfare (the product of item impacts). — Cornell page; ACM DL
- Tunable λ in every re-ranker. MMR, xQuAD and calibration all trade relevance against diversity or fairness through one parameter (sources above). The Cost of Fairness is measured as the utility lost. — Singh & Joachims
Inferences (comparison with "one turn per owner before any gets a second, soft per-owner ceiling past 10 cards")
- What it is in this literature's terms. It is a group-level, equal-share proportional allocation. In PM-2's language, every owner-party has the same vote weight, so seats are dealt in rotation; ordering within an owner's turn still follows merit (story size, recency). The soft ceiling is a hard cap that relaxes when no other owner has a candidate, close to Google's "two per site, unless especially relevant".
- Strengths for a non-personalised page. It is deterministic, explainable and checkable by readers ("every owner gets a turn"). It needs no relevance model or click data, so it has no feedback loop. It enforces fairness at every prefix of the list, the property FA*IR enforces for a protected group, rather than only on average. And it cannot be bought by adding feeds, the attack the volume literature worries about.
- Trade-offs and failure modes the literature predicts:
- It is demographic-parity-style (equal exposure regardless of merit). Singh & Joachims note this can cost a lot of utility when groups differ greatly in relevance. Here: when one owner genuinely broke many stories, an equal-turn rule pushes its lesser items below others' marginal items.
- Single-ranking position effects remain. Even with equal turns, whoever goes first in each round gets more attention (1/log(1+j)). Biega et al.'s answer is to amortise: rotate the order of owners over time or across refreshes, so that accumulated attention is even. A fixed tie-break order (alphabetical, by outlet count) would give one owner a lasting position advantage.
- Small owners with one marginal item get a full turn. This is the "dominance over uniform" question in reverse. Saito & Joachims suggest judging fairness by impact, not raw slots.
- Group definition decides the outcome. Counting by owner rather than by feed is the decisive design choice: a pure per-feed round-robin is exactly what volume flooding exploits.
- MMR/xQuAD versus round-robin. MMR penalises redundant content, not over-represented sources, so it would not stop one owner filling a list with distinct stories. xQuAD or PM-2 with owner as the aspect is the closest formal analogue of the round-robin.
- Calibration (Steck) and OFAiR are personalised; their user-specific parts do not transfer to a same-for-everyone page, but their re-ranking mechanics do, with a fixed target distribution.
Gaps
- No paper found that evaluates exactly "round-robin by owner plus a soft cap" on news; the comparison above is inferred from the general frameworks.
- The full texts of PM-2 and Steck could not be fetched (403), so no specific numbers from them are given.
3. Equal exposure versus exposure proportional to merit: the debate
Takeaway
Yes, it is an explicit and unresolved debate. Singh & Joachims and Biega et al. favour merit-proportional exposure (disparate treatment / equity of attention) and note that demographic parity can be costly; Saito & Joachims question whether "proportional" itself has any principled basis; the surveys treat the choice as value-laden. News-diversity scholars add a normative, democratic argument for deliberately not following merit or popularity alone.
Cited Findings
- Singh & Joachims define three constraints. Demographic parity: equal average exposure across groups, which they say "can cause substantial utility losses, particularly when groups differ significantly in relevance". Disparate treatment: exposure "proportional to its average utility" (Disparate Treatment Ratio = 1 when satisfied). Disparate impact: expected click-through proportional to average utility, with P(click) = exposure × relevance. — arXiv:1802.07281
- Their motivating point for merit-based fairness: under pure relevance ranking, small relevance differences turn into large exposure differences. — arXiv:1802.07281; illustrated with news in Morik et al. 2020 (2% relevance gap → "vastly less exposure").
- Biega et al. take attention proportional to relevance as the individual-fairness target. — arXiv:1805.01788
- Saito & Joachims: no principled justification exists for linear exposure-to-merit; an exposure-fair policy can still make an item worse off than random ranking. — KDD 2022
- Zehlike, Yang & Stoyanovich: interventions reflect underlying value frameworks, and fairness definitions are "contingent on the chosen value framework". — arXiv:2103.14000
- Burke: P-fairness may require going beyond personal preference. — arXiv:1707.00093
- On the user–provider tension: every re-ranking method reports a λ trade-off between accuracy and long-tail/provider coverage. — Abdollahpouri et al. 2019; OFAiR
- Dagstuhl manifesto: diversity is "an inherently normative concept", news diversity serves democratic functions, and "one actor should not be able to dominate public discourse". — Bernstein et al., Dagstuhl Manifestos 9(1), 2021
Inferences
- For a news aggregator, "merit" is itself contested: is it story size, recency, originality? Volume of output is not merit, so equal exposure per owner can be defended as correcting a production-capacity advantage (the freshness effect in the Top Stories audit) rather than ignoring relevance.
- A middle path from the literature is to keep merit ordering inside an equal-turn structure (what the soft ceiling does), or to aim for exposure proportional to a merit signal that is not volume, such as the number of distinct owners carrying a story.
Gaps
- I found no user study measuring whether news readers perceive equal-per-owner pages as lower quality.
4. Group-level (owner, conglomerate) versus individual-provider fairness; grouping by ownership
Takeaway
The literature distinguishes group fairness (Singh & Joachims; FA*IR; Burke's protected provider groups) from individual fairness (Biega et al.; Saito & Joachims). Provider groups are usually defined by a protected attribute (e.g. geography, minority ownership), and only rarely by corporate ownership. I found no CS ranking paper that groups news outlets by owner for fairness. Ownership-concentration measurement (HHI) appears in news audits, not in re-ranking algorithms.
Cited Findings
- Group exposure constraints: Singh & Joachims 2018; ranked group fairness for a protected group at every prefix: FA*IR 2017.
- Individual, amortised: Biega et al. 2018; individual impact-based: Saito & Joachims 2022.
- Burke's examples of provider groups include minority-owned businesses in job recommendation, i.e. groups defined by who owns the provider. — arXiv:1707.00093
- Gómez, Boratto, Salamó and colleagues study provider groups defined by geography (continent/country): "The Winner Takes it All: Geographic Imbalance and Provider (Un)fairness in Educational Recommender Systems" (SIGIR 2021) and "Provider fairness across continents in collaborative recommender systems" (IP&M 2022); a 2024 UMAP paper addresses "coarse and fine-grained provider groups" together. — ResearchGate; Boratto; ACM DL 2024
- OFAiR handles multiple provider attributes at once ("multi-aspect"). — arXiv:2005.12974
- News audits measure concentration with HHI and Gini across outlets. — arXiv:2410.23842
Inferences
- Grouping by owner is a group-fairness definition in which the "group" is defined by economic control rather than a protected attribute. The coarse/fine-grained provider-group work (UMAP 2024) is the closest formal analogue: fairness among owners (coarse) and among outlets within an owner (fine).
- One risk of owner-level grouping: if ownership data is wrong or out of date, the fairness guarantee silently fails. Ownership facts need sources and dates.
- Individual-level fairness (each outlet) and group-level fairness (each owner) can conflict: a 10-outlet owner's individual outlets each get less than a 1-outlet owner's single outlet. That is the intended result, but it should be stated openly.
Gaps
- No peer-reviewed CS paper found that uses media ownership or conglomerate membership as the grouping for fair ranking in news aggregation.
5. News-specific fairness and diversity work
Takeaway
News recommender research mostly treats this as normative diversity rather than provider fairness. RADio (Vrijenhoek et al., RecSys 2022) gives rank-aware metrics grounded in democratic theory, and the Dagstuhl manifesto (2021) sets the agenda. These metrics are distributional and rank-aware, so they apply to a non-personalised list as well.
Cited Findings
- RADio (Vrijenhoek, Bénédict, Gutierrez Granada, Odijk, de Rijke; RecSys 2022, best paper runner-up). A rank-aware Jensen–Shannon divergence framework with five normative metrics: calibration, fragmentation, activation, representation and alternative voices. It is rank-aware because attention declines down the list; it compares full distributions; tested on six neural recommenders with the MIND dataset. — arXiv:2209.13520. Extended version "RADio*" in ACM TORS. — ACM DL
- Earlier: Vrijenhoek et al., "Recommenders with a Mission: Assessing Diversity in News Recommendations" (CHIIR 2021). — ResearchGate
- Dagstuhl manifesto, "Diversity in News Recommendation" (Bernstein, de Vreese, Helberger, Schulz, Zweig et al.; Dagstuhl Perspectives Workshop 19482; Dagstuhl Manifestos 9(1), 2021). Diversity is "heterogeneity of media content in terms of one or more specified characteristics"; current models "do not capture the multidimensionality of diversity"; calls for "quantifiable, meaningful, and relevant definitions" while noting "there cannot be a single one"; argues "one actor should not be able to dominate public discourse". — PDF; Maastricht record
- Kaminskas & Bridge, "Diversity, Serendipity, Novelty, and Coverage: A Survey and Empirical Analysis of Beyond-Accuracy Objectives in Recommender Systems", ACM TiiS 7(1), 2016 (2017 issue). — ACM DL
- The 2% → large-exposure-gap example in Morik et al. is explicitly a news-polarisation scenario. — Morik et al. 2020
- User studies on diverse news recommendations, e.g. "Benefits of Diverse News Recommendations for Democracy: A User Study" (Digital Journalism, 2022). — Taylor & Francis
Inferences
- RADio's "representation" and "alternative voices" metrics, computed over owners rather than individual outlets, would be a principled way to report how balanced a non-personalised edition is (a count, not a judgement). Its rank-aware discount matches the position-bias concern above.
- "Calibration" and "fragmentation" in RADio are about personalisation and so are largely moot for a same-for-everyone page; "representation" and "alternative voices" apply directly.
Gaps
- I did not find RecSys news-track papers specifically on provider (publisher) fairness, as opposed to content diversity.
6. Gaming and manipulation: content farms, volume flooding and defences
Takeaway
Volume flooding is a known way to game ranking. The standard defences are near-duplicate detection and clustering (shingling, SimHash), per-site caps, and spam policies against mass-produced content. Counting by owner and clustering into stories are the aggregator's equivalents.
Cited Findings
- Near-duplicate detection for web crawling with SimHash fingerprints (Manku, Jain & Das Sarma, WWW 2007). — ACM DL; probabilistic SimHash extension (CIKM 2011). — ACM DL
- Benchmarking near-duplicate detection on pay-walled news (WWW 2025 companion). — ACM DL
- Google's per-site cap (two per site in top results; subdomains counted as one site) addresses one site occupying a results page. — Search Engine Roundtable
- Google introduced new spam policies with its March 2024 core update, including "scaled content abuse". — Google Search Central blog, March 2024 (the fetch returned only the archive index, so the policy wording and any claimed reduction figures are not verified here)
- The freshness advantage in Google Top Stories (83.5% of articles under 24 h) shows how continuous high-volume publishing turns into exposure. — CJR
Inferences
- Treating subdomains as one site (Google) corresponds to treating one owner's outlets as one; both close the "add more feeds" loophole.
- Story clustering (pg_trgm grouping) is a near-duplicate defence: agency copy syndicated across one owner's outlets collapses into one story, so it cannot fill several slots. Counting distinct owners per story, rather than distinct outlets, keeps syndication from inflating "N reports".
- The remaining attack surface: owners whose links are hidden (outlets that look separate but share ownership), and junk desks or advertorials, which need classification and junk filtering.
Gaps
- I did not find peer-reviewed measurements of how much ownership-unaware aggregators are skewed by syndicated or wire copy.
- The text of Google's 2024 "scaled content abuse" definition could not be retrieved.
Corrected since
- Search Engine Roundtable: Google’s site-diversity change was reported to cover only the main web listings, not other search features; the article doesn’t mention Top Stories.
- Search Engine Journal: A site’s subdomains are generally counted as the site, not always.
- Singh and Joachims: They offer exposure in proportion to merit as an alternative to equal exposure.
- Biega and colleagues: The target is attention in proportion to relevance, summed over many rankings, not attention evened out.
- Zehlike and colleagues, FA*IR: The anthology link is dead; the paper is on arXiv. FA*IR keeps a minimum share for a protected group at every point in the list.