How Huduku fetches
Last updated 2026-10-11.
Huduku is an independent media-literacy project. It shows headlines from chosen outlets, each with a link to the publisher’s own page. This page says what its fetcher, HudukuBot, asks publishers for, how often, and how to make it stop.
Who asks
Every request Huduku makes to a publisher’s site carries one name:
HudukuBot/1.0 (+https://huduku.org/fetching)
HudukuBot asks under no other name. It uses no proxy, no browser pretending to be a person, and no copy of a site kept by someone else. It never reads Google News or another aggregator in a publisher’s place.
What it takes
- From the publisher’s own feed (RSS or Atom), its news sitemap, or, for a few sites, the list of headlines on its own home page: each item’s headline, its link, when it was published, and which outlet published it.
- Where a feed gives them, the category a piece is filed under (so Huduku can say “Opinion” only when the publisher does) and, for magazines, the author line, which Huduku doesn’t show.
- A summary in a feed is read only to file the headline under a topic, while the item is in hand, and is not kept. Huduku doesn’t fetch article text, images or logos.
- To check whether an article has been withdrawn, it asks each public headline’s link once, with a single request for its headers and nothing else, when the headline is two days old and again at sixteen days (see Takedowns).
How often
- Each feed or sitemap is read about once an hour.
- A site’s robots.txt is read about once a day, and kept for that day.
- It asks a site for one thing at a time, never several at once.
The rules it keeps
- robots.txt first. Before any request, HudukuBot reads the site’s robots.txt and keeps
the rules for
HudukuBotif there are any, otherwise those for*. It doesn’t ask for a path they disallow. A content signal ofsearch=nocounts as a refusal: a headline and a link are closest to search’s use. If robots.txt itself answers 401 or 403, the whole site is treated as refused. robots.txt is obeyed as it stands each day: change it, and HudukuBot follows. - A refusal is final. If a request answers 401, 403 or 451, or with a challenge page, HudukuBot stops reading that source. It doesn’t try again another way, and the source returns only when the publisher gives permission or lifts the block, with the reason recorded.
- Asked to wait, it waits. On 429 or 503 it waits as long as the
Retry-Afterheader asks, and at least an hour. Three such answers in a row stop it until someone at Huduku looks.
Opting out
- Write one email to contact@huduku.org from an address at your publication’s domain, naming the outlet or the site. Within 48 hours Huduku stops fetching it and takes its headlines off every public page.
- Or disallow
HudukuBotin your robots.txt: it takes effect within a day, when the file is next read. - More for publishers, including how to give or withhold permission: For publishers. To have a single headline removed: Takedowns.
Videos
Huduku lists videos from a few YouTube channels through YouTube’s own API, with Huduku’s key; HudukuBot doesn’t fetch from YouTube. The privacy page says what is kept.