The record, to replay
Four times a day Argus collects the news, groups it into stories and ranks them by one published rule. Since 23 Sep 2026 every run has been kept. These files hold each run's ranking with the score split into its parts, and every article's publisher, channel and address, so anyone can check a number on this site or rank a past run by a different rule.
What is in each file
runs-DATE.jsonl.gz: one line per story per run. run_at, cluster (the first 16 characters of its id, stable across runs and midnight), rank, category, score and its parts outlet_points (3 per outlet), volume_points (2 × ln(1 + articles)) and recency_points, outlets, articles (their short ids), first_seen (the run that first saw the story) and story (its page here, when it has one).
articles-DATE.jsonl.gz: one line per article collected that day. id (short), outlet (the publisher, spellings folded), via (newsroom feed, Google News, Reddit or GDELT), published_at, fetched_at, url.
Left out: headlines, summaries and article text. Each channel has its own terms about redistributing its content in bulk; the numbers and the addresses are the record's own. A story's headline is on its page; an article's is at its address.
JSON lines, gzip-compressed, UTF-8. Times are UTC. GDELT's times are when it crawled a page, not when it was published. Files are rewritten with each publish; a day's file grows until the day is over.