EXPERIMENTS · 10 THAT HAVE RUN

What the machine has tested about itself

Each experiment asks one question about how this site works, answers it from the site's own data, and says what was decided. The results are rebuilt after every run; the editor test runs weekly. Where an answer rests on labelled examples, the page says who labelled them, including how many were labelled by an AI model. Results that do not flatter the site stay in. What the record cannot see gives the denominators behind every number here, and the record itself is released to replay.

  1. How much of What Moved is the rule itself?

    Stories enter and leave the top 25 every run. Some of that is news; some is the ranking rule acting on its own.

    Question
    How much of what moves in the top 25 between runs is the rule rather than the news?
    Result
    Over 34 run pairs, 23% of 238 moves were mechanical: 34 from outlets ageing out of the window, 20 from the clock, 1 from GDELT crawl times. 40% were decided by less than one outlet's weight; the gap between #25 and #26 was under one outlet in 74% of runs.
    Decision
    GDELT crawl times no longer feed recency (from 29 Sep 2026). Each What Moved page now says how many of its moves were mechanical. The window's share is the next thing to fix in the rule.

    LAST BUILT · THE NUMBERS →

  2. How often does a newcomer last a day?

    Stories enter the top 25 from nowhere every day or so. The record keeps count of how many are still there a day later.

    Question
    Of the stories that enter the top 25 from nowhere, how many are still there a day later?
    Result
    16 of 23 newcomers were still in the top 25 a day later (70%); about 0.64 a run enter from outside the top 100.
    Decision
    Record only. No story page shows a probability. The rate by entry band is not read until each band has 30 newcomers.

    LAST BUILT · THE NUMBERS →

  3. Does Argus state forecasts as fact?

    The site's own stories sometimes say what is "set to" happen or "could" follow. How often without saying who expects it, and did it happen?

    Question
    How often does Argus's own writer state a forecast as fact, and does the record later show the event?
    Result
    87 forecast phrases in the last 30 days; 41% attributed, the rest stated bare. Of speculation (not schedules), 56% was attributed.
    Decision
    Measure first. If bare speculation stays above half of speculation, the editor gets a rule to attribute or cut it, and this study records the run the rule shipped.

    LAST BUILT · THE NUMBERS →

  4. Which earthquakes did the news under-cover?

    USGS lists every quake whatever the news does. Matching a year of them to coverage shows which got less than quakes of their size.

    Question
    Which earthquakes got more or less coverage than quakes of their size?
    Result
    19 of 138 quakes had any coverage in the record. 15 got less than quakes like them, 5 more. The matcher's hand-checked precision is 0.947 and recall 0.818.
    Decision
    Published as a replication under a public rule, not a discovery: disaster coverage is known to follow size and news-day crowding. The matcher is re-checked by hand on new quakes; openFDA recalls and FRED series follow with the same spine.

    LAST BUILT · THE NUMBERS →

  5. Is the ranking just the clock?

    Every story loses score as it ages. How much of what leaves the top 25 is that, rather than news?

    Question
    How much of the top 25's churn is the clock rather than the news?
    Result
    The clock causes 0% of turnover between six-hourly runs and 0% day to day.
    Decision
    Keep the 12-hour half-life: recency is not what moves the top 25.

    LAST BUILT · THE NUMBERS →

  6. Does shared copy inflate a story's reach?

    Some outlets run the same copy as others. If each such group counted once, would the top stories change?

    Question
    If the outlets of one syndication ring counted once, how would the top 25 change?
    Result
    Rings of one owner, network or public broadcaster hold 17% of a top-25 story's outlets; counting each once changes about 4% of the top 25. Counting wire and mixed rings once too: 33% of the outlets, about 8% of the top 25.
    Decision
    Production ranking unchanged: ownership rings barely move it, and whether a wire pickup is a newsroom's own choice is a question of definition the data cannot settle. A ring is a shared content pattern, not proof of common ownership.

    LAST BUILT · THE NUMBERS →

  7. Which stories are the same event?

    An edition can carry several stories about one thing that happened. How well does the rule that groups them work?

    Question
    Which of an edition's stories are the same real-world event?
    Result
    Precision 0.92, recall 0.338 on 265 labelled pairs. Latest edition: 105 stories shown as 91 events.
    Decision
    Keep the rule: it is precise, and no threshold in the sweep raises recall by more than a few points while keeping precision above 0.9. The pairs it misses are different angles on one event, which similarity alone does not see; catching them needs another signal.

    LAST BUILT · THE NUMBERS →

  8. Does the editor catch planted errors?

    A second model checks each story against its sources. Stories with a known error planted in them test it.

    Question
    Does the second model catch errors the writer makes, without changing clean stories?
    Result
    Recall 100% on 30 seeded stories; it changed 3 of 6 clean ones.
    Decision
    Keep the editor in the run. Watch the false-alarm rate on clean stories.

    LAST BUILT · THE NUMBERS →

  9. Is the clustering cut in the right place?

    Articles are grouped into stories at one similarity cut. Moving it shows what splits and what merges.

    Question
    Is the clustering cut in the right place?
    Result
    At the current cut 0.42, 10% of groups survive whole on headlines alone; the largest cluster holds 2% of articles.
    Decision
    Keep the cut: production keeps Google's groups whole, and a higher cut starts swallowing topics.

    LAST BUILT · THE NUMBERS →

  10. Can a model spot loaded language better than word lists?

    Word lists flag loaded wording. A model trained on labelled sentences is the check on them.

    Question
    Where do the framing word lists over- or under-count?
    Result
    Model F1 0.179 against the word lists' 0.462, on 121 labelled sentences (28 loaded).
    Decision
    The word lists stay: on these labels the model does worse than they do. With 28 loaded examples it has too few to learn from; it needs more labels, especially loaded ones.

    LAST BUILT · THE NUMBERS →

Claims, and how they came out

A study can say in advance what a number will be by a given date. The claim is recorded the first time a build sees it, and scored on its due date against the study's own result. A claim recorded after its due date, or changed after it was recorded, is never scored. Failed claims stay on this list.

ClaimStudyRegisteredDueOutcome
The earthquake matcher's hand-checked precision is still at least 0.9 after the next three months of quakes are checked. validation.precision >= 0.9 Which earthquakes did the news under-cover? 2026-09-29 2026-12-31 pending
In the four weeks after the crawl-time fix, at least 15% of top-25 moves are still made by the rule rather than the news. after_fix.mechanical_share >= 0.15 How much of What Moved is the rule itself? 2026-09-29 2026-10-27 pending
After the crawl-time fix, no top-25 move comes from GDELT crawl times. after_fix.causes.crawl == 0 How much of What Moved is the rule itself? 2026-09-29 2026-10-13 pending