Does the editor catch planted errors?
A second model checks each story against its sources. Stories with a known error planted in them test it.
- Question
- Does the second model catch errors the writer makes, without changing clean stories?
- Method
- Plant one known error per story (number, name, quote, inference, negation) and send it to the editor as a run does; unseeded copies are the controls. Weekly.
- Result
- Recall 100% on 30 seeded stories; it changed 3 of 6 clean ones.
- Decision
- Keep the editor in the run. Watch the false-alarm rate on clean stories.
How it was measured
One language model writes each story from the outlets' own articles, and a second model edits it: it may cut, attribute or correct against the same sources, and nothing else. To test it, stories it had passed as written get one known error planted in each: a number changed, a name swapped, a quote altered, an inference stated as fact, or a negation flipped. Each seeded story goes to the editor as a run would send it.
- Recall: the share of planted errors the editor fixed or flagged.
- False alarms: the same stories unseeded are the controls; this is how often it changed a clean one anyway.
The same planted errors are used each time, so a change between tests is the editor, not the draw.
The numbers
Latest test · 2026-09-27 · 30 seeded stories, 6 controls
| Error | Planted | Fixed | Flagged | Missed | Recall |
|---|---|---|---|---|---|
| number | 6 | 6 | 0 | 0 | 100% |
| name | 6 | 6 | 0 | 0 | 100% |
| quote | 6 | 6 | 0 | 0 | 100% |
| inference | 6 | 6 | 0 | 0 | 100% |
| negation | 6 | 6 | 0 | 0 | 100% |
Overall recall 100%. The editor changed 3 of 6 clean stories.
Over time · recall by kind of error, one point per test
On real stories, by what the writer had
How often the editor revised or held a real story, by how much of the writer's input carried the publisher's own text, and by how many independent newsrooms it came from (each syndication ring counted once). 1846 stories.
| Inputs with text | Stories | Editor intervened | 95% interval |
|---|---|---|---|
| none | 777 | 35% | 32–38% |
| some | 866 | 33% | 30–37% |
| majority | 203 | 47% | 40–54% |
| Independent newsrooms | Stories | Editor intervened | 95% interval |
|---|---|---|---|
| 1 | 20 | 65% | 43–82% |
| 2-3 | 439 | 28% | 24–33% |
| 4+ | 1387 | 37% | 35–40% |
The latest test had only 6 clean control stories, too few to say how often the editor changes a clean story (3 of 6); from October the weekly test uses 40. Until then these rates cannot be separated from that noise. More text or more newsrooms also means more sentences to check, so a higher rate there is not by itself evidence the editor finds more; a lower rate on headline-only stories may be what it cannot check.
Limits
- Planted errors are cleaner than real ones: a real mistake can be subtler, or can agree with a wrong source.
- 30 seeded stories per test is a small sample; one story more or less moves recall by several points.
- This tests the editor against its sources. It is not fact-checking: if every source is wrong, so is the story.