EXPERIMENTS · BUILT

Can a model spot loaded language better than word lists?

Word lists flag loaded wording. A model trained on labelled sentences is the check on them.

Question
Where do the framing word lists over- or under-count?
Method
Train TF-IDF with logistic regression on sentences labelled loaded or plain, cross-validated, and compare it with the lists on the same sentences.
Result
Model F1 0.179 against the word lists' 0.462, on 121 labelled sentences (28 loaded).
Decision
The word lists stay: on these labels the model does worse than they do. With 28 loaded examples it has too few to learn from; it needs more labels, especially loaded ones.

How it was measured

The site flags loaded wording with six word lists: loaded verbs, unattributed claims, asserted motives, consequences forecast as fact, hedges dressed as forecasts, and intensifiers. Word lists are easy to explain and wrong in places. The check on them is a small model, TF-IDF with logistic regression, trained on sentences from the week's headlines and summaries labelled loaded (wording that carries a judgement the facts do not) or plain (reports what happened).

Both are scored on the same labelled sentences, the model by five-fold cross-validation, so it is never tested on a sentence it learned from. F1 balances how many loaded sentences each finds against how many it flags wrongly.

The numbers

AccuracyF1
Model0.620.179
Word lists0.5950.462

121 labelled sentences: 28 loaded, 93 plain. The two agree on 56% of them.

Who labelled, and limits

  • 117 of the 121 sentences were labelled by Claude, an AI model, and 4 by the site's author. Mostly one labeller's judgement of what is loaded.
  • 28 loaded examples are too few for a model to learn from; accuracy is flattered by the many plain sentences.
  • Scores for individual outlets are not published: a model this weak is not a fair judge of a named newsroom.