How We Score
Every claim is inspected the same way: five labs score it blind, then two masters drawn at random write the report. The scoring model is the arithmetic that turns those five numbers into the one you see. It has a version, every Inspection Report says which one it was reached under, and the current model is v2.
When a model changes, the claims already published are not silently re-numbered. Each is re-inspected under the new model and reviewed before the new verdict replaces the old one. The report as it was stays readable, marked as an earlier version, and any card made from it says so when followed.
The scale everyone scores on
A Baloney Score runs from 1 to 99. Low is true, high is false, and the number is published under the band it falls in:
| Score | Band | Meaning |
|---|---|---|
| 1–5 | Super Fresh Truth | The claim checks out. |
| 6–20 | Fishy Baloney | Something smells off… |
| 21–40 | Stinky Baloney | That’s not the whole story… |
| 41–60 | Rotting Baloney | Most of this claim has spoiled. |
| 61–98 | Radioactive Baloney | Dangerously misleading. |
| 99 | Zombie Baloney | DEAD. BURIED. STILL SHAMBLING AROUND. |
On a Super Fresh Truth claim the page shows the truth percentage instead — a score of 2 is displayed as 98% true — but the stored number is the same number.
What every model has in common
- Five labs, blind, in parallel. Each writes evidence for, evidence against, sources, a confidence and a recommended score. None sees another.
- The published number is a median, never a mean. The middle of the lab scores. One lab wildly out moves it not at all, which is the property a panel of five needs.
- The masters own the words, not the number. The Primary Master writes the report from the five anonymised lab reports. The Blind Master scores the same reports without knowing the Primary exists. Their scores are published beside the panel median and used to check it; if either master sits more than 20 points from the median, the verdict is held back for re-evaluation rather than published.
- Everything is on the report. Every lab score, both masters, and a line-by-line account of how the number was reached.
Model v1 superseded
The model every claim was scored under before September 2026. It is kept unchanged so that what it published can always be reproduced.
- The labs were told one sentence about the scale: 1 is essentially true, 100 is pure misinformation. The six bands, their names and their ranges were given to the two masters but never to the labs.
- Every lab counted. The published score was the median of all five recommended scores, whatever each lab had written in words.
- The zombie flag was recorded but never changed the number. The Primary Master could flag a long public debunk history; the score stayed the median. Publishing Zombie Baloney needed three labs to independently write 99 on a scale that never named it.
Worked example. Five labs score a debunked claim 3, 85, 90, 95 and 99. The 3 came from a lab whose evidence was entirely against the claim at confidence 94 — it read the scale backwards. v1 publishes the median of all five: 90, Radioactive Baloney. The zombie flag is true and does nothing.
What was wrong with it. A lab that thought a claim was “substantially true with a caveat” would reasonably write 12 or 15, which on this scale is Fishy Baloney — it was scoring a ruler it had never seen. And a lab saying “false” in words and “true” in its number counted the same as one that had read the scale correctly.
Model v2 current
Four changes to v1, and nothing else. The panel, the masters, the prompts’ evidence questions and the report are identical.
- 1. The labs get the scale. The same six-band guide the masters have always had — names, ranges, “score the statement as written” — goes to every lab word for word, with one explicit check: a claim you have just shown to be false cannot carry a number under 20, and a claim you have just shown to be true cannot carry a number over 80.
- 2. A lab that contradicts itself is not counted. A report whose only evidence is against the claim, with confidence 80 or above, and a score of 20 or below has read the scale backwards. The mirror — only evidence for, confidence 80 or above, score 80 or above — is the same error the other way. That lab’s report is still shown, marked NOT COUNTED with the reason; its number is left out of the median. If fewer than two labs survive, all are counted and the report says so.
- 3. Zombie is a call. When the Primary Master flags a long, well-documented debunk history spanning years, AND the counted median is 61 or above, AND the Blind Master is 61 or above, the published score is 99. All three are required. The flag on a claim the panel does not already have in the false bands changes nothing. The rule can only ever raise a number already at Radioactive.
- 4. Fresh is a call. The mirror of rule 3. When the Primary Master scores 5 or below, AND the Blind Master scores 5 or below, AND the counted median is 20 or below, the published score is the higher of the two masters’ numbers — inside Super Fresh Truth, and the more cautious of the two. Any gate fails, the median stands. The rule can only ever lower a number already on the true side.
Worked example, the false side. The same five scores: 3, 85, 90, 95, 99. The lab that wrote 3 with all its evidence against the claim at confidence 94 is not counted. The median of the four counted is 92.5. The Primary Master flags a long debunk history and the Blind Master scores 92, so the Zombie rule applies: 99, Zombie Baloney. The report shows all five labs, marks one NOT COUNTED, and prints every step.
Worked example, the true side. Five labs score a true claim 2, 4, 8, 12 and 3. Nothing is excluded; the median is 4. The Primary Master scores 3 and the Blind Master 4, so the Fresh rule applies and the published score is the higher master, 4, Super Fresh Truth — shown as 96% true. Had the Blind Master written 9, the rule would not apply and the median of 4 would stand on its own.
Why the numbers moved. Claims re-inspected under v2 often land in a different band from their v1 report. That is mostly rule 1: the labs are now scoring the ruler the site publishes, so their numbers cluster in named bands instead of a vague middle. It is a calibration change, not a change in what the labs think about the claim.
Reading a report
- The line above the claim names the model: “Scoring model v2” links here.
- “How the number was reached” lists every decision the model made on that claim, including any lab not counted and any rule applied.
- If the claim has been re-inspected, the same line offers “Earlier report, model v1”: the report as it was, marked as superseded and pointing back at the live one.
The prompts themselves, verbatim, are on How We Inspect.
