How scoring works

Every factual claim in a debate is researched against the web and measured on two separate axes: whether it is true, and what it left out. Verdicts are evidence-based estimates — not a claim to absolute truth — and every one links the sources it relied on so you can judge for yourself. This page explains both axes and why they are never blended.

These are the full instructions. Too long, didn't read? Here's the quick version →

Axis 1

Veracity — is the claim true?

Every factual claim is researched against the web and scored 0–100, with the sources it relied on attached. That one score is read two ways — absolutely (true exactly as worded) and substantively (true in its gist). Two readings of this axis, not two axes.

92% 64% 12%
Axis 2

Omissions — what was left out?

A claim can be entirely true and still mislead, by leaving out the one fact that changes what it means. So claims that check out get a second, separate look, and each finding names a missing fact and cites a source. There is no score on this axis at all.

⚠ 2 omissions Context checked

The two never mix — an omission is never subtracted from a score. Everywhere on this site, purple means omissions: a purple badge, rule or heading is always about what a claim left out, never about how accurate it was.

Axis 1 — the veracity score (0–100)

Each checked claim gets a score from 0 (clearly false) to 100 (clearly true), shown as a coloured badge. The colour is a quick read of which band the score falls into — and the same scale carries both readings of it below. The second axis has no scale of its own at all.

0–19
20–49
50–79
80–100
12% False / very low confidence 35% Mostly false / low 64% Mixed / leaning true 92% True / high confidence Still checking, or not a scorable claim
Verdict labels

Alongside the number, each claim carries a plain-language verdict on the same colour bands:

True95% Mostly true82% Mixed58% Mostly false32% False8% Unverifiable(not scored)

Unverifiable claims (no reliable source either way) are shown but excluded from a speaker's average — they neither help nor hurt the score.

Two ways to read that score: Absolute vs Substantive

The veracity score is reported two ways, because "wrong" can mean two very different things — a precise figure being slightly off, versus the underlying point being untrue. Both are readings of the same axis: still is this claim true?, asked strictly and then generously. The second axis asks something else entirely, and comes further down.

Absolute strict & literal

Is the statement true exactly as worded — every number, date and name? This is the canonical reading: it decides a statement's status and is the value reused across debates.

Substantive the gist

Is the underlying point true, even if a figure is rounded or imprecise? It allows reasonable approximation. When a checker doesn't distinguish the two, the substantive score simply mirrors the absolute one.

Worked examples (made-up claims)

Claim Absolute Substantive Why they differ
“Our plan created exactly three million jobs.” 30% · Mostly false 85% · Mostly true Official data show ~2.6M, so “exactly three million” is literally off — but the gist (millions of jobs added) holds.
“Crime has tripled since 2010.” 8% · False 10% · False Reported crime rose ~12%, nowhere near threefold. Wrong literally and in substance, so both readings are low.
“The law saves households about $500 a year.” 88% · Mostly true 95% · True Independent estimates land near $480. The hedge “about” keeps even the literal reading accurate.

This split is why a claim that is only literally imprecise (high substantive, low absolute) is not treated as a “persistent falsehood” on a speaker's profile.

When was it true? Every verdict carries a date

Facts move. Unemployment figures, prices, who holds which office — a claim that was plainly true when it was made can be plainly false a year later, and marking it false would be scoring a speaker on events they could not have known.

So every session records the date it was held, and each claim is researched as of that date. Verdicts say so: open Evidence & sources on any claim below and you will find a line reading Judged as of 28 May 2026. Where that line is absent, the claim was judged against the present.

A claim already checked in an earlier debate can be reused instead of re-researched, and the checker itself sets the limit: alongside each verdict it reports how long that answer holds — a single day for a monthly statistic, effectively forever for a historical date. A stored verdict is reused only where its own window reaches the new claim's date, and it keeps the date it was originally judged against, so no chain of reuses can walk an answer away from the day it was really checked. Reused verdicts are labelled as reused.

From claims to a speaker's average

A speaker's (and a debate's) overall score is the mean of their scored claims only, computed separately for each reading and rounded to the nearest whole percent. Opinions, unverifiable claims, and errors are left out entirely.

Example — one speaker with four scored claims (Absolute)

100%25%50%33% 208 ÷ 4 = 52%

The same four claims, scored substantively

100%70%60%90% 320 ÷ 4 = 80%

See it in action

Below is a short made-up debate run through the very same components used elsewhere in the app. Expand Evidence & sources on any scored claim to see both summaries and the cited sources. The scoreboard's averages are computed live from the scored statements.

Statements
  • AlexThe United States declared independence in 1776.
    Abs 100% · TrueSub 100% · True
    Evidence & sources

    Absolute: The Declaration of Independence was adopted on July 4, 1776 — confirmed by multiple primary sources.

    Context checked — no material omissions found.

    Detected by xAI grok-3-mini · 06/01/2026 15:00
    Checked by xAI grok-4.3 · 06/01/2026 15:00
    Judged as of 28 May 2026
    Context checked by xAI grok-4.3
  • AlexExports to Europe have grown 40% since 2019.
    “Since 2019, our exports to Europe have grown by 40%.”
    Abs 100% · TrueSub 100% · True
    ⚠ Context: 1 omission
    Evidence & sources

    Absolute: Trade figures confirm a 40% rise in exports to Europe between 2019 and the present.

    ⚠ Context: 1 omission What you weren't told — this claim is accurate; these are omissions that change what it means.

    Missing baseline
    2019 was the weakest export year in over a decade.
    Measuring growth from an unusually low year makes an ordinary recovery look like a boom. Against 2018 the same figure is a 3% rise.
    Office for National Statistics — Trade in goods — “Exports to the EU fell to a decade low in 2019.”

    The claim is accurate. The chosen starting point is what carries the impression.

    Detected by xAI grok-3-mini · 06/01/2026 15:00
    Checked by xAI grok-4.3 · 06/01/2026 15:00
    Judged as of 28 May 2026
    Context checked by xAI grok-4.3
  • AlexThe plan created exactly three million jobs.
    “Our plan created exactly three million jobs.”
    Abs 30% · Mostly falseSub 85% · Mostly true
    Evidence & sources

    Absolute: Official figures show roughly 2.6 million jobs, so the precise claim of “exactly three million” is literally inaccurate.

    Substantive: Directionally sound: on the order of three million jobs were added, so the broad point stands.

    Detected by xAI grok-3-mini · 06/01/2026 15:00
    Checked by xAI grok-4.3 · 06/01/2026 15:00
    Judged as of 28 May 2026
  • JordanCrime has tripled since 2010.
    Abs 8% · FalseSub 10% · False
    Evidence & sources

    Absolute: Reported crime rose by roughly 12% over the period — far from a threefold increase — so the claim is false both literally and in substance.

    Detected by xAI grok-3-mini · 06/01/2026 15:00
    Checked by xAI grok-4.3 · 06/01/2026 15:00
    Judged as of 28 May 2026
  • JordanThis is the best budget our country has ever seen.
    Not a factual claim
  • AlexThe opponent privately said they would raise taxes.
    “My opponent told me privately he would raise taxes.”
    Unverifiable
    Evidence & sources

    Absolute: No public record can confirm or deny a private conversation, so this claim can't be checked.

    Detected by xAI grok-3-mini · 06/01/2026 15:00
    Checked by xAI grok-4.3 · 06/01/2026 15:00
    Judged as of 28 May 2026
Speaker scores
Alex
Absolute
77%
Substantive
95%
3 scored / 4 claims
⚠ 1 omission · 2 checked
For the motion
Jordan
Absolute
8%
Substantive
10%
1 scored / 2 claims
Against the motion

⚠ Axis 2 — omissions: what was left out?

Everything above this line was axis one. Everything below it is the second axis, and it is purple throughout — on this page, on a statement row, and on a speaker's profile.

Everything above answers one question — is this claim true? A claim can pass that test completely and still leave a listener with a false impression, by omitting the one fact that changes what it means. Cherry-picking, missing baselines, a timeframe chosen because it flatters. The most skilled operators live entirely in this space and never say anything false.

So claims that check out as substantially true get a second, separate look: what was the listener not told? Findings appear beside the accuracy score and are never blended into it — a cherry-picked claim is still 100% factually accurate, and its veracity score still says so. The interesting case is exactly that pairing: a green accuracy badge next to a flagged context badge.

Which claims earn that look is decided on the Substantive reading, not the Absolute one, and that choice is the whole feature. “Unemployment fell 2%” when it really fell 2.1% reads as Mixed taken literally and True taken for its gist — and a claim like that is exactly the one worth examining for what it leaves out. Gating on the strict reading would skip it. A claim needs a substantive score of at least 70 to be examined; below that it is already flagged as inaccurate, and needs no omission analysis to say so.

There is deliberately no context score. "Misleading" is a judgement people can reasonably disagree about, and a number like "64% integrity" would give false precision to it. Instead you get specific omissions, each with a source you can open and check for yourself.

Arguing your side is not dishonesty

In a debate, arguing one side is the format. A speaker is not obliged to make their opponent's case, and treating that as an omission would flag every statement in every debate and make the whole thing meaningless. So each candidate omission has to pass one test: does it change what the speaker's own claim means?

That allowance belongs to a debate, because an opponent is in the room to make the missing case. In an interview or a speech nobody is, so it is withdrawn — a fact left out there is simply one the audience never hears. The motion and each speaker's side are recorded either way, and are optional either way: they are what lets us say which direction an omission favours.

Motion This house believes remote work harms productivity.
An opponent of the motion says: “Company X reported a 15% productivity increase after going remote.” The figure checks out — it is true.

What was left outFlagged?Why
Company X also cut its least productive division that quarter. ⚠ Flagged The stated 15% no longer means what it appears to — it reflects who left, not who improved. This changes the speaker's own claim.
Other companies saw productivity decreases. Not flagged True, and relevant to the motion — but it is the other side's argument to make. It does not change what Company X's own figure means.
Three rules every omission must pass
  1. It has to matter. Would knowing it change a reasonable listener's conclusion about this claim? Every true statement omits something; almost none of it matters.
  2. It has to be cited. Every omission names a specific fact and links a source that was actually found when researching it. Anything the app cannot trace to a real page is discarded rather than shown — an uncited accusation is worse than no finding at all.
  3. It has to undercut the speaker's own claim, not merely support the other side, as above.

Context checked Nothing survives all three? Then the claim is recorded as checked with no findings — which is the common result, and carries that badge. It is deliberately quieter than a flagged one, and it is shown differently from a claim that was never examined at all.

Did the missing fact damage their own case?

Every finding above passes one test: the missing fact changes what the claim means.

Where we know which side the speaker was arguing, we ask one more thing — would that fact have damaged their own case? In the Company X example above it would: the missing detail deflates the very figure the speaker was leaning on.

Those are the ones a speaker's profile counts. One on its own proves nothing — omit enough facts by accident and some will land against you by chance. So the profile reports what share of someone's omissions were the damaging kind, and reports nothing at all until there are enough of them to be worth reading.

A speaker who says nothing false, but whose omissions keep turning out to be the damaging ones, is what this app calls Accurate · One-sided. That is a pattern, not a motive — whether it is selection or coincidence is for you to decide.

The two answers read together

Side by side, the accuracy score and the context findings sort speakers into three groups. A profile names the group its speaker falls into, and Speakers ranks each:

  • Accurate · Forthright — Accurate, and checked for what they left out
  • Accurate · One-sided — Say nothing false, and what they leave out always damages their own case
  • Inaccurate · Incomplete — Inaccurate even read generously, and leaving things out as well

Most speakers are in none of them. Each group has a bar to clear — enough claims examined, and a clear enough result — and Speakers states what that bar is. A name missing from a list is not a verdict about that person.

The kinds of omission, and what each is called

Every finding is filed under how the missing fact did its work. You will meet these labels on a claim, on a speaker's profile and on the home page — always in the omissions purple, never on the green-to-red veracity ramp, because these are not ranked. Cherry-picking is not worse than a missing baseline; they are different, not ordered.

  • Cherry-picking The favourable data points presented as the picture.
  • Missing baseline A change quoted with no baseline, so its size cannot be judged.
  • Selective timeframe A window chosen because it flatters — the trend reverses just outside it.
  • Subgroup selection True of a narrow subgroup, stated as though true generally.
  • Count vs rate A raw count where a rate is the meaningful measure, or the reverse.
  • Quote mining A quotation stripped of the passage that gave it its meaning.
  • Missing context Context missing in some other way.
What counts toward a score

Not every line in a transcript is scored. Here is the lifecycle of a single statement:

Detected Checking… Scored ✓ counts toward the average Not a factual claim excluded — opinion / question Unverifiable excluded — no reliable source Error excluded — the check failed
An unhandled error has occurred. Reload 🗙

Rejoining the server...

Rejoin failed... trying again in seconds.

Failed to rejoin.
Please retry or reload the page.

The session has been paused by the server.

Failed to resume the session.
Please reload the page.