Sept. 2, 2026Methodology v3.0 · State metrics: 12 source systems, data through 2023–2025 by metric
Political Grades
government — just the stats

Methodology · Integrity Record rulebook v1.0

Honesty is not numerically scored on this site

For two methodology versions, 25% of every legislator grade was reserved for an “Honesty” score that would publish “only with evidence trails and two agreeing human reviews.” That gate was designed for a newsroom and staffed by one person. It published nothing, and it never would have. We owed readers either the staff or the truth, and this page is the truth: honesty is not a number here. In its place, three factual ledgers — each recomputable from cited public records — and one set of preregistered grade caps triggered only by findings other institutions made on the record. The grade’s former honesty weight is retired for all 531 members identically (a changelogged formula change), not silently redistributed per member.

Why not just automate the score? We tried, adversarially, four different ways, before writing a line of pipeline. Every route to a numeric honesty score died as one of three things: degenerate (nearly everyone scores 100%, so all variance is error), purchasable (stuffable with safe pledges, gameable by strategic silence), or model-laundered (a language model’s judgment about a named official, dressed as arithmetic). The full design record, including the probe that showed bill text changes between endorsement and vote on essentially every bill that reaches the floor, is in the project repository.

The invariants (every tier)

Tier A — the Adjudicated Record

The only tier that can touch a letter grade, and only downward. An event scores only when the deciding body is evenly bipartisan by rule — the House Ethics Committee (5R/5D), the Senate Select Committee on Ethics (3–3), the FEC (3–3 by statute) — or an Article III court, on a deception-map offense (false statements, perjury, fraud, obstruction, falsification), with affirmative finality (mandate issued or appeal dismissed; docket silence is never finality; a vacated judgment is removed by append-only correction). Simple-majority floor discipline never scores alone: a floor majority can censure anyone, which is exactly the party-line judgment this tier exists to exclude — it scores only when paired with a same-Congress ethics-committee report recommending that discipline.

Caps (frozen): Tier 1 (conviction or accepted plea; expulsion-with-report) caps the overall grade at F; Tier 2 (censure or reprimand with report; personal knowing-and-willful FEC finding) at C-; Tier 3 (letter of reproval; admission-bearing FEC conciliation) at B. Caps do not stack; the lowest applies; no decay curve, because a curve is a knob. No member is ever labeled “clear” — every panel prints exactly which sources were searched, and 1 of 6 are connected today:

Fact-checker verdicts never score, in either direction: which statements get checked is an editorial choice no reweighting can make symmetric, and the verdicts are another outlet’s composed prose about named persons. Scored adjudication events site-wide today: 0 — with zero events the caps bind no one, and this page says so rather than implying a search that found innocence.

Tier B — Name-Vote Consistency

A formal endorsement — sponsorship, cosponsorship, a discharge signature — is a public, dated, signed act. This ledger pairs each one with the member’s final-passage vote on the same measure, and a pair qualifies only when the text at the vote is essentially the text they signed: no adopted amendment and no new text version between endorsement and vote. Suspension-calendar votes are excluded in both directions (consensus traffic that pads numerators); a Nay on concurrence never produces an adverse row (it addresses the other chamber’s changes); endorsements filed after floor scheduling are excluded (applause, not commitment); one pair per member-bill; an adverse row additionally requires confirmation from a second official record before it publishes. This ledger never feeds the composite grade, and a percentage renders only at five or more qualifying pairs, always beside the chamber median.

Run against all 163 qualifying passage votes of the 119th Congress (160 bills): 27 qualifying pairs, 0 adverse rows. The gates ate nearly everything — which is the finding, not a failure: on the modern floor, the text members endorse and the text they vote on are almost never the same document. Exclusions, by reason:

Tier C — Checkable Commitments

The corpus is spoken floor proceedings from the Congressional Record daily edition, harvested identically for every member over a disclosed window — currently 2026-03-01 to 2026-08-31 (98 session days). Extensions of Remarks are excluded by section class: members can insert unspoken text there, and counting insertions rewards corpus-stuffing. Speaker attribution is exact-string against GPO’s own per-granule tagging; ambiguous segments are dropped and counted.

A pledge-shaped sentence becomes a ratable commitment only if the deterministic compiler can build its test entirely from the sentence itself: a first-person commissive (“I will…”), a named bill (“H.R. 25”) in the same sentence, and an unambiguous direction verb. Everything else is excluded with a reason code. The kill-lists are deliberately over-broad — a missed promise costs coverage; a mis-rated one is a false accusation, and those are not the same size. From the current window, 1,895 pledge-shaped sentences compiled to 1 ratable commitment:

Outcomes are computed by code, never judged. A matching final-passage vote after the pledge → Kept. A contrary final-passage vote that passed, with no amendment adopted between pledge and vote → Contradicted by record — the only adverse label, always rendered beside the roll call and the verbatim quote. No qualifying vote yet → Not tested. Sponsorship and outcome pledges stay Open until the Congress ends; a window that closes empty reads No qualifying record found and costs delivery credit only — absence of evidence is never an honesty penalty, because a void does not say which of its explanations is true. The lead line of every panel is the delivery count, because delivery is what a record can prove.

Known v1 limits, stated rather than implied away: FEC and stock-trading pledge predicates are not connected; the window covers part of the 119th Congress and extends with each run. Attributed floor-speech volume in this window: D 0 words, R 0 words — the corpus inherits the floor’s asymmetries (the majority controls the calendar), which is disclosed here and audited per mechanism rather than tuned.

Tier C-extended — the model-assisted tier, measured

The core tier above uses no language model. The extended tier asks whether a language model, kept strictly untrusted, can safely add the commitments the core compiler misses — chiefly pledges that name a bill by its short title (“the SAVE Act”) rather than its number. The design keeps the model on the far side of a deterministic gate: two independent model passes must agree and an adversarial refuter must fail before a proposal is even considered, and then code re-derives every field from official bytes — the verbatim quote must be a byte-exact substring of the archived record, the government’s own title for the proposed bill number must contain the words the member said, and the for/against direction is recomputed by the same rules as the core tier. The model resolves a title; it never originates an adverse claim, picks a label, or types a number that publishes.

And then we measured it, because the rulebook said to before shipping it. Of 27 short-title candidates the core tier could not resolve, the dual-model proposer classified every one as something other than a first-person vote-position commitment (27 not vote position) — they were pledges to speak, to introduce, to fight for passage, or to describe a bill, none of which is a testable statement of how the speaker will vote. 0 items were published from the extended tier. On this corpus the model added nothing the model-free core did not already have. The verifier is retained and unit-certified against deliberate fabrications (a hallucinated quote, a wrong bill number, a flipped direction — each is caught at its own gate), so the tier can be re-measured as the corpus grows; but it does not publish a single reader-facing item today, and the honest reason is printed here rather than the machinery being quietly shelved.

The governors’ edition

A governor casts no roll-call votes and leaves no uniform machine- readable speech record, so the Name-Vote Consistency and Checkable Commitments tiers do not apply to the office — a stated boundary, not a silent gap. The Adjudicated Record does transfer, because its admission test is about the structure of the deciding body and generalizes to the states unchanged: a finding may cap a governor’s grade only from a court of record or a state body that is evenly bipartisan by rule. Legislative impeachment — a legislative majority acting — is excluded by rule, as is any governor-appointed board with a working partisan majority; those are the party-line judgments the whole test exists to keep out.

Whether a given state’s ethics commission meets the evenly-bipartisan-by-rule bar is a per-state question of statutory composition, and a wrong “yes” would let a partisan body damage a grade — so no state ethics body is connected on a machine’s say-so. Those determinations are researched with the composition statute cited, then held as a proposal for a person to confirm before anything is wired. Until then each governor’s panel connects the courts of record only, wires no event feed yet, and prints the same coverage line the congressional tier does — which sources were searched, how many events, and never the word “clear.” The rule is published now, before any event exists.

What changed in the grade

The alignment block’s 25% is retired, not refilled: the surviving weights (state record, bills, fiscal, attendance) each scaled by exactly 4/3, so every member’s grade changed by the same multiplicative constant and no rank moved except at the winsorization bound. Filling the slot with any of the numbers above would put a degenerate or purchasable statistic where readers were told honesty lived; leaving it hatched forever would reproduce the never-publishes failure this rulebook removes. Adjudicated findings, when final, cap the grade under ADJ-7 — that is honesty’s entire remaining contact with the number.