Methodology › Full text
Scoring Methodology v2.1, Verbatim
The binding rulebook, published unedited — every rule here was fixed before any politician was scored. The methodology page is a plain-English summary of this document; where they differ, this document wins.
# Political War Room — Scoring Methodology v1.0
*The complete rubric for the Honesty Score, the Performance Score, and the attribution rules that determine who gets scored on what. Designed to be published verbatim on the site: the methodology IS the product's credibility.*
---
## Part 0: The Five Fairness Principles
Every rule below derives from these. When an edge case isn't covered, resolve it by these principles and log the decision publicly.
1. **Preregistration.** All metrics, weights, and rules are fixed *before* any politician is scored. Changes require a new methodology version, applied retroactively to everyone, with a public changelog. You can never add or drop a metric after seeing how it affects someone's score.
2. **Symmetry.** Every rule is written without knowing which party it helps. Test: swap the party labels in any example — the ruling must not change.
3. **Trend, not level; peers, not perfection.** Politicians are scored on the *change* during their tenure relative to *comparable peers over the same window* — never on the absolute state of what they inherited.
4. **Honesty ≠ Delivery.** Failing to achieve something you genuinely fought for is not a lie. These are two different virtues and get two different scores.
5. **Show the work.** Every score decomposes, on click, into the individual promises, votes, and data series that produced it, each with a source link. Uncertainty is displayed, never hidden.
---
## Part 1: The Honesty Score
*"Did they do what they said, or try to?"* — Applies to ALL officials: executives and legislators alike.
### 1.1 What counts as a promise
A ratable promise must be **specific enough that a reasonable person could later say whether it happened**. Extraction rules:
| Ratable | Not ratable (logged as "Aspiration," excluded) |
|---|---|
| "I will veto any income tax increase" | "I will fight for working families" |
| "I will hire 500 more police officers" | "I will make our streets safer" |
| "I will sign a bill legalizing X in my first year" | "I support X" (position, not commitment) |
| "I will not cut Medicaid" | "Medicaid is important to me" |
Each promise record stores: verbatim quote, source URL (with Wayback archive), date, venue (ad / debate / platform / speech), and a category tag. Both an LLM pass and a human reviewer must independently agree it's ratable; disagreement defaults to Aspiration.
### 1.2 Promise prominence weights
Not all promises are equal. Weight by how central the candidate made it:
- **Signature (×3):** appeared in paid advertising, convention/kickoff speech, or repeated in 3+ distinct venues
- **Major (×2):** in the official platform or repeated in 2 venues
- **Minor (×1):** stated once in a ratable form
Prominence is assigned at extraction time, before any outcome is known (preregistration principle).
### 1.3 The six outcomes
At rating time (see 1.5), every promise lands in exactly one bucket:
| Outcome | Definition | Honesty credit | Delivery credit |
|---|---|---|---|
| **Kept** | The promised action/result occurred substantially as described | 1.0 | 1.0 |
| **Partial** | Meaningful movement in the promised direction, materially short of the commitment (promised 500 officers, funded 250) | 0.5 | 0.5 |
| **Blocked** | Official made a *documented genuine effort* (see 1.4) and was stopped by an actor outside their control — legislature, courts, referendum, Congress | 1.0 | 0.0 |
| **Abandoned** | No documented genuine effort; the promise was simply dropped | 0.0 | 0.0 |
| **Betrayed** | Official actively did the *opposite* of the promise (promised to veto a tax increase, signed one) | 0.0 (and flagged ⚠) | 0.0 |
| **In Progress / Unratable** | Term ongoing and outcome genuinely open, or events made the promise moot (facts changed, office lacks the power and candidate couldn't have reasonably claimed it) | excluded | excluded |
The **Honesty Score** = weighted average of honesty credits over all rated promises.
The **Delivery Score** = weighted average of delivery credits over the same set.
This split is the single most important fairness feature. A governor with a hostile legislature can be perfectly honest (fought for everything promised) and low on delivery. A governor with a friendly supermajority who delivers 90% has shown less about their honesty under adversity — the scores let readers see both truths. Betrayed is displayed distinctly from Abandoned because acting contrary to your word is categorically worse than dropping it.
### 1.4 The "genuine effort" test (what separates Blocked from Abandoned)
Blocked requires documented evidence of at least two of the following, proportional to the office's powers: introduced/formally proposed it (bill, budget line, executive order attempt); publicly advocated for it *while in office* (not just re-promising at re-election); expended political capital (veto, special session call, lobbying record, floor speech); the blocking event is identifiable (failed vote, court injunction, federal preemption). One tweet does not constitute genuine effort. Reviewers must cite the evidence in the rating record.
### 1.5 When ratings happen
- Promises are rated at **end of term** (the fair deadline the politician themselves set) with provisional "In Progress" status shown during the term.
- A promise with an explicit internal deadline ("in my first 100 days") is rated when its deadline passes.
- Re-elected officials: unkept promises re-made in the next campaign re-enter the ledger; quietly dropped ones are rated on the first term.
### 1.6 Review pipeline
LLM proposes a rating with cited evidence → trained human reviewer confirms or overrides → a second reviewer signs off on any Broken/Betrayed rating (the reputationally damaging ones) → published with full evidence trail → subject dispute window (officials' offices can submit evidence; disputes and resolutions are public).
---
## Part 2: The Performance Score
*"Is the place they run actually getting better?"* — Applies to **executives only** (see Part 3).
### 2.1 The metric basket
Six domains, equal default weights (users can re-weight sliders on the site; the official score uses the published defaults). Preregistered metrics per domain:
| Domain | Core metrics (jurisdiction-level) | Primary sources |
|---|---|---|
| **Economy** | Unemployment rate, median household income (real), labor-force participation, GDP per capita growth | BLS, Census ACS, BEA |
| **Public Safety** | Violent crime rate, property crime rate, traffic fatalities | FBI NIBRS, NHTSA |
| **Health** | Age-adjusted mortality, overdose deaths, infant mortality, uninsured rate | CDC WONDER, Census |
| **Education** | HS graduation rate, NAEP 4th/8th grade scores (states), chronic absenteeism | NCES/EDFacts, state DOEs |
| **Environment & Infrastructure** | Days of unhealthy air (AQI), drinking-water violations per system, % roads in poor condition | EPA AQS, EPA SDWIS, FHWA |
| **Fiscal Governance** | Credit rating change, rainy-day fund as % of spending, pension funded ratio, structural balance | Ratings agencies, NASBO, Pew |
Population-level affordability (median rent/income ratio, cost-of-living-adjusted income) is included in Economy for v1.1 consideration — flagged now so adding it later isn't post-hoc.
### 2.2 The computation (per metric)
1. **Tenure window with lag.** Metrics are assigned a preregistered lag reflecting how fast policy can plausibly move them: employment/crime = 1 year; health = 1–2 years; education = 2 years; fiscal = 0 (budgets are directly theirs). An official's window for a lag-2 metric starts 2 years after inauguration and ends 2 years after departure. First-year outcomes on lagged metrics belong to the predecessor. This kills the "inherited a recovery/collapse" problem at the boundary.
2. **Compute the annualized trend** over the window (log change per year for rates).
3. **Subtract the peer median trend** over the identical calendar window. Peers: governors → other 49 states; mayors → cities in the same population tier (5 tiers); president → the state-median trend plus OECD comparison shown as context. This differencing removes national shocks (recessions, pandemics, fentanyl waves) that hit everyone — the score captures *relative* movement.
4. **Standardize** to a z-score within the peer set; winsorize at ±2.5 to stop one freak metric from dominating.
5. **Average metrics → domain score; average domains → composite; convert to percentile rank** among peers ("62nd percentile of governors").
### 2.3 Fairness machinery
- **Minimum tenure:** no Performance Score before 2 full years in office (provisional badge shown instead).
- **Coverage badges:** every metric shows its data coverage (e.g., "FBI data covers 71% of this state's agencies"). Metrics below 60% coverage are shown but excluded from the composite.
- **Shock flags:** jurisdiction-specific catastrophes (hurricane, plant closure) are annotated on the affected series — never silently adjusted, because adjustment is where bias hides. Peer-differencing already handles nationwide shocks.
- **Small-area suppression:** for towns where CDC/FBI suppress small counts, the domain is marked "insufficient data" rather than scored on noise.
- **No endpoint cherry-picking:** trends use every year of the window via regression slope, not first-vs-last year.
- **Term-by-term scoring:** multi-term executives get a score per term plus a career composite, so late collapses or improvements are visible.
---
## Part 3: The Attribution Split
The central question: *who owns outcomes?* The answer differs by how much unilateral power the office holds.
### 3.1 The principle
**Attribution follows the pen.** Executives propose and sign budgets, run agencies, hire, veto, and issue orders — they own outcome trends, adjusted by the machinery above. Legislators cast one vote among many — they own their *votes, words, and work product*, not the district's water quality.
### 3.2 Executives — full scoring
**Presidents, governors, county executives, strong mayors** receive: Honesty Score + Delivery Score + Performance Score.
One adjustment: a **divided-government context badge** (not a score modifier) showing what share of tenure the executive faced an opposition legislature. It contextualizes Delivery without excusing it — voters can weigh it themselves. Making it a badge instead of a multiplier is deliberate: any numerical "adjustment for opposition" would be gamed and disputed forever.
Presidents are the hard case (peer group of one). Solution: score U.S. trends against (a) the pre-tenure 8-year trend and (b) the G7/OECD median over the same window, presented as two separate views rather than falsely merged into one number.
### 3.3 Legislators — record scoring, not outcome scoring
**Members of Congress, state legislators, council members in council-manager or weak-mayor cities** receive:
1. **Honesty Score** (identical rubric — promises vs. votes/actions; a legislator who promised to vote against X and voted for X is Betrayed)
2. **Alignment Score:** % of relevant floor votes consistent with their stated platform positions, including positions too vague to be promises. This catches the politician who never technically promised but campaigns as one thing and votes as another
3. **Effectiveness Score:** modeled on the academic Legislative Effectiveness Score (Center for Effective Lawmaking): bills sponsored weighted by how far they progress (introduced → committee → passed chamber → law), plus substantive amendments and secured district funding, normalized within chamber and majority/minority status — a minority-party member is compared to minority-party peers
4. **District Dashboard (unscored):** the same outcome metrics shown as context, clearly labeled "outcomes reflect many actors; this legislator is one vote of [N]"
The one exception: where a legislative body holds genuinely executive power (council-manager cities, county commissions), the *body* gets a collective Performance Score, and each member's page shows it alongside their individual vote alignment with the majority — so a dissenting member who voted against the failing policies is visibly distinct from the majority that enacted them.
### 3.4 The gray zone table
| Office | Honesty | Delivery | Alignment | Effectiveness | Performance |
|---|---|---|---|---|---|
| President / Governor | ✔ | ✔ | — | — | ✔ (individual) |
| Strong mayor / county executive | ✔ | ✔ | — | — | ✔ (individual) |
| Weak mayor | ✔ | ✔ | — | — | ✔ (shared w/ council, labeled) |
| Legislator (Congress/state) | ✔ | — | ✔ | ✔ | context only |
| Council member (council-manager) | ✔ | — | ✔ | ✔ | ✔ (collective, with dissent tracking) |
| Elected sheriff / DA / school board | ✔ | ✔ | — | — | ✔ on domain-specific metrics only |
Domain-specific offices are scored only on what they control: a sheriff on crime/jail metrics, a school board on education metrics — never on the general basket.
---
## Part 4: Anti-Bias Infrastructure
Methodology alone doesn't create trust; auditable process does.
1. **Annual symmetry audit:** regress published scores on party after controlling for the underlying data inputs. Any residual party effect triggers a public investigation of which rules produced it. Publish the audit either way.
*Amendment, July 2026 (logged, not silent): while only mechanically computed scores are published (v1.0 covers one metric with no judgment-based inputs to control for), the audit ships as a two-sided permutation test on the party gap in published scores, run on every data refresh and published on the methodology page. The regression form activates when judgment-based ratings (the Honesty ledger) begin publishing.*
2. **Adversarial review panel:** a standing panel with disclosed members from across the spectrum reviews a random sample of Broken/Betrayed ratings each quarter and every disputed rating. Their dissents are published.
3. **Dispute process with SLA:** any subject's office may submit evidence; response within 30 days; the dispute, evidence, and ruling are permanently public on the rating page.
4. **Versioned everything:** scores display their methodology version; historical scores are never silently recomputed.
5. **Inter-rater reliability:** track reviewer agreement rates (Cohen's kappa) on promise ratings and publish them. If humans can't agree on the rubric, the rubric — not the politicians — gets fixed.
6. **The unratable is visible:** every politician's page shows how many statements were Aspirations and how many promises were Unratable. A politician who promises nothing testable earns no high Honesty Score by default — they earn a badge: "Made few verifiable commitments."
---
## Part 5: Worked Example (illustrative)
**Governor X, one term, opposition legislature for 4 of 4 years.**
- 31 ratable promises extracted (weighted N = 52 after prominence). End of term: 14 Kept, 4 Partial, 7 Blocked (genuine-effort documented), 4 Abandoned, 1 Betrayed, 1 mooted.
- **Honesty = (Kept + Partial×0.5 + Blocked) / rated = 82%** — high; they fought for what they said.
- **Delivery = (Kept + Partial×0.5) / rated = 57%** — middling; divided-government badge shown.
- Performance: economy +0.4σ vs. peer states, safety +0.9σ, health −0.2σ, education −0.6σ (2-yr lag applied), environment +0.1σ, fiscal +1.2σ → **68th percentile of governors**.
- Page verdict a reader takes away: honest, fiscally strong, education slipping — which is exactly the kind of textured, non-tribal conclusion the site exists to produce.
---
*v1.0 — Every rule above is open to challenge before launch. After launch, changes only by public versioned amendment.*
---
## Amendment v1.1 — Basket expansion (July 2026, logged)
Additive only; every v1.0 rule stands unchanged. New preregistered metrics join
the Performance basket, each registered with its direction and lag FIXED BEFORE
any scores were computed, and each shipping only after two-stage independent
verification of its source data (self-verification against published figures,
then an adversarial recheck by a second reviewer):
| Metric | Domain | Direction | Lag | Source (keyless public data) |
|---|---|---|---|---|
| NAEP 8th-grade math score | education | higher is better | 2 | NAEP Data Service |
| Median daily air-quality index | environment | lower is better | 1 | EPA AirData annual files |
| Drinking-water systems with health violations | environment | lower is better | 1 | EPA SDWIS/ECHO |
| Adults without health coverage | health | lower is better | 1 | CDC BRFSS |
| Teen birth rate | health | lower is better | 1 | CDC WONDER natality |
| Poor mental-health days (mean, past 30) | health | lower is better | 1 | CDC BRFSS |
| Homicide death rate | public_safety | lower is better | 1 | CDC WONDER mortality |
| State tax collections per resident | fiscal | lower is better | 0 | Census STC + population estimates |
| Residential electricity price | economy | lower is better | 1 | EIA |
| 4th-graders below basic reading | education | lower is better | 2 | NAEP Data Service |
| Adult obesity prevalence | health | lower is better | 2 | CDC BRFSS |
| Firearm death rate | public_safety | lower is better | 1 | CDC WONDER mortality |
Context-only data (displayed, never graded): school-shooting deaths and
injuries (K-12 School Shooting Database) — most state-years are zero, and
rare-event counts cannot support fair trend-vs-peers scoring. Requested but
unmeasurable pending a defensible source: school food quality (no national
state-by-state quality measure exists); adult illiteracy (no annual series
exists — the 4th-grade below-basic reading share is the preregistered proxy,
labeled as what it is).
Aggregation (unchanged in form, now meaningful): domain score = mean of the
term's scored metric z's in that domain; composite = mean of domain scores.
**Composite gating:** a composite publishes only with >= 3 scored categories;
it is thin-flagged when every underlying category is thin. Ordinal composite
ranks go only to non-thin composites.
Value-judgment disclosure: treating lower taxes and lower electricity prices
as "better" is a preregistered editorial choice, applied identically to every
governor and disclosed here; the trend-vs-peers form means the score measures
whether the burden GREW slower than peers, not its level. Two requested
categories are deliberately NOT scored in v1.1: Honesty (requires two agreeing
human reviews per Part 1 before anything publishes) and spending discipline /
"waste" (no defensible preregistered rubric yet). Their card slots show status,
not numbers, until those bars are met.
---
## Amendment v1.2 — Weighting, fair adjustments, transparency (July 2026, logged)
Additive; every v1.0/v1.1 rule stands. Directions, lags, weights and the
transparency penalty were all fixed before any v1.2 score was computed.
**1. Tax burden as share of income.** The fiscal domain's tax metric changes
from state taxes per resident to **state taxes ÷ residents' personal income**
(Census STC numerator, BEA SAINC1 denominator). Per-capita punished wealthy
states by arithmetic; share-of-income does not. The old per-capita metric is
retired to unscored context.
**2. Social-services domain.** Two preregistered metrics: **SNAP Program Access
Index** (USDA FNS — participants relative to eligible people in need; higher =
easier access) and **unemployment-insurance recipiency rate** (US DOL — share
of unemployed receiving UI; higher = easier access). Physically-impossible
source artifacts (UI recipiency > 100%, one FNS outlier) were dropped and
documented in the staged files, never interpolated.
**3. Citizen-importance weights.** The overall score is published in two forms:
equal-weight, and weighted by the share of Americans calling each domain a "top
priority" in the Pew Research Center January 2024 policy-priorities survey
(economy 73%, health 60%, education 60%, public safety 58%, deficit 54%,
environment 45%, poverty/social services 44%; normalized to sum to 1). Weights
are cited on the site and applied identically to all governors.
**4. Data-transparency penalty.** A flat **0.1σ per state-controlled
data-collection year missed** during a governor's tenure (capped at 0.5σ),
subtracted from the overall score, with missed years disclosed per state. Scope
is strict: it applies ONLY to reporting a state administers and fails to
deliver (e.g. a state's own BRFSS survey year that CDC excluded for
non-participation or quality failure). It is NEVER applied to federal
small-count suppression or any withholding outside the state's control — a
low-count state is not opaque. Flat-per-miss (not share-based) so the same
lapse costs the same regardless of tenure length.
---
## Amendment v2.0 — Senators are scored on their state (July 2026, logged)
**This amends a core rule of v1.0.** It is recorded here rather than applied
silently, and applies identically to every senator, past and present.
**What changed.** v1.0 Part 3 held that Performance Scores go *only* to
executives, and that legislators receive district/state outcomes as UNSCORED
context ("one vote of N"). That blanket exemption is replaced.
**The new rule.** An official is scored on the constituency they actually
answer to, compared only against others holding the same job:
- **Senators** answer to an entire state. They are therefore scored on the same
verified state outcomes as the governor — identical categories, identical
peer math (their state's trend vs. the other 49 over the senator's own
current term), identical citizen-weighted OVR — but **percentiled and ranked
only against other senators**. Mixing offices with different jobs on one
yardstick would be meaningless.
- **House members** answer to a *district*, not a state. They are explicitly
NOT given their state's numbers, which would credit or blame them for the
other 434 districts. District-level scoring awaits district-level data.
- **The data-transparency penalty stays with the executive.** The governor's
administration runs the state's data collections; a senator has no control
over them and is never charged for a missed collection.
**Why this is defensible.** The original concern — misattribution — is
addressed by construction rather than by exemption: the score is explicitly
*relative movement of the state during their term vs. peer states*, never a
claim of sole causation; influence-sharing is disclosed on every senator card
and board; and the peer set is same-office-only. Voters do hold senators
accountable for their state; refusing to measure it was a choice, not a
statistical necessity.
**Unchanged.** Trend-not-level, the lags, peer-differencing, winsorization at
±2.5σ, the citizen weights, the ≥3-category composite gate, thin-data flags,
provisional status, and the promise-rating human-review requirement all stand
exactly as published.
---
## Amendment v2.1 — The blended legislator grade (July 2026, logged)
Preregistered before any grade was computed. Applies identically to every member.
LEG = 0.30·ρ_R·R + 0.35·ρ_B·B + 0.25·ρ_A·A + 0.10·ρ_P·P
**Weights never renormalize.** A block we cannot measure enters at ρ=0 — the
peer mean, which is what "we don't know" means — and its weight is RESERVED,
not spread over the others. Renormalizing would score a freshman 100% on bills
and rank them against a veteran scored on a blend: two yardsticks, one
leaderboard. It also means that when alignment ships, no weight moves.
**Region (R) — 30%, a ceiling.** A state's economy is shaped by a governor, a
legislature, the other senator, the President and a world economy. 30% is the
most we will attribute to one senator.
R = [(1−φ)·OVR_trend_z + φ·OVR_level_z] / SD_cohort
ρ_R = ρ_window × ρ_attr
**φ — the tenure ramp.** "I inherited this" is a true claim with an expiry
date. φ = 0.40 × clamp((T−6)/18, 0, 1): a senator in their first term carries
ZERO level weight; at 24+ years the level is 40% of their region block (12% of
the whole grade — a hard cap). T is continuous **Senate** service: House years
represent a district, not a state, so a House→Senate switcher is not credited
for them. Note the trend was always the *future* term — a slope is a
derivative, the best unbiased estimate of where a state is heading. φ adds the
*present*, priced by ownership.
**Bills (B) — 35%, the largest block**, because a senator's bills are theirs
alone. LES is Blom rank-normalized within (chamber × majority-status), then
winsorized. Two measured findings force this: (1) a log transform fails at the
FLOOR, not the ceiling — the worst lawmaker reaches only −1.57 while the best
clips at +2.5, making legislative failure nearly unpunishable; (2) raw z is
inert — the long tail inflates the SD until 86% of the chamber sits inside
|z|<0.7. Blom is monotone: it cannot change anyone's order, only the spacing.
Within-cell is mandatory, not cosmetic: in the 118th House the majority median
LES was 1.17 vs 0.36 for the minority, a 3.2× gap; scoring across it would
measure who held the gavel.
**Participation (P) — 10%, one-sided.** Showing up is a hygiene factor, not an
achievement: P = −min(2.5, max(0, (missed% − p75)/5)). It can only subtract.
Max damage to any grade: 0.25. *Currently reserved at ρ=0* — our missed-votes
figure is lifetime, and the rule requires per-Congress matched to the LES
congress. Lifetime is a blocker, not a caveat.
**Ideology is excluded entirely.** It measures where a member stands, not
whether they did their job, and no direction sign can be written for it without
knowing which party it helps — it fails Symmetry on its face. Retained for
symmetry audits only, exactly like terms.party.
**House members** carry ρ_R = 0: no defensible district trend exists.
Redistricting leaves at most three annual points on current lines (2020 was
cancelled), and 57 districts across NY, NC, GA, AL and LA were redrawn
mid-decade — Georgia districts swung +60%/−27% on boundary changes alone while
the national median moved +5.3%. Scoring that would publish cartography as
performance. They are graded on what they control until district history is
rebuilt onto current lines from census tracts.