July 20, 2026Methodology v2.1 · Data: BLS, 20102025
Political Grades
government — just the stats

Methodology v2.1

How We Score

Every rule below was fixed before any politician was scored. Changing a rule requires a new methodology version, applied to everyone, with a public changelog. Nobody gets a custom rulebook.

Two scores, never blended

Honesty / Delivery measures whether an official did what they said: campaign promises matched against votes, bills, orders and budgets. It applies to every official. A governor blocked by a hostile legislature after genuinely trying keeps full honesty credit and zero delivery credit — being stopped is not lying. “Genuinely trying” has a preregistered evidence test, not a vibe test: a Blocked rating requires at least two documented actions on the record (a bill formally proposed, public advocacy while in office, political capital spent, an identifiable blocking event) — one tweet does not count. Doing the opposite of a promise (Betrayed) is tracked separately from quietly dropping it (Abandoned), and neither publishes without an evidence trail and two agreeing human reviews. The Honesty ledger is in human review and is not yet published on this site.

Performance measures whether the place got better on the official’s watch. Every official who answers to a constituency is scored on how that place moved during their tenure — but only against others holding the same job, and always as relative movement rather than sole credit. Influence is shared; the score says how the place moved, not that one person moved it.

Who gets which score

The fairness rule: an official is scored on the place they actually answer to, compared only against others holding the same job — and never handed numbers for a constituency that isn’t theirs.

  • Governors (and other executives) sign budgets and run agencies, so they get the full performance score — outcome trends vs. peer states, the card you see on each state page.
  • Mayors of the 50 largest cities are executives too, scored the same way against peer cities — but only where they hold the pen. A strong mayor gets a full score; a council-manager city runs day-to-day through an unelected manager, so the mayor’s score is shared and labeled that way. City outcome metrics are in verification.
  • Senators represent an entire state, so they are scored on the same verified state outcomes as the governor — same categories, same peer math, same citizen-weighted OVR — but ranked against other senators, over their own current term. They share influence with 99 colleagues and the governor, so the score measures relative movement, never sole credit. Alongside it, every senator carries their own record: legislative effectiveness (how far their bills move, benchmarked by the Center for Effective Lawmaking to a chamber average of 1.0), voting statistics (missed votes, ideology position from GovTrack), and — once it publishes — honesty (promises vs. their actual votes). One thing they are never charged for: the state’s data-transparency penalty, because the governor’s administration runs those collections, not a senator.
  • House members represent a district, not a state, so they are not given the state’s numbers — that would credit or blame them for 434 other districts. District-level outcome scoring is in development (almost no federal source publishes by congressional district; it requires Census district tables plus county-to-district apportionment). Until it lands they are scored on their own record, with state outcomes shown as labeled context.
  • The President is an executive but a peer group of one. Rather than force a false comparison, presidential performance is scored two honest ways — U.S. trends vs. the pre-tenure trend, and vs. the OECD median — presented side by side, not merged into a single number. That view is in development.

The Performance Score, step by step

  1. Window with a lag. Each metric carries a preregistered lag reflecting how fast policy can plausibly move it (employment: 1 year). A governor’s window opens one year after inauguration — the first year runs on the predecessor’s budget — and only full calendar years count.
  2. Trend, not level. We fit a log-linear trend to the metric inside the window. Inheriting a bad number costs nothing; only the direction and speed of change during tenure counts.
  3. Subtract the peers. The median trend of the other 49 states over the identical calendar window is subtracted. This removes the shared national shock and the shared recovery — not state-to-state differences in how hard a shock hit, which remain part of each state’s record (see Known limitations).
  4. Standardize and cap. The differenced trend is converted to a z-score — its distance from the peer pack, measured in standard deviations (σ) — then winsorized (extreme values capped at ±2.5σ, not deleted) so no single freak series dominates.
  5. Aggregate and rank. Metric scores average into domain scores, domains into a composite, and the composite becomes a percentile rank among scored governors.

What’s live today

Version v2.1 publishes 16 preregistered categories, each from a federal statistical source and each shipped only after two independent verification passes against published figures:

Educationeducation · higher is better · lag 2 yr · NAEP (Nation's Report Card)
Air qualityenvironment · lower is better · lag 1 yr · EPA AirData
Water qualityenvironment · lower is better · lag 1 yr · EPA SDWIS/ECHO
Unemploymenteconomy · lower is better · lag 1 yr · BLS LAUS
Health coveragehealth · lower is better · lag 1 yr · CDC BRFSS
Tax burdenfiscal · lower is better · lag 0 yr · Census STC ÷ BEA personal income
Energy costeconomy · lower is better · lag 1 yr · EIA
Cost of livingeconomy · lower is better · lag 1 yr · BEA
Crimepublic_safety · lower is better · lag 1 yr · CDC WONDER
Teen birthshealth · lower is better · lag 1 yr · CDC WONDER
Wellbeinghealth · lower is better · lag 1 yr · CDC BRFSS
School shootingscontext only, never graded — public_safety · lower is better · lag 1 yr · K-12 School Shooting Database (CHDS)
Obesityhealth · lower is better · lag 2 yr · CDC BRFSS
Early literacyeducation · lower is better · lag 2 yr · NAEP (Nation's Report Card)
Gun deathspublic_safety · lower is better · lag 1 yr · CDC WONDER
Food-aid accesssocial_services · higher is better · lag 1 yr · USDA FNS
Jobless-benefit accesssocial_services · higher is better · lag 1 yr · US DOL ETA

Unfilled card categories show their honest status instead of a number: Honesty is in human review (nothing publishes without two agreeing reviewers), spending discipline and school food quality await defensible preregistered measures, and any category whose data fails verification stays “pending” rather than shipping shaky numbers. The overall rating (OVR) averages a governor’s scored categories by domain and requires at least 3 of them. Adding a category is additive: published methodology versions are never silently rewritten.

Because every sitting governor’s term is still in progress, every published score is provisional. Scores computed on fewer than three annual data points additionally carry a thin data flag — they are real numbers, honestly computed, on not-much-data-yet.

The legislator grade

A senator’s grade blends three things they do, into one number ranked among senators: the bills they move (35% — a senator’s bills are theirs alone), how their state has fared on their watch (30%, a ceiling — a state’s economy has a governor, a legislature, the other senator and a world economy in it), attendance (10%), and promise-keeping (25%, once human review ships). We never renormalize: a part we can’t measure enters at the peer average and its weight is reserved, so nothing is filled with a guess and the card says how much of the grade is real.

Within the state block, we mostly score the direction the state is moving — a trend is the best estimate of where it’s heading. But the longer someone serves, the more the state’s current condition is genuinely theirs: a first-termer owns none of it, a 24-year veteran owns the most (capped at 12% of the grade). Bills are scored by the Legislative Effectiveness Score, rank-normalized within chamber and majority status so the grade reflects lawmaking, not who holds the gavel. Ideology is excluded entirely — it measures where a member stands, not whether they did their job. House members get the same grade minus the region block: no honest district trend exists yet (redistricting), so it is reserved rather than faked.

Bill-spending vote pies

A member’s votes are part of their record, so each member’s page shows two pies: what they voted to fund and what they voted against, sliced by category, with every dollar quoted from the bill. The dollars are extracted verbatim from the enacted bill text (and, where a landmark law’s real dollars live there, the nonpartisan CBO cost estimate) and cited to the exact section — click a slice to read the line-items.

Two rules keep this honest. Bill selection is mechanical, never editorial: a bill enters only if it became law or got a recorded floor vote, moved at least $1 billion, and has dollars verifiable to a public source — both parties’ signature laws enter, or don’t, by the same numeric test. And not every dollar in a bill is spending: a pie counts appropriations only. Loan-guarantee ceilings (the government backstopping private loans, no cash out), trust-fund obligation limits, authorizations (a ceiling on futurespending) and pay-fors are shown separately and never summed in — a naive “add every dollar” would overstate a single bill by more than tenfold. We extract the largest verified line-items, so a pie total is a floor labeled “tracked,” never presented as the bill’s complete total. See the bills & the full extraction rules →

How the overall score is weighted

The overall rating combines a governor’s domain scores. We publish two versions of it, side by side, so you can see exactly what the weighting does:

  • Equal-weight — every domain counts the same.
  • Citizen-weighted (the headline number) — each domain is weighted by how many Americans call that issue a “top priority,” so the score reflects what people actually say matters most. The weights aren’t our opinion; they come from a national survey, cited below.
economy19% of the overall weight
health15% of the overall weight
education15% of the overall weight
public safety15% of the overall weight
fiscal14% of the overall weight
environment11% of the overall weight
social services11% of the overall weight

Source: Pew Research Center, “Americans’ Top Policy Priorities” (fielded Jan. 16–21, 2024) — the % calling each issue a top priority, normalized to sum to 100%. See the survey. Weights were fixed before any weighted score was computed and apply identically to every governor; changing them will be a logged methodology change.

Fair-comparison adjustments

  • Taxes are measured as a share of income, not dollars per person — otherwise a high-income state would look high-tax by arithmetic alone. We divide each state’s tax collections (Census) by its residents’ total personal income (BEA). Treating lower taxes as “better” is a preregistered editorial choice, applied to every governor; and because we score the trend vs. peers, the number really measures whether the burden grew slower than other states, not whether a state is low-tax.
  • Rates, not raw counts, wherever population matters — homicides per 100,000, teen births per 1,000, and so on — so big states and small states are compared on equal footing.

The transparency penalty

A government that doesn’t collect and report data on its own residents is dodging accountability, and the score says so. When a state fails to deliver a data collection it runs — for example, skipping or flunking the CDC health survey it administers — the governor’s overall score is docked a flat 0.1σ for each missed year, capped at 0.5σ, with the exact missed years listed on the state’s card.

The line we hold carefully: this penalty is only for reporting failures the government itself controls — whether a state, a city, or a federal agency. It is never applied when a statistic is withheld for a reason outside that government’s control — a state with too few homicides to compute a stable rate is safe, not secretive, and is never penalized for it. The penalty is flat per missed year (not a percentage) so a short-tenured official and a long-tenured one pay the same for the same lapse.

Fairness by construction

  • Symmetry. Every rule reads identically with party labels swapped. The scoring code never reads party at all — party is recorded solely so we can audit the output for skew, and that audit is published below with every data refresh.
  • No silent adjustments. A state-specific catastrophe (a hurricane, a plant closure) gets an annotation on the affected series, never a quiet numerical correction — adjustment is where bias hides.
  • Reproducibility. Every score row stores the exact inputs used to compute it. The “Show the Work” table on each state page is that snapshot, rendered.
  • Disputes. Any office named on this site can file a dispute; rulings are published. Corrections get their own changelog entry — never a silent edit.

The symmetry audit, published

Does the output favor a party? We check on every data refresh, and publish the result unedited. As of the current refresh:

Democratic governors scored47 · mean score -0.052σ
Republican governors scored53 · mean score -0.04σ
Gap between party means-0.012σ
Chance of a gap this big if party carried no information89%
Same test, solid-data governors onlygap -0.012σ · 89% chance

Method: two-sided permutation test (10,000 label shuffles, fixed seed 42 for reproducibility). A high percentage means the party gap is indistinguishable from random noise. The scoring pipeline itself never reads party; this audit checks the output anyway, and publishes whatever it finds.

Known limitations

Trust means naming what the score can’t do, not just what it can:

  • One metric so far. The Economy domain currently holds only the unemployment rate — and unemployment can fall for unwelcome reasons (people leaving the labor force) as well as good ones. More preregistered metrics join as their sources clear verification.
  • COVID-era windows. Most scored windows open during the pandemic recovery. Peer-differencing removes the shared shock and shared recovery, but a state whose economy was hit harder (a tourism collapse, say) will also show a steeper rebound — part of what some top and bottom scores measure is that exposure, not gubernatorial skill alone.
  • Different windows in one list. Each governor is compared with all 49 peer states over their own tenure window, so the ranked list pools scores measured over different years and lengths. That is the price of scoring every tenure on its own terms; the scored window is printed beside every number.
  • Trends partly revert. States that entered a window unusually high or low tend to drift back toward the pack. Scoring the change, not the level, limits this — it does not eliminate it.
  • Influence, not control. Governors shape their state’s economy at the margin. The score measures relative movement on their watch; it cannot isolate their causal contribution.

The rulebook, verbatim

This page is a plain-English summary. The binding document is the full preregistered rulebook, published unedited at Methodology v2.1 — full text. Rule changes require a new version, applied to everyone, and are recorded in the changelog & corrections log. The complete dataset behind every page is downloadable at /pwr.json.

Why isn’t my governor scored?

12 governors took office between 2024 and 2026. With per-category attribution lags and annual data, their first scorable year hasn’t been published yet. They appear the moment the data does; an empty score means nothing about their performance.

← Back to the index