Two scores, never blended
Honesty is not numerically scored on this site — a change made openly in v3.0. The promised score (“promises vs. actions,” publishing only after two agreeing human reviews) was never staffable and never published; scoring rules we cannot staff are scores we cannot honestly compute. In its place, the Integrity Record publishes three factual ledgers — checkable spoken commitments tested by code against roll calls, formal endorsements beside the votes on them, and findings by courts and evenly-bipartisan bodies — and adjudicated findings, when final, cap the letter grade under a preregistered rule. The distinction the old design got right survives in the ledgers: being stopped is not lying, so an unkept commitment costs delivery credit beside the quote, never a claimed verdict on intent.
Performance measures whether the place got better on the official’s watch. Every official who answers to a constituency is scored on how that place moved during their tenure — but only against others holding the same job, and always as relative movement rather than sole credit. Influence is shared; the score says how the place moved, not that one person moved it.
Who gets which score
The fairness rule: an official is scored on the place they actually answer to, compared only against others holding the same job — and never handed numbers for a constituency that isn’t theirs.
- Governors (and other executives) sign budgets and run agencies, so they get the full performance score — outcome trends vs. peer states, the card you see on each state page.
- Mayors of the 50 largest cities are executives too, scored the same way against peer cities — but only where they hold the pen. A strong mayor gets a full score; a council-manager city runs day-to-day through an unelected manager, so the mayor’s score is shared and labeled that way. City scoring is live: three economic outcomes — unemployment, real median household income and poverty — against the 50-largest peer group, using the identical trend-vs-peers math as governors. Crime, housing and health are not yet included, so a city score is narrower than a governor’s, and every city page says so.
- Senators represent an entire state, so they are scored on the same verified state outcomes as the governor — same categories, same peer math, same citizen-weighted OVR — but ranked against other senators, over their own current term. They share influence with 99 colleagues and the governor, so the score measures relative movement, never sole credit. Alongside it, every senator carries their own record: legislative effectiveness (how far their bills move, benchmarked by the Center for Effective Lawmaking to a chamber average of 1.0), voting statistics (missed votes and party-line agreement, each counted from the chambers’ own published roll calls and printed with its denominator), with their Integrity Record beside the grade — facts, not a score; adjudicated findings cap the grade. One thing they are never charged for: the state’s data-transparency penalty, because the governor’s administration runs those collections, not a senator.
- House members represent a district, not a state, so they are not given the state’s numbers — that would credit or blame them for 434 other districts. District-level outcome scoring is in development (almost no federal source publishes by congressional district; it requires Census district tables plus county-to-district apportionment). Until it lands they are scored on their own record, with state outcomes shown as labeled context.
- The President is an executive but a peer group of one. We built a presidential grade, tested it, and refused to ship it: the only honest whole-country benchmark — its own prior trajectory — turned out to be dominated by crisis timing, so whoever inherits a bad year looks good regardless of what they did. A single ranking would have measured luck. The president pages therefore publish the national outcomes during each term next to the OECD median, as data, not a grade, and no composite is computed.
The Performance Score, step by step
- Window with a lag. Each metric carries a preregistered lag reflecting how fast policy can plausibly move it, in years: 0 for tax burden; 2 for education, obesity and early literacy; 1 for the other 12. A metric’s window opens that many years after inauguration, so a lag-1 metric’s first year runs on the predecessor’s budget while a lag-0 metric starts the year they take office — a budget is directly theirs. Only full calendar years count, and every card prints the window it used.
- Standing first — 60% of every category score. How does the state compare with the other 49 right now, at the end of the scoring window? Below average is a negative, above average a positive, with the sign flipped for the metrics where low is good (homicide, obesity, cost of living). This is the larger half deliberately: through v2.3 the score was mostly trend, and Oklahoma ranked No. 2 while standing below the national average on 11 of its 16 categories — teen births, wellbeing, early literacy, education and health coverage among them — because it is cheap and improving. A board that calls that second-best in the country fails the reader.
- Direction second — the other 40%. The fitted trend across the window, minus the peer-median trend over the identical years. Improving fast still moves a score a long way, and it is the fastest thing a governor can actually change. The weight is the same 60/40 for every governor, because a ranked list whose members were scored by different formulas is not a ranking.
The trade-off, stated plainly: this board therefore ranks state outcomes, with the sitting governor named — not a governor’s personal performance. A governor two years in inherited nearly all of their state’s standing. Nobody is scored before two years in office, and both halves are published separately on every card so neither hides inside the other. - Subtract the peers. The median trend of the other 49 states over the identical calendar window is subtracted. This removes the shared national shock and the shared recovery — not state-to-state differences in how hard a shock hit, which remain part of each state’s record (see Known limitations).
- Standardize and cap. The differenced trend is converted to a z-score — its distance from the peer pack, measured in standard deviations (σ) — then winsorized (extreme values capped at ±2.5σ, not deleted) so no single freak series dominates.
- Aggregate and rank. Metric scores average into domain scores, domains into a composite, and the composite becomes a percentile rank among scored governors.
What’s live today
Version v3.0 publishes 16 graded categories, each from a federal statistical program and each shipped only after two independent verification passes against published figures — plus one context-only series not from a federal statistical program (the row below names the source), shown but never graded:
| Education (scale score) | education · higher is better · lag 2 yr · NAEP (Nation's Report Card) data through 2024 · source checked Sept. 3, 2026 |
| Air quality (AQI) | environment · lower is better · lag 1 yr · EPA AirData data through 2024 · source checked Sept. 3, 2026 |
| Water quality (% of systems with a violation) | environment · lower is better · lag 1 yr · EPA SDWIS/ECHO data through 2025 · updated by hand· EPA ECHO bulk file, refreshed by hand on release |
| Unemployment (%) | economy · lower is better · lag 1 yr · BLS LAUS data through 2025 · source checked Sept. 3, 202611-month average — October 2025 was never collected (federal shutdown) |
| Health coverage (% of adults 18-64 uninsured) | health · lower is better · lag 1 yr · CDC BRFSS data through 2024 · source check failing since Sept. 3, 2026 |
| Tax burden (% of personal income) | fiscal · lower is better · lag 0 yr · Census STC ÷ BEA personal income data through 2024 · source checked Sept. 3, 2026 |
| Energy cost (cents/kWh) | economy · lower is better · lag 1 yr · EIA data through 2024 · source checked Sept. 3, 2026 |
| Cost of living (index, US=100) | economy · lower is better · lag 1 yr · BEA data through 2024 · source checked Sept. 3, 2026 |
| Homicide (age-adjusted deaths per 100,000) | public_safety · lower is better · lag 1 yr · CDC WONDER data through 2024 · updated by hand· CDC WONDER blocks automated queries; updated by hand on release |
| Teen births (per 1,000 females 15-19) | health · lower is better · lag 1 yr · CDC WONDER data through 2024 · source checked Sept. 3, 2026 |
| Wellbeing (mean poor-mental-health days of 30) | health · lower is better · lag 1 yr · CDC BRFSS data through 2024 · updated by hand· CDC blocks cloud-hosted runners (403); refreshed from a local run or by hand |
| School shootings (people killed or wounded) | context only, never graded — public_safety · lower is better · lag 1 yr · K-12 School Shooting Database (CHDS) data through 2024 · source checked Sept. 3, 2026 |
| Obesity (% of adults) | health · lower is better · lag 2 yr · CDC BRFSS data through 2024 · source checked Sept. 3, 2026 |
| Early literacy (% of 4th-graders below basic) | education · lower is better · lag 2 yr · NAEP (Nation's Report Card) data through 2024 · source checked Sept. 3, 2026 |
| Gun deaths (age-adjusted deaths per 100,000) | public_safety · lower is better · lag 1 yr · CDC WONDER data through 2024 · updated by hand· CDC WONDER blocks automated queries; updated by hand on release |
| Food-aid access (index (higher = easier access)) | social_services · higher is better · lag 1 yr · USDA FNS data through 2023 · source checked Sept. 3, 2026 |
| Jobless-benefit access (% of unemployed receiving UI) | social_services · higher is better · lag 1 yr · US DOL ETA data through 2024 · source checked Sept. 3, 2026 |
Unfilled card categories show their honest status instead of a number: Honesty is not numerically scored (v3.0 — see the Integrity Record rulebook; each governor now carries the record’s Adjudicated tier, with its admission rule published and its event feed not yet wired), spending discipline and school food quality await defensible preregistered measures, and any category whose data fails verification stays “pending” rather than shipping shaky numbers. The overall rating (OVR) averages a governor’s scored categories by domain and requires at least 3 of them. Adding a category is additive: published methodology versions are never silently rewritten.
Because every sitting governor’s term is still in progress, every published score is provisional. Scores computed on fewer than three annual data points additionally carry a thin data flag — they are real numbers, honestly computed, on not-much-data-yet.
The legislator grade
A senator’s grade blends what they do, into one number ranked among senators: the bills they move (30% — a senator’s bills are theirs alone), how their state has fared on their watch (25%, a ceiling — a state’s economy has a governor, a legislature, the other senator and a world economy in it; since v2.6 the formula inside is identical for every senator: trend fitted over their own Senate service, blended 40/60 with current condition — the same blend the governor board uses — with the tenure ramps that once varied the formula per senator removed, for the reason v2.4 published: a fixed weight is the only honest basis for an ordinal list), fiscal votes (10%, added in v2.5 — the deficit footprint of the laws they voted for, from verified CBO estimates; a member needs recorded votes on at least three scored laws to be measured, and the scored laws today are dollar-dominated by the largest two — one from each party’s trifecta — which the changelog states plainly), attendance (10%), and promise-keeping (25%, once human review ships). We never renormalize: a part we can’t measure enters at the peer average and its weight is reserved, so nothing is filled with a guess and the card says how much of the grade is real.
What that costs today, stated plainly: are reserved for every graded member and state outcomes additionally for the whole House. So a typical House member’s grade rests on bills, fiscal votes and attendance — 0.67 of the weight actually measured — and a typical senator’s on state outcomes, bills, fiscal votes and attendance, 1.00. For 82 of the 531 graded members no block is measured at all; the card prints 0.00 coverage rather than a grade dressed up as complete.
Within the state block, current condition and direction are blended in a single published ratio, the same 60/40 for every senator — 60% where the state stands now, 40% which way it moved — with the trend fitted over that senator’s own Senate service. Nothing in it varies with how long they have served: the tenure ramp that did that was removed in v2.6, because it scored different senators with different formulas inside one ranked list. The state block is 33% of the grade, so the most a state’s standing can ever move a senator’s grade is 20% of it. Bills are scored by the Legislative Effectiveness Score, rank-normalized within chamber and majority status so the grade reflects lawmaking, not who holds the gavel. Ideology is excluded entirely — it measures where a member stands, not whether they did their job. House members get the same grade minus the region block: no honest district trend exists yet (redistricting), so it is reserved rather than faked.
Bill-spending vote pies
A member’s votes are part of their record, so each member’s page shows two pies: what they voted to fund and what they voted against, sliced by category, with every dollar quoted from the bill. The dollars are extracted verbatim from the enacted bill text (and, where a landmark law’s real dollars live there, the nonpartisan CBO cost estimate) and cited to the exact section — click a slice to read the line-items.
Two rules keep this honest. Bill selection is preregistered, never editorial — and an audit of our own earlier sentence is why it now says more. The old test (“at least $1 billion”) admits roughly 160 laws while this site extracts a handful, which means an unstated “landmark” judgment was doing the real selecting; an unstated filter is exactly where a preference can hide, so we replaced it with a stated one. The rule: every public law since the 113th Congress (2013) whose enacted text carries a single dollar figure of $50 billion or more is a candidate — one figure, never a sum, because summing every dollar string overstates a bill more than tenfold — and a candidate enters the corpus once hand extraction verifies at least $100 billion of budget authority. Both thresholds apply identically to every Congress and both parties, and may only ever be lowered for everyone, never waived for one bill. A recorded floor vote in either chamber qualifies; where a chamber passed a law by voice vote, that fact is disclosed on every member’s row and positions are never invented. The queue of qualifying laws not yet extracted is published — a coverage gap this site owns, not a choice it hides — and so is the other half of the rule’s work: every law that was read and did not qualify, with the reason. That is where the famous laws are — defence authorization acts that appropriate nothing, budget acts that are caps and pay-fors, a debt-limit statute — and publishing the verdicts is the only way a reader can check that the threshold was applied rather than invoked — and the extractions made before this rule existed remain, stated as grandfathered, because the corpus is append-only. One entry is a deliberate exception, and named as one: H.R. 4820 — Reported in House (RH). It never became law, so it is joined to no member’s votes and no member is credited or charged for a dollar in it; it is on file only as the plainest illustration of the mechanism rule that follows. The 30 bills that do feed a member’s pies are all enrolled public laws. And not every dollar in a bill is spending: a pie counts real outlays — appropriations plus mandatory (direct) spending, the two mechanisms that actually move cash. 15 of the 31 laws on the site carry mandatory spending in their pie, and on 5 of those there are no appropriations in it at all. Loan-guarantee ceilings (the government backstopping private loans, no cash out), trust-fund obligation limits, authorizations (a ceiling on futurespending) and pay-fors are shown separately and never summed in — a naive “add every dollar” would overstate a single bill by more than tenfold. We extract the largest verified line-items, so a pie total is a floor labeled “tracked,” never presented as the bill’s complete total. See the bills & the full extraction rules →
Comparing dollars from different years
A dollar in a 2013 bill and a dollar in a 2025 bill are not the same quantity, and for a long time every figure on this site was nominal — so the charts carried a warning (“a longer bar can be higher prices rather than more spending”) that was honest and left the reader nowhere to go. Where two years are now compared, they are put on one footing using the composite outlay deflator published in OMB’s Historical Tables, indexed to FY2017 = 1.000.
Deliberately not a consumer price index. CPI measures a household’s shopping basket; the federal government does not buy that basket — it buys military pay, procurement, medical care, grants to states and interest. The composite outlay deflator is weighted to that actual mix, and it is the index OMB deflates its own constant-dollar table with, which means a constant-dollar figure here can be checked against the government’s own published one rather than resting on a choice made on this site. Before the series is allowed to ship, three checks must pass: OMB’s Table 1.3 must publish the same deflator year for year; constant-dollar outlays times the deflator must reproduce current-dollar outlays in every year (the check that catches an index applied upside down, which passes every other test and is wrong on every year); and the outlays must agree with the figures this site already publishes from a different OMB table, themselves cross-checked against the Treasury’s Monthly Treasury Statement.
Two limits, stated rather than buried. The deflator is applied at the fiscal year of a law’s action date, which is an approximation for money that spends across a decade — so the nominal figure, the one quoted from the bill text, is always printed beside the constant-dollar one rather than replaced by it. And a second yardstick that needs no deflator at all is printed next to both: the tracked total against one fiscal year of total federal outlays. It is not a claim that a law was that share of that year’s budget; it is a measure whose meaning does not drift between 1962 and 2025, which is exactly what nominal dollars are not. Federal spending by year →
How the overall score is weighted
The overall rating combines a governor’s domain scores. We publish two versions of it, side by side, so you can see exactly what the weighting does:
- Equal-weight — every domain counts the same.
- Citizen-weighted (the headline number) — each domain is weighted by how many Americans call that issue a “top priority,” so the score reflects what people actually say matters most. The weights aren’t our opinion; they come from a national survey, cited below.
| economy | 19% of the overall weight |
| health | 15% of the overall weight |
| education | 15% of the overall weight |
| public safety | 15% of the overall weight |
| fiscal | 14% of the overall weight |
| environment | 11% of the overall weight |
| social services | 11% of the overall weight |
Source: Pew Research Center, “Americans’ Top Policy Priorities” (fielded Jan. 16–21, 2024) — the % calling each issue a top priority, normalized to sum to 100%. See the survey. Weights were fixed before any weighted score was computed and apply identically to every governor; changing them will be a logged methodology change. Two honest caveats about that survey: it is one poll by one organization, and its question asked what the president and Congress should prioritize — federal priorities, applied here to governors. That wording gap is why the equal-weight score is published beside the weighted one on every card: if you distrust the weights, the unweighted answer is one line away.
Choices both sides hate
Every scored metric needs a preregistered direction — which way is “better” — and several of those directions are value judgments that draw fire from opposite ends. We list them here with the attack each one invites, because a choice you have to discover is a choice that looks hidden. 13 of the 16 live metrics score lower-is-better, 3 higher-is-better — the full list with each direction sits in the metric table above, and every direction was fixed before any state was scored.
- Tax burden and cost of living, lower is better — readable from the left as a prize for austerity and low wages. It is a preregistered editorial choice: the site scores what a dollar of income endures, not what a government spends it on.
- Food-aid access and jobless-benefit access, higher is better — readable from the right as a prize for enrollment. Both are access measures (how easily an eligible person can use a program that exists), not caseload counts — and they are the mirror image of the tax choice above: one direction each side dislikes, kept on the same board.
- Gun deaths, lower is better — attacked as scoring gun ownership, since most gun deaths are suicides. The metric is age-adjusted deaths; the direction judgment is only that fewer deaths are better. (The board was re-run without this metric during a red-team check: no governor moves more than four places, 14 of 29 do not move at all — and the top two swap. One metric of sixteen is not carrying the board, and the exact movement is stated here rather than an adverb.)
- Wellbeing (self-reported poor-mental-health days) — a phone survey answer, vulnerable to regional differences in what people tell a surveyor. It stays because it is the only state-level mental-health series we found with a consistent method, and its weight is one metric among 16.
Public safety, monthly (context, never scored)
City pages carry a monthly hate/bias-crime module under one preregistered rule: any of the 50 roster cities whose own government or police department publishes machine-readable hate/bias-crime data through an official open-data portal, currently maintained, gets the identical module — monthly incident counts by the city’s own published bias motives, compared same months year over year, rises and falls printed alike. Third-party aggregators and federal roll-ups never qualify; a city without a qualifying feed has no module and no inferred number. The counts are context: they enter no grade, carry no rank, and are shown with their source, their data-through month, and the standing caveat that an incident count is a report, not a conviction, and moves with reporting as well as incidence. Trailing stray records past the last real month are excluded by a stated threshold and the exclusion is counted on the page rather than made silently.
What keeps these choices honest is not that they are neutral — no direction is — but that they are symmetric and inspectable: fixed in advance, identical for every state, published beside an equal-weight score, and stress-tested by the party-symmetry audit below. If you would flip a direction, the data files are public — recompute the board your way.
Fair-comparison adjustments
- Taxes are measured as a share of income, not dollars per person — otherwise a high-income state would look high-tax by arithmetic alone. We divide each state’s tax collections (Census) by its residents’ total personal income (BEA). Treating lower taxes as “better” is a preregistered editorial choice, applied to every governor; and because we score the trend vs. peers, the number really measures whether the burden grew slower than other states, not whether a state is low-tax.
- Rates, not raw counts, wherever population matters — homicides per 100,000, teen births per 1,000, and so on — so big states and small states are compared on equal footing.
The transparency penalty
A government that doesn’t collect and report data on its own residents is dodging accountability, and the score says so. When a state fails to deliver a data collection it runs — for example, skipping or flunking the CDC health survey it administers — the governor’s overall score is docked a flat 0.1σ for each missed year, capped at 0.5σ, with the exact missed years listed on the state’s card.
The line we hold carefully: this penalty is only for reporting failures the government itself controls — whether a state, a city, or a federal agency. It is never applied when a statistic is withheld for a reason outside that government’s control — a state with too few homicides to compute a stable rate is safe, not secretive, and is never penalized for it. The penalty is flat per missed year (not a percentage) so a short-tenured official and a long-tenured one pay the same for the same lapse.
Fairness by construction
- Symmetry. Every rule reads identically with party labels swapped. The scoring code never reads party at all — party is recorded solely so we can audit the output for skew, and that audit is published below with every data refresh.
- No silent adjustments. A state-specific catastrophe (a hurricane, a plant closure) gets an annotation on the affected series, never a quiet numerical correction — adjustment is where bias hides.
- Reproducibility. Every score row stores the exact inputs used to compute it. The “Show the Work” table on each state page is that snapshot, rendered.
- Disputes. Any office named on this site can file a dispute; rulings are published. Corrections get their own changelog entry — never a silent edit.
The symmetry audit, published
Does the output favor a party? We check on every data refresh, and publish the result unedited. As of the current refresh:
| Democratic governors scored | 19 · mean score -0.07σ |
| Republican governors scored | 19 · mean score -0.109σ |
| Gap between party means | 0.039σ |
| Chance of a gap this big if party carried no information | 69% |
| Same test, solid-data governors only | 29 governors · gap 0.065σ · 54% chance |
| The legislator grade — the number the Congress board ranks on, tested for the first time | 220 D / 227 R · gap 0.15σ toward Democrats · 0.02% chance if party carried no information |
| Same test, House only | gap 0.117σ · 0.98% chance |
| Same test, Senate only | gap 0.282σ · 1.07% chance |
| Senators, tested separately | 28 D / 34 R · gap -0.038σ · 66% chance |
| Everyone pooled (governors + senators) | 47 D / 53 R · gap -0.013σ · 83% chance |
| Everyone pooled, solid data only | 26 D / 33 R · gap 0.109σ · 11% chance |
Method: two-sided permutation test (10,000 label shuffles, fixed seed 42 for reproducibility). A high percentage means the party gap is indistinguishable from random noise. The scoring pipeline itself never reads party; this audit checks the output anyway, and publishes whatever it finds — all five cohorts, including the least flattering. One limit, stated: with 19 governors a side, only a large skew would be detectable; a subtle one would pass this test. The pooled rows run over larger samples for exactly that reason.
Known limitations
Trust means naming what the score can’t do, not just what it can:
- Thin domains. 5 of the 7 scored domains rest on one or two metrics — fiscal on 1, education on 2, environment on 2, public safety on 2, social services on 2. A domain built from one or two series inherits whatever those series measure and whatever they miss. The extreme case is fiscal: one series — tax burden, Census STC ÷ BEA personal income — so the whole domain moves with that one number. Separately, a headline metric can move for more than one reason: the unemployment rate can fall because people left the labor force as well as because more of them found work. More preregistered metrics join as their sources clear verification.
- COVID-era windows. Most scored windows open during the pandemic recovery. Peer-differencing removes the shared shock and shared recovery, but a state whose economy was hit harder (a tourism collapse, say) will also show a steeper rebound — part of what some top and bottom scores measure is that exposure, not gubernatorial skill alone.
- Different windows in one list. Each governor is compared with all 49 peer states over their own tenure window, so the ranked list pools scores measured over different years and lengths. That is the price of scoring every tenure on its own terms; the scored window is printed beside every number.
- Trends partly revert. States that entered a window unusually high or low tend to drift back toward the pack. Scoring the change, not the level, limits this — it does not eliminate it.
- Influence, not control. Governors shape their state’s economy at the margin. The score measures relative movement on their watch; it cannot isolate their causal contribution.
The UK edition (beta)
The UK pages republish Parliament’s own records — the Commons and Lords rosters, recorded division (vote) results, and bill progress — fetched from Parliament’s official APIs and shown with the same derive-don’t-type discipline as the US pages. Nothing in the UK edition is graded yet: no MP, peer or party carries a score, and the participation figures state their own window inline (a few weeks of divisions is a snapshot, and each page says exactly which weeks). Grades and outcome scorecards, if they come, will get their own preregistered methodology section here first. Parliamentary data is used under the Open Parliament Licence v3.0; devolved-nation statistics carry Office for National Statistics attribution on the page that uses them.
The rulebook, verbatim
This page is the plain-English statement of the rules in force (v3.0). The original rulebook — the v1.0 text with its logged amendments — is published unedited at Methodology v1.0 — full text, as amended; the amendments since are summarized here and dated in the changelog. Rule changes require a new version, applied to everyone, and are recorded in the changelog & corrections log. The complete dataset behind every page is downloadable at /governors.json.
Why isn’t my governor scored?
12 governors took office between 2024 and 2026. With per-category attribution lags and annual data, their first scorable year hasn’t been published yet. They appear the moment the data does; an empty score means nothing about their performance.