What is a risk-limiting audit, and why does it matter?
In a race decided by 369 votes, officials had nothing independent to recount — because the machines were paperless. That gap still exists in elections near you.
It is 2 a.m. on election night in Sarasota County, Florida, and the tally for Florida's 13th Congressional District is finishing its crawl across the screen. The race ends with a margin of 369 votes. But buried in the same numbers is something stranger: roughly 18,000 ballots in the county recorded no choice at all in that race. Not a protest. Not an accident of demographics. One in eight voters, apparently, skipped the most contested contest on their ballot.
Nobody could explain it. And nobody could check it.
The county's touchscreen machines — the ES&S iVotronic — left no independent paper record behind. When the U.S. Government Accountability Office later tested the systems, it could not identify a machine malfunction that caused the undervote. It also noted, with the quiet understatement of a federal audit, that a voter-verified paper trail "could have provided independent confirmation that the touchscreens recorded votes correctly." There was no such trail. The GAO's full report is here.
In a race decided by 369 votes, the audit method available was: trust the machine.
That was 2006. The question for today is whether the infrastructure we now have — paper ballots, post-election audits, the statistical tools to use them — is actually strong enough to replace that kind of blind trust with something a skeptical citizen can independently verify. The answer is: sometimes. In some places. Under some conditions. And the gap between "sometimes" and "always" is the whole problem.
Why audits exist at all — and why most of them are weak
Every election produces a number. That number comes from a machine, or a hand count, or both. The question an audit tries to answer is: did the reported winner actually win?
That sounds obvious. But the mechanics matter enormously, because most of the world's post-election audits are not actually designed to answer that question with statistical confidence. They are designed to look like they do.
The most common type is the fixed-percentage audit: after the election, officials pull a set percentage of ballots — say, 1% or 5% — examine them by hand, and compare the result to the machine tally. If they match closely enough, everyone goes home. This approach is intuitive. It feels rigorous. It is not.
Here is the problem. A 1% sample in a landslide race and a 1% sample in a race decided by 0.1% of votes are doing completely different statistical work. In the landslide, even a moderately flawed count would be so obvious a small sample would catch it. In the razor-thin race — the one where catching an error actually matters — that same 1% sample might have almost no chance of detecting a problem. The closer the race, the more checking you need. A fixed percentage inverts the logic.
And yet fixed-percentage audits are still standard practice in the majority of democracies. The gap between the audit's apparent rigor and its actual statistical power is precisely where a miscounted election can disappear unnoticed.
What a risk-limiting audit actually does differently
A risk-limiting audit, or RLA, starts from the opposite direction.
Instead of asking "what percentage should we check?", it asks: "what is the maximum probability we are willing to accept that a wrong outcome slips through?" That probability — the risk limit — is set before the audit begins. Colorado, which pioneered this approach, started at 9% and later tightened it to 3%. A 5% risk limit means: if the reported winner actually lost, there is at most a 1-in-20 chance this audit would fail to catch it.
The math then works backward from that guarantee to determine how many ballots need to be hand-examined. And the result is counterintuitive in the best possible way: the closer the race, the larger the sample the RLA demands. A blowout might need only a few hundred ballots checked by hand. A race decided by a fraction of a percent might require a full hand recount of every ballot in the jurisdiction — which is exactly what Georgia had to do in 2020.
The other critical ingredient is a trustworthy paper record. An RLA cannot audit a digital ledger against itself. It needs physical, voter-marked ballots — the kind a voter actually touched — to serve as the ground truth. Sarasota County in 2006 had no such thing. An RLA conducted there would have been mathematically sound and physically impossible at the same time.
Colorado, 2017: the first state to actually do it
On November 22, 2017, Colorado's Secretary of State announced that the state had completed the first statewide risk-limiting audit in American history.
This is not a trivial claim. Plenty of states have conducted audits. Colorado was the first to conduct one that came with a mathematically defined confidence level — one that anyone, in principle, could verify by checking the statistical work. The procedure is public. The FAQ the Secretary of State published explains the methodology in plain language. The sample sizes, the risk limit, the ballots selected — all of it documented.
The reason this matters is not that Colorado's 2017 election was controversial. It wasn't. The reason is that a method proven in an uncontroversial election is the method you want waiting for you when the controversial one arrives. You do not want to be inventing your audit infrastructure the week after a disputed result. Colorado built its before it needed it.
The key ingredients Colorado assembled: mandatory paper ballots (the physical ground truth), a public risk limit set before the election, and a random sample drawn in a transparent, documented way. Take any one of those away and the statistical guarantee collapses.
Georgia, 2020: when the math demands a full hand count
Colorado's RLA in 2017 was, by statistical luck, efficient. The margins were wide enough that relatively small samples provided the required confidence.
Georgia in November 2020 had no such luck.
The presidential margin in Georgia was narrow enough that when officials applied the risk-limiting audit framework, the math gave them no shortcut: to meet the required confidence level, they had to hand-examine every ballot. All roughly five million of them. Across 159 counties. Examining 41,881 batches. In less than six days.
They did it. Georgia Public Broadcasting reported that the hand count affirmed the machine-tabulated outcome, with the variation between the original count and the hand tally coming in at about a tenth of one percent. NPR's reporting confirmed the result.
Here is what the Georgia case actually proved: paper ballots plus a sound audit method lets anyone check a close result, instead of accepting the machine's word for it. Five million ballots. An observable, documented process. A result that any county-level observer could cross-check against their own jurisdiction's totals.
That is not nothing. That is enormously more than what Sarasota County had in 2006 with its paperless touchscreens and its 369-vote margin and its 18,000 missing choices.
But Georgia's case also contains a subtler lesson. The variation of about a tenth of a percent sounds small. It is small. And in a race of five million ballots it did not affect the outcome. But a tenth of a percent of five million is five thousand votes. In a different race — a state legislative contest, a county commission seat, a race decided by hundreds rather than thousands — a similar rate of hand-count variation could matter. Even the gold standard of hand counting carries irreducible human error. Tired counters misread ballots. Stray marks get interpreted differently by different people. Adjudication calls — what counts as a valid mark? — vary by table.
This is not a knock on Georgia's process. It is a statement about the limits of any human-operated system.
Michigan, 2020: what happens when there's no safeguard before publication
Antrim County, Michigan is a reliably Republican county. When the first unofficial results on the night of November 4, 2020 showed Joe Biden ahead by roughly 3,000 votes with most ballots counted, the number drew immediate scrutiny — not because of any internal alarm system, but because anyone who knew the county's politics knew the result was implausible.
The cause turned out to be human error. County clerk Sheryl Guy had reprogrammed some but not all of the cards used in the tabulators after a last-minute change to a local race, so some precinct data combined incorrectly when the unofficial totals were compiled. Guy said the error was her own. FactCheck.org's detailed account makes clear this was an administrative mistake, not a software conspiracy. The county corrected the totals on November 5. The certified outcome was unaffected.
On December 17, to settle the matter definitively, officials conducted a hand audit of every presidential ballot in the county. The Michigan Department of State announced that the hand count — 9,759 for Trump, 5,959 for Biden — affirmed the accuracy of the election results.
Now stop there. Read those two numbers back. The hand count and the machine tabulation differed by about a dozen votes out of roughly 15,700 cast. The Michigan Department of State presented this as confirmation. The Secretary of State called it affirmation of accuracy.
But what does a 12-vote spread actually mean?
In Antrim County's presidential race, with its large margin, 12 votes is noise. It doesn't change the result, and it probably reflects exactly what experts would predict: a mix of human miscounts during the hand tally, ambiguous ballot marks, and adjudication judgment calls that went slightly differently by different counters at different moments. Hand counting is not a perfect oracle. It is a highly reliable but imperfect human process, and its error rate — small as it is — becomes the conversation in any race decided by a similarly small number.
The deeper problem that Antrim County illustrates is not the 12 votes. It is the moment before the correction. Nothing in the published account of Antrim County credits an internal safeguard with catching the wrong numbers before they were posted. The error surfaced because the result was glaringly implausible — a county-wide political profile that anyone could observe made the published numbers absurd. A subtler error, in a close statewide race, shifting a few hundred votes without flipping a county's obvious partisan lean, might have passed the plausibility check entirely. You cannot audit for what you cannot see.
That is the argument for verifiability that does not depend on implausibility as the alarm.
The gap between "we ran an audit" and "you can verify the audit"
An RLA sounds like a complete solution. In many ways, it is the best post-election verification tool available at scale. But it is worth being precise about what it does and does not provide.
What a well-run RLA provides:
- A defined, quantified confidence level that the reported winner actually won
- A sample drawn by a process that is transparent and reproducible
- A comparison against physical ballots that voters actually marked
What a well-run RLA still requires you to trust:
- That the paper ballots being audited have not been altered since they were cast
- That the chain of custody for those ballots — the sealed bags, the locked rooms, the transfer logs — held at every step
- That the jurisdiction conducting the audit did not make errors in its own audit process
The last one is not hypothetical. Antrim County ran a hand audit and got a number that differed by 12 from the machine count. Georgia ran an RLA and got a variation of about a tenth of a percent. Those are good results. They are not zero. And the entities verifying the audit results were, in both cases, the same government bodies whose work was being scrutinized.
"The audit confirmed it" and "the Secretary of State says it's fine" are reassurances, not proof. They may well be true. But reassurance is not verifiability. Verifiability means an independent party — a researcher, a journalist, a skeptical citizen with a spreadsheet — can check the work using the same inputs and get the same answer.
The gap between "we ran an audit" and "you can verify the audit independently" is exactly the gap that software-and-cryptography-based approaches exist to close. Techniques like end-to-end verifiable voting systems allow voters to confirm their own ballot was counted without revealing how they voted, and allow any observer to verify the aggregate result without trusting any single authority. These are not theoretical: they exist, and some jurisdictions use versions of them already.
The question is not whether audits matter. They do. The question is whether "we audited it" is the endpoint, or the beginning of a higher standard.
What to look for in your own jurisdiction
Before the next election in your jurisdiction, three questions are worth knowing the answers to:
First: does your jurisdiction use voter-verified paper ballots? If not, an RLA is mathematically impossible. There is nothing to audit against. This is the Sarasota situation. If your jurisdiction still uses paperless touchscreens or internet voting without an independent paper record, the audit conversation cannot even start.
Second: does your jurisdiction conduct a risk-limiting audit, or a fixed-percentage audit? The difference matters most in the closest races — precisely the ones that generate the most public concern. A fixed-percentage audit that inspects 2% of ballots in a race decided by 0.1% may have almost no statistical power to catch a wrong outcome. An RLA in the same race would either catch the problem or explicitly expand the sample until it did.
Third: are the audit results — the sampled ballot IDs, the hand-count tallies, the comparison against machine counts — published in full, in a form that an independent observer can check? This is where most jurisdictions, even those running good audits, fall short. Publishing a summary press release is not the same as publishing the audit data.
The answers to those three questions tell you more about the real integrity of your local election than any official statement ever will.
See how common the 'no mandatory audit' and 'weak audit method' gaps are across the world →
Read the 2-minute version of election integrity gaps →
Explore specific gaps in your region →
The shareable truth: an election that can only be trusted — but not independently verified — is not a secure election. It is a well-intentioned one. Those are not the same thing, and in a close race, the difference is the whole story.
Sources
- Michigan Department of State — Final numbers from Antrim County audit affirm accuracy of election results
- FactCheck.org — Audit in Michigan County Refutes Dominion Conspiracy Theory
- Georgia Public Broadcasting — Risk-Limiting Audit Confirms Biden Won Georgia
- NPR — Georgia Releases Hand Recount Results, Affirming Biden's Lead
- Colorado Secretary of State — A new kind of election audit: Colorado is first to complete it
- Colorado Secretary of State — Risk-Limiting Audit (RLA) FAQs
- U.S. GAO (GAO-08-97T) via GovInfo — Testing of Voting Systems in Florida's 13th Congressional District