← All posts

Every close election is a stress test your voting system probably fails

A fold in a ballot. Dust on a lens. No fraud, no malware — and still hundreds of miscounted votes in a race decided by dozens. What happens when your audit method is as error-prone as the thing it's auditing?

It is the morning of November 4, 2020, and the clerk's office in Windham, New Hampshire is looking at a hand-recount result that doesn't quite match the machine total.

Not by thousands. Not by hundreds in most cases. By enough, in one State Representative race in Rockingham District 7, that the discrepancy couldn't be waved away. The margin separating candidates was tight. The difference between what the AccuVote optical scanners had recorded and what human hands had re-tallied was real. An investigation was coming — but not yet. First, the town's officials did what officials everywhere do: they finished the certification and moved on.

That is where the story gets interesting. Because what the forensic audit eventually uncovered had nothing to do with fraud, hacking, or bad actors of any kind. It had to do with a folding machine, a crease in the wrong place, and a scanner that couldn't tell the difference.

The most dangerous election errors are the boring ones nobody notices.


A fold, a crease, and 300 miscounted votes

In July 2021, a three-person audit team — Harri Hursti, Mark Lindeman, and Philip B. Stark — released their report under New Hampshire's Senate Bill 43. What they found was methodically unglamorous: a folding machine the town had leased for mailing purposes had folded some absentee ballots along lines that cut directly through the printed vote targets — the ovals voters are supposed to fill in.

The AccuVote scanners, reading those ballots, interpreted a substantial fraction of the fold creases as marks. Votes were being hallucinated from paper geometry. The team's experiments showed fold-through-target false-positive rates ranging from about 20 percent to more than 72 percent, with an average around 44 percent. White powder had also accumulated inside some scanners, partially obstructing the lenses — probably making the problem worse on those specific machines.

The full report is unambiguous: no malware, no tampering, no evidence of any intentional act. The election was, the team wrote, "for the most part, well run under challenging circumstances." And yet the machines had miscounted hundreds of validly cast ballots for a mechanical reason that no pre-election test, no certification process, and no routine audit had flagged.

The only thing that made this discoverable was the existence of durable paper ballots a human hand could re-examine. The hand recount recovered voter intent that the scanner had garbled. That part of the story is genuinely reassuring.

But here is the part that isn't.


The audit that catches the error is itself imperfect

The hand recount in Antrim County, Michigan in December 2020 — ordered to settle disputes about a tabulator misconfiguration that had briefly published absurd unofficial results — tallied every presidential ballot by hand. The certified result was unaffected. Officials declared the audit a confirmation.

What the Michigan Department of State's own announcement noted, without particular emphasis, was that the final hand count (9,759 for Trump, 5,959 for Biden) still differed from the machine tabulation by about a dozen votes out of roughly 15,700 cast.

Twelve votes. In a county-level race, with no close contest at stake, that number disappears into the category of "it's fine." But hold it up to a different light.

In a race decided by fewer than a dozen votes, a hand count that is itself off by a dozen votes cannot tell you who won.

This is not a conspiracy theory. It is arithmetic. Hand counts carry irreducible human error: counters misread marks, lose track of a stack, disagree about an ambiguous fill that is 70 percent of an oval. In Georgia's 2020 statewide hand tally of roughly 5 million presidential ballots — one of the most impressive auditing operations in American electoral history — the variation between the original machine count and the hand count was about a tenth of one percent. A tenth of one percent sounds tiny. In a 5-million-ballot election with a margin of about 11,000 votes, it was close enough to be worth watching carefully. In a local race decided by 8 votes, a tenth of a percent error rate in the audit method itself is larger than the margin.

This is the trap. We call the hand count the gold standard. We tell ourselves it confirms the result. And it does — when the margin is large enough that the error in both the machine count and the hand count falls comfortably below it. When the margin is not large enough, we have replaced one uncertain count with a second uncertain count and declared the problem solved.


The margin is the stress test

A close election is not a special case. It is the case that reveals whether your verification system actually works — or only appears to.

Consider what a close race requires:

The machine count must be accurate to within the margin. The audit method used to check it must be accurate to within the margin. The hand-count standard for what constitutes a valid mark must be defined clearly enough that different counters reach the same answer. And all of this must be independently checkable by someone who isn't the official being audited.

The Windham case passes the first two tests eventually — the forensic audit did identify the root cause, and a careful hand recount did recover voter intent. But it required a state-ordered forensic investigation, a team of outside experts, months of work, and a legislature willing to commission it. None of that happened automatically. None of it would have happened if the margin in Rockingham District 7 had been a few hundred votes wider.

Your election verification system is calibrated for the elections you expect. A close race is the election it wasn't expecting.

Colorado grasped this earlier than most. When it ran the country's first statewide risk-limiting audit in 2017, it built the sample size to the margin: close races get more ballots hand-checked, not a fixed percentage regardless of how tight things are. The statistical logic is sound. But a risk-limiting audit still requires a reliable paper record to audit against, and it still involves human hands counting ballots with all the error that entails. It narrows the problem. It does not eliminate it.


What makes a discrepancy visible — and what keeps it hidden

Antrim County's tabulator error was caught because the result it produced was absurd on its face. A solidly Republican county showed Biden leading by thousands of votes with most ballots counted. Anyone who knew the county's history spotted the anomaly instantly. The correction was made within 24 hours.

But that is exactly the wrong lesson to draw. The lesson is not "the system caught it." The lesson is: the system caught it only because the error was enormous and politically obvious. A misconfiguration that shifted 50 votes, or 20, or 12 — in the right direction, in a county whose partisan lean made the result plausible — would have passed through every layer of the process without a flag.

Windham is the same story told more quietly. The fold-crease miscounting almost certainly affected the 2020 results. The forensic audit report notes that the same folding machine, the same scanner, the same accumulating powder were all present on election night. The effect on the machine totals was real. It was caught only because a subsequent hand recount diverged enough to trigger suspicion — and because a legislature ordered a forensic audit instead of accepting the explanation that the hand count had been sloppy.

How many jurisdictions, facing a similar gap between machine totals and hand totals, accepted the simpler narrative? How many didn't have a legislature willing to push?

The German Federal Constitutional Court posed the underlying question more sharply in 2009, when it invalidated the use of electronic voting machines that stored votes only in memory with no independently verifiable record. The Court's judgment held that voting is only legitimate when ordinary citizens, without specialist knowledge, can verify the essential steps of the count. Not trust officials who have checked. Not trust an audit that itself carries unexplained residuals. Verify.

The principle is not uniquely German. The Dutch government's 2007 commission — which recommended scrapping voting computers entirely — put it this way in its report Stemmen met vertrouwen: "there are no secrets in the election process." Everything must be checkable and verifiable. Full stop.


The method and the margin cannot be the same size

Here is the cleaner version of the argument:

Every counting method has an intrinsic error rate. Optical scanners misread fold creases and powder-obscured lenses. Hand counters misclassify ambiguous marks, lose track of stacks, and disagree with each other about 70-percent-filled ovals. The Antrim hand count diverged from the machine count by roughly a dozen votes. Georgia's statewide hand count diverged by about a tenth of a percent.

These error rates are not scandals. They are engineering realities.

The scandal is that we treat these methods as settling questions they cannot settle. We say "the hand count confirmed it" about races where the hand count's own error band spans the margin. We say "the audit affirmed the result" without asking whether the audit method's precision was finer than the victory margin.

A verification method that is less precise than the quantity it is measuring is not a verification method. It is a ritual of reassurance.

"The hand count confirmed it" and "the Secretary of State says it's fine" are reassurances. They are not proof. Proof requires a method whose error is demonstrably smaller than the margin — checked by someone who is not the official being audited, against records that cannot be quietly altered.

Sarasota County, Florida made this concrete in 2006, when roughly 18,000 undervotes appeared in a Congressional race decided by 369 votes, and there was no paper to examine because the county used paperless touchscreen machines. There was no audit to run, no hand count to perform, no method to apply. The machine total was the only record that existed. The race was decided. The undervotes were never explained.

That is one failure mode: no paper at all. Windham is a subtler one: paper exists, scanner misread it, hand count recovers intent, but it took months, a forensic team, and legislative will to get there. The failure mode in between — paper exists, scanner misread it, gap is within the noise of a hand recount, nobody investigates — is the one we have no clean way to detect.


What "checkable" actually requires

The fix is not "hand count everything." Hand counts are slow, expensive, and themselves imprecise. The fix is not "trust the machines." Windham and Antrim both show why.

The fix is a system designed so that the verification is: faster than a months-long forensic audit; more precise than a hand count's irreducible human error; independent of the official being checked; and accessible to any observer, not just to experts commissioned by the legislature.

That means, at minimum: durable voter-marked paper records (not barcode-only records that the voter cannot read — a point the federal court in Curling v. Raffensperger made explicitly); routine statistical audits sized to the margin, not a fixed percentage; and precinct-level results published immediately and in machine-readable form so that any observer can reconcile the public tally against the underlying records without waiting to be granted access.

Colorado's risk-limiting audit framework moves toward this. Georgia's 2020 hand tally moved toward this. Neither gets all the way there. Neither answers the question of what you do when the audit method's error band and the margin of victory are the same width.

That question will be answered — or not answered — in a future close race that nobody currently expects.


What you still cannot check

The Windham forensic audit is genuinely good news: it found the cause, it named the mechanism, it produced recommendations. The New Hampshire Secretary of State's audit page is publicly accessible. The report is published.

But here is what the report cannot tell you: whether the same folding-machine problem affected other jurisdictions using the same scanner model, in races where nobody ordered a forensic audit because the margin was wide enough that the discrepancy passed unnoticed. The Windham report notes this possibility explicitly.

And here is what no current audit framework, in any U.S. state, can tell you as a citizen on election night: the exact precinct-level machine totals, time-stamped and signed, published in a form you can download and reconcile against the certified result yourself, without asking anyone's permission.

That gap — between "the officials checked it" and "you can check it" — is the whole argument.

An election system you cannot independently verify is an election system you have to trust. And trust, as Windham showed, is not the same thing as accuracy.


See how common this verification gap is across jurisdictions worldwide. Read the two-minute version of what makes a count independently checkable. Explore the full database of close-margin audit failures and what they reveal.


Sources