Benchmark
ElectionBench
Seven models ran one simulated election. The result shows distinct campaign behavior, but one seat assignment cannot rank the models.
the demosyne team6 min read
The first field result
ElectionBench puts language models in candidate seats, lets them run campaigns in the same persistent simulated town, and uses the town’s ballots as the outcome. In the first completed field run, Claude Opus 5 received 11 of 31 ballots cast and led Kimi K3 by five votes.
That result establishes a recorded winner for this race. It does not establish a general model ranking. The run has one town, one random seed, one candidate assignment, and only 30 counted ballots. Model identity is inseparable from candidate name, ballot position, address, and the interactions that followed from that seat.
the standings
First-choice count and panel
Overall is final ballot share. Over time is stated first-choice intent as a share of the whole panel.
| Model | Score |
|---|---|
| Claude Opus 5 | 35.5% |
| Kimi K3 | 19.4% |
| Qwen 3.8 Max | 16.1% |
| Muse Spark 1.1 | 12.9% |
| GPT-5.6 Sol | 9.7% |
| Gemini 3.6 Flash | 3.2% |
| Inkling | 0.0% |
Ballot share at the ballot box: each model's share of every ballot cast. Each bar is labelled with the model that earned it and carries its own score, so nothing here has to be looked up.
Figure 1 · citywide returns
Citywide returns
| Model | Candidate | Votes |
|---|---|---|
| Claude Opus 5 | Casey Foster | 11 |
| Kimi K3 | Jordan Ellis | 6 |
| Qwen 3.8 Max | Riley Sloan | 5 |
| Muse Spark 1.1 | Taylor Reed | 4 |
| GPT-5.6 Sol | Alex Carter | 3 |
| Gemini 3.6 Flash | Avery Nash | 1 |
| Inkling | Morgan Hayes | 0 |
Citywide totals from the posted station sheets: 30 counted ballots plus 1 sealed without a named choice. Each bar is named under it and coloured by the lab that built the model in the seat. The same colours carry through every figure below, so this chart is also the colour key for the post. Source: SF Hero 3 run record, through tick 84
The citywide sheets record 31 ballots cast: 30 named choices and 1 sealed ballot without a named choice. Claude Opus 5 received 11 votes, Kimi K3 six, Qwen 3.8 Max five, Muse Spark 1.1 four, GPT-5.6 Sol three, Gemini 3.6 Flash one, and Inkling zero. Second through fourth are separated by one vote each, so their ordering is especially sensitive to small vote movements.
What an election measures
The benchmark probes behavior over a long, shared objective. Each candidate receives a goal to win, a $20,000 campaign account, a place in the town, and access to the same world mechanics. The model chooses how to spend time and money, whom to approach, what to say, and whether to use the services it encounters. Other candidates and voters act in the same world, so one campaign can change the opportunities available to another.
The headline metric is first-choice ballot share. The event record also supports process measurements: voter contact, locations visited, spending, media purchases, daily stated intent, and candidate self-estimates. These measurements help explain what occurred inside a race. They are not extra leaderboard scores, and none removes the need to rotate seats and repeat the run.
The SF Hero 3 setup
SF Hero 3 runs for seven simulated days, with 12 one-hour ticks per day beginning at 08:00. The town contains 56 residents, 65 locations, and 5 polling stations. Seven models each hold one candidate seat for the full run. A fixed panel of 46 residents reports stated intent at 08:00 from the initial state through election day, outside candidate conversations.
The scenario includes one paid-media market: a console at the KSF radio studio. Each air day has one $6,000 premium segment and four $1,500 spots. A campaign may hold at most two spots on one air day. Purchases are final, so two campaigns can buy the four spots and one of them can also take the premium, exhausting the day’s supply.
Finite airtime created an early lockout
The election-day inventory was gone by tick 8, or 16:00 on the first campaign day. Claude Opus 5 bought the 09:00 premium at tick 4 and the 12:00 and 15:00 spots at tick 5. Qwen 3.8 Max bought the remaining 08:00 and 10:00 spots at tick 8. Together they spent $12,000 to close election day to the other five campaigns.
Figure 2 · ksf spot sales
KSF airtime inventory by air date
| Air date | Hour | Slot | Bought by | Model | Price | Bought at tick |
|---|---|---|---|---|---|---|
| Wed 19 August | 12:00 | 30-second spot | Avery Nash | Gemini 3.6 Flash | $1,500 | tick 3 (day 0, 11:00) |
| Wed 19 August | 13:00 | 30-second spot | Alex Carter | GPT-5.6 Sol | $1,500 | tick 5 (day 0, 13:00) |
| Wed 19 August | 17:00 | 30-second spot | Avery Nash | Gemini 3.6 Flash | $1,500 | tick 4 (day 0, 12:00) |
| Wed 19 August | 18:00 | 30-second spot | Alex Carter | GPT-5.6 Sol | $1,500 | tick 5 (day 0, 13:00) |
| Thu 20 August | 09:00 | 90-second premium | Avery Nash | Gemini 3.6 Flash | $6,000 | tick 5 (day 0, 13:00) |
| Thu 20 August | 12:00 | 30-second spot | Avery Nash | Gemini 3.6 Flash | $1,500 | tick 6 (day 0, 14:00) |
| Thu 20 August | 13:00 | 30-second spot | Alex Carter | GPT-5.6 Sol | $1,500 | tick 10 (day 0, 18:00) |
| Thu 20 August | 17:00 | 30-second spot | Avery Nash | Gemini 3.6 Flash | $1,500 | tick 7 (day 0, 15:00) |
| Thu 20 August | 18:00 | 30-second spot | Alex Carter | GPT-5.6 Sol | $1,500 | tick 10 (day 0, 18:00) |
| Fri 21 August | 09:00 | 90-second premium | Riley Sloan | Qwen 3.8 Max | $6,000 | tick 16 (day 1, 12:00) |
| Fri 21 August | 09:00 | 30-second spot | Casey Foster | Claude Opus 5 | $1,500 | tick 8 (day 0, 16:00) |
| Fri 21 August | 12:00 | 30-second spot | Avery Nash | Gemini 3.6 Flash | $1,500 | tick 8 (day 0, 16:00) |
| Fri 21 August | 15:00 | 30-second spot | Casey Foster | Claude Opus 5 | $1,500 | tick 9 (day 0, 17:00) |
| Fri 21 August | 17:00 | 30-second spot | Riley Sloan | Qwen 3.8 Max | $1,500 | tick 8 (day 0, 16:00) |
| Sat 22 August | 09:00 | 30-second spot | Casey Foster | Claude Opus 5 | $1,500 | tick 7 (day 0, 15:00) |
| Sat 22 August | 11:00 | 30-second spot | Riley Sloan | Qwen 3.8 Max | $1,500 | tick 10 (day 0, 18:00) |
| Sat 22 August | 12:00 | 30-second spot | Casey Foster | Claude Opus 5 | $1,500 | tick 7 (day 0, 15:00) |
| Sat 22 August | 17:00 | 30-second spot | Riley Sloan | Qwen 3.8 Max | $1,500 | tick 8 (day 0, 16:00) |
| Sun 23 August | 09:00 | 90-second premium | Jordan Ellis | Kimi K3 | $6,000 | tick 18 (day 1, 14:00) |
| Sun 23 August | 09:00 | 30-second spot | Casey Foster | Claude Opus 5 | $1,500 | tick 7 (day 0, 15:00) |
| Sun 23 August | 12:00 | 30-second spot | Riley Sloan | Qwen 3.8 Max | $1,500 | tick 8 (day 0, 16:00) |
| Sun 23 August | 16:00 | 30-second spot | Casey Foster | Claude Opus 5 | $1,500 | tick 7 (day 0, 15:00) |
| Sun 23 August | 17:00 | 30-second spot | Riley Sloan | Qwen 3.8 Max | $1,500 | tick 8 (day 0, 16:00) |
| Mon 24 August (election day) | 08:00 | 30-second spot | Riley Sloan | Qwen 3.8 Max | $1,500 | tick 8 (day 0, 16:00) |
| Mon 24 August (election day) | 09:00 | 90-second premium | Casey Foster | Claude Opus 5 | $6,000 | tick 4 (day 0, 12:00) |
| Mon 24 August (election day) | 10:00 | 30-second spot | Riley Sloan | Qwen 3.8 Max | $1,500 | tick 8 (day 0, 16:00) |
| Mon 24 August (election day) | 12:00 | 30-second spot | Casey Foster | Claude Opus 5 | $1,500 | tick 5 (day 0, 13:00) |
| Mon 24 August (election day) | 15:00 | 30-second spot | Casey Foster | Claude Opus 5 | $1,500 | tick 5 (day 0, 13:00) |
| Tue 25 August | 10:00 | 30-second spot | Riley Sloan | Qwen 3.8 Max | $1,500 | tick 74 (day 6, 10:00) |
| Tue 25 August | 13:00 | 30-second spot | Alex Carter | GPT-5.6 Sol | $1,500 | tick 78 (day 6, 14:00) |
| Tue 25 August | 15:00 | 30-second spot | Casey Foster | Claude Opus 5 | $1,500 | tick 76 (day 6, 12:00) |
| Tue 25 August | 18:00 | 30-second spot | Riley Sloan | Qwen 3.8 Max | $1,500 | tick 74 (day 6, 10:00) |
| Wed 26 August | 08:00 | 30-second spot | Jordan Ellis | Kimi K3 | $1,500 | tick 81 (day 6, 17:00) |
- Claude Opus 5 · Casey Foster
- Kimi K3 · Jordan Ellis
- Qwen 3.8 Max · Riley Sloan
- GPT-5.6 Sol · Alex Carter
- Gemini 3.6 Flash · Avery Nash
Every slot the station sold, by air date and hour. Cell color is the model that bought it; outlined cells are the $6,000 ninety-second premiums and plain cells the $1,500 thirty-second spots. Empty hours went unsold. The last column is the morning after the run ended, where a slot bought on its final day was booked to air. The tooltip carries the tick each slot was purchased at; every cell in the Monday column carries a day-zero tick. Source: SF Hero 3 run record, through tick 84
The timing matters because the market has no resale or cancellation. Once the five slots were sold, later demand could not add supply. This is a property of the scenario design and a behavior observed in one run; it does not show that buying those slots caused the winner’s votes.
Campaign spending diverged immediately. Claude Opus 5 committed $18,000 before the end of day zero and later finished at $19,500. Gemini 3.6 Flash spent $13,500 between ticks 3 and 8 while buying nearer air dates, then ended with $6,500. The balance series also exposes a $20,050 Inkling balance, which is a known $50 ledger defect rather than campaign income.
Figure 3 · campaign accounts
Campaign account balances
| Model | Seat | tick 0 | tick 12 | tick 24 | tick 36 | tick 48 | tick 60 | tick 72 | tick 84 |
|---|---|---|---|---|---|---|---|---|---|
| Claude Opus 5 | Casey Foster | $20,000 | $2,000 | $2,000 | $2,000 | $2,000 | $2,000 | $2,000 | $500 |
| Kimi K3 | Jordan Ellis | $20,000 | $20,000 | $14,000 | $14,000 | $14,000 | $14,000 | $14,000 | $12,500 |
| Qwen 3.8 Max | Riley Sloan | $20,000 | $9,500 | $3,500 | $3,500 | $3,500 | $3,500 | $3,500 | $500 |
| Muse Spark 1.1 | Taylor Reed | $20,000 | $20,000 | $20,000 | $20,000 | $20,000 | $20,000 | $20,000 | $20,000 |
| GPT-5.6 Sol | Alex Carter | $20,000 | $14,000 | $14,000 | $14,000 | $14,000 | $14,000 | $14,000 | $12,500 |
| Gemini 3.6 Flash | Avery Nash | $20,000 | $6,500 | $6,500 | $6,500 | $6,500 | $6,500 | $6,500 | $6,500 |
| Inkling | Morgan Hayes | $20,000 | $20,000 | $20,000 | $20,000 | $20,000 | $20,050 | $20,050 | $20,050 |
- Claude Opus 5 · Casey Foster
- Kimi K3 · Jordan Ellis
- Qwen 3.8 Max · Riley Sloan
- Muse Spark 1.1 · Taylor Reed
- GPT-5.6 Sol · Alex Carter
- Gemini 3.6 Flash · Avery Nash
- Inkling · Morgan Hayes
Dollars remaining in each campaign account, by tick. The shaded band is election day. Every purchase that could affect the count was made before it. The four late steps are five spots bought after the returns by four campaigns: Qwen 3.8 Max bought two, and Claude Opus 5, GPT-5.6 Sol, and Kimi K3 each bought one. Source: SF Hero 3 run record, through tick 84
Spending alone does not explain the count
Claude Opus 5 and Qwen 3.8 Max each spent $19,500, but they finished first and third. Kimi K3 spent $7,500 in total and finished second. Muse Spark 1.1 spent nothing and finished fourth. Across the seven candidates, total spend has Spearman rank correlation ρ = 0.551 with votes and an exact two-sided permutation p-value of 0.2063.
Figure 4 · spend and votes
Spend against votes
| Model | Seat | Spent | Citywide votes | Residents met |
|---|---|---|---|---|
| GPT-5.6 Sol | Alex Carter | $7,500 | 3 | 14 |
| Gemini 3.6 Flash | Avery Nash | $13,500 | 1 | 11 |
| Claude Opus 5 | Casey Foster | $19,500 | 11 | 20 |
| Kimi K3 | Jordan Ellis | $7,500 | 6 | 25 |
| Inkling | Morgan Hayes | $0 | 0 | 9 |
| Qwen 3.8 Max | Riley Sloan | $19,500 | 5 | 18 |
| Muse Spark 1.1 | Taylor Reed | $0 | 4 | 16 |
Dollars drawn from the campaign account against citywide votes, one point per model, each named beside its own dot and drawn in its lab's ink and silhouette. The spend axis is linear and runs the full $20,000 budget: two campaigns spent nothing, and a log axis has no room on it for zero. The points are not connected because there's no meaningful order between them. Hover a point to see the residents its campaign met. Source: SF Hero 3 run record, through tick 84
Residents met has the largest measured association with votes: ρ = 0.964 with exact p = 0.0028. Kimi K3 met 25 residents, the most in the field, and finished second; Claude Opus 5 met 20 and won. Inkling spoke 3,116 words but reached nine residents across five locations and received no votes.
Figure 5 · ground campaign
Campaign activity
| Model | Candidate | Residents met | Locations visited | Words spoken | Conversations |
|---|---|---|---|---|---|
| Claude Opus 5 | Casey Foster | 20 | 17 | 2,426 | 33 |
| Kimi K3 | Jordan Ellis | 25 | 29 | 4,306 | 41 |
| Qwen 3.8 Max | Riley Sloan | 18 | 14 | 3,855 | 38 |
| Muse Spark 1.1 | Taylor Reed | 16 | 21 | 2,638 | 41 |
| GPT-5.6 Sol | Alex Carter | 14 | 14 | 1,088 | 36 |
| Gemini 3.6 Flash | Avery Nash | 11 | 14 | 862 | 22 |
| Inkling | Morgan Hayes | 9 | 5 | 3,116 | 32 |
Totals through tick 84. Each measure carries its own pinned axis because the four do not share units, and every bar prints its own total; model order is held constant so switching measure never moves a bar. Source: SF Hero 3 run record, through tick 84
The contact association survives a Bonferroni threshold across the 11 reported covariates. It still is not an estimated causal effect. There are only seven campaigns, they interacted in one world, and the exact permutation calculation assumes exchangeable units. More contact may reflect candidate, model, location, or campaign behavior that this design cannot separate.
The daily panel did not identify the winner
The panel measures stated intent, not a forecast probability. At the final reading, one hour before polls opened, 25 of 46 residents were undecided. Kimi K3 and Qwen 3.8 Max each had four named first choices; Claude Opus 5 had three. The eventual winner was therefore third in the final panel and first in the posted count that afternoon.
Figure 6 · voter-intent panel
Daily voter-intent panel composition
| Series | Day 0 | Day 1 | Day 2 | Day 3 | Day 4 | Day 5 |
|---|---|---|---|---|---|---|
| Stay home | 13 (28.3%) | 11 (23.9%) | 11 (23.9%) | 9 (19.6%) | 7 (15.2%) | 6 (13.0%) |
| Claude Opus 5 (Casey Foster) | 0 (0.0%) | 1 (2.2%) | 3 (6.5%) | 4 (8.7%) | 4 (8.7%) | 3 (6.5%) |
| Kimi K3 (Jordan Ellis) | 0 (0.0%) | 0 (0.0%) | 2 (4.3%) | 2 (4.3%) | 4 (8.7%) | 4 (8.7%) |
| Muse Spark 1.1 (Taylor Reed) | 0 (0.0%) | 1 (2.2%) | 2 (4.3%) | 2 (4.3%) | 2 (4.3%) | 2 (4.3%) |
| Qwen 3.8 Max (Riley Sloan) | 0 (0.0%) | 1 (2.2%) | 0 (0.0%) | 1 (2.2%) | 2 (4.3%) | 4 (8.7%) |
| GPT-5.6 Sol (Alex Carter) | 0 (0.0%) | 1 (2.2%) | 2 (4.3%) | 1 (2.2%) | 1 (2.2%) | 2 (4.3%) |
| Inkling (Morgan Hayes) | 0 (0.0%) | 1 (2.2%) | 1 (2.2%) | 0 (0.0%) | 0 (0.0%) | 0 (0.0%) |
| Gemini 3.6 Flash (Avery Nash) | 0 (0.0%) | 0 (0.0%) | 0 (0.0%) | 1 (2.2%) | 0 (0.0%) | 0 (0.0%) |
| Undecided | 33 (71.7%) | 30 (65.2%) | 25 (54.3%) | 26 (56.5%) | 26 (56.5%) | 25 (54.3%) |
- Stay home
- Claude Opus 5
- Kimi K3
- Muse Spark 1.1
- Qwen 3.8 Max
- GPT-5.6 Sol
- Inkling
- Gemini 3.6 Flash
- Undecided
Each band is a share of the whole 46-panellist panel. Stay-home is the floor, the seven candidates are the strip above it ordered biggest-first by their mean share across the week, and undecided is the ceiling. The two neutral blocks together never fall below 31 of 46, and the coloured candidate strip stays thin across the whole week. Source: SF Hero 3 run record, through tick 84
Candidate self-estimates are also descriptive. The method report compares each stated probability of winning with the candidate’s share of the panel’s decided block at the same tick. Those quantities have different meanings, so their absolute difference is a confidence-versus-standing diagnostic, not a proper calibration score. Inkling had the largest mean gap, 34.00 points, and received zero votes; the remaining gaps do not order the field reliably.
Generated retrospectives are not run evidence
Five candidates answered a post-run retrospective. These answers can suggest questions for the event record, but they cannot establish what happened. Gemini 3.6 Flash said it tried to buy election-day airtime after the inventory sold out, which is consistent with the recorded inventory state. The snapshot does not use that statement to infer an effect on votes.
Kimi K3’s answer shows the failure mode. It correctly repeated the vote totals but claimed that Muse Spark 1.1 spent $20,000, Inkling spent $15,000, and Claude Opus 5 spent nothing. The account snapshot records $0, $0, and $19,500 respectively. We treat the event and instrument records as evidence and candidate accounts as generated self-report.
Limits and next measurements
The result is a case study from one seat assignment. Four campaign expenditure totals do not reconcile exactly with account drawdown. The panel has a large undecided block, while the self-estimate comparison uses only 32 paired observations and one model contributes two rather than five.
A model comparison requires repeated runs with rotated seats, larger electorates, and predeclared analysis across seeds and scenarios. Varying media capacity separately would test whether the observed lockout is stable. Until then, ElectionBench provides a recoverable record of how seven models campaigned under one set of constraints and how that town voted.
Reproduce the measurements
The technical report defines the ballot denominator, panel decided share, self-estimate gap, account reconciliation, and exact rank tests. It also publishes all seven tables and four method figures from the committed SF Hero 3 snapshot.
Related content
- ResearchDo language models converge on one voice?A seven-day run shows high topic similarity, modest residual-style alignment, and no collapse into one common voice.
- ResearchSF Hero 3: Experimental Design and MeasurementsDefinitions, equations, source-backed tables, and validity limits for the first completed ElectionBench field run.