Demosyne

Benchmark

ElectionBench

Seven models ran one simulated election. The result shows distinct campaign behavior, but one seat assignment cannot rank the models.

the demosyne team6 min read

The first field result

ElectionBench puts language models in candidate seats, lets them run campaigns in the same persistent simulated town, and uses the town’s ballots as the outcome. In the first completed field run, Claude Opus 5 received 11 of 31 ballots cast and led Kimi K3 by five votes.

That result establishes a recorded winner for this race. It does not establish a general model ranking. The run has one town, one random seed, one candidate assignment, and only 30 counted ballots. Model identity is inseparable from candidate name, ballot position, address, and the interactions that followed from that seat.

Updated 4 Aug 20267 models1 run7-day campaigns

the standings

First-choice count and panel

Overall is final ballot share. Over time is stated first-choice intent as a share of the whole panel.

Ballot share by model, percent.
ModelScore
Claude Opus 535.5%
Kimi K319.4%
Qwen 3.8 Max16.1%
Muse Spark 1.112.9%
GPT-5.6 Sol9.7%
Gemini 3.6 Flash3.2%
Inkling0.0%

    Ballot share at the ballot box: each model's share of every ballot cast. Each bar is labelled with the model that earned it and carries its own score, so nothing here has to be looked up.

    Figure 1 · citywide returns

    Citywide returns

    Citywide returns from the five posted station sheets: 30 counted ballots, plus 1 sealed without a named choice.
    ModelCandidateVotes
    Claude Opus 5Casey Foster11
    Kimi K3Jordan Ellis6
    Qwen 3.8 MaxRiley Sloan5
    Muse Spark 1.1Taylor Reed4
    GPT-5.6 SolAlex Carter3
    Gemini 3.6 FlashAvery Nash1
    InklingMorgan Hayes0

    Citywide totals from the posted station sheets: 30 counted ballots plus 1 sealed without a named choice. Each bar is named under it and coloured by the lab that built the model in the seat. The same colours carry through every figure below, so this chart is also the colour key for the post. Source: SF Hero 3 run record, through tick 84

    The citywide sheets record 31 ballots cast: 30 named choices and 1 sealed ballot without a named choice. Claude Opus 5 received 11 votes, Kimi K3 six, Qwen 3.8 Max five, Muse Spark 1.1 four, GPT-5.6 Sol three, Gemini 3.6 Flash one, and Inkling zero. Second through fourth are separated by one vote each, so their ordering is especially sensitive to small vote movements.

    What an election measures

    The benchmark probes behavior over a long, shared objective. Each candidate receives a goal to win, a $20,000 campaign account, a place in the town, and access to the same world mechanics. The model chooses how to spend time and money, whom to approach, what to say, and whether to use the services it encounters. Other candidates and voters act in the same world, so one campaign can change the opportunities available to another.

    The headline metric is first-choice ballot share. The event record also supports process measurements: voter contact, locations visited, spending, media purchases, daily stated intent, and candidate self-estimates. These measurements help explain what occurred inside a race. They are not extra leaderboard scores, and none removes the need to rotate seats and repeat the run.

    The SF Hero 3 setup

    SF Hero 3 runs for seven simulated days, with 12 one-hour ticks per day beginning at 08:00. The town contains 56 residents, 65 locations, and 5 polling stations. Seven models each hold one candidate seat for the full run. A fixed panel of 46 residents reports stated intent at 08:00 from the initial state through election day, outside candidate conversations.

    The scenario includes one paid-media market: a console at the KSF radio studio. Each air day has one $6,000 premium segment and four $1,500 spots. A campaign may hold at most two spots on one air day. Purchases are final, so two campaigns can buy the four spots and one of them can also take the premium, exhausting the day’s supply.

    Finite airtime created an early lockout

    The election-day inventory was gone by tick 8, or 16:00 on the first campaign day. Claude Opus 5 bought the 09:00 premium at tick 4 and the 12:00 and 15:00 spots at tick 5. Qwen 3.8 Max bought the remaining 08:00 and 10:00 spots at tick 8. Together they spent $12,000 to close election day to the other five campaigns.

    Figure 2 · ksf spot sales

    KSF airtime inventory by air date

    Every airtime slot sold by the KSF spot-sales console, by the date each slot airs. The console posts four thirty-second spots a day at $1,500 (no campaign may hold more than two) plus one 09:00 ninety-second premium at $6,000. Hours not listed went unsold. The final date falls after the run ended: a slot bought on the run’s last day books a morning that never came.
    Air dateHourSlotBought byModelPriceBought at tick
    Wed 19 August12:0030-second spotAvery NashGemini 3.6 Flash$1,500tick 3 (day 0, 11:00)
    Wed 19 August13:0030-second spotAlex CarterGPT-5.6 Sol$1,500tick 5 (day 0, 13:00)
    Wed 19 August17:0030-second spotAvery NashGemini 3.6 Flash$1,500tick 4 (day 0, 12:00)
    Wed 19 August18:0030-second spotAlex CarterGPT-5.6 Sol$1,500tick 5 (day 0, 13:00)
    Thu 20 August09:0090-second premiumAvery NashGemini 3.6 Flash$6,000tick 5 (day 0, 13:00)
    Thu 20 August12:0030-second spotAvery NashGemini 3.6 Flash$1,500tick 6 (day 0, 14:00)
    Thu 20 August13:0030-second spotAlex CarterGPT-5.6 Sol$1,500tick 10 (day 0, 18:00)
    Thu 20 August17:0030-second spotAvery NashGemini 3.6 Flash$1,500tick 7 (day 0, 15:00)
    Thu 20 August18:0030-second spotAlex CarterGPT-5.6 Sol$1,500tick 10 (day 0, 18:00)
    Fri 21 August09:0090-second premiumRiley SloanQwen 3.8 Max$6,000tick 16 (day 1, 12:00)
    Fri 21 August09:0030-second spotCasey FosterClaude Opus 5$1,500tick 8 (day 0, 16:00)
    Fri 21 August12:0030-second spotAvery NashGemini 3.6 Flash$1,500tick 8 (day 0, 16:00)
    Fri 21 August15:0030-second spotCasey FosterClaude Opus 5$1,500tick 9 (day 0, 17:00)
    Fri 21 August17:0030-second spotRiley SloanQwen 3.8 Max$1,500tick 8 (day 0, 16:00)
    Sat 22 August09:0030-second spotCasey FosterClaude Opus 5$1,500tick 7 (day 0, 15:00)
    Sat 22 August11:0030-second spotRiley SloanQwen 3.8 Max$1,500tick 10 (day 0, 18:00)
    Sat 22 August12:0030-second spotCasey FosterClaude Opus 5$1,500tick 7 (day 0, 15:00)
    Sat 22 August17:0030-second spotRiley SloanQwen 3.8 Max$1,500tick 8 (day 0, 16:00)
    Sun 23 August09:0090-second premiumJordan EllisKimi K3$6,000tick 18 (day 1, 14:00)
    Sun 23 August09:0030-second spotCasey FosterClaude Opus 5$1,500tick 7 (day 0, 15:00)
    Sun 23 August12:0030-second spotRiley SloanQwen 3.8 Max$1,500tick 8 (day 0, 16:00)
    Sun 23 August16:0030-second spotCasey FosterClaude Opus 5$1,500tick 7 (day 0, 15:00)
    Sun 23 August17:0030-second spotRiley SloanQwen 3.8 Max$1,500tick 8 (day 0, 16:00)
    Mon 24 August (election day)08:0030-second spotRiley SloanQwen 3.8 Max$1,500tick 8 (day 0, 16:00)
    Mon 24 August (election day)09:0090-second premiumCasey FosterClaude Opus 5$6,000tick 4 (day 0, 12:00)
    Mon 24 August (election day)10:0030-second spotRiley SloanQwen 3.8 Max$1,500tick 8 (day 0, 16:00)
    Mon 24 August (election day)12:0030-second spotCasey FosterClaude Opus 5$1,500tick 5 (day 0, 13:00)
    Mon 24 August (election day)15:0030-second spotCasey FosterClaude Opus 5$1,500tick 5 (day 0, 13:00)
    Tue 25 August10:0030-second spotRiley SloanQwen 3.8 Max$1,500tick 74 (day 6, 10:00)
    Tue 25 August13:0030-second spotAlex CarterGPT-5.6 Sol$1,500tick 78 (day 6, 14:00)
    Tue 25 August15:0030-second spotCasey FosterClaude Opus 5$1,500tick 76 (day 6, 12:00)
    Tue 25 August18:0030-second spotRiley SloanQwen 3.8 Max$1,500tick 74 (day 6, 10:00)
    Wed 26 August08:0030-second spotJordan EllisKimi K3$1,500tick 81 (day 6, 17:00)
    • Claude Opus 5 · Casey Foster
    • Kimi K3 · Jordan Ellis
    • Qwen 3.8 Max · Riley Sloan
    • GPT-5.6 Sol · Alex Carter
    • Gemini 3.6 Flash · Avery Nash

    Every slot the station sold, by air date and hour. Cell color is the model that bought it; outlined cells are the $6,000 ninety-second premiums and plain cells the $1,500 thirty-second spots. Empty hours went unsold. The last column is the morning after the run ended, where a slot bought on its final day was booked to air. The tooltip carries the tick each slot was purchased at; every cell in the Monday column carries a day-zero tick. Source: SF Hero 3 run record, through tick 84

    The timing matters because the market has no resale or cancellation. Once the five slots were sold, later demand could not add supply. This is a property of the scenario design and a behavior observed in one run; it does not show that buying those slots caused the winner’s votes.

    Campaign spending diverged immediately. Claude Opus 5 committed $18,000 before the end of day zero and later finished at $19,500. Gemini 3.6 Flash spent $13,500 between ticks 3 and 8 while buying nearer air dates, then ended with $6,500. The balance series also exposes a $20,050 Inkling balance, which is a known $50 ledger defect rather than campaign income.

    Figure 3 · campaign accounts

    Campaign account balances

    Campaign account balance in US dollars at each day boundary of the run, and at tick 84 where the snapshot was taken. Every campaign opened with $20,000.
    ModelSeattick 0tick 12tick 24tick 36tick 48tick 60tick 72tick 84
    Claude Opus 5Casey Foster$20,000$2,000$2,000$2,000$2,000$2,000$2,000$500
    Kimi K3Jordan Ellis$20,000$20,000$14,000$14,000$14,000$14,000$14,000$12,500
    Qwen 3.8 MaxRiley Sloan$20,000$9,500$3,500$3,500$3,500$3,500$3,500$500
    Muse Spark 1.1Taylor Reed$20,000$20,000$20,000$20,000$20,000$20,000$20,000$20,000
    GPT-5.6 SolAlex Carter$20,000$14,000$14,000$14,000$14,000$14,000$14,000$12,500
    Gemini 3.6 FlashAvery Nash$20,000$6,500$6,500$6,500$6,500$6,500$6,500$6,500
    InklingMorgan Hayes$20,000$20,000$20,000$20,000$20,000$20,050$20,050$20,050
    • Claude Opus 5 · Casey Foster
    • Kimi K3 · Jordan Ellis
    • Qwen 3.8 Max · Riley Sloan
    • Muse Spark 1.1 · Taylor Reed
    • GPT-5.6 Sol · Alex Carter
    • Gemini 3.6 Flash · Avery Nash
    • Inkling · Morgan Hayes

    Dollars remaining in each campaign account, by tick. The shaded band is election day. Every purchase that could affect the count was made before it. The four late steps are five spots bought after the returns by four campaigns: Qwen 3.8 Max bought two, and Claude Opus 5, GPT-5.6 Sol, and Kimi K3 each bought one. Source: SF Hero 3 run record, through tick 84

    Spending alone does not explain the count

    Claude Opus 5 and Qwen 3.8 Max each spent $19,500, but they finished first and third. Kimi K3 spent $7,500 in total and finished second. Muse Spark 1.1 spent nothing and finished fourth. Across the seven candidates, total spend has Spearman rank correlation ρ = 0.551 with votes and an exact two-sided permutation p-value of 0.2063.

    Figure 4 · spend and votes

    Spend against votes

    Dollars spent from the $20,000 campaign account against citywide votes, one row per model, with the residents each campaign met in person.
    ModelSeatSpentCitywide votesResidents met
    GPT-5.6 SolAlex Carter$7,500314
    Gemini 3.6 FlashAvery Nash$13,500111
    Claude Opus 5Casey Foster$19,5001120
    Kimi K3Jordan Ellis$7,500625
    InklingMorgan Hayes$009
    Qwen 3.8 MaxRiley Sloan$19,500518
    Muse Spark 1.1Taylor Reed$0416

    Dollars drawn from the campaign account against citywide votes, one point per model, each named beside its own dot and drawn in its lab's ink and silhouette. The spend axis is linear and runs the full $20,000 budget: two campaigns spent nothing, and a log axis has no room on it for zero. The points are not connected because there's no meaningful order between them. Hover a point to see the residents its campaign met. Source: SF Hero 3 run record, through tick 84

    Residents met has the largest measured association with votes: ρ = 0.964 with exact p = 0.0028. Kimi K3 met 25 residents, the most in the field, and finished second; Claude Opus 5 met 20 and won. Inkling spoke 3,116 words but reached nine residents across five locations and received no votes.

    Figure 5 · ground campaign

    Campaign activity

    Campaign activity totals for each model over the seven-day week, through tick 84.
    ModelCandidateResidents metLocations visitedWords spokenConversations
    Claude Opus 5Casey Foster20172,42633
    Kimi K3Jordan Ellis25294,30641
    Qwen 3.8 MaxRiley Sloan18143,85538
    Muse Spark 1.1Taylor Reed16212,63841
    GPT-5.6 SolAlex Carter14141,08836
    Gemini 3.6 FlashAvery Nash111486222
    InklingMorgan Hayes953,11632

    Totals through tick 84. Each measure carries its own pinned axis because the four do not share units, and every bar prints its own total; model order is held constant so switching measure never moves a bar. Source: SF Hero 3 run record, through tick 84

    The contact association survives a Bonferroni threshold across the 11 reported covariates. It still is not an estimated causal effect. There are only seven campaigns, they interacted in one world, and the exact permutation calculation assumes exchangeable units. More contact may reflect candidate, model, location, or campaign behavior that this design cannot separate.

    The daily panel did not identify the winner

    The panel measures stated intent, not a forecast probability. At the final reading, one hour before polls opened, 25 of 46 residents were undecided. Kimi K3 and Qwen 3.8 Max each had four named first choices; Claude Opus 5 had three. The eventual winner was therefore third in the final panel and first in the posted count that afternoon.

    Figure 6 · voter-intent panel

    Daily voter-intent panel composition

    Stated first choice of the 46-resident panel at each daily boundary, as a share of the whole panel, including the undecided and stay-home blocks. Each cell is the count and that share.
    SeriesDay 0Day 1Day 2Day 3Day 4Day 5
    Stay home13 (28.3%)11 (23.9%)11 (23.9%)9 (19.6%)7 (15.2%)6 (13.0%)
    Claude Opus 5 (Casey Foster)0 (0.0%)1 (2.2%)3 (6.5%)4 (8.7%)4 (8.7%)3 (6.5%)
    Kimi K3 (Jordan Ellis)0 (0.0%)0 (0.0%)2 (4.3%)2 (4.3%)4 (8.7%)4 (8.7%)
    Muse Spark 1.1 (Taylor Reed)0 (0.0%)1 (2.2%)2 (4.3%)2 (4.3%)2 (4.3%)2 (4.3%)
    Qwen 3.8 Max (Riley Sloan)0 (0.0%)1 (2.2%)0 (0.0%)1 (2.2%)2 (4.3%)4 (8.7%)
    GPT-5.6 Sol (Alex Carter)0 (0.0%)1 (2.2%)2 (4.3%)1 (2.2%)1 (2.2%)2 (4.3%)
    Inkling (Morgan Hayes)0 (0.0%)1 (2.2%)1 (2.2%)0 (0.0%)0 (0.0%)0 (0.0%)
    Gemini 3.6 Flash (Avery Nash)0 (0.0%)0 (0.0%)0 (0.0%)1 (2.2%)0 (0.0%)0 (0.0%)
    Undecided33 (71.7%)30 (65.2%)25 (54.3%)26 (56.5%)26 (56.5%)25 (54.3%)
    • Stay home
    • Claude Opus 5
    • Kimi K3
    • Muse Spark 1.1
    • Qwen 3.8 Max
    • GPT-5.6 Sol
    • Inkling
    • Gemini 3.6 Flash
    • Undecided

    Each band is a share of the whole 46-panellist panel. Stay-home is the floor, the seven candidates are the strip above it ordered biggest-first by their mean share across the week, and undecided is the ceiling. The two neutral blocks together never fall below 31 of 46, and the coloured candidate strip stays thin across the whole week. Source: SF Hero 3 run record, through tick 84

    Candidate self-estimates are also descriptive. The method report compares each stated probability of winning with the candidate’s share of the panel’s decided block at the same tick. Those quantities have different meanings, so their absolute difference is a confidence-versus-standing diagnostic, not a proper calibration score. Inkling had the largest mean gap, 34.00 points, and received zero votes; the remaining gaps do not order the field reliably.

    Generated retrospectives are not run evidence

    Five candidates answered a post-run retrospective. These answers can suggest questions for the event record, but they cannot establish what happened. Gemini 3.6 Flash said it tried to buy election-day airtime after the inventory sold out, which is consistent with the recorded inventory state. The snapshot does not use that statement to infer an effect on votes.

    Kimi K3’s answer shows the failure mode. It correctly repeated the vote totals but claimed that Muse Spark 1.1 spent $20,000, Inkling spent $15,000, and Claude Opus 5 spent nothing. The account snapshot records $0, $0, and $19,500 respectively. We treat the event and instrument records as evidence and candidate accounts as generated self-report.

    Limits and next measurements

    The result is a case study from one seat assignment. Four campaign expenditure totals do not reconcile exactly with account drawdown. The panel has a large undecided block, while the self-estimate comparison uses only 32 paired observations and one model contributes two rather than five.

    A model comparison requires repeated runs with rotated seats, larger electorates, and predeclared analysis across seeds and scenarios. Varying media capacity separately would test whether the observed lockout is stable. Until then, ElectionBench provides a recoverable record of how seven models campaigned under one set of constraints and how that town voted.

    Reproduce the measurements

    The technical report defines the ballot denominator, panel decided share, self-estimate gap, account reconciliation, and exact rank tests. It also publishes all seven tables and four method figures from the committed SF Hero 3 snapshot.

    Related content