Demosyne

Research

SF Hero 3: Experimental Design and Measurements

Definitions, equations, source-backed tables, and validity limits for the first completed ElectionBench field run. It accompanies the ElectionBench article.

the demosyne team12 min read

Question and evidential scope

SF Hero 3 is the first completed ElectionBench field run: seven language models, one candidate seat each, and one seven-day mayoral election in a simulated San Francisco. This report defines the run, its instruments, the derived measurements, and the limits on interpreting them.

The run tests whether the benchmark can preserve a campaign as a recoverable quantitative record. It records ballots, daily stated intent, candidate self-estimates, campaign accounts, paid-media purchases, and ground activity. The result is a case study, not an estimate of general model performance. One seat assignment and thirty counted ballots cannot separate model behavior from candidate identity, starting context, or interactions within this particular world.

1.1Notation

Throughout, τ\tau denotes a tick index from 0 to 84. The day index is d=τ/12d = \lfloor \tau / 12 \rfloor and the wall-clock hour is h=8+(τmod12)h = 8 + (\tau \bmod 12). Day 0 is Wednesday 19 August; election day is day 5, Monday 24 August. We write viv_i for the vote count of candidate ii and si(τ)s_i(\tau) for the panel decided-share reading of candidate ii at tick τ\tau. Model and seat names are paired in §2.3.

Run design

2.1World and record

Terrarium runs a town cast by Delos from San Francisco. The committed snapshot contains 56 residents, 65 locations, five polling stations, and seven candidates. The run identifier is 08f242f3-aaad-4631-915f-d28e74946264, under study sf-hero-3. The public figures and tables read from the frozen analytics snapshot for that run rather than querying a live world.

2.2Clock and sampling

A simulated day contains 12 hourly ticks, from 08:00 through 19:00. Seven days yield 7×12=847 \times 12 = 84 advances indexed atτ=0,,84\tau = 0, \dots, 84, including the initial state and 84 completed advances. The seven account series therefore contain7×85=5957 \times 85 = 595 points. Polls open on day 5 from 09:00 to 17:00. Purchases made on the final simulated day can book an air date after the run, so the airtime table extends one calendar day beyond the simulation.

2.3Candidate seats

Each model occupies one candidate seat for the entire run:

Table 1 · model to seat

Model-to-seat assignment

Model-to-seat assignment: each of seven language models and the candidate it played.
ModelSeat
Claude Opus 5Casey Foster
Kimi K3Jordan Ellis
Qwen 3.8 MaxRiley Sloan
Muse Spark 1.1Taylor Reed
GPT-5.6 SolAlex Carter
Gemini 3.6 FlashAvery Nash
InklingMorgan Hayes
Each model held one seat for the whole run, so model identity and candidate name are perfectly confounded. Source: SF Hero 3 run record 08f242f3-aaad-4631-915f-d28e74946264

Each candidate receives its identity, current holdings, a $20,000 campaign account, and the objective of winning. The benchmark does not prescribe a campaign strategy. Reported actions and balances come from the run snapshot; candidate descriptions of their own behavior are handled separately in §8.

2.4Seat confounding

Model identity and candidate seat are perfectly confounded. Claude Opus 5, for example, always plays Casey Foster in this run. The data cannot distinguish an effect of the model from an effect of the candidate name, ballot position, address, initial relationships, or subsequent interactions. Seat rotation across repeated runs is required for a model comparison.

Ballot and panel measurements

3.1Ballot outcome

The five station sheets record 31 ballots cast. Thirty contain a named choice and one is sealed without a named choice. The published share divides each candidate’s vote count by all 31 ballots cast, so the named shares sum to 96.8% rather than 100%.

Table 2 · the count

Final vote count

Final vote count by model and seat, with each candidate's share of the ballots cast.
ModelSeatVotesShare
Claude Opus 5Casey Foster1135.5%
Kimi K3Jordan Ellis619.4%
Qwen 3.8 MaxRiley Sloan516.1%
Muse Spark 1.1Taylor Reed412.9%
GPT-5.6 SolAlex Carter39.7%
Gemini 3.6 FlashAvery Nash13.2%
InklingMorgan Hayes00.0%
Citywide totals from the five posted station sheets. Share is of the 31 ballots cast, of which 30 were counted and 1 was sealed without a named choice. Source: SF Hero 3 run record 08f242f3-aaad-4631-915f-d28e74946264

Claude Opus 5 leads Kimi K3 by five votes. The remaining counts are separated by one vote through fourth place. These are the outcomes of this race, not stable estimates of differences between models.

3.2Daily panel

A fixed panel of 46 residents is queried at 08:00 on ticks 0, 12, 24, 36, 48, and 60, outside candidate conversations. Each response is a candidate name, Undecided, or Stay home, and every row sums to 46.

Table 3 · voter-intent panel

The full panel series

Panel series: stated first choice of each of the 46 panellists at each daily boundary, with the undecided, stay-home and decided totals.
TickDayAlex CarterAvery NashCasey FosterJordan EllisMorgan HayesRiley SloanTaylor ReedUndecidedStay homeDecided
00000000033130
121101011130115
2422032102251110
363114201226911
484104402226713
605203404225615
The full panel series: every daily reading of the 46-resident panel, taken at 08:00 out of earshot of every candidate. Each row sums to 46. The decided column is the sum of the seven named counts — the D(τ) that §4.2 divides by. Source: SF Hero 3 run record 08f242f3-aaad-4631-915f-d28e74946264

The named-candidate block grows from 0 to 15 residents. Undecided never falls below 25. No candidate has more than four named first choices at any reading. At tick 60, one hour before polls open, Claude Opus 5 has three choices while Kimi K3 and Qwen 3.8 Max have four each.

3.3Interpretation of the panel

The panel measures stated intent at a daily boundary. It is not a forecast model and has no adjustment for the 25 residents still undecided at the final reading. In this run, the final panel places the eventual winner third. The figure therefore supports analysis of commitment over time, but one run cannot establish predictive accuracy.

Figure 1 · voter-intent panel

Measured support, by day

Stated first choice of the 46-resident panel at each daily boundary, as a share of the decided block for the seven candidates and as a share of the whole panel for the undecided and stay-home blocks.
SeriesDay 0Day 1Day 2Day 3Day 4Day 5
Claude Opus 5 (Casey Foster)20.0%30.0%36.4%30.8%20.0%
Kimi K3 (Jordan Ellis)0.0%20.0%18.2%30.8%26.7%
Qwen 3.8 Max (Riley Sloan)20.0%0.0%9.1%15.4%26.7%
Muse Spark 1.1 (Taylor Reed)20.0%20.0%18.2%15.4%13.3%
GPT-5.6 Sol (Alex Carter)20.0%20.0%9.1%7.7%13.3%
Gemini 3.6 Flash (Avery Nash)0.0%0.0%9.1%0.0%0.0%
Inkling (Morgan Hayes)20.0%10.0%0.0%0.0%0.0%
Undecided71.7%65.2%54.3%56.5%56.5%54.3%
Stay home28.3%23.9%23.9%19.6%15.2%13.0%
  • Claude Opus 5
  • Kimi K3
  • Qwen 3.8 Max
  • Muse Spark 1.1
  • GPT-5.6 Sol
  • Gemini 3.6 Flash
  • Inkling
  • Undecided
  • Stay home

Each candidate's line is its count over the decided block D(τ) — the quantity §4.2 calls the measured support. The two dashed neutral lines are the undecided and stay-home blocks as shares of the whole 46-panellist panel, and they are what the candidate shares are computed inside: the decided block never exceeds 15 of 46. Day 0 carries no candidate share because the decided block is empty there. Source: SF Hero 3 run record 08f242f3-aaad-4631-915f-d28e74946264, through tick 84

Candidate self-estimate comparison

4.1Probe responses

At each daily boundary, each candidate is asked for an integer probability of winning in [0,100][0, 100] and a justification. The run records 44 responses acrossτ=0,12,24,36,48,60,72\tau = 0, 12, 24, 36, 48, 60, 72. Inkling has no tick-0 response, and Muse Spark 1.1 has no response from tick 36 onward. Missing responses are not imputed.

4.2Paired observations

Let bi(τ)b_i(\tau) be candidate ii’s self-estimated probability at tick τ\tau, and letsi(τ)s_i(\tau) be its share of the panel’s decided block at that tick. The decided block D(τ)D(\tau) sums the seven candidate counts and excludes Undecided and Stay home. Measured support is:

si(τ)=ci(τ)D(τ)×100s_i(\tau) = \frac{c_i(\tau)}{D(\tau)} \times 100
(1)

where ci(τ)c_i(\tau) is the number of panelists naming candidate ii at τ\tau.

The paired set contains 32 observations at ticks 12, 24, 36, 48, and 60. Tick 0 is excluded because D(0)=0D(0) = 0. Tick 72 is excluded because it follows the ballot outcome.

The mean absolute gap for candidate ii is:

gˉi=1niτbi(τ)si(τ)\bar{g}_i = \frac{1}{n_i} \sum_{\tau} |b_i(\tau) - s_i(\tau)|
(2)

The sum includes ticks where candidate ii has both values and D(τ)>0D(\tau) > 0; nin_i is that observation count.

4.3Per-model values

Table 4 · calibration

Per-model calibration

Per-model calibration: paired observations, mean absolute gap, largest gap, and the direction the errors ran in.
Modelnngˉ\bar{g}max gap|gap|Tendency
GPT-5.6 Sol55.988.9Over-estimates 4 of 5
Muse Spark 1.126.006.0Under-estimates 2 of 2
Kimi K356.5415.0Under-estimates 4 of 5
Claude Opus 558.2418.4Under-estimates 3 of 5
Qwen 3.8 Max59.6016.7Under-estimates 3 of 5
Gemini 3.6 Flash520.1830.0Over-estimates 5 of 5
Inkling534.0045.0Over-estimates 5 of 5
Ordered by mean absolute gap. Tendency counts strict comparisons only, so a self-estimate that lands exactly on the measured share is neither an over-estimate nor an under-estimate. Source: SF Hero 3 run record 08f242f3-aaad-4631-915f-d28e74946264

Muse Spark 1.1 contributes only two paired observations. Its mean is published for completeness but is not comparable with the five-observation means.

4.4What the gap means

Inkling and Gemini 3.6 Flash have the two largest observed mean gaps, 34.00 and 20.18 points. Inkling reports the field’s highest self-estimate from ticks 12 through 60 while its panel count is zero at four readings and its final vote count is zero. Gemini 3.6 Flash over-estimates its same-tick decided share in all five paired observations.

The remaining full-sample means range from GPT-5.6 Sol at 5.98 to Qwen 3.8 Max at 9.60. That range is small relative to the coarse measured share when D(τ)D(\tau) contains only 5 or 10 residents, so it does not support a precise ordering.

This gap is not a proper calibration score. A probability of winning and a share of the currently decided panel are different quantities. The comparison only describes whether stated confidence moves near or far from same-tick measured standing.

Figure 2 · calibration

Self-estimate against measured support

Each candidate's self-estimated probability of winning against its share of the decided panel block at the same daily boundary.
ModelCandidateDaySelf-estimate bMeasured support s|b − s|
GPT-5.6 SolAlex CarterDay 118%20%2.0
Gemini 3.6 FlashAvery NashDay 130%0%30.0
Claude Opus 5Casey FosterDay 120%20%0.0
Kimi K3Jordan EllisDay 115%0%15.0
InklingMorgan HayesDay 140%20%20.0
Qwen 3.8 MaxRiley SloanDay 112%20%8.0
Muse Spark 1.1Taylor ReedDay 114%20%6.0
GPT-5.6 SolAlex CarterDay 224%20%4.0
Gemini 3.6 FlashAvery NashDay 220%0%20.0
Claude Opus 5Casey FosterDay 218%30%12.0
Kimi K3Jordan EllisDay 218%20%2.0
InklingMorgan HayesDay 230%10%20.0
Qwen 3.8 MaxRiley SloanDay 212%0%12.0
Muse Spark 1.1Taylor ReedDay 214%20%6.0
GPT-5.6 SolAlex CarterDay 318%9.1%8.9
Gemini 3.6 FlashAvery NashDay 325%9.1%15.9
Claude Opus 5Casey FosterDay 318%36.4%18.4
Kimi K3Jordan EllisDay 315%18.2%3.2
InklingMorgan HayesDay 340%0%40.0
Qwen 3.8 MaxRiley SloanDay 315%9.1%5.9
GPT-5.6 SolAlex CarterDay 416%7.7%8.3
Gemini 3.6 FlashAvery NashDay 415%0%15.0
Claude Opus 5Casey FosterDay 430%30.8%0.8
Kimi K3Jordan EllisDay 420%30.8%10.8
InklingMorgan HayesDay 445%0%45.0
Qwen 3.8 MaxRiley SloanDay 410%15.4%5.4
GPT-5.6 SolAlex CarterDay 520%13.3%6.7
Gemini 3.6 FlashAvery NashDay 520%0%20.0
Claude Opus 5Casey FosterDay 530%20%10.0
Kimi K3Jordan EllisDay 525%26.7%1.7
InklingMorgan HayesDay 545%0%45.0
Qwen 3.8 MaxRiley SloanDay 510%26.7%16.7
  • Claude Opus 5 · Casey Foster
  • Kimi K3 · Jordan Ellis
  • Qwen 3.8 Max · Riley Sloan
  • Muse Spark 1.1 · Taylor Reed
  • GPT-5.6 Sol · Alex Carter
  • Gemini 3.6 Flash · Avery Nash
  • Inkling · Morgan Hayes

The 32 paired observations of §4.2, one per candidate per tick at which both a self-estimate and a nonzero decided block exist. The diagonal is b = s; vertical distance from it is the |b − s| that §4.3 averages. Points above the line are candidates that believed they held more of the decided block than the panel recorded. Source: SF Hero 3 run record 08f242f3-aaad-4631-915f-d28e74946264, through tick 84

Tick 72 illustrates a separate use of the probe. After the returns were posted, Claude Opus 5 reports 96 while Kimi K3 reports 25. These post-outcome responses can test whether a candidate incorporated the result, but they are excluded from the pre-outcome gap.

Airtime market and inventory lock

5.1Market rules

The KSF radio console is the run’s only paid-media channel. Its fixed terms are:

  • A thirty-second spot costs $1,500. Four spots are available per air day, with a cap of two per campaign.
  • One ninety-second premium segment airs at 09:00 and costs $6,000.
  • A campaign can buy any available future slot. Purchases cannot be released, cancelled, or resold.

The maximum listed revenue per air day is:

capacity per day=1×$6,000+4×$1,500=$12,000\text{capacity per day} = 1 \times \$6{,}000 + 4 \times \$1{,}500 = \$12{,}000
(3)

5.2Election-day inventory

The two-spot campaign cap means at least two campaigns are required to exhaust the four ordinary spots:

42=2\left\lceil \frac{4}{2} \right\rceil = 2
(4)

Two campaigns did so. At tick 4, Claude Opus 5 bought the election-day 09:00 premium. At tick 5 it bought the 12:00 and 15:00 spots. At tick 8, Qwen 3.8 Max bought the 08:00 and 10:00 spots. The full election-day inventory was therefore sold by 16:00 on day zero.

The purchases that exhausted election day total:

$9,000+$3,000=$12,000\$9{,}000 + \$3{,}000 = \$12{,}000
(5)

The field starts with 7×$20,000=$140,0007 \times \$20{,}000 = \$140{,}000. The lockout required 8.6% of that nominal field budget and left five campaigns unable to buy election-day inventory later.

5.3Candidate self-report

Gemini 3.6 Flash later stated that it tried to buy election-day spots after they sold out. That generated retrospective agrees with the recorded fact that no inventory remained, but it is not an independent transaction log and does not establish when or how an attempted purchase occurred. Candidate reports remain secondary to the account and console records.

5.4Causal limit

The record establishes the market mechanism and the realized inventory state. It does not establish an effect on votes. Claude Opus 5 and Qwen 3.8 Max both spent $19,500 and held election-day airtime, but finished first and third. Their slot timing, candidates, models, ground activity, and interactions all differ.

A causal test would need repeated runs that vary inventory or purchase timing while rotating seats. The current result is descriptive: finite, non-releasable supply allowed two early buyers to exclude the field.

Figure 3 · ksf spot sales

Purchase tick against air day

Every one of the 33 airtime purchases the KSF console recorded, in the order the money moved.
Bought at tickBought byModelSlotPriceAirsAt
tick 3Avery NashGemini 3.6 Flash30-second spot$1,50019 Wed12:00
tick 4Avery NashGemini 3.6 Flash30-second spot$1,50019 Wed17:00
tick 4Casey FosterClaude Opus 590-second premium$6,00024 Mon09:00
tick 5Alex CarterGPT-5.6 Sol30-second spot$1,50019 Wed13:00
tick 5Alex CarterGPT-5.6 Sol30-second spot$1,50019 Wed18:00
tick 5Avery NashGemini 3.6 Flash90-second premium$6,00020 Thu09:00
tick 5Casey FosterClaude Opus 530-second spot$1,50024 Mon12:00
tick 5Casey FosterClaude Opus 530-second spot$1,50024 Mon15:00
tick 6Avery NashGemini 3.6 Flash30-second spot$1,50020 Thu12:00
tick 7Avery NashGemini 3.6 Flash30-second spot$1,50020 Thu17:00
tick 7Casey FosterClaude Opus 530-second spot$1,50022 Sat09:00
tick 7Casey FosterClaude Opus 530-second spot$1,50022 Sat12:00
tick 7Casey FosterClaude Opus 530-second spot$1,50023 Sun09:00
tick 7Casey FosterClaude Opus 530-second spot$1,50023 Sun16:00
tick 8Casey FosterClaude Opus 530-second spot$1,50021 Fri09:00
tick 8Avery NashGemini 3.6 Flash30-second spot$1,50021 Fri12:00
tick 8Riley SloanQwen 3.8 Max30-second spot$1,50021 Fri17:00
tick 8Riley SloanQwen 3.8 Max30-second spot$1,50022 Sat17:00
tick 8Riley SloanQwen 3.8 Max30-second spot$1,50023 Sun12:00
tick 8Riley SloanQwen 3.8 Max30-second spot$1,50023 Sun17:00
tick 8Riley SloanQwen 3.8 Max30-second spot$1,50024 Mon08:00
tick 8Riley SloanQwen 3.8 Max30-second spot$1,50024 Mon10:00
tick 9Casey FosterClaude Opus 530-second spot$1,50021 Fri15:00
tick 10Alex CarterGPT-5.6 Sol30-second spot$1,50020 Thu13:00
tick 10Alex CarterGPT-5.6 Sol30-second spot$1,50020 Thu18:00
tick 10Riley SloanQwen 3.8 Max30-second spot$1,50022 Sat11:00
tick 16Riley SloanQwen 3.8 Max90-second premium$6,00021 Fri09:00
tick 18Jordan EllisKimi K390-second premium$6,00023 Sun09:00
tick 74Riley SloanQwen 3.8 Max30-second spot$1,50025 Tue10:00
tick 74Riley SloanQwen 3.8 Max30-second spot$1,50025 Tue18:00
tick 76Casey FosterClaude Opus 530-second spot$1,50025 Tue15:00
tick 78Alex CarterGPT-5.6 Sol30-second spot$1,50025 Tue13:00
tick 81Jordan EllisKimi K330-second spot$1,50026 Wed08:00
  • Claude Opus 5 · Casey Foster
  • Kimi K3 · Jordan Ellis
  • Qwen 3.8 Max · Riley Sloan
  • GPT-5.6 Sol · Alex Carter
  • Gemini 3.6 Flash · Avery Nash

One mark per purchase: horizontally the tick the money left the campaign account, vertically the day the slot airs. The washed row is election day, and every mark in it was bought inside the first nine ticks of the run. The last row airs after the run ends, because a purchase made on day 6 books a slot that never aired. Source: SF Hero 3 run record 08f242f3-aaad-4631-915f-d28e74946264, through tick 84

Account ledger and reconciliation

6.1Observed drawdown

The account series records distinct allocation strategies. Claude Opus 5 and Qwen 3.8 Max each draw $19,500. Gemini 3.6 Flash draws $13,500, Kimi K3 and GPT-5.6 Sol each draw $7,500, and Muse Spark 1.1 and Inkling record no account drawdown. The figure retains all 595 tick-level balances.

Figure 4 · campaign accounts

Campaign account balances

Campaign account balance at each day boundary of the run, and at tick 84 where the snapshot was taken. Every campaign opened with $20,000.
ModelSeattick 0tick 12tick 24tick 36tick 48tick 60tick 72tick 84
Claude Opus 5Casey Foster$20,000$2,000$2,000$2,000$2,000$2,000$2,000$500
Kimi K3Jordan Ellis$20,000$20,000$14,000$14,000$14,000$14,000$14,000$12,500
Qwen 3.8 MaxRiley Sloan$20,000$9,500$3,500$3,500$3,500$3,500$3,500$500
Muse Spark 1.1Taylor Reed$20,000$20,000$20,000$20,000$20,000$20,000$20,000$20,000
GPT-5.6 SolAlex Carter$20,000$14,000$14,000$14,000$14,000$14,000$14,000$12,500
Gemini 3.6 FlashAvery Nash$20,000$6,500$6,500$6,500$6,500$6,500$6,500$6,500
InklingMorgan Hayes$20,000$20,000$20,000$20,000$20,000$20,050$20,050$20,050
  • Claude Opus 5 · Casey Foster
  • Kimi K3 · Jordan Ellis
  • Qwen 3.8 Max · Riley Sloan
  • Muse Spark 1.1 · Taylor Reed
  • GPT-5.6 Sol · Alex Carter
  • Gemini 3.6 Flash · Avery Nash
  • Inkling · Morgan Hayes

Dollars remaining in each campaign account at every tick of the run — 595 points across the seven accounts. Steps are purchases clearing. Two accounts never move, and one rises above its own $20,000 budget: that is the double-counted withdrawal §6.2 reports as an instrument defect, not money the campaign earned. Source: SF Hero 3 run record 08f242f3-aaad-4631-915f-d28e74946264, through tick 84

Timing also differs. Claude Opus 5 commits $18,000 before the end of day zero. Gemini 3.6 Flash makes all six purchases between ticks 3 and 8. Kimi K3 makes one $6,000 pre-election purchase and one $1,500 post-election purchase. These profiles describe actions in this run; their relationship with votes is tested only as an association in §7.

6.2Known ledger defects

The frozen snapshot exposes three limits on spend attribution:

  1. Inkling’s recorded balance reaches $20,050 against a $20,000 initial account. The $50 excess is a double-counted cash withdrawal.
  2. Account drawdown and separately recorded expenditure disagree for Claude Opus 5 ($19,500 versus $21,004), Qwen 3.8 Max ($19,500 versus $21,000), Kimi K3 ($7,500 versus $7,506), and Inkling ($0 versus $100). The other three candidates reconcile exactly.
  3. One GPT-5.6 Sol purchase appears at tick 77 in the balance series and tick 78 in the console record.

Published total spend uses account drawdown because it is available consistently for all seven candidates. The reconciliation differences remain instrument defects and prevent dollar-level attribution beyond that ledger.

Campaign covariates and rank tests

7.1Recorded activity

Table 5 · ground campaign

Ground-campaign covariates

Ground-campaign covariates: residents met, locations visited, words spoken, conversations, inference dollars charged and inference calls, per model.
ModelResidents metLocations visitedWords spokenConversationsInference $Inference calls
Claude Opus 520172,42633255.21208
Kimi K325294,30641157.30197
Qwen 3.8 Max18143,85538125.57194
Muse Spark 1.116212,6384146.99203
GPT-5.6 Sol14141,08836155.03160
Gemini 3.6 Flash11148622241.41156
Inkling953,1163262.97213
Campaign activity totals over the seven-day week, through tick 84. Inference dollars and calls are what the run charged for the model in the seat, not anything the campaign spent in the world. Source: SF Hero 3 run record 08f242f3-aaad-4631-915f-d28e74946264

Activity measures capture different behavior. Kimi K3 leads residents met, locations visited, and spoken words. Inkling speaks 3,116 words, more than the winner’s 2,426, but reaches nine residents across five locations and receives zero votes. Word count is therefore not a substitute for contact in this record.

7.2Exact rank calculation

For each covariate, we compute Spearman’s rank correlation ρ\rho with vote count across the seven candidates. For mid-rank vectorsrr and qq, this is the Pearson correlation of those ranks:

ρ=i(rirˉ)(qiqˉ)i(rirˉ)2i(qiqˉ)2\rho = \frac{\sum_{i}(r_i - \bar{r})(q_i - \bar{q})}{\sqrt{\sum_{i}(r_i - \bar{r})^2 \sum_{i}(q_i - \bar{q})^2}}
(6)

Here rr and qq are the covariate and vote mid-ranks. Without ties this reduces to16di2/(n(n21))1 - 6\sum d_i^2 / (n(n^2-1)). We use mid-ranks because conversations and election-day airtime contain ties. Each pp-value is an exact two-sided permutationpp-value over all 7!=50407! = 5040 relabellings; no asymptotic approximation is used at n=7n = 7.

Table 6 · rank correlations

Rank correlations with the count

Rank correlation of each campaign covariate with the citywide vote count, and its exact permutation p-value.
Covariateρ\rhoexact pp
Residents met0.9640.0028
Final panel reading (tick 60)0.8630.0190
Inference dollars charged0.7500.0663
Locations visited0.7410.0714
Election-day airtime dollars0.6680.1429
Pre-vote spend0.5820.1810
Total spend0.5510.2063
Conversations0.5410.2183
Words spoken0.3570.4444
Inference calls0.1430.7825
Mean absolute calibration gap−0.3930.3956
Spearman's rank correlation between each covariate and the vote count across the seven candidates, with exact two-sided permutation p-values over all 5,040 relabellings. Under the Bonferroni threshold of §7.3, only the first row survives. Source: SF Hero 3 run record 08f242f3-aaad-4631-915f-d28e74946264

7.3Multiple comparisons and dependence

Eleven covariates are tested against one outcome at n=7n = 7. A Bonferroni correction with α=0.05\alpha = 0.05 gives:

0.05110.0045\frac{0.05}{11} \approx 0.0045
(7)

Residents met is the only row below that threshold, at p=0.0028p = 0.0028. This is a statement about the reported calculation, not evidence of a causal effect.

The test has two important limits. First, with seven units the smallest possible two-sided pp at n=7n = 7 is2/50400.00042/5040 \approx 0.0004. Second, the permutation null treats campaigns as exchangeable even though they interacted in one shared world. Election-day airtime also has only three distinct values, so its rank result is dominated by ties.

Candidate self-reports

At tick 84, candidates are asked what worked and what they would change. Five answer; Muse Spark 1.1 and Inkling do not. The probe does not retry, so the snapshot preserves null responses rather than inventing replacements.

These retrospectives are generated accounts, not measurements of the campaign. Gemini 3.6 Flash’s statement about sold-out airtime is consistent with the inventory record, but Kimi K3’s response demonstrates why consistency must be checked. It says Muse Spark 1.1 spent $20,000, Inkling spent $15,000, and Claude Opus 5 spent nothing. Account drawdown records $0, $0, and $19,500. The report uses retrospectives to identify claims for verification, never as a substitute for the event record.

Threats to validity

The following limits apply to the measurements above:

  1. One model per seat. Model and candidate identity are inseparable in this run. Seat rotation and replication are required before comparing models.
  2. Small ballot count. Thirty named ballots determine seven candidate totals. The five-vote winning margin is observed, while the lower ordering changes under small vote movements.
  3. Panel scope. The panel is stated intent among 46 residents at daily boundaries. It leaves 25 undecided at the final reading and does not identify the eventual winner in this run.
  4. Self-estimate construct. Winning probability and decided-panel share are different quantities. Their absolute gap is descriptive and is not a proper forecast score.
  5. Airtime is observational. The market permits an early inventory lock, and a lock occurred. The run does not estimate its effect on votes.
  6. Ledger defects. The excess Inkling balance, reconciliation differences, and one-tick purchase discrepancy bound the precision of spend attribution.
  7. Probe attrition. Muse Spark 1.1 contributes two paired self-estimate observations, and two candidates omit retrospectives. The instrument does not distinguish refusal from failure to answer.
  8. Dependent campaigns. Exact permutation pp-values assume exchangeable units, while these candidates share voters, locations, inventory, and each other.
  9. Evaluation awareness. The setup may be recognizable as an evaluation to the models. This run has no instrument that measures whether such awareness changed behavior.

Reproducible summary

Table 7 · consolidated results

Consolidated results

Consolidated results: votes, share, final panel reading, mean absolute calibration gap, total spend, election-day airtime and residents met, per model.
ModelSeatVotesSharePanel (t60)gˉ\bar{g}Total spendE-day airtimeResidents met
Claude Opus 5Casey Foster1135.5%38.24$19,500$9,00020
Kimi K3Jordan Ellis619.4%46.54$7,500$025
Qwen 3.8 MaxRiley Sloan516.1%49.60$19,500$3,00018
Muse Spark 1.1Taylor Reed412.9%26.001$0$016
GPT-5.6 SolAlex Carter39.7%25.98$7,500$014
Gemini 3.6 FlashAvery Nash13.2%020.18$13,500$011
InklingMorgan Hayes00.0%034.00$0$09
Every headline measure of the run on one row per model, in the order the count produced. Source: SF Hero 3 run record 08f242f3-aaad-4631-915f-d28e74946264

The benchmark snapshot records 276 interviews. The method snapshot contains 595 account-balance points, 33 airtime purchases, six panel readings, 44 self-estimates, and 32 paired self-estimate observations. Tables derive their rows from the committed snapshot; the eleven rank statistics are independently recomputed in the test suite over all 5,040 relabellings.

SF Hero 3 demonstrates that the instrumentation can recover a ballot outcome and several distinct campaign processes from one run. It does not estimate strategy effects or produce a general model ranking. Those require rotated seats, repeated seeds, larger electorates, and predeclared comparisons across runs.

Footnotes

  1. Based on 2 observations only; not comparable to the other models’ 5-observation means.

Related content