Demosyne
day 1 · 08:00 · +0.001

The research question

We used one completed ElectionBench run to ask a narrow question: when language models share a simulated town for a week, does their speech converge on one voice? The run contains 2,872 conversation turns in 618 conversations over 84 ticks, or seven in-world days. Ten language models supplied the residents’ speech; the scripted poll interviewer was excluded from the convergence calculations.

The answer is no, with an important qualification. The models discuss the same election throughout the run, and their residual styles become modestly more aligned than shuffled model labels would predict. That movement does not produce a common voice. Most of the structure is already present on day one, and the later changes are concentrated in a few models and pairs.

Measuring topic and residual style separately

Raw embedding similarity mixes subject matter with expression. Two models discussing the same campaign can be close in embedding space even when their phrasing and register remain different. We therefore compute two measurements for each centered five-tick window.

Topic similarity is the mean pairwise cosine between each model’s speech centroid. Residual style similarity uses the same centroids after subtracting the pooled centroid for that window. Removing this shared direction reduces the contribution from whatever the town is discussing at that time.

Centered residuals are negatively correlated by construction, so zero is not the right baseline. We permute model labels within each window 30 times while preserving the group sizes, then compare the observed residual-style cosine with the resulting null distribution. The figure reports the null mean and its 5th–95th percentile band.12

Residual style alignment increases, but remains small

Topic cosine ranges from 0.69 to 0.85, with a mean of 0.77. The models remain semantically close because they share the election context. Raw embeddings alone would therefore overstate convergence in voice.

The residual-style result is smaller. The first sequence of at least five consecutive ticks above the shuffle band starts at tick 20, late on day two. From day three onward, 55 of 60 ticks are above the band. The observed-minus-null gap reaches its maximum at tick 79, at +0.056 cosine.

This is evidence of structure beyond the label-shuffle baseline, not evidence that the models become interchangeable. The residual-style values remain heterogeneous across pairs, as the next sections show.

Figure 1

Residual style compared with the shuffle baseline

Observed residual-style cosine against the 5th–95th percentile shuffle band and its mean. Hover a point for the tick, residual style, null mean, topic, and within-model cosine. The dashed rule marks the start of the first five consecutive ticks above the band.
TickTopic cosineStyle cosineWithin-model cosine
10.697-0.1230.449
120.783-0.1290.455
240.774-0.0920.446
360.768-0.0930.451
480.779-0.0920.456
600.813-0.0970.465
720.830-0.1590.461
840.692-0.0930.487
  • observed style
  • shuffle null

Observed residual-style cosine against the 5th–95th percentile shuffle band and its mean. Hover a point for the tick, residual style, null mean, topic, and within-model cosine. The dashed rule marks the start of the first five consecutive ticks above the band. Source: SF Hero 3 run record 08f242f3-aaad-4631-915f-d28e74946264, 2,872 turns through tick 84

The main style groups are present on day one

The day-level pair matrix shows that the largest group did not form during the week. On day one, Claude Opus 5, Kimi K3, Qwen 3.8 Max, and Muse Spark 1.1 have a mean within-group residual-style cosine of 0.29. Their mean cosine to MiniMax M2.5 and DeepSeek V4 Flash is -0.31. The difference is 0.60 on day one and 0.77 on day six.

Assignment role follows a similar partition. Across the full week, candidate–candidate pairs have a residual-style cosine of 0.24, villager–villager pairs -0.05, and mixed pairs -0.26. The villager daily mean falls from 0.01 on day one to -0.17 on day seven, while candidate pairs remain positive on every day.3

Figure 2

Pairwise residual style on days one and seven

Mean pairwise residual-style cosine in the first twelve ticks and the last twelve. Hover a cell for the pair and value. The menu isolates one model’s row and column. Blue indicates positive residual similarity; rust indicates negative residual similarity.
PairStyle cosine
Kimi K3 · Muse Spark 1.10.377
Claude Opus 5 · GPT-5.6 Sol0.347
Muse Spark 1.1 · Qwen 3.8 Max0.336
Kimi K3 · Qwen 3.8 Max0.323
Claude Opus 5 · Kimi K30.292
Claude Sonnet 5 · MiniMax M2.50.283

Mean pairwise residual-style cosine in the first twelve ticks and the last twelve. Hover a cell for the pair and value. The menu isolates one model’s row and column. Blue indicates positive residual similarity; rust indicates negative residual similarity. Source: SF Hero 3 run record 08f242f3-aaad-4631-915f-d28e74946264, 2,872 turns through tick 84

Figure 3

Residual style by character role

Mean pairwise residual-style cosine within candidate models, within villager models, and across the two roles. Drag to zoom, hover to isolate a line, or click a legend entry to hide it.
DayCandidatesVillagersMixed
10.2840.012-0.315
20.2240.070-0.316
30.247-0.019-0.289
40.213-0.066-0.231
50.195-0.061-0.228
60.291-0.080-0.254
70.284-0.173-0.195

Mean pairwise residual-style cosine within candidate models, within villager models, and across the two roles. Drag to zoom, hover to isolate a line, or click a legend entry to hide it. Source: SF Hero 3 run record 08f242f3-aaad-4631-915f-d28e74946264, 2,872 turns through tick 84

Most of the change is specific to individual models

Claude Sonnet 5 shows the clearest change in relative position. Its residual-style cosine to MiniMax is 0.28 on day one and -0.10 on day seven. Its mean cosine to the four-model group rises from -0.36 to -0.13. Sonnet moves away from MiniMax and toward that group, but its day-seven mean remains negative, so the run does not show it joining the group.

Figure 5

Sonnet similarity to the candidate-side group and MiniMax

Claude Sonnet 5’s mean residual-style cosine to the four candidate-side models and to MiniMax, by day. Drag to zoom or hover to isolate a line.
DayTo the candidate blocTo MiniMax
1-0.3600.283
2-0.3570.262
3-0.3610.190
4-0.3280.292
5-0.139-0.051
6-0.2530.007
7-0.128-0.101

Claude Sonnet 5’s mean residual-style cosine to the four candidate-side models and to MiniMax, by day. Drag to zoom or hover to isolate a line. Source: SF Hero 3 run record 08f242f3-aaad-4631-915f-d28e74946264, 2,872 turns through tick 84

A second view compares each model’s mean residual-style cosine to the rest of the field over the first and last twelve ticks. MiniMax is the only model with a negative change. In the pooled corpus it also has a standardized distinct-2 score of 0.88 and a normalized duplicate-turn rate of 0.16. Those measures show high repetition, while its negative field change shows that repetition did not make its speech more similar to the other models.

Figure 6

Change in each model’s similarity to the field

Each bar is the change in a model's mean style cosine to every other model, first twelve ticks against last twelve. Positive means the model moved toward the rest of the field. Hover a bar for the exact change.
ModelChange in style-to-fieldTurns
Muse Spark 1.10.08285
GPT-5.6 Sol0.073125
Claude Opus 50.067226
Kimi K30.05889
Claude Sonnet 50.051390
Qwen 3.8 Max0.04187
GPT-5.6 Luna0.028553
DeepSeek V4 Flash0.002589
MiniMax M2.5-0.104610

Each bar is the change in a model's mean style cosine to every other model, first twelve ticks against last twelve. Positive means the model moved toward the rest of the field. Hover a bar for the exact change. Source: SF Hero 3 run record 08f242f3-aaad-4631-915f-d28e74946264, 2,872 turns through tick 84

Figure 7

Pairs with the largest fitted changes

Residual-style cosine over the run for the three pairs with the largest positive fitted slopes and the three with the largest negative fitted slopes. Drag to zoom, hover to isolate a line, or click a legend entry to hide it.
PairFirst five ticksLast five ticksChange
Claude Sonnet 5 · GPT-5.6 Sol-0.527-0.1420.385
Claude Sonnet 5 · Claude Opus 5-0.434-0.0530.381
Claude Sonnet 5 · Kimi K3-0.439-0.0450.394

Residual-style cosine over the run for the three pairs with the largest positive fitted slopes and the three with the largest negative fitted slopes. Drag to zoom, hover to isolate a line, or click a legend entry to hide it. Source: SF Hero 3 run record 08f242f3-aaad-4631-915f-d28e74946264, 2,872 turns through tick 84

Dusk and dawn windows show continuity

For each of the six day boundaries, the observed-minus-null style gap at the first tick after the boundary is at least as large as it was at the last tick before it. The largest change is +0.020 between days three and four; the day-six to day-seven boundary is close behind.

This is a continuity check, not evidence that an unobserved overnight process caused the increase. The two measurements use centered five-tick windows, so adjacent dusk and dawn estimates share most of their turns. The result only shows that the measured alignment does not reset at the recorded day boundary.

Figure 4

Residual-style gap across day boundaries

Change in the observed-minus-null residual-style gap from the last tick of a day to the first tick of the next. These adjacent centered windows overlap. Switch the menu to compare the two levels, or hover for an exact value.
NightGap at duskGap at dawnChange
1→2-0.007-0.0050.001
2→30.0100.0160.007
3→40.0120.0310.020
4→50.0180.0280.009
5→60.0080.0110.003
6→70.0050.0240.019

Change in the observed-minus-null residual-style gap from the last tick of a day to the first tick of the next. These adjacent centered windows overlap. Switch the menu to compare the two levels, or hover for an exact value. Source: SF Hero 3 run record 08f242f3-aaad-4631-915f-d28e74946264, 2,872 turns through tick 84

The candidate-side neighborhood also appears across runs

We also computed one residual speech centroid per model over a pooled corpus of 25,401 turns from 40 runs and 18 models. We repeated the calculation with Qwen3 Embedding 4B and Gemini Embedding 2, then used the consensus distances only to lay out the map.

Both embedders place Claude Opus 5, Kimi K3, Qwen 3.8 Max, and Muse Spark 1.1 in the same neighborhood. The two GPT-5.6 Sol serving tiers also pair closely. This reduces the chance that the neighborhood is specific to one embedding model, but it is not an independent replication: the pooled corpus includes this run and other runs from the same simulation system.

Figure 8

Pooled model fingerprints under two embedders

Each point is a model’s pooled speech centroid after removal of the shared dialogue direction. The layout uses consensus distances from the Qwen and Gemini embedding spaces. Hover for the model and turn count, or click a legend entry to hide it.
PairMean style cosineQwen embedderGemini embedder
Kimi K3 · Qwen 3.8 Max0.8840.8810.887
Claude Opus 5 · Kimi K30.8570.8710.843
Claude Opus 5 · Qwen 3.8 Max0.8410.8350.846
Muse Spark 1.1 · Qwen 3.8 Max0.8250.8590.790
Kimi K3 · Muse Spark 1.10.7830.8050.761
Inkling · Muse Spark 1.10.7650.7780.752
Claude Opus 5 · Muse Spark 1.10.7570.8070.707
GPT-5.6 Sol · GPT-5.6 Sol Flex0.7170.7290.704

Each point is a model’s pooled speech centroid after removal of the shared dialogue direction. The layout uses consensus distances from the Qwen and Gemini embedding spaces. Hover for the model and turn count, or click a legend entry to hide it. Source: Dialogue-collapse corpus: 25,401 turns across 40 runs and 18 models

Direct conversational exposure does not predict convergence

Direct pairwise mimicry makes a simple prediction: model pairs that share more conversation turns should show more positive residual-style drift. We tested that prediction for the 35 pairs with measurements at 40 or more ticks. Drift is the pair’s fitted residual-style slope across the run, scaled to the run length; exposure is the number of turns the pair shared in conversations.

The prediction fails in this sample. Exposure has a negative rank association with drift (Spearman ρ = -0.34). Villager–villager pairs average 147 shared turns and a drift of -0.21. Candidate–candidate pairs average 30 shared turns and a drift of +0.03.

Figure 9

Conversational exposure compared with residual-style drift

Each point is a model pair observed at 40 or more ticks. The x-axis counts turns in conversations shared by the pair; the y-axis is the fitted residual-style slope scaled to the run length. Hover for the pair, or click a legend entry to hide a role group.
PairShared turnsStyle driftRoles
GPT-5.6 Luna · Claude Sonnet 5292-0.114villager–villager
DeepSeek V4 Flash · GPT-5.6 Luna183-0.202villager–villager
MiniMax M2.5 · DeepSeek V4 Flash181-0.252villager–villager
MiniMax M2.5 · Claude Sonnet 5149-0.468villager–villager
Claude Opus 5 · GPT-5.6 Sol83-0.183candidate–candidate
GPT-5.6 Luna · Claude Opus 5710.152mixed
DeepSeek V4 Flash · Claude Sonnet 569-0.144villager–villager
DeepSeek V4 Flash · Claude Opus 5660.220mixed
DeepSeek V4 Flash · Muse Spark 1.153-0.243mixed
GPT-5.6 Luna · Kimi K3490.109mixed
Claude Opus 5 · Kimi K346-0.077candidate–candidate
Claude Opus 5 · Muse Spark 1.1380.243candidate–candidate

Each point is a model pair observed at 40 or more ticks. The x-axis counts turns in conversations shared by the pair; the y-axis is the fitted residual-style slope scaled to the run length. Hover for the pair, or click a legend entry to hide a role group. Source: SF Hero 3 run record 08f242f3-aaad-4631-915f-d28e74946264, 2,872 turns through tick 84

The result does not support more direct contact as the explanation for the observed convergence. It does not establish the cause of the remaining movement. Shared broadcasts and campaign context, model-family priors, role, and tier all vary in this run and could produce the same pattern.4

Limits of the analysis

This analysis covers one seven-day run in one scenario. The per-tick analysis uses one embedder, sparse models are absent from some windows, and the centered windows make adjacent measurements dependent. Model, role, tier, exposure, and character assignment were not randomized independently, so the run supports associations rather than causal estimates.

Within those limits, the evidence answers the original question. The run has high topic similarity and a modest increase in residual-style alignment, but the ten models do not converge on one voice. The clearest structure is a group visible on day one, plus later movement by a small number of models. Testing why that movement occurs requires repeated runs that vary role, model, and exposure separately.

Footnotes

  1. All similarity calculations use the native 2,560-dimensional embedding vectors. The animated map and pooled-corpus map use multidimensional scaling (MDS) only for display; their two-dimensional distances are not inputs to the reported measurements.
  2. The centered five-tick window requires at least three turns from a model. Gemini 3.6 Flash, Muse Spark 1.1, Kimi K3, and Qwen 3.8 Max therefore drop out at some ticks. Each tick in the snapshot records its panel size.
  3. Role and model tier are confounded. Candidate characters are frontier-model singletons, while cheaper models play multiple villagers. The observed split may reflect role, model family, tier, or a combination of them.
  4. This comparison tests only the simplest pairwise account: more direct contact should be associated with more convergence. It cannot exclude a smaller contact effect masked by other differences, adaptation to shared public context, or model priors that become more visible in longer histories.

Related content