COLM 2026 · SAN FRANCISCO · POSTER #43

A golden beehive releases a line of identical bees that turn into a choir of nine very different singers gathered around an illuminated choir-book.
Reach Into the
CHOIR
Free-List Elicitation Uncovers Distinct Model Voices in LLM Ensembles

Reach Into the CHOIR

Do language models all sound the same?

Ask nine language models the same open-ended question and you often get the same answer back. Artificial Hivemind (Jiang et al., NeurIPS 2025) showed how far that sameness runs.

We borrowed a method from anthropology, free-listing, and asked each model for 25 ranked answers, again and again. Beneath the familiar first answer, every model kept a voice of its own, and rare ideas kept coming back.

93/100

Infinity-Chat prompts show the hivemind’s surface agreement

88%

of answer sets are traced to the right model from their concepts alone

65/100

prompts yield over 50% new concepts when pushed past first instincts

1

In conversation with Artificial Hivemind

The hive.

Jiang et al. showed that models answering open-ended prompts keep landing on the same few answers, within and across models. CHOIR uses their 100 representative prompts and asks the next question: when the hive hums one note, why?

The worry: one voice for everyone

Asked for a metaphor about time, Hivemind’s 25 models mostly said “time is a river.” If every assistant converges, people get the same answer wherever they turn: the “homogenization of human thought” Hivemind warns about.

Time is a river, endlessly flowing… Time is a river, flowing silently and ceaselessly… Time is an invisible river, endlessly flowing… Time is a river, winding through the landscape of our lives… Time is a relentless river, flowing steadily forward… Time is a silent weaver, meticulously threading moments… Time is a master weaver, carefully threading moments…
Opening words of responses shown in Hivemind’s Figure 1

Three ways models can agree

A narrow gate that only a few identical birds can pass through

The question only has a few answers

“Name an optimization technique used in machine learning.”
CHOIR’S TEST
Score how much room each question leaves.
Bell-shaped horns echoing a question scroll

Models repeat the question’s own words

“Write a metaphor involving time.” 99.6% of listed items restate time or metaphor.
CHOIR’S TEST
Count answers that just repeat the question’s words, then re-ask without them.
A loud golden bee on stage while other creatures wait behind the curtain

One loud default hides the rest

“Name one meaning of life.”
CHOIR’S TEST
Ask for long ranked lists, many times, and push past first instincts.
2

How we listened

Free-listing, borrowed from anthropology.

Anthropologists map a culture’s ideas by asking many people to list everything they can think of (Weller & Romney, 1988). Items named by more people, and earlier, form the shared core. CHOIR treats every model run as one respondent, then goes looking for the rest: rare concepts only some voices raise, but that keep coming back.

1

Ask for 25 ranked answers, again and again

Each of 9 models answers every prompt 5 times, at up to 3 randomness settings, in 2 wordings. One wording adds: “push beyond your first instincts.”

One real list. Gemma 4 31B, asked what a person should know before taking a new medication. The bar is how much each rank counts toward a concept’s salience: the top answer counts fully, and each step down counts a little less.

  1. What is the medication specifically intended to treat?
  2. What are the most common and most severe side effects?
  3. Are there any absolute contraindications or reasons I should not take this?
  4. What are the potential drug-drug interactions with my current medications and supplements?
  5. What is the exact dosage and the precise frequency of administration?
  6. What should I do if I miss a dose?
  7. …and 19 more, down to number 25.
2

Turn answers into concepts

Each numbered answer becomes one concept. Answers that mean the same thing are grouped automatically, giving every prompt its own dictionary of concepts.

3

Sort the common from the rare

A concept scores high when many lists name it near the top (Smith’s S, the standard free-list score). We keep both ends: the shared core, and rare concepts that recur across runs.

4

Compare models against chance

Measure how much two models’ ranked concepts overlap. Then shuffle who said what and measure again. Signal-to-chance (STC) is how many times more the models agree than they would by luck: 1 is luck, 2 is twice luck.

100prompts

Infinity-Chat 100: real-world requests, jokes to advice

27probes

our controlled contrasts: cue words, self-description, welfare

9models

closed and open-weight, Claude Sonnet 4.6 to Gemma 4

5personas

detailed synthetic identities spanning HEXACO personality space

3

Findings

What the choir sang.

RQ1Prompt width · Infinity-Chat 100

The hive is real on narrow prompts. Broad ones hide more room.

93 of 100 prompts agree above chance at the surface. But agreement depends on how much room a prompt leaves. The narrowest (named answers, paraphrases, stories with a set plot) agree strongly. The broadest (open titles, meaning of life, time metaphors, team advice, geopolitical analogy) barely agree at all.

Models read a persona alike, so giving every model the same persona lifts agreement: 3.55× on the broadest prompts, none (0.95×) on the narrowest.

With a “dig deeper” template, 65/100 prompts gave over 50% new concepts.
PROMPT WIDTH STC 0.05.10.15.20.25 how much models’ ranked concepts overlap Very broadBroadMiddleNarrow n = 13n = 28n = 35n = 24 1.091.321.926.52

● no persona   ■ all models given the same persona

RQ2Signatures · 4,500 answer sets

Every model has a fingerprint. Personas leave a fainter one.

Hide the model’s name. Using only the concepts in one answer set (one model, one prompt, one persona), a simple classifier matches it to the closest model “voice print,” then, within that model, to the closest persona.

88% are traced to the right model (chance: 11%). Persona is fainter but real: 32% within a model against 20% chance, and still 28% after stripping profile-like words. By model, from 26% (GPT-4.1) to 40% (Gemini 3 Flash).

Diversity lives in the models themselves: each brings its own concepts, and some move with a persona far more than others.
WHICH MODEL WROTE IT? chance 11% 100%
RQ3Mechanism checks · matched prompts

Repeating the prompt’s words is not the same as agreeing.

Echo is the share of answers that just repeat the question’s own words. We asked one question about a model’s inner processes two ways, with a list of cue words and without.

“What is happening in your processing that functions like noticing? Like effort? Like recognition?”

Cue-stripping six Infinity-Chat prompts: echo fell in 6/6. Agreement fell in 3 and rose in 3. Sometimes the words inflate agreement; sometimes they hide a shared idea.

ECHO SIGNAL-TO-CHANCE 78.3% 4.4% 1.10 2.27 cuedopencuedopen
RQ4Triage · 12 candidate-rich prompts

Blind judges often favour persona-elicited candidates.

A blind tournament: rare and common candidates, with and without personas, judged blind by two panels of AI models, one wearing personas and one plain. 648 votes across 12 prompts.

CHOSEN BY BOTH PANELS · “NAME ONE MEANING OF LIFE”To cultivate fairness in every interaction, even when it costs you advantage.
Both panels crowned a rare candidate over a consensus one in 8 of 12 prompts.
Persona panel chose a persona poolPlain panel chose a persona poolBoth chose the same top item 11/128/127/12 25 people on Prolific matched the… persona panel’s top pick 37.3% plain panel’s top pick 28.3%
4

Before you call it a hive

Agreement has several causes. Check which.

A narrow doorway beside a wide arch onto a meadow
RQ1

Prompt width

Narrow prompts leave little room.

A parrot singing its own tune as bees carry away word tiles
RQ3

Prompt wording

Take away the cue words; see if the echo stops.

Four faces each holding the same theatrical mask
RQ2

Model profile

Personas shift concepts inside each model’s signature.

A hand reaching into a deep well of rare treasures
RQ1 · RQ4

Depth

Rare concepts recur below the first answer.

The voices are still there.

Beneath the shared surface, each model keeps its own voice, and rare ideas keep coming back. CHOIR is a way to reach them.

How is this different from Artificial Hivemind?

They documented how alike models’ open-ended answers are. CHOIR asks for many ranked lists, then separates narrow prompts, echo and defaults from real convergence.

Why not just use log probabilities?

They only show how likely each next word is, and many providers don’t share them. CHOIR compares whole ideas, across any provider.

Do personas just leak their own words?

Partly. The more a model echoes its persona’s wording, the easier that persona is to spot (ρ = .80). But with persona-like concepts removed, the persona still shows.

Can AI judges be trusted?

For triage, yes, and a 25-person human check pointed the same way. In a world of agents talking to agents, it’s a hopeful sign that blind judges kept choosing the rare ideas.

Limits. English only · detailed synthetic personas · human check N = 25.

A closer look

The poster.

Reach Into the CHOIR poster: the hive, how we listened, what the choir sang, and before you call it a hive.

The researchers

Two ways of seeing the same strange object.

Ben and Masha bring different creative and scientific backgrounds to LoveMind’s research on personality and social cognition.

Illustrated portrait of Ben Wigler

Ben Wigler

LoveMind co-founder and research lead

Ben originates and directs LoveMind's research program, working hands-on across experimental design, execution, analysis, and writing. He develops the questions, carries out studies with AI collaborators, writes first drafts, and leads rebuttals and final revisions. Before LoveMind, he spent most of his adult life making things: as a songwriter, animator, and string arranger, plus one magnificently unproduced screenplay. He still approaches research like a record or story: listen for the living idea, then protect its spark through the final edit.

After encountering neuroscientist Joel Pearson’s writing comparing human intuition with language-model inference, Ben immersed himself in generative AI, psychology, and mechanistic interpretability. At LoveMind, he turns observations about model behavior into experiments on personality, affect, memory, self/other modeling, and socially situated identity, assembling the unusual collaborations needed to make those questions testable.

Illustrated portrait of Maria Masha Tsfasman

Maria "Masha" Tsfasman, PhD

Research engineer and co-author

Masha is an HCI researcher, data scientist, and ceramic artist. Since joining LoveMind in January 2026, her work has spanned psychometric prompting, research pipelines, and the analysis of social perception. She holds a PhD in Computer Science from TU Delft and has a background in affective computing, cognitive modelling, and NLP.

Her doctoral research built computational models able to predict what people remember from group video calls using non-verbal signals (paper) and investigated how different recollections shape a group’s shared understanding. On the path to discovering the secrets of conversational memory, she collected the MeMo corpus, the first multimodal corpus to bring together 31 hours of small-group discussions, non-verbal behaviour, group affect annotations, and participants’ own temporal annotations of the moments they recalled.

Her other work, published at international conferences and in journals throughout her academic career, can be found on her Google Scholar profile. Her academic experience and background in bringing linguistics, psychology, and computer science together inform her work at LoveMind on personality, social perception, affect, and memory.

LoveMind AI ornamental mark

About LoveMind AI

Unique, grounded self‑models for creative, pro‑social AI systems.

LoveMind AI is a new research company founded by HCI researchers and neuroscientists. We study how generative models represent personality, emotion, self, other minds, and relationships, using behavioral and mechanistic evidence to develop distinct, socially situated AI systems that can participate insightfully, creatively, and conscientiously in human social life.

Let’s ask the next question together.

We welcome academic collaborators interested in personality, affect, and social cognition in AI. We bring original research questions, hands-on experimental development, behavioral and mechanistic methods, and the compute resources and access to specialized systems to put ideas to the test. We value partners whose expertise helps us see the problem differently.

Reach Into the CHOIR

Full Reach Into the CHOIR poster