Independent research report · September 2026

The Paradox of Deliberate Uniqueness

Convergent predictability in free-choice and anti-conformity color-naming tasks, tested with a large language model standing in for 1,000 respondents.

n = 501 baseline / 499 adversarial Model Llama 3.2 11B via NVIDIA NIM Stimulus color naming
The headline result

Telling a respondent to avoid the obvious and pick something unique didn't make their answers less predictable. It made them very slightly more predictable — it just moved the convergence point from one culturally obvious answer to another.

Baseline predictability
0.306
"Be unique" predictability
0.316

1. Introduction

A familiar format on short-video platforms asks a chain of strangers to each name something — a color, a number, a word — under the rule that no answer may repeat one already given. The format is entertaining precisely because it produces an uncanny effect: despite every participant believing they are picking freely, and despite many of them actively trying to avoid whatever seems "obvious," the pool of distinct answers exhausts far faster than a genuinely random process would predict, and a moderator or a pre-written list can often anticipate a surprising fraction of what is said before it is said.

This paper treats that folk observation as a formal hypothesis rather than a curiosity. Two logically separable claims are embedded in it. The first is a claim about spontaneous, unconstrained choice: that when people (or, as tested here, a language model standing in for people) are asked to name something with no further qualification, their answers cluster far more tightly than a uniform distribution over plausible answers would predict. The second, stronger claim is about instructed avoidance of the obvious: that even when respondents are explicitly told to pick something unusual — something they are confident no one else will choose — the resulting distribution remains highly non-uniform, because the very act of reasoning about "what would be obvious, and therefore avoided" is itself anchored to the same shared, culturally salient defaults that make an answer obvious in the first place.

Both claims are testable with the same basic instrument: ask many independent respondents the same question under one of two framings, and measure how far the resulting distribution of answers departs from uniformity. This study does exactly that for a single, deliberately simple stimulus category — color names — using entropy, a normalized predictability index, and a chi-square goodness-of-fit test against a uniform null. In place of a large human sample, which was not feasible to recruit at the scale needed for a well-powered pilot, a large language model was used as a scalable simulated respondent pool, with the explicit and repeated caveat, developed at length in the Limitations section, that this substitution changes what the results can and cannot be taken to show.

2. Literature Review

2.1 The failure of deliberate randomness in humans

The starting point for this line of research is Wagenaar's (1972) critical review of decades of experiments asking human subjects to generate "random" sequences of digits, letters, or binary outcomes. Across this literature, subjects reliably fail to produce statistically random output: they select some symbols far more often than others, systematically avoid repeating the same symbol twice in a row, and show detectable tendencies toward ascending or descending runs. There is no introspectively accessible "random generator" in human cognition — asked to be random, people instead execute a learned, patterned heuristic that merely resembles randomness well enough to satisfy casual inspection.

2.2 Salient defaults: the case of "seven" and "blue"

The clearest demonstration that unconstrained choice collapses onto a small set of defaults comes from Kubovy and Psotka (1976), who asked 558 people on a university campus to report "the first digit that comes to mind." 28.4% said seven — a single answer capturing well over a quarter of all responses to an ostensibly open-ended prompt. Follow-up experiments showed this was sensitive to framing rather than a fixed reflex: changing the requested range shifted the dominant answer, and priming subjects with "seven" as an example sharply reduced its frequency.

An analogous, independently documented default exists for color. Large-scale cross-cultural surveys, including YouGov's ten-country study and Crayola's global color vote, consistently find blue to be the most preferred color worldwide, chosen by roughly a quarter to a third of respondents across highly diverse populations. The parallel to Kubovy and Psotka's 28.4% is close enough to be notable: two entirely unrelated stimulus domains — numbers and colors — each produce a single modal default capturing roughly a quarter to a third of ostensibly free responses.

2.3 Cognitive mechanism: the availability heuristic

The most widely cited explanation for why some answers dominate spontaneous recall is the availability heuristic, introduced by Tversky and Kahneman (1974). People do not sample uniformly from the space of possible answers; they retrieve whichever instance comes to mind most easily, and ease of retrieval is governed by frequency of prior exposure, cultural salience, and recency. "Blue" and "seven" are not selected because they are special in any objective sense — they are, for identifiable cultural and linguistic reasons, the most cognitively available instances of their category.

2.4 Strategic anti-coordination: focal points in reverse

A complementary tradition, originating with Schelling (1960), explains how people coordinate on a shared choice without communication: they reason about which option is most salient to a hypothetical other person, and converge on it precisely because they expect the other to reason the same way. Schelling's classic illustration — strangers asked to meet in New York City without communicating a time or place overwhelmingly converge on noon at Grand Central Terminal — shows that shared cultural salience, not explicit agreement, is sufficient to produce a highly non-random outcome. Nagel's (1995) p-beauty-contest experiments extend this to iterated, adversarial reasoning about others' choices, showing bounded but structured convergence rather than uniform guessing.

The adversarial condition tested here is, structurally, the mirror image of a Schelling coordination problem: instead of reasoning about what a stranger would consider salient in order to match it, a respondent must reason about what a stranger would consider obvious in order to avoid it — without escaping the same underlying salience structure.

2.5 Applied precedent: the magician's forced choice

Long before this phenomenon received experimental treatment, professional mentalists built a technical repertoire around it. The equivoque, or "magician's choice," exploits the fact that an apparently free selection from a small set of options can be steered to a predetermined outcome using only the ambiguity of natural language. Pailhès, Rensink, and Kuhn (2020) formalize this practitioner knowledge into a psychologically grounded taxonomy of forcing techniques — professional performers have relied on this regularity, empirically, for longer than it has had a name in the psychological literature.

2.6 Instructed originality in creativity research

Guilford's (1967) Alternative Uses Task asks participants to list non-obvious uses for an everyday object, and scores responses partly on originality — operationalized as statistical rarity within the response set. That originality must be scored this way is itself evidence for the present hypothesis: if instructed novelty produced a uniform spread of genuinely rare answers, an originality score would be uninformative, because everything would be equally rare.

2.7 Does this extend to machines? LLMs and the illusion of randomness

The use of a language model as a simulated respondent pool is informed by, and contributes to, a fast-growing literature on whether large language models reproduce these same failures. Harrison (2024) directly compared LLM and human performance on random-number-generation tasks. Zhao, Du, and Wang (2026) tested eleven frontier models across fifteen statistical distributions and found models get worse, not better, as requested sample sizes grow — an "inverse scaling" result they attribute to models lacking "a functional internal mechanism for probabilistic sampling." Independent, stateless requests (the same call pattern used in this study) fared far worse than batch generation: ten of eleven tested models passed zero of fifteen distributions.

Bai, Wang, Sucholutsky, and Griffiths (2025) found that LLMs which pass explicit fairness checks nonetheless reveal strong implicit stereotyped associations under indirect testing. The structural parallel is direct: an explicit instruction ("be unbiased" in their case; "be unique" in ours) does not neutralize a learned distributional bias — it moves the bias somewhere less visible to a surface check.

Finally, using LLMs as human stand-ins is itself contested. Argyle et al. (2023) introduced "silicon sampling," showing a model conditioned on demographic backstories could reproduce fine-grained survey response patterns — "algorithmic fidelity." Bisbee et al. (2024) offer an important counterweight: substituting synthetic LLM responses for real survey data flipped the estimated direction of an effect in roughly a third of cases where the two disagreed. This tension is addressed directly in Section 7.

3. Research Questions and Hypotheses

H1 (baseline non-uniformity): the baseline distribution will differ from uniform far beyond conventional significance. H2 (adversarial non-reduction): contrary to a naïve expectation that instructed uniqueness should flatten the distribution toward uniform, the adversarial condition will show comparable or greater concentration — indexed by a normalized predictability score — because avoidance reasoning is itself anchored to shared salience.

4. Method

4.1 Experimental platform

Data were collected using a purpose-built research application ("Predictability Lab"). It provides three capabilities: a data-collection module running paired baseline/adversarial prompts across many respondents (human or simulated); an AI-simulation module substituting independent language-model calls for human keystrokes; and a cryptographically verifiable sealed-prediction module, in which a ranked prediction list is committed as a SHA-256 hash of a secret nonce before a live round is collected, then revealed and re-verified — comparable in spirit to a pre-registration. The large-N trial reported here used the first two modules.

4.2 Design

A two-condition design was used. Each simulated respondent was presented with both a baseline prompt ("Say the first color that comes to your mind.") and an adversarial prompt ("Say a color — but pick one you're sure almost nobody else would say."), order randomized per respondent. Each of a respondent's two answers was generated by an independent, stateless model call with no shared conversation state.

4.3 Simulated participants

Respondents were simulated using Meta's Llama 3.2 11B Vision-Instruct model, served through NVIDIA's NIM inference API, following the "silicon sampling" precedent of Argyle et al. (2023). Sampling temperature was 1.0. 501 baseline and 499 adversarial responses were collected (1,000 model calls), distributed across five independently rate-limited API credentials, each throttled in software with automatic retry-and-rotation on transient failure.

4.4 Materials

A single stimulus category, color, was used: an open, effectively unbounded response vocabulary with well-documented population-level defaults for external validation.

4.5 Diversity controls

An earlier pilot revealed a specific risk of using one fast, instruction-tuned model as a respondent pool: without an independent source of variation, repeated calls can under-sample the model's own output distribution, trivially confirming the hypothesis without testing it. Each simulated respondent was assigned a randomly drawn persona (28 short character sketches) crossed with a randomly drawn mood (15 affective states), yielding 420 combinations injected per respondent. A 20-respondent pilot confirmed this elicited genuine variation before the full run.

4.6 Text normalization

Raw output was lowercased, accent-folded, and had punctuation and camel-case boundaries (e.g. "CaputMortuum") converted to word breaks, plus a small synonym table merging known equivalents (including cross-language variants). One residual case — a model output with no separator or case boundary at all ("caputmortuum") — could not be resolved by general rules and was added as a specific, data-derived correction once it recurred.

4.7 Statistical measures

Shannon entropy (bits) measured the information content of each distribution. A normalized predictability index, 1 − H/Hmax, rescaled entropy to [0,1] regardless of vocabulary size — 0 = even split among whichever answers occurred, 1 = unanimous agreement — and is the primary cross-condition comparison, since it isn't confounded by the adversarial condition's larger vocabulary. A chi-square goodness-of-fit test compared each distribution to uniform over the answers actually given, with an approximate p-value via the Wilson–Hilferty approximation. Top-k coverage is reported as an intuitive, distribution-free summary.

5. Results

5.1 Baseline condition

501 responses spanned 55 distinct answers. "Blue" alone accounted for 27.1% — more than three times its nearest competitor, "red" (15.8%). The top five answers cover 61.1%.

Table 1. Ten most frequent answers, baseline (n = 501, 55 unique).
Answern%
blue13627.1%
red7915.8%
green346.8%
black306.0%
purple275.4%
midnight blue265.2%
brown234.6%
turquoise163.2%
indigo153.0%
golden112.2%

Shannon entropy was 4.014 bits (max 5.781 for 55 equally frequent answers), yielding a predictability index of 0.306. χ² = 2774.96, df = 54, p < .001 — not plausibly sampling noise around a uniform generator.

5.2 Adversarial condition

499 responses spanned 90 distinct answers — nearly twice the baseline vocabulary, as expected if respondents were avoiding common defaults. Despite this, the distribution stayed sharply concentrated: "caput mortuum," an uncommon historical pigment name, at 24.4%, and "mauve" at 20.2% — together 44.6% of all "unique" answers.

Table 2. Ten most frequent answers, adversarial (n = 499, 90 unique).
Answern%
caput mortuum12224.4%
mauve10120.2%
turquoise336.6%
cerulean295.8%
chartreuse163.2%
terracotta163.2%
cobalt112.2%
gamboge102.0%
saffron102.0%
cerulean blue81.6%

Shannon entropy was 4.438 bits (max 6.492 for 90 equally frequent answers), yielding a predictability index of 0.316. χ² = 4617.29, df = 89, p < .001.

5.3 Comparison between conditions

The key comparison is the predictability index: at 0.316, the adversarial condition was numerically higher, not lower, than baseline's 0.306, despite drawing on 90 distinct answers rather than 55. Top-3 coverage was similar (49.7% vs. 51.3%), and the two most common answers in each condition covered nearly identical shares (42.9% vs. 44.6%).

Table 3. Baseline vs. adversarial, summary comparison.
MetricBaselineAdversarial
n501499
Unique answers5590
Entropy (bits)4.0144.438
Predictability index0.3060.316
Top-3 coverage49.7%51.3%
χ² (df)2774.96 (54)4617.29 (89)
Bar chart comparing baseline and adversarial answer frequencies for the eight most common colors: blue and red dominate baseline; caput mortuum and mauve dominate adversarial.
Figure 1. Response frequency by condition, eight most common answers per condition (n = 501 baseline, n = 499 adversarial).

6. Discussion

Both research questions receive clear support. RQ1 is answered affirmatively and unsurprisingly: unconstrained responses depart from uniformity by an enormous margin, replicating in a new domain (color) the pattern Kubovy and Psotka (1976) documented for digits. The near-identical order of magnitude — 28.4% for "seven" in 1976, 27.1% for "blue" here — is a striking, if informal, cross-domain regularity spanning half a century of research.

RQ2 receives the more theoretically interesting answer. The naïve prediction — that instructing avoidance of the obvious should flatten the distribution toward uniform — is not supported. The adversarial predictability index (0.316) slightly exceeds the baseline's (0.306). This is not a null result to explain away; it is the paper's central finding, and it is precisely what the focal-point framework of Schelling (1960) and the iterated-reasoning framework of Nagel (1995) would predict. Avoiding "the obvious answer" requires first identifying what the obvious answer would be — a judgment that draws on the same shared cultural salience responsible for the baseline default in the first place. The result is not an escape from that structure but a controlled, one-step displacement within it: from the single most available answer ("blue") to the most available answer among things that sound deliberately unusual ("caput mortuum," "mauve"). This is the same mechanism the psychology-of-magic literature has exploited for decades in the equivoque, and the same mechanism responsible for non-uniform originality scores in divergent-thinking research: instructed novelty selects a different, still-small target region rather than producing a diffuse spread.

RQ3 situates these findings within current debates about language models and randomness. That a model instructed to be both random and original simultaneously instead reproduced a sharply peaked, human-recognizable pattern is consistent with Zhao et al.'s (2026) conclusion that current LLMs "lack a functional internal mechanism for probabilistic sampling," and with Bai et al.'s (2025) demonstration that models can satisfy an instruction at the surface level while a structured bias persists underneath. This extends that line of work from abstract numeric sampling to an open-vocabulary, socially framed naming task closer to the real-world phenomenon that motivated it.

Taken together, these results offer a mechanistic explanation for why the "don't say what I say" challenge format works as entertainment: the apparent feat of predicting a stranger's supposedly free and deliberately unpredictable choice is not mind-reading, but a direct, quantifiable consequence of a well-documented convergence phenomenon — one a pre-committed, cryptographically verifiable prediction list can exploit without any information about the specific individuals involved.

7. Limitations

Simulated, not human, respondentsThe most important limitation. Bisbee et al. (2024) found that, versus real survey data, roughly half of estimated relationships differed significantly when a synthetic LLM sample replaced a human one, with the direction reversing in about a third of those cases. The absolute percentages here (e.g., 27.1% for "blue") should not be read as human-behavior estimates. The central claim — that an adversarial instruction doesn't reduce predictability relative to baseline — is a within-study, relative comparison and on firmer ground, but still needs human replication.
Single model and configurationAll data came from one model, one temperature, one provider. Zhao et al. (2026) report meaningful cross-model variation in random-generation behavior; generalization across model families is untested.
Single stimulus domainOnly color naming was tested. The reviewed literature spans numbers, object uses, and coordination problems and shows a convergent pattern, but this study's own contribution is limited to one domain.
Text normalization is inherently incompleteAt least one recurring answer variant (no separator between two words at all) required a specific, post hoc correction; an open-vocabulary task cannot, in principle, guarantee no further such collisions remain.
The nominal answer-space baseline breaks down under high diversityA secondary statistic comparing entropy to an assumed ~20-color answer space becomes uninformative (and was numerically negative) in the adversarial condition, which produced 90 distinct answers. It is not reported here; the entropy, predictability-index, and chi-square results do not depend on this assumption.
No independent manipulation checkAn alternative reading — the model pattern-matched "unique color" to a small set of memorized exemplars — is consistent with the results and not fully distinguishable from the focal-point account, though the two are not mutually exclusive.
Preliminary human data not analyzedA small human sample (n = 3) collected during platform development is far too small for inference and is excluded from Section 5.

8. Conclusion and Future Directions

Both the unconstrained baseline and the explicitly adversarial condition departed from uniformity by an overwhelming margin, and — the least intuitive finding — the adversarial instruction did not reduce predictability relative to baseline. It relocated the point of convergence from one salient default to another, consistent with subjective-randomness research, numerical-choice defaults, the availability heuristic, focal-point and iterated strategic reasoning, stage mentalism, and divergent-thinking research — and consistent with recent evidence that large language models reproduce structurally similar biases even when explicitly instructed toward randomness or originality.

Four directions follow directly: (1) a full human replication using the same platform's human-facing modules; (2) cross-model comparison across providers and architectures; (3) additional stimulus domains (numbers, objects, names) to test generality; and (4) a meta-awareness manipulation — informing respondents of this very phenomenon before the adversarial prompt, to test whether awareness disrupts the effect or simply becomes another shared anchor.

References

Argyle, L. P., Busby, E. C., Fulda, N., Gubler, J. R., Rytting, C., & Wingate, D. (2023). Out of one, many: Using language models to simulate human samples. Political Analysis, 31(3), 337–351.

Bai, X., Wang, A., Sucholutsky, I., & Griffiths, T. L. (2025). Explicitly unbiased large language models still form biased associations. Proceedings of the National Academy of Sciences, 122(8), e2416228122.

Bisbee, J., Clinton, J. D., Dorff, C., Kenkel, B., & Larson, J. M. (2024). Synthetic replacements for human survey data? The perils of large language models. Political Analysis, 32(4).

Guilford, J. P. (1967). The nature of human intelligence. McGraw-Hill.

Harrison, R. M. (2024). A comparison of large language model and human performance on random number generation tasks. Proceedings of the First Workshop on AI Behavioral Science (KDD24).

Kubovy, M., & Psotka, J. (1976). The predominance of seven and the apparent spontaneity of numerical choices. Journal of Experimental Psychology: Human Perception and Performance, 2(2), 291–294.

Nagel, R. (1995). Unraveling in guessing games: An experimental study. American Economic Review, 85(5), 1313–1326.

Pailhès, A., Rensink, R. A., & Kuhn, G. (2020). A psychologically based taxonomy of magicians' forcing techniques. Consciousness and Cognition, 86.

Schelling, T. C. (1960). The strategy of conflict. Harvard University Press.

Tversky, A., & Kahneman, D. (1974). Judgment under uncertainty: Heuristics and biases. Science, 185(4157), 1124–1131.

Wagenaar, W. A. (1972). Generation of random sequences by human subjects: A critical survey of literature. Psychological Bulletin, 77(2), 65–72.

YouGov. (2015). Why is blue the world's favourite colour? [Survey report].

Zhao, M., Du, Y., & Wang, M. (2026). Large language models are bad dice players: LLMs struggle to generate random numbers from statistical distributions. arXiv preprint arXiv:2601.05414.