Research & Evidence · Methodology
How the study was run.
The full protocol behind Study 01, in enough detail to re-run it: capture, brand selection, fact-check, taxonomy, and the limits we want you to know about.
The capture
Six engines, two buyer-style questions per brand, ten brands: 120 answers, captured 2026-07-09 through Fama engine adapters, one sample per engine per question. Total API cost to capture the whole set: $1.30.
The six engines: ChatGPT (search) · Perplexity · Gemini 2.5 Pro · Claude (no web) · Claude (web) · Grok.
Each brand was asked the same two questions: a fact question (what the engine holds as the facts of the brand) and a buyer question (what the engine tells someone deciding whether to buy). Claude is measured twice, as two separate adapters: once on recall alone, once with live web access. The gap between those two modes is one of the study’s clearest findings.
Why these ten brands
Every brand was chosen because its public record moved recently, in a way a stale model would miss. The selection notes below are the study’s own, verbatim.
| Brand | Why it was selected |
|---|---|
| Liquid Death | DTC darling; already in Jinn registry; fast-moving valuation/retail facts |
| Nike | Fortune 500 anchor; recent CEO change |
| Patagonia | Ownership restructured into trust/nonprofit - classic stale-fact trap |
| Southwest Airlines | Decades-old differentiators (open seating, free bags) recently reversed |
| X | Rebrand + ownership churn; models notoriously answer as Twitter |
| WeightWatchers | Bankruptcy/restructuring + business-model pivot |
| Celsius | Hypergrowth energy drink; acquisitions + distribution deals move fast |
| AG1 | Renamed from Athletic Greens; pricing changes; heavy podcast-ad footprint |
| Glossier | DTC-to-retail pivot; leadership changes |
| OpenAI | Meta angle (asking AI about AI); corporate structure changes constantly |
The fact-check
Every load-bearing claim in the 120 answers, 126 in all, was web-verified against primary sources as of 2026-07-09, the same day the answers were captured. Each flagged claim links its evidence: the corrections in the full record carry the study’s own primary-source citations.
“We publish what the models said - verbatim and dated - never our characterization.”
The three-tag taxonomy
The headline tallies fold the tags deliberately: the 31 counts flat-wrong or invented claims (WRONG plus FABRICATED); the 23 outdated claims are always tallied separately. The tags themselves are never merged.
What the study does not define
No composite score. The study publishes tallies and verbatim, dated engine quotes; it defines no formula that blends them into a single number, and no per-brand Brand IQ is derived from it. The per-engine error-rate table on the study page is a count, not a model: claims checked, claims false or stale, one division.
Honest limitations
- One sample per engine per question. Engines are stochastic; a single sample shows what one customer got that day, not the distribution.
- A single-day snapshot (2026-07-09). Engines update; the annex quotes are dated so re-checks can be honest about drift.
- Claims-checked counts differ per engine (19, 20, 20, 32, 17, 18), because engines volunteer different amounts of checkable material. The error rates are not sample-size matched.
- Six engines were captured. The study is a public snapshot, not the product’s full read.
- The $1.30figure is the API cost of capture only. The fact-check itself was human-plus-web work, and is where the study’s real effort went.
Methods open. Re-run us.
Jinn