The bar we will not publish below.
Before Jinn publishes a number about how AI sees your brand, that number has to clear a floor we set for ourselves. This page is the method, not a result: the floor, the gates, and why we would rather publish nothing than publish a number we do not yet trust.
There is no accuracy figure on this page, and that is deliberate. The discipline below is why.
What this is
The method behind measuring how AI sees a brand: the floor a number has to clear, and the separate gates that catch the ways measurement goes wrong. It is method, not a score.
The floor
Our grader’s verdicts must agree with a human-labeled golden set at least 85% of the time before any score is allowed to publish. Eighty-five percent is the bar to clear, not a result we are claiming.
Why no number
The method is real code; the grader has not yet cleared its own floor, so no result publishes - not on this page, and not in the product. That refusal is the point.
The bar we hold ourselves to
The publish floor is straightforward: our grader’s verdicts must agree with a human-labeled golden set at least 85% of the time, claim by claim, before any score is allowed to publish. The same code path computes that agreement in the test suite and in the live gate, so the bar cannot quietly drift from the check that enforces it.
Eighty-five percent is the bar, not a result. It is the line the method has to clear before a number is fit to show you - never a claim about where the grader stands today. We may raise it; we will not quietly publish under it.
Each gate catches a different failure
Agreement alone can be fooled, so the discipline is more than a single number. Each gate below catches a different way measurement goes wrong.
- agreement
- Our verdicts must match a human-labeled golden set at least 85% of the time before any score publishes.
- recall
- A separate gate checks how many known-bad claims the first stage flagged for verification. AI engines fail in correlated ways - a consensus can agree on the same wrong fact - so high agreement with low recall is a named failure the harness surfaces, not a pass.
- coverage
- A brand is scored only when at least five distinct engines answered every question. Below that, the run is re-queued, never scored on thin data.
- confidence
- A published score carries a 95% confidence interval. When there are too few confident observations to estimate honestly, the method reports no estimate rather than a shaky one.
- no accusation
- Calling an answer "wrong" requires an authority-class source. Without one, the verdict is downgraded to unverified disagreement, so the method never accuses a brand’s critics on a hunch.
These gates are deliberately separate. A run can pass one and fail another, and a single strong number never buys its way past the rest.
This is the discipline behind the numbers. The place to start is your own brand. Read your brand free
Publishing nothing, on purpose
There is no accuracy figure on this page by design. The method above is real code; the grader has not yet cleared the floor it sets, so no result publishes - not on this page, and not in the product.
That refusal is the whole point. When the number is good enough to trust, you will see it. Until then, publishing nothing is the honest answer, and the more defensible one.
Keep going
This page is the measurement discipline. The full tour shows the mechanism it sits inside - how Jinn learns, wears, and works a brand.
- The full tourThe mechanism end to end: how Jinn learns, wears, and works a brand.
- The recordThe 346-signal Brand DNA record, group by group.
- Models & routingWhich model does which job, and the gateway rules underneath.
- Guardrails & spendThe deterministic checks that guard your money and your name.
- Check it yourselfWhat “verified” means here: human approvals, recorded provenance, and facts that age on purpose.
Measurement, answered
- Why doesn’t this page show an accuracy number?
- Because the number is not ready to trust. Our grader must agree with a human-labeled golden set at least 85% of the time before any score publishes; until it clears that floor, we publish nothing rather than a figure we cannot stand behind.
- Is 85% your accuracy?
- No. 85% is the floor the method has to clear before a result is fit to show, not a measurement of where the grader stands today. It is a bar, not a score.
- How do you know the AI engines are not agreeing on the same wrong answer?
- That exact failure is why agreement is not the only gate. A separate recall gate checks how many known-bad claims the first stage flagged, because engines fail in correlated ways - high agreement with low recall is treated as a failure, not a pass.
- Can I measure my brand’s AI visibility in Jinn today?
- This page is about the measurement method and the bar it has to clear, not a live product surface. The free brand read shows what Jinn extracts from your brand today; the visibility scoring described here publishes only once it clears its floor.
We would rather publish nothing than a number we do not trust.
That is the whole discipline. Start with the free brand read - the part of Jinn that is ready for you today.