How Often Do AI Models Actually Agree?
Satcove asks up to six independent AI models the same question and scores how strongly they converge. The public figure is withheld until enough runs have accumulated to be meaningful.
Not enough runs yet
218 of 300 consensus runs recorded. The aggregate publishes automatically once the floor is crossed. This page is excluded from search until then.
What this number is — and is not
It is a usage-transparency figure
The mean agreement score across real consensus runs on Satcove. It describes how often independent models line up on the questions people actually bring.
It is not an accuracy benchmark
Agreement is not correctness. A reproducible accuracy evaluation, with a versioned public prompt set, is tracked separately on the Accuracy Index.
It exposes nothing private
Only counts and a mean. No prompt text, no per-model answers, no user or account data is part of this aggregate.
Frequently asked questions
How often do AI models agree with each other?
Satcove scores how strongly up to six independent models (Claude, GPT, Gemini, Mistral, Perplexity, Grok) converge in meaning and direction on every consensus run. It is a measure of alignment, not of truth.
Why does this figure only cover recent runs?
Scores from the first months after launch differ sharply from later ones, so averaging them together would misstate how often models agree today. The figure covers recent runs only and is recomputed from the live data.
Does a high agreement score mean the answer is correct?
No. Models can share the same training-data blind spot and agree on something wrong. Agreement is a routing and confidence signal; it is never presented as accuracy or probability of truth. Sourced evidence is reported separately on the /proof page.
What should I do when the models disagree?
Treat the disagreement as the most useful part of the result: read which claims differ, check the decisive ones against a primary source, and do not act on the majority view alone. Satcove names the disagreement instead of averaging it away.
How is the agreement score calculated?
It blends how close the answers are in meaning with whether they point the same direction (recommend / caution / against). The score is bucketed: strong (75-100), moderate (45-74), divergent (under 45). The exact blend is part of Satcove's consensus contract.
What data is this based on?
Aggregate counts over Satcove's own consensus runs — no prompts, no per-model outputs and no user identifiers are published. The figure updates as more runs accumulate and only appears publicly once at least 300 runs have been recorded.
Related: where the panel diverges · consensus vs. sourced proof · Accuracy Index (research preview)
Satcove — A product by Abyssal Group