Factual disagreement
They differ on something with a ground truth: a number, a date, a status, who said what. This is the easiest kind and the one people handle worst — usually by counting votes.
The instinct when three platforms disagree is to treat it as a problem — noise to be resolved so you can get to the answer. That instinct is backwards. Agreement is cheap and weakly informative. Disagreement is a map of where the genuine uncertainty lives, and it is the most valuable thing a session produces.
What matters is not how much the platforms disagree but what kind of disagreement it is. Two responses differing on a date is a different event from two responses differing on what the question is fundamentally about, and treating them the same way is how sessions go wrong.
Four kinds, and each one has a different next move.
They differ on something with a ground truth: a number, a date, a status, who said what. This is the easiest kind and the one people handle worst — usually by counting votes.
They're answering different questions. One read your brief as being about cost, another about risk, another about timing. What looks like conflict is three reasonable readings of an ambiguous prompt.
They agree on the facts and weigh them differently — different risk tolerance, different priority among competing goods, different read on what matters. Nothing is wrong here. The question is genuinely contested.
They differ because their information is dated differently — different training cutoffs, and one retrieved live sources while another answered from memory. This masquerades as factual disagreement and is far more common than people realize.
Comparing three essays to each other is slow and unreliable — the differences hide inside prose that mostly overlaps. So don't compare essays. Split each response into atomic claims, line them up in a row each, and mark the row. What you keep at the end of a session isn't one platform's answer. It's a marked-up ledger.
Work at claim level, not document level. A three-page answer you'd describe as "broadly agreeing" will routinely contain two claims that flatly conflict, and those two are the entire value of having run three platforms.
| Claim | Claude | ChatGPT | Grok | State |
|---|---|---|---|---|
| The effect direction is positive | Positive | Positive | Positive | Unflagged |
| Effect size in the pivotal trial | 0.4 SD | 0.4 SD | 0.2 SD | Flagged |
| Pivotal trial sample size | n≈1,200 | n≈1,200 | n≈1,200 | Verified |
| Result replicated independently | Unclear | Yes | Unclear | Flagged |
| Effect holds in the older subgroup | Not established | Yes | Not established | Flagged |
The tempting version of this ledger has two states: agreed and flagged. Resist it. A two-state ledger quietly teaches you that agreement is the finish line, so the agreed rows ship unexamined — and those are precisely the rows where three platforms can be wrong together.
The third state is what makes the ledger honest. Agreement moves a claim out of the "needs adjudication" pile; only verification moves it out of the "might be wrong" pile. Most rows in a real session end at unflagged, and the discipline is saying so rather than rounding up.
Once you're keeping ledgers, the share of rows that come back flagged becomes a genuinely useful number — and it moves in an informative direction. A high rate means the question is contested or your brief was loose. A very low rate on a question you know is hard is not a clean bill of health; it's the blind-spot signal arriving as a statistic.
Across sessions it also profiles your ensemble. If one platform is almost never the odd one out, it has stopped earning its chair — it's an echo, and you should audition a replacement that fails differently.
Fast, confident convergence on a question your own experience tells you is genuinely hard is not a reliable answer. It is the blind-spot signal.
If the ensemble is wrong and the error is structural, no amount of further deliberation will surface it. Asking a fourth time produces a fourth version of the same mistake. Only your outside perspective, or an external source, can catch it. When three platforms agree instantly on something hard, ask: what would have to be true for this answer to be wrong? — and then go check that thing specifically.
Three questions that sort a disagreement into its type. The widget is not the point — the questions are, and they're the ones worth asking before you pick a side.
Answer all three to get a verdict.
Deadlock is when the responses genuinely conflict and you cannot determine which is more reliable. The recovery sequence, in order:
A session that produces only agreement is not necessarily a session that found the truth. Before accepting consensus, run three checks: reread the brief for loading, ask your friction platform explicitly for the dissenting case, and search for counterevidence rather than confirmation. If the agreement survives all three, it is probably genuine — and now you have a reason to believe it beyond the fact that it was unanimous.
"The disagreement is the information — it tells you where the genuine uncertainty lives and where you need to look harder."