Divergence ← Home

The disagreement is the finding. Don't average it away.

The instinct when three platforms disagree is to treat it as a problem — noise to be resolved so you can get to the answer. That instinct is backwards. Agreement is cheap and weakly informative. Disagreement is a map of where the genuine uncertainty lives, and it is the most valuable thing a session produces.

Quality over quantity

What matters is not how much the platforms disagree but what kind of disagreement it is. Two responses differing on a date is a different event from two responses differing on what the question is fundamentally about, and treating them the same way is how sessions go wrong.

Four kinds, and each one has a different next move.

Checkable

Factual disagreement

They differ on something with a ground truth: a number, a date, a status, who said what. This is the easiest kind and the one people handle worst — usually by counting votes.

DoGo to the primary source. Two platforms agreeing against one tells you nothing about which is right; correlated error makes majority a poor estimator here.
Structural

Framing disagreement

They're answering different questions. One read your brief as being about cost, another about risk, another about timing. What looks like conflict is three reasonable readings of an ambiguous prompt.

DoFix the brief and re-run. Adjudicating between them now just launders your own ambiguity into a conclusion.
Substantive

Judgment disagreement

They agree on the facts and weigh them differently — different risk tolerance, different priority among competing goods, different read on what matters. Nothing is wrong here. The question is genuinely contested.

DoReport the split as your finding. Name what would decide between the positions, and say which way the evidence leans and why.
Temporal

Time-skew

They differ because their information is dated differently — different training cutoffs, and one retrieved live sources while another answered from memory. This masquerades as factual disagreement and is far more common than people realize.

DoAsk each one directly: as of what date is this true, and did you retrieve it or recall it? A confident stale answer should lose to a sourced fresh one every time.

The ledger is the deliverable

Comparing three essays to each other is slow and unreliable — the differences hide inside prose that mostly overlaps. So don't compare essays. Split each response into atomic claims, line them up in a row each, and mark the row. What you keep at the end of a session isn't one platform's answer. It's a marked-up ledger.

Work at claim level, not document level. A three-page answer you'd describe as "broadly agreeing" will routinely contain two claims that flatly conflict, and those two are the entire value of having run three platforms.

A worked ledger for one research question briefed to all three. Five claims, three flagged — and note that the two unflagged rows are not finished. They are merely unopposed.
Claim Claude ChatGPT Grok State
The effect direction is positive Positive Positive Positive Unflagged
Effect size in the pivotal trial 0.4 SD 0.4 SD 0.2 SD Flagged
Pivotal trial sample size n≈1,200 n≈1,200 n≈1,200 Verified
Result replicated independently Unclear Yes Unclear Flagged
Effect holds in the older subgroup Not established Yes Not established Flagged
Flagged They disagree. Classify it using the four types above, then resolve or report it. These rows get your attention first because they're the ones announcing themselves.
Unflagged They agree — and that is all it means. Not checked, not settled. This is where correlated error lives, so unflagged is a holding state, never a conclusion.
Verified You took it to a source yourself and the source held. Only these rows are actually finished, and only these should carry weight in what you write.
Two states is one state too few

The tempting version of this ledger has two states: agreed and flagged. Resist it. A two-state ledger quietly teaches you that agreement is the finish line, so the agreed rows ship unexamined — and those are precisely the rows where three platforms can be wrong together.

The third state is what makes the ledger honest. Agreement moves a claim out of the "needs adjudication" pile; only verification moves it out of the "might be wrong" pile. Most rows in a real session end at unflagged, and the discipline is saying so rather than rounding up.

Divergence rate

Once you're keeping ledgers, the share of rows that come back flagged becomes a genuinely useful number — and it moves in an informative direction. A high rate means the question is contested or your brief was loose. A very low rate on a question you know is hard is not a clean bill of health; it's the blind-spot signal arriving as a statistic.

Across sessions it also profiles your ensemble. If one platform is almost never the odd one out, it has stopped earning its chair — it's an echo, and you should audition a replacement that fails differently.

Where the ledger earns its keep

  • Evidence and literature review. Citations and mechanisms are where models improvise most. A ledger turns a plausible reference list into a list with the disputed entries marked — and every named reference belongs in the verified column or nowhere.
  • Contracts and compliance. Obligations, deadlines, and carve-outs are exactly the claims a single platform states too smoothly. Claim-level comparison across three is the cheapest first pass available.
  • Numbers and derivations. Three independent derivations landing on the same figure are worth far more than one arriving with a confident explanation attached. Reconcile first, then trust.
  • Clinical reading. Never a diagnosis. A second and third reading of the same summary, with every disagreement surfaced for the clinician rather than averaged away.
Unanimous confidence is a warning, not a comfort

Fast, confident convergence on a question your own experience tells you is genuinely hard is not a reliable answer. It is the blind-spot signal.

If the ensemble is wrong and the error is structural, no amount of further deliberation will surface it. Asking a fourth time produces a fourth version of the same mistake. Only your outside perspective, or an external source, can catch it. When three platforms agree instantly on something hard, ask: what would have to be true for this answer to be wrong? — and then go check that thing specifically.

Triage a conflict

Three questions that sort a disagreement into its type. The widget is not the point — the questions are, and they're the ones worth asking before you pick a side.

Question 1

Reading the three responses closely, did they all actually answer the same question?

Question 2

Does the point they disagree on have a checkable ground truth — a number, date, document, or record that exists somewhere?

Question 3

Does the answer depend on recent events, current status, or anything that has changed in the last year?

Answer all three to get a verdict.

When it won't resolve

Deadlock is when the responses genuinely conflict and you cannot determine which is more reliable. The recovery sequence, in order:

  1. Reframe. Restate the question more precisely and re-brief. A surprising share of deadlocks are framing problems wearing a disguise.
  2. Decompose. Break the question into smaller components and find the specific sub-claim where the paths diverge. Deadlocks are rarely total; usually one hinge is doing all the work.
  3. Second-order pass. Show each platform the competing positions — after the independent round, never before — and ask it to evaluate the reasoning rather than restate its answer.
  4. Ground it. Route the factual hinge to a grounded source. Many "irreconcilable" splits collapse the moment one fact is nailed down.
  5. Accept it. If it persists, treat unresolvability as information. Some questions are not settled, and reporting that honestly is a better result than a confident answer you manufactured.

The opposite problem

A session that produces only agreement is not necessarily a session that found the truth. Before accepting consensus, run three checks: reread the brief for loading, ask your friction platform explicitly for the dissenting case, and search for counterevidence rather than confirmation. If the agreement survives all three, it is probably genuine — and now you have a reason to believe it beyond the fact that it was unanimous.

Field note

"The disagreement is the information — it tells you where the genuine uncertainty lives and where you need to look harder."