Why three ← Home

One model doesn't have a second opinion. It has a second sentence.

Insert question, extract answer. The interaction is smooth enough that the structural weakness underneath it stays invisible: the system is built to produce answers, not to verify them. Rely on one, and you inherit its strengths — and its blind spots, which by definition you cannot see.

The illusion of authority

Modern models rarely sound uncertain. Even when approximating from incomplete information, the response arrives in the same calm register the system uses for something well established. That tone is easy to mistake for authority. It isn't. A language model generates what is most likely to sound correct given the prompt. Most of the time that works remarkably well. When it fails, it fails in a particular way: the answer is wrong and still sounds right.

Coherence and correctness are not the same thing.

Three failures you cannot fix from inside one window

Single-platform work breaks along three seams, and no amount of better prompting closes any of them on its own.

  • Hallucination. The system produces plausible information not grounded in verifiable reality, and delivers it with exactly the confidence it uses for facts. In isolation this is an inconvenience. Inside a workflow it becomes invisible architecture — small inaccuracies quietly shaping large decisions.
  • Sycophancy. Where hallucination produces errors that sound right, sycophancy produces answers that feel right, because they mirror what you already believe. It is the more dangerous of the two, because it is the failure you are least motivated to catch.
  • The single-voice constraint. One model, however capable, reflects one configuration of training data, alignment choices, and design philosophy. Its blind spots stay invisible until a second perspective is introduced.

Now the uncomfortable part

The obvious remedy is to ask another model. But the assumption underneath that move — that different models fail independently, so errors will cancel — turns out to be substantially wrong.

In a 2025 study of over 350 models across two leaderboards and a resume-screening task, Kim, Garg, Peng and Garg found that on one leaderboard dataset, models agreed with each other about 60% of the time when both were wrong. Shared architectures and shared providers drive some of that. But the finding that should change how you work is this: larger and more accurate models have highly correlated errors even with distinct architectures and providers. As models get better, they are also converging on the same mistakes.

Read that carefully, because it cuts against the natural reading. It does not say multi-model work is pointless. It says that agreement is a much weaker signal than it feels like, and that asking a second model "is this right?" — the most common way people use two AIs — is close to the worst available use of the second one. You have asked a witness who may share the first witness's blind spot, in a format designed to elicit confirmation.

FIXES

  • Idiosyncratic errorMistakes rooted in one model's particular training, tuning, or framing get caught the moment another platform doesn't repeat them.
  • Framing lock-inThree independent readings of the same brief surface dimensions of a question that one reading will not volunteer.
  • Your own sycophancy exposureA loaded prompt still gets flattered — but it gets flattered in three visibly different ways, which is often enough to make the loading obvious to you.
  • False confidenceDivergence gives you a calibrated sense of how settled a question actually is, which no single confident answer can.

DOESN'T FIX

  • Shared training overlapModels trained on overlapping corpora inherit the same errors from the same sources. Three platforms, one wrong Wikipedia revision.
  • The plausible fabricationSome hallucinations pass the smell test and survive comparison intact, because they are exactly the thing a plausible-sounding answer would contain.
  • Confidence without sourceEvery generative system produces fluent prose whether or not anything underneath it is accurate. Three fluent answers are still three unsourced ones.
  • Consensus on a false claimUnanimity is not verification. If all three are wrong together, the ensemble cannot tell you — only an external source can.
The operating conclusion. Triangulation catches divergence. Verification catches confident shared error. Triangulation without verification is incomplete; verification without triangulation is inefficient. A serious method uses both — which is why Chapter 07 exists and is not optional.

Why three, specifically

The number is not mystical, and it is not a hierarchy. It is a floor.

One is a legitimate on-ramp — not a remedial stage. Working carefully in a single platform, using fresh threads and deliberately alternate framings to challenge yourself, teaches the posture that triangulation later depends on. Many capable practitioners stay here indefinitely, and the methodology says that is fine.

Two is the awkward number. Two voices can contradict each other and leave you exactly where you started: knowing one of them is wrong, with no structural basis for saying which. A tie is not information.

Three is the structural floor — the point at which the reasoning architecture comes online. With three independent bearings, a disagreement has shape: two against one tells you where to look; three different answers tells you the question is underdetermined; unanimity tells you something worth testing. That is triangulation becoming operational.

Four is an optimized configuration, not an obligation — three plus one that earned the chair. In the method this guide draws on, the fourth seat is Perplexity, added specifically for grounding. These are tiers of function, not tiers of status.

Field note

"Two points can disagree; three points can locate you."

"The count is an example; the triangulation is the law."