Two points can disagree. Three points can locate you.
One AI gives you an answer. Two give you a contradiction you can't settle. Three give you a position — but only if you run them so their errors don't line up. This is the discipline of getting research-grade work out of Claude, ChatGPT, and Grok: identical briefs, independent outputs, and a human who does the deciding.
Asking a second AI "is this right?" is weaker than it feels.
The intuitive case for using several platforms is that errors will cancel out. The research says otherwise: models fail in correlated ways, and the correlation is strongest exactly where you'd least expect it — among the largest, most accurate models, across different providers.
That doesn't make triangulation worthless. It makes naive triangulation worthless. Everything in this guide is about engineering the independence back in.
On one leaderboard dataset, when two models were both wrong, they were wrong in the same way about 60% of the time. A second opinion that shares your first one's blind spot is not a second opinion.
Kim, Garg, Peng & Garg, arXiv:2506.07962The same study found that shared architecture and shared provider drive correlation — but that larger, more accurate models have highly correlated errors even across distinct architectures and providers.
Two leaderboards + a resume-screening taskTwo voices can contradict each other without telling you which is wrong. Three is where triangulation becomes operational — the point at which disagreement starts carrying usable information.
The count is an example; the triangulation is the lawNine chapters, in playing order.
Read straight through like a score, or jump to the movement you need. Each chapter stands on its own and ends where the next one picks up.
Why three
Hallucination, sycophancy, and the single-voice constraint — plus what the correlated-error research actually shows.
Read → 02The seats
Claude, ChatGPT, and Grok by temperament rather than benchmark — and the fourth seat that grounds them.
Read → 03The protocol
How a session actually runs, from framing the question to closing it without deciding.
Read → 04The brief
Identical inputs, three components, neutral phrasing — and why a loaded prompt is self-inflicted.
Read → 05Routing
Breadth, depth, or grounding? Who to call first, when to call everyone, and when to stop polishing the prompt.
Read → 06Divergence
Reading disagreement as a map of genuine uncertainty — the claim ledger, plus a triage tool for conflicts you can't settle.
Read → 07Verification
Triangulation catches divergence. Verification catches confident shared error. You need both.
Read → 08Failure modes
Deadlock, false consensus, session drift, fatigue — fifteen ways the method breaks, each with its fix.
Read → 09Reference
Copy-ready briefs, the session checklist, the log format, and the rules in compressed form.
Read →