Verification ← Home

Triangulation catches divergence. It cannot catch confident shared error.

This is the chapter people skip, and skipping it is what turns a rigorous-feeling process into an elaborate way of being wrong with more confidence than before. Running three platforms does not make a claim true. It makes a claim triply asserted, which is a different thing.

Why hallucination survives a good session

Three risks persist no matter how cleanly you ran the protocol:

  • Shared training overlap. Platforms trained on overlapping corpora inherit the same errors from the same upstream sources. If a wrong figure propagated widely enough on the open web, three independent systems will reproduce it independently — and their agreement is a measure of how widely the error spread, not of whether it's true.
  • The plausible fabrication. Some hallucinations survive comparison precisely because they are what a correct answer would look like. A citation with the right author, plausible journal, and plausible year. A statistic in the range you'd expect. Comparison doesn't catch these, because there's nothing anomalous to catch.
  • Confidence without source. Generative systems produce fluent prose regardless of whether anything underneath is accurate. Three fluent answers are three unsourced answers, and fluency is not evidence.

This is the correlated-error finding from Chapter 01 arriving in practice. It is also why the fourth seat earns its chair: a platform whose native job is retrieval fails differently from three platforms whose native job is generation.

The protocol, in one sentence. Before any session output containing specific factual claims is used for a real decision, those claims pass through a grounded, source-citing check — and if it cannot confirm a claim, the claim is unverified, not merely unconfirmed.

What warrants a check

Not everything. Verification is a budget, and spending it evenly is the same as not having one. Four categories earn it:

  • Specific numbers. Figures, percentages, dollar amounts, dates, counts. Anything a reader could look up and find wrong.
  • Named references. Papers, cases, statutes, books, people, quotes. The single highest-yield category — fabricated citations are the classic failure and the easiest to catch.
  • Current-status claims. Anything phrased as "currently," "as of," "now offers," "no longer." Training cutoffs make this category structurally unreliable.
  • Anything harmful if wrong. The catch-all. If being wrong would be expensive, embarrassing, or irreversible, it gets checked regardless of which category it falls in.

Note what isn't on the list: the platforms' reasoning, their framing, their recommendations. Those you evaluate with judgment, not verification. Verification is for claims about the world.

Unanimity does not grant an exemption

The most under-verified claims in any session are the ones all three platforms agreed on, because agreement feels like it did the checking for you. It didn't. Route the claim anyway. In practice this is where verification most often pays for itself.

Verification brief · grounded seat
VERIFICATION REQUEST:
Please verify the following specific claims before I act on them. Treat
each independently. I am not asking whether they sound reasonable — I am
asking what the source says.

CLAIMS TO VERIFY:
1. [Specific claim — number, date, reference, current status]
2. [Second claim]
3. [Third claim]

For each claim:
· Confirm or correct it
· Provide the source, with enough detail that I can find it myself
· State the date the source reflects
· Flag anything that has changed recently or may be jurisdiction-specific

If you cannot find a source for a claim, say so explicitly rather than
reasoning toward a plausible answer. "Unverified" is a useful result.

Run it before synthesis closes, not after

Verification is not a final proofread. Run it while the session is still open, because a corrected fact frequently changes the shape of the conclusion — and discovering that after you've written your synthesis means rewriting it, which creates pressure to keep the conclusion and adjust the fact. Order protects you from yourself here.

  • Every named reference resolves. You found the paper, case, or quote yourself — not a platform's assurance that it exists.
  • Every number has a source and a date. A figure without a date is not a fact; it's a fact-shaped object.
  • Current-status claims were checked against something published recently. Not against a second model's memory.
  • Unverifiable claims are labelled as such in the synthesis. Not quietly dropped, and not silently promoted to established.
  • Anything the check corrected got fed back to see whether it changes the conclusion.

High-stakes domains

Medical, legal, and financial questions share a structure: the information moves faster than training data, jurisdiction changes the answer, and being wrong has consequences that don't unwind. In these domains the ensemble is a powerful tool for education, framing, and preliminary analysis. It is not a substitute for domain expertise, and no amount of triangulation converts it into one.

Use the ensemble to arrive at the specialist's door with a good question. That is a genuine and substantial contribution — and it is a different thing from an answer.

When the ensemble has hit its ceiling

Three signals that you need a human specialist rather than another round:

  • All three platforms are returning similar hedges that end with some version of "consult a professional."
  • The question turns on current, jurisdiction-specific, or highly technical information that isn't reliably in any training set.
  • The consequences of being wrong are irreversible.

The handoff has three stages. Preparation — use the ensemble to do everything it can, so you arrive with a clear situation statement, the specific questions it couldn't resolve, and an explicit flag on what couldn't be verified. Consultation — bring the synthesis as context, not as a conclusion to be ratified. Integration — return to the ensemble with what you learned and ask each platform to update its analysis.

Redact identifying information before sharing sensitive material. And note that specialists are not infallible either; for genuinely high-stakes decisions, a second expert opinion is often worth pursuing for exactly the reasons this whole guide exists.

Field note

"Triangulation catches divergence. Verification loops catch confident shared errors. The council uses both."

Triangulation without verification is incomplete. Verification without triangulation is inefficient.