The seats ← Home

Look at them the way you'd look at players in a band.

Not model sizes, not benchmark scores — those change quarterly and mean less in practice than you'd expect. What matters is what each one tends to do well, where it wobbles, how it thinks, and when to call on it. These are tendencies, not cages.

The three cognitive modes

Useful AI work falls into three modes, and having a name for them is what turns "I'll ask a few AIs" into a method. Every platform can operate in all three. Each one leans differently — the way players lean toward groove, harmony, or melody.

  • Exploratory Mode — breadth. Options, variations, angles, "what if." Generating the space of possible answers.
  • Analytical Mode — depth. Weighing tradeoffs, examining structure, reasoning through implications to a coherent argument.
  • Grounded Mode — contact with reality. Facts, sources, documents, external verification. What can actually be confirmed.

The most common gap in a self-assembled ensemble is Grounded Mode. Most people start with generative platforms and discover only later that they have no reliable way to check what those systems are telling them. The fix is not a fourth platform — it is briefing one of your three explicitly for retrieval, and treating that grounding pass as its own step rather than something the session picks up on the way past. If it cannot confirm a claim, treat the claim as unverified. The second most common gap is friction — a member that will push back on your plan when explicitly asked.

Analytical cornerstone

Claude

Anthropic · the thoughtful collaborator

The member that listens to the whole chart before offering a structured, considered response. Strongest when the question is dense or ambiguous, when implications run long-term or ethically layered, or when raw material from the others needs organizing into a coherent argument.

Leans
Analytical first; comfortable in Exploratory when guided
Characteristic failure
Hedging. The instinct to note edge cases can feel like the rhythm section slowing the soloist down — a strength at high stakes, a drag when you need momentum.
Research surface
Sustained work over documents you supply. Large context is now common across frontier models; what varies is fidelity to the material — staying inside what the source actually says instead of drifting toward what it probably says
Pair with
A grounding pass whenever external facts matter
Generative range

ChatGPT

OpenAI · the fluent generalist

The versatile session player — usually the fastest way to get a fluent, well-formed answer onto the page. Lives comfortably in Exploratory Mode: drafting in a specific tone, generating examples and analogies, brainstorming with minimal friction. Give it structure and it moves into Analytical work just as readily.

Leans
Exploratory; Analytical when asked to show its steps
Characteristic failure
Fluency bias. It is good enough at making sentences sound smooth that invented details and shaky reasoning can read as convincing — the polished version of the vending-machine trap.
Research surface
Agentic web investigation — multi-step search returning a cited report, with browser-driven retrieval for pages a plain fetch can't reach (Deep Research / agent mode, as of July 2026)
Pair with
Grounded checks whenever specifics matter more than style
Friction surface

Grok

xAI · the bold improviser

The member most willing to challenge a comfortable assumption. Insight with edge: commentary that doesn't just restate the obvious, unexpected angles, tensions and contradictions surfaced rather than smoothed. It pushes the frame instead of filling it — which is precisely what you want before committing to a plan.

Leans
Exploratory, with a bias toward challenge and contradiction
Characteristic failure
The improvisational instinct can take you farther from the chart than intended. Useful while exploring; less useful when locking a final decision.
Research surface
Live retrieval across web and social data, fanning out parallel searches and reporting what it could not find alongside what it could — the partial-failure report is the useful part (DeepSearch, as of July 2026)
Pair with
Grounded follow-through when moving from ideas to decisions
Capabilities as of July 2026 and changing constantly. Treat this as a starting hypothesis, not a verdict — architecture sets the reachable range, and reputation is not a reliable signal for current capability.
  Claude ChatGPT Grok
Dominant mode Analytical Exploratory Exploratory / challenge
Call first when The question is ambiguous, layered, or needs a structured document You need a first draft, variations, or fast exploration You need a plan stress-tested before you commit
Characteristic failure Over-hedging; slow to commit Fluent fabrication; confident smoothness Drift; edge for its own sake
Long documents Strongest — most faithful to supplied material over long inputs Good; multimodal in one thread Adequate; better at live material than archives
Live / recent information Weakest of the three without retrieval Deep Research reads and cites sources Strongest — live web plus X data
Will it push back? When asked, politely Least likely without explicit instruction Most likely, sometimes unprompted
What it contributes to the fix Depth and structure Breadth and range Friction and reframing

What earns a seat

A member earns its place by doing at least one of three things consistently: seeing the problem from a useful angle the others miss, catching what the others smooth over, or grounding the conversation when the others float.

A system that merely confirms what the others say is not a member of your ensemble. It is an echo. This is the practical test for adding a fourth or fifth platform, and it is also why swapping Grok for a model that shares more training lineage with the other two would weaken the ensemble even if that model scored higher on every benchmark. You are not assembling the three best models. You are assembling three bearings that fail differently.

Audition, don't assume

When you evaluate a new platform, do it with real work rather than demos. Run it through three kinds of session: a question you already know the answer to, a question you don't, and a question where you hold a strong opinion. The third is the most revealing — it is where you find out whether the platform will tell you something you don't want to hear.

And when a result disappoints, diagnose before you iterate. Ask whether the problem is the prompt, the platform, the task, or the routing choice. Prompting can improve access to a system's strengths; it cannot create strengths the system does not have. Architecture sets the reachable range; prompting determines how much of that range you can actually use. If the limitation is architectural, stop polishing the prompt and change the instrument.

Field note

"Let evidence, not familiarity or symbolism, decide who gets a seat."

Audition the platform in the task, not in your memory of its brand.

The field guides

Eight guides, one method. Roughly in the order a reader meets them: how the field got here, how to work across models, then the tools themselves.