The seats ← Home

Look at them the way you'd look at players in a band.

Not model sizes, not benchmark scores — those change quarterly and mean less in practice than you'd expect. What matters is what each one tends to do well, where it wobbles, how it thinks, and when to call on it. These are tendencies, not cages.

The three cognitive modes

Useful AI work falls into three modes, and having a name for them is what turns "I'll ask a few AIs" into a method. Every platform can operate in all three. Each one leans differently — the way players lean toward groove, harmony, or melody.

  • Exploratory Mode — breadth. Options, variations, angles, "what if." Generating the space of possible answers.
  • Analytical Mode — depth. Weighing tradeoffs, examining structure, reasoning through implications to a coherent argument.
  • Grounded Mode — contact with reality. Facts, sources, documents, external verification. What can actually be confirmed.

The most common gap in a self-assembled council is Grounded Mode. Most people start with generative platforms and discover only later that they have no reliable way to check what those systems are telling them. The second most common gap is friction — a member that will push back on your plan when explicitly asked.

Analytical cornerstone

Claude

Anthropic · the thoughtful collaborator

The member that listens to the whole chart before offering a structured, considered response. Strongest when the question is dense or ambiguous, when implications run long-term or ethically layered, or when raw material from the others needs organizing into a coherent argument.

Leans
Analytical first; comfortable in Exploratory when guided
Characteristic failure
Hedging. The instinct to note edge cases can feel like the rhythm section slowing the soloist down — a strength at high stakes, a drag when you need momentum.
Research surface
Long-context document work (1M-token window on current frontier models as of July 2026); strong faithfulness to supplied material
Pair with
A Grounded partner whenever external facts matter
Generative range

ChatGPT

OpenAI · the fluent generalist

The versatile session player — usually the fastest way to get a fluent, well-formed answer onto the page. Lives comfortably in Exploratory Mode: drafting in a specific tone, generating examples and analogies, brainstorming with minimal friction. Give it structure and it moves into Analytical work just as readily.

Leans
Exploratory; Analytical when asked to show its steps
Characteristic failure
Fluency bias. It is good enough at making sentences sound smooth that invented details and shaky reasoning can read as convincing — the polished version of the vending-machine trap.
Research surface
Deep Research runs multi-step agentic web investigation and returns a cited report; agent mode adds a visual browser (as of July 2026)
Pair with
Grounded checks whenever specifics matter more than style
Friction surface

Grok

xAI · the bold improviser

The member most willing to challenge a comfortable assumption. Insight with edge: commentary that doesn't just restate the obvious, unexpected angles, tensions and contradictions surfaced rather than smoothed. It pushes the frame instead of filling it — which is precisely what you want before committing to a plan.

Leans
Exploratory, with a bias toward challenge and contradiction
Characteristic failure
The improvisational instinct can take you farther from the chart than intended. Useful while exploring; less useful when locking a final decision.
Research surface
DeepSearch and a deep-research mode that fans out parallel agents across web and live X data, returning citations plus what it could not find (as of July 2026)
Pair with
Grounded follow-through when moving from ideas to decisions
The fourth seat

Perplexity

Optional · the grounded researcher

Not required for the floor of three — but this is the seat that earns its place fastest. Where the others start by generating prose, Perplexity starts by looking outward: retrieving, summarizing, and synthesizing from external sources with an evidentiary discipline the generative systems rarely match.

Leans
Grounded first, then Analytical
Characteristic failure
Under-explores creative options because it stays close to sources; may surface material that reflects the limits of what is indexed. Your judgment about source quality still matters.
Why add it
It is the member that anchors deliberation to the external record — and the one the Verification Protocol routes to
Rule of thumb
If it cannot confirm a claim, treat the claim as unverified.
Capabilities as of July 2026 and changing constantly. Treat this as a starting hypothesis, not a verdict — architecture sets the reachable range, and reputation is not a reliable signal for current capability.
  Claude ChatGPT Grok
Dominant mode Analytical Exploratory Exploratory / challenge
Call first when The question is ambiguous, layered, or needs a structured document You need a first draft, variations, or fast exploration You need a plan stress-tested before you commit
Characteristic failure Over-hedging; slow to commit Fluent fabrication; confident smoothness Drift; edge for its own sake
Long documents Strongest — very large context, faithful to supplied material Good; multimodal in one thread Adequate; better at live material than archives
Live / recent information Weakest of the three without retrieval Deep Research reads and cites sources Strongest — live web plus X data
Will it push back? When asked, politely Least likely without explicit instruction Most likely, sometimes unprompted
What it contributes to the fix Depth and structure Breadth and range Friction and reframing

What earns a seat

A member earns its place by doing at least one of three things consistently: seeing the problem from a useful angle the others miss, catching what the others smooth over, or grounding the conversation when the others float.

A system that merely confirms what the others say is not a council member. It is an echo. This is the practical test for adding a fourth or fifth platform, and it is also why swapping Grok for a model that shares more training lineage with the other two would weaken the ensemble even if that model scored higher on every benchmark. You are not assembling the three best models. You are assembling three bearings that fail differently.

Audition, don't assume

When you evaluate a new platform, do it with real work rather than demos. Run it through three kinds of session: a question you already know the answer to, a question you don't, and a question where you hold a strong opinion. The third is the most revealing — it is where you find out whether the platform will tell you something you don't want to hear.

And when a result disappoints, diagnose before you iterate. Ask whether the problem is the prompt, the platform, the task, or the routing choice. Prompting can improve access to a system's strengths; it cannot create strengths the system does not have. Architecture sets the reachable range; prompting determines how much of that range you can actually use. If the limitation is architectural, stop polishing the prompt and change the instrument.

Field note

"Let evidence, not familiarity or symbolism, decide who gets a seat."

Audition the platform in the task, not in your memory of its brand.