What earns a seat
A member earns its place by doing at least one of three things consistently: seeing the problem from
a useful angle the others miss, catching what the others smooth over, or grounding the conversation
when the others float.
A system that merely confirms what the others say is not a council member. It is an echo.
This is the practical test for adding a fourth or fifth platform, and it is also why swapping Grok for
a model that shares more training lineage with the other two would weaken the ensemble even if that
model scored higher on every benchmark. You are not assembling the three best models. You are
assembling three bearings that fail differently.
Audition, don't assume
When you evaluate a new platform, do it with real work rather than demos. Run it through three kinds
of session: a question you already know the answer to, a question you don't, and a question where you
hold a strong opinion. The third is the most revealing — it is where you find out whether the platform
will tell you something you don't want to hear.
And when a result disappoints, diagnose before you iterate. Ask whether the problem is the prompt, the
platform, the task, or the routing choice. Prompting can improve access to a system's strengths; it
cannot create strengths the system does not have. Architecture sets the reachable range;
prompting determines how much of that range you can actually use. If the limitation is
architectural, stop polishing the prompt and change the instrument.
Field note
"Let evidence, not familiarity or symbolism, decide who gets a seat."
Audition the platform in the task, not in your memory of its brand.