The composition is not arbitrary. The ensemble covers the complete deliberative arc — mapping, validity, integrity, adversarial pressure, synthesis, and application. Remove any one role and the process has a structural gap. Every agent operates under a typed output contract defining exactly what it can produce, what it must produce, and what it is architecturally forbidden from doing.
| # | Agent | Role | Decoding | Active phases | Primary output |
|---|---|---|---|---|---|
| 01 | Topic Architect | Session orchestration | Near-deterministic | Init · transitions · delivery | Session params · final delivery |
| 02 | Cartographer | Evidence mapping | Exploratory | Phases 1 · 2 · 3 | Five-component terrain map |
| 03 | Methodologist | Validity assessment | Balanced | Phases 1 · 2 · 3 | Per-node confidence bounds |
| 04 | Guardian | Integrity layer | Near-deterministic | Phase boundaries only | Flag classification · halt auth |
| 05 | Contrarian | Adversarial pressure | Maximum variation | Phases 1 · 2 · 3 | Steelmanned objections |
| 06 | Synthesizer | Integration & finding | Deterministic · locked | All phases | Evidence-anchored finding |
| 07 | Pragmatist | Application notes | Low variation | Phase 3 only | Actionable recommendations |
Decoding behavior reflects the functional requirement of each role — adversarial pressure runs with maximum variation; deterministic finding production runs locked. Dispatch order is fixed within every phase and cannot be altered by configuration or instruction.
Each specification includes the agent's role, its decoding rationale, what the agent must produce, and what it is architecturally forbidden from doing. Output contracts are not guidelines — they are hard constraints enforced at the system level.
The Guardian is architecturally isolated from the research deliberation loop. It evaluates integrity — it does not contribute to findings. Its model identity is withheld from all user-facing surfaces, including the other agents, to prevent anchoring effects on research reasoning. It runs near-deterministically to produce consistent integrity classifications: the same integrity event must produce the same flag classification across sessions for the audit trail to be reliable.
Near-deterministic decoding produces consistent output. Integrity classification must be repeatable — a given citation verification outcome must always produce the same flag classification. Stochastic integrity behavior would undermine the reliability of the audit trail as an accountability artifact.
The only agent with a direct user-facing interface. Fires once at session initialization to parse the research question, set depth tier, configure Guardian integrity mode, and queue the first dispatch. Manages all phase transitions and surfaces the final delivery package. Does not participate in research discourse between initialization and final delivery.
Low-variation decoding allows modest flexibility in how session parameters are framed and communicated to the user, while keeping the orchestration logic predictable. The Topic Architect's job is routing — not reasoning — so near-deterministic behavior is appropriate.
Dispatched first in every phase — every deliberation round begins with the Cartographer mapping or updating the evidence terrain. Produces a five-component landscape output that becomes the structured foundation every subsequent agent builds from. Exploratory decoding ensures broad, creative evidence retrieval and reduces the risk that a narrow initial framing constrains the entire deliberation.
Exploratory decoding maximises evidence retrieval breadth. A Cartographer that consistently produces the same narrow landscape would systematically miss relevant evidence at the edges of the research question. High variation increases the probability of surfacing non-obvious evidence nodes and cross-domain connections.
Issues formal confidence bounds on every evidence node across four validity dimensions. These bounds propagate as hard constraints on the Synthesizer — the Synthesizer cannot produce a conclusion at higher confidence than the minimum bound across its supporting evidence nodes. Also issues [GRADE CHALLENGE] flags when the Synthesizer's preliminary conclusion exceeds its evidentiary warrant, triggering a mandatory revision loop.
Balanced decoding trades off methodological consistency against sensitivity to domain-specific validity considerations. Too rigid and the Methodologist applies a fixed template regardless of domain context. Too loose and confidence bounds become inconsistent across sessions, undermining their reliability as constraints.
Maximum-variation decoding to maximize output diversity and reduce sycophantic convergence. The Contrarian is required to steelman every claim before challenging it — adversarial, not contrarian for its own sake. Each objection must specify a resolution condition and a strength grade. Unresolved Strong objections at Phase 3 surface verbatim in the final output — not summarized, not softened.
Maximum-variation decoding diversifies objections across sessions. If the Contrarian produced the same objections every time, it would be exploitable — researchers would learn to expect and pre-answer those objections. High variation ensures the adversarial pressure is genuinely unpredictable. In Deep sessions, the strongest available model is used for the highest-quality steelmanning.
Decoding is locked to deterministic — a hard architectural requirement, not a configuration choice. The same evidence base must produce the same finding across sessions: this is the Finding Invariance Requirement. The Synthesizer anchors exclusively to the structured evidence nodes registry — not the discourse thread — to prevent reasoning contamination from the deliberation history. Subject to three inviolable constraints that no other agent or instruction can override.
Deterministic decoding is a hard requirement for the Finding Invariance property. If the Synthesizer produced stochastic findings, the corpus calibration pipeline would be invalid — the same session with the same evidence could produce different findings, undermining the reproducibility the reasoning corpus depends on.
Fires in Phase 3 only — the final agent in the dispatch sequence. Converts the Synthesizer's finding into context-specific actionable output. Inherits the Synthesizer's confidence ceiling as an absolute constraint: it cannot produce recommendations more confident than the synthesis permits. Gap-graded findings cannot be converted to directional recommendations — the Pragmatist must surface the gap and propose follow-on sessions for the unresolved question.
Low-variation decoding lets the Pragmatist tailor actionable output to the user's specific context with modest flexibility, while remaining close enough to deterministic that the recommendations are reliably bounded by the Synthesizer's finding. Higher variation would risk recommendations drifting outside the confidence ceiling.
The ensemble deliberately runs on four independent frontier models rather than a single model wearing different hats. No one model's training distribution, capability profile, or systematic biases get to set the terms of the deliberation. When the models produce incompatible outputs, the disagreement is structurally preserved — surfaced and recorded, not averaged away into a false consensus.
Each model is selected strategically — matched to the kind of reasoning it does best, so the ensemble draws on the distinct strengths of four frontier systems rather than the limits of any one.
Each model is trained on a different data distribution, so their blind spots don't line up. Where one is systematically weak, another is likely to be strong — and the ensemble sees both.
The four models bring genuinely different reasoning profiles. The composition is chosen so coverage is broad — not one vendor's view repeated four times over.
When independent models fail, they tend to fail differently. Correlated error is far less likely than when a single model is left to check its own work.
Conflicting outputs are treated as signal. The architecture records the disagreement and carries it through to the finding rather than smoothing it into a single confident answer.
"The composition is not arbitrary. The ensemble covers the complete deliberative arc. Remove any one role and the process has a structural gap."
SVS mechanics, flag taxonomy, halt authority, domain-specific integrity modes, and the reasoning behind hidden model identity.
Read more →How the three phases constrain each other — what carries forward, what gets locked, and how the constraint propagation chain works.
Read more →The four confidence grades, the Grade Challenge mechanism, how Contrarian objections are classified by strength, and what the Ledger records.
Read more →Join waitlist and run a session — every agent dispatched in sequence,
every constraint enforced, every finding auditable.