The Guardian is not a research agent. It doesn't produce findings, participate in deliberation, or recommend conclusions. It operates exclusively at phase boundaries — authenticating sources, classifying integrity events, and holding permanent halt authority over every session it monitors.
Its model identity is hidden from all user-facing surfaces, including the other agents. This is not a privacy feature. It is an architectural decision to prevent anchoring effects — the well-documented tendency of agents and users to anchor reasoning toward known model capabilities.
The Guardian operates outside the research deliberation loop entirely. It evaluates the integrity of what the ensemble produces — it does not contribute to producing it. This separation is architectural, not procedural.
"An integrity agent that participates in deliberation is not an integrity agent. It is a participant with audit responsibilities — a fundamentally different and weaker guarantee."
Running at temperature 0.1, the Guardian is near-deterministic — the same input produces the same integrity classification across sessions. This is essential for calibration: if the Guardian's behavior were stochastic, the flag record would be unreliable as an audit artifact.
The Guardian classifies every integrity event into one of three severity tiers. The tier determines what happens next — whether deliberation is blocked, surfaced for acknowledgment, or logged silently. There is no discretion in the consequence: the tier determines the response deterministically.
Deliberation cannot proceed. The Topic Architect halts dispatch immediately. The user must resolve the condition before the next round can fire. Critical flags indicate a fundamental integrity failure — not a data quality issue.
Session halted. User notified. If a Hard Block terminates the session permanently, full credit refund is issued. The flag is written to the session audit trail with full resolution detail.
Surfaces as a banner before the next round fires. User acknowledgment is required but does not block dispatch — the user can proceed after reviewing the flag. The condition is recorded in the session audit trail regardless of the user's choice.
User acknowledgment required before next round. Dispatch is not blocked. Full flag record written to audit trail with evidence node ID, verification source, and resolution detail.
Logged in the session flag registry and included in the session summary output. Does not surface interactively during the session. Informational flags represent conditions the Guardian has noted but assessed as non-material to the deliberation's integrity.
Logged silently. Included in final session summary. Available in the exportable audit trail. No user acknowledgment required.
A Guardian Hard Block terminates the session permanently when the Guardian determines the deliberation cannot be certified under any conditions. This typically occurs due to compound integrity violations or a fundamental research question framing problem that cannot be corrected mid-session. Hard blocks trigger a full refund of session credits. The condition is logged with full detail in the session audit trail.
The SVS runs as a background service between the evidence extraction layer and the Methodologist agent dispatch. It intercepts citation hallucinations — the most damaging failure mode in AI reasoning systems — before they enter the evidence base and propagate downstream.
The SVS detects the citation type — URL, DOI, arXiv identifier, or ISBN — and routes to the appropriate protocol handler. Each type has a distinct resolution pathway and content match verification method.
The SVS resolves the identifier against the live resource. A resource that doesn't exist at the claimed location returns SVS_NOT_FOUND. A resource that exists but has been retracted is flagged accordingly.
The SVS compares the agent's citation claim against the metadata returned by the resolved resource — title, author, publication year. A resource that exists but doesn't match the claim receives verification_status: unverified, not verified. Existence alone is not sufficient.
Based on the verification outcome, the SVS applies tiered confidence downgrade rules to the evidence node and raises a typed integrity flag into the session flag registry. The evidence node is preserved — not removed — with its verification status recorded.
The verification_status, verification_source, verification_timestamp, and verification_url are written to the evidence node record. The Methodologist receives the updated registry — never the pre-verification state.
| SVS outcome | Prior grade | Downgrade applied | Flag raised |
|---|---|---|---|
| Verified | Any | None | — |
| Unverified | Established | Cap to Probable | SVS_UNVERIFIED · Informational |
| Unverified | Probable / Contested / Gap | No change | SVS_UNVERIFIED · Informational |
| Not found | Any | Hard downgrade to Contested | SVS_NOT_FOUND · Moderate |
| Pending | Established | No change at session time | SVS_UNVERIFIED · Informational |
Removing an evidence node that fails verification would silently improve the apparent quality of the session output while concealing a potential hallucination. The SVS preserves every node, downgrades its confidence bound, and records the failure in the flag registry. This is a transparency obligation — not a data retention policy.
The Guardian's core behaviour — SVS authentication, flag taxonomy, halt authority — is constant across all modes. What changes is the domain-specific integrity ruleset applied on top of that foundation.
Applied in Letters & Science sessions for research, dissertation, and grant work. Enforces academic source standards and retraction database checks.
Applied in Letters & Science sessions for case analysis, expert evidence review, and regulatory applicability questions.
Applied in Letters & Science sessions for drug interactions, clinical trial design, and healthcare coverage decisions.
Applied in sessions for VC, PE, financial services, and enterprise strategy work.
Applied in Letters & Science sessions for journalism, media analysis, and science reporting work.
The Guardian writes a structured record of every integrity decision to the session audit trail — not a summary, a full provenance record. Every SVS verification outcome, every flag raised, every confidence downgrade applied, every halt decision made.
The audit trail is exportable in full at session close. For regulated industries — clinical, financial, legal — this record is the accountability artifact that demonstrates due diligence on the evidence base used to support a decision.
What the audit trail records
Pashler et al. (2008) · DOI resolved · content match PASS · confidence_bound: Established → unchanged
Martinez et al. (2024) · arXiv:2403.12847 · resolution_result: 404 · confidence_bound: Probable → Contested · flag written to registry
Chen & Liu (2023) · conference proceedings · content match FAIL (title mismatch) · confidence_bound: Established → capped at Probable
Phase 1 integrity evaluation complete · 2 flags raised · no Critical conditions · Phase 2 dispatch authorised
Phase 2 integrity evaluation complete · 0 new flags · no Critical conditions · Phase 3 dispatch authorised
Full audit trail available for export · 3 flags total · 0 Critical · 1 Moderate · 2 Informational · session certified
The full multi-agent ensemble explained — dispatch order, phase architecture, output contracts, and session modes.
Read more →How the three phases constrain each other — what carries forward, what gets locked, and how the structured output protocol works.
Read more →The four confidence grades, how the Methodologist issues them as hard constraints, and how unresolved Contrarian objections are preserved in the final output.
Read more →Join waitlist and run a session — every source authenticated,
every flag logged, every decision auditable.