Augle doesn't produce a single confidence number. It produces four typed confidence grades — one per evidence node — issued by the Methodologist as hard constraints, propagated downstream, and enforced architecturally. The Synthesizer cannot exceed them. The Pragmatist inherits them. No agent and no user input can override them.
Every unresolved Contrarian objection is preserved verbatim alongside the finding that triggered it. The complete record — grades, objections, resolution conditions — is what you get at the end of every session.
These are not labels a researcher applies to summarize their judgment. They are typed outputs issued by the Methodologist after a four-dimension validity assessment, encoded into the evidence nodes registry, and propagated as constraints to every downstream agent.
The highest confidence grade. Requires replication across independent studies using distinct methodologies — convergence from a single research group or a single methodology does not qualify. The Contrarian is required to issue a formal objection with a resolution condition to trigger a downgrade from Established.
An Unresolved Strong Contrarian objection with a specified resolution condition. The Methodologist reassesses and may downgrade to Probable or Contested if the objection identifies a previously unconsidered validity gap. Downgrade is not automatic — it requires a valid methodological challenge.
The most common ceiling in real-world research sessions. Best available evidence supports the claim, but the evidentiary base has limitations — sample size, methodology constraints, limited replication across contexts, or recency. The Synthesizer may assert moderate confidence but cannot use language implying certainty.
Single-study support with sound methodology. Multiple converging studies with methodological variation but no independent replication. Strong evidence base with an external validity constraint identified by the Methodologist. The majority of evidence nodes in substantive research sessions carry Probable grades.
The Contested grade reflects genuine evidential uncertainty — either because the literature itself is split, the methodology is materially weak, or the Contrarian has raised an objection so strong that it cannot be resolved within the current evidence base. Once Contested, the Synthesizer cannot produce a directional finding on that node.
A Contested finding tells you precisely where the limits of the evidence are. It names the dispute, preserves the competing positions, and surfaces the resolution condition that would move the grade. This is more useful than a confident wrong answer — and more honest about what the evidence actually supports.
The Gap grade is not a failure state — it is a first-class output. It tells you what isn't knowable from the current evidence base. A finding that the evidence doesn't exist is valuable information. The Pragmatist is architecturally forbidden from converting a Gap-graded finding into any directional recommendation. Instead, it produces a follow-on session proposal.
Named knowledge gap entered into the session Ledger with the specific evidence that would be required to resolve it. A structured follow-on session proposal from the Pragmatist specifying the research question that would fill the gap. The gap record is exported as part of the full session audit trail.
The propagation chain — Methodologist → Synthesizer → Pragmatist — is enforced at the orchestration layer, not through prompt instruction. The Topic Architect's round transition logic detects Grade Challenge violations and halts dispatch until the Synthesizer revises. There is no mechanism for any agent or user to override this chain.
Evaluates each evidence node across four dimensions and issues a grade — Established, Probable, Contested, or Gap — as a hard upper constraint. Also issues [GRADE CHALLENGE] flags when the Synthesizer violates the constraint.
For each claim k supported by evidence nodes Eₖ, the Synthesizer's grade cannot exceed the minimum Methodologist grade across all supporting nodes. Operates at T=0.0 — deterministic. Any violation triggers a mandatory [GRADE CHALLENGE] revision loop.
The Pragmatist's recommendation confidence cannot exceed the Synthesizer's conclusion grade. Gap-graded findings cannot become directional recommendations under any circumstances. Fires in Phase 3 only — after Guardian clearance.
These are hard constraints, not guidelines or prompt instructions. They are enforced through the session orchestration layer — the Topic Architect's round transition logic detects open [GRADE_CHALLENGE] flags and halts dispatch until the Synthesizer produces a compliant revision.
The Synthesizer operates at T=0.0 and anchors its verdict exclusively to the evidence nodes registry — not the discourse thread. This produces a deterministic relationship between evidence and finding: the same question with the same evidence base must always produce the same confidence verdict. This invariance is essential for the calibration corpus — it ensures each confidence grade represents a deterministic function of the evidence, enabling meaningful comparison against ground truth resolution outcomes.
If the Synthesizer produces a preliminary claim graded above the minimum confidence bound of its supporting evidence nodes, the Methodologist issues a [GRADE CHALLENGE] flag. This is not an advisory note — it triggers a mandatory revision loop enforced at the orchestration layer.
The Topic Architect detects the open [GRADE CHALLENGE] flag and halts dispatch of all subsequent agents — including the Contrarian, Guardian, and Pragmatist — until the Synthesizer produces a revised output in which all claims satisfy the propagation constraint.
Both the challenge and the revision are written to the session audit trail verbatim. The loop repeats until no Grade Challenge violations remain, or a maximum iteration count is reached — at which point remaining violations are surfaced to the user as blocking items.
Example trigger
Methodologist grades a research paper's construct validity as Probable. Synthesizer produces a preliminary claim labeled Established using that paper as the sole supporting evidence node. Constraint 1 is violated. [GRADE CHALLENGE] fires. Synthesizer must revise the claim to Probable before any agent dispatches.
The Contrarian issues typed objections — not general criticism. Each objection carries a strength grade and a resolution condition. The strength grade determines what happens to the objection when Phase 3 closes: Strong unresolved objections surface verbatim in the final output. The resolution condition tells you exactly what evidence would change the finding.
The most consequential objection grade. Issued when the Contrarian identifies a fundamental methodological flaw, a construct validity problem, or an evidentiary gap that materially undermines a claim. A Strong objection that remains unresolved at Phase 3 forces the Synthesizer to reflect it in the confidence grade — the affected node cannot stay at Established or Probable if a Strong objection is unresolved.
Issued when the Contrarian identifies a real concern — a replication limitation, an external validity constraint, or a methodology-claim mismatch — that is material but does not fundamentally invalidate the claim. Moderate objections that are unresolved at Phase 3 surface in the final output alongside the finding, with their resolution conditions intact.
Issued when the Contrarian identifies a concern that is plausible but not directly supported by evidence in the current session. Speculative objections are logged and included in the session audit trail, but do not affect confidence grades and do not surface in the main finding output.
"The assignment of 'weight regain follows discontinuation' to Settled Ground overstates the case: while regain after discontinuation is well-documented, its completeness varies between individuals, and the ≤2-year evidence cannot establish whether structured tapering or lifestyle transition sustains any of the loss."
"The lifestyle-alone maintenance evidence rests on a single trial paradigm across all studies — convergence without independent methodological variation."
"Follow-up gap — all primary evidence nodes are limited to ≤104 weeks. Long-term (>2yr) off-drug maintenance cannot be inferred from current trials."
Long-term off-drug maintenance claim graded Probable by Synthesizer — no controlled evidence beyond ~2 years. Claim split required.
The following is an illustrative Standard-depth Letters & Science session — the same confidence and dissent record you receive at the end of every session.
The session demonstrates all five architectural properties of the confidence system: Strong objection successfully amended a terrain classification; evidence ceiling propagated correctly across all three phases; Grade Challenge mechanism fired and resolved correctly; Pragmatist declined to produce a directional recommendation for the Gap-graded finding; and the Unresolved Moderate objection surfaced verbatim in final delivery.
Research question: Does the current evidence base support long-term weight maintenance without continued GLP-1 dosing?
Depth: Standard · Mode: Letters & Science · Guardian: Active
Illustrative session · representative of the record structure
SVS mechanics, flag taxonomy, halt authority, domain integrity modes, and the reasoning behind hidden model identity.
Read more →How the three phases constrain each other — what carries forward, what gets locked, and how the Grade Challenge loop works across rounds.
Read more →Full specifications for each agent — including the Methodologist's four validity dimensions and the Synthesizer's three inviolable constraints.
Read more →Join waitlist and run a session — every grade enforced,
every objection logged, every finding auditable.