Confidence + dissent scoring

Not a score.
A structured
evidence record.

Augle doesn't produce a single confidence number. It produces four typed confidence grades — one per evidence node — issued by the Methodologist as hard constraints, propagated downstream, and enforced architecturally. The Synthesizer cannot exceed them. The Pragmatist inherits them. No agent and no user input can override them.

Every unresolved Contrarian objection is preserved verbatim alongside the finding that triggered it. The complete record — grades, objections, resolution conditions — is what you get at the end of every session.

Four confidence grades
EstablishedMultiple independent replications. Contrarian challenge required to downgrade.
ProbableBest available evidence supports. Replication is limited.
ContestedActive dispute or unresolved Strong objection. Cannot grade higher.
GapEvidence insufficient. Named explicitly. Cannot become a recommendation.
Propagation constraint · formal
C_S(k) ≤ min{ C_M(eᵢ) : eᵢ ∈ Eₖ }
C_P ≤ C_S(k*)
The four grades

Each grade has a precise
definition and a hard consequence.

These are not labels a researcher applies to summarize their judgment. They are typed outputs issued by the Methodologist after a four-dimension validity assessment, encoded into the evidence nodes registry, and propagated as constraints to every downstream agent.

EstablishedGrade 1 of 4 · highest

Multiple independent replications at Moderate or better across distinct methodological contexts.

The highest confidence grade. Requires replication across independent studies using distinct methodologies — convergence from a single research group or a single methodology does not qualify. The Contrarian is required to issue a formal objection with a resolution condition to trigger a downgrade from Established.

What it requires for downgrade

An Unresolved Strong Contrarian objection with a specified resolution condition. The Methodologist reassesses and may downgrade to Probable or Contested if the objection identifies a previously unconsidered validity gap. Downgrade is not automatic — it requires a valid methodological challenge.

Synthesizer may use unqualified assertive claim language for Established-grade nodes
ProbableGrade 2 of 4

At least one Moderate-confidence study, or multiple Weak-confidence studies converging. Replication is limited.

The most common ceiling in real-world research sessions. Best available evidence supports the claim, but the evidentiary base has limitations — sample size, methodology constraints, limited replication across contexts, or recency. The Synthesizer may assert moderate confidence but cannot use language implying certainty.

Common causes of Probable assignment

Single-study support with sound methodology. Multiple converging studies with methodological variation but no independent replication. Strong evidence base with an external validity constraint identified by the Methodologist. The majority of evidence nodes in substantive research sessions carry Probable grades.

Synthesizer must include hedging language — cannot assert with certainty on Probable-grade nodes
ContestedGrade 3 of 4

Active dispute exists, methods are weak, or an Unresolved Strong Contrarian objection applies. Cannot grade higher regardless of discourse content.

The Contested grade reflects genuine evidential uncertainty — either because the literature itself is split, the methodology is materially weak, or the Contrarian has raised an objection so strong that it cannot be resolved within the current evidence base. Once Contested, the Synthesizer cannot produce a directional finding on that node.

Why Contested is a complete finding, not a failure

A Contested finding tells you precisely where the limits of the evidence are. It names the dispute, preserves the competing positions, and surfaces the resolution condition that would move the grade. This is more useful than a confident wrong answer — and more honest about what the evidence actually supports.

Synthesizer must present alternative positions — cannot produce a directional claim on Contested nodes
GapGrade 4 of 4 · Insufficient Evidence

Evidence is insufficient to evaluate the claim. Named explicitly. Cannot be converted to a directional recommendation by any downstream agent.

The Gap grade is not a failure state — it is a first-class output. It tells you what isn't knowable from the current evidence base. A finding that the evidence doesn't exist is valuable information. The Pragmatist is architecturally forbidden from converting a Gap-graded finding into any directional recommendation. Instead, it produces a follow-on session proposal.

What a Gap finding produces

Named knowledge gap entered into the session Ledger with the specific evidence that would be required to resolve it. A structured follow-on session proposal from the Pragmatist specifying the research question that would fill the gap. The gap record is exported as part of the full session audit trail.

Pragmatist is forbidden from producing any directional recommendation — must generate a follow-on session proposal
Unidirectional confidence propagation

Confidence flows one direction.
No exceptions.

The propagation chain — Methodologist → Synthesizer → Pragmatist — is enforced at the orchestration layer, not through prompt instruction. The Topic Architect's round transition logic detects Grade Challenge violations and halts dispatch until the Synthesizer revises. There is no mechanism for any agent or user to override this chain.

Methodologist
Issues confidence bounds

Evaluates each evidence node across four dimensions and issues a grade — Established, Probable, Contested, or Gap — as a hard upper constraint. Also issues [GRADE CHALLENGE] flags when the Synthesizer violates the constraint.

Constraint issued
C_M(eᵢ) = grade for node eᵢ · Hard ceiling for all downstream agents
Synthesizer
Inherits as hard ceiling

For each claim k supported by evidence nodes Eₖ, the Synthesizer's grade cannot exceed the minimum Methodologist grade across all supporting nodes. Operates at T=0.0 — deterministic. Any violation triggers a mandatory [GRADE CHALLENGE] revision loop.

Constraint enforced
C_S(k) ≤ min{ C_M(eᵢ) : eᵢ ∈ Eₖ } · Revision loop fires on violation
Pragmatist
Inherits Synthesizer ceiling

The Pragmatist's recommendation confidence cannot exceed the Synthesizer's conclusion grade. Gap-graded findings cannot become directional recommendations under any circumstances. Fires in Phase 3 only — after Guardian clearance.

Constraint enforced
C_P ≤ C_S(k*) · Gap → follow-on session, not recommendation
Formal propagation constraints · from AUGLE-001P patent application
Let C_M(eᵢ) denote the Methodologist's confidence bound for evidence node eᵢ
Let C_S(k) denote the Synthesizer's grade for claim k supported by node set Eₖ
Let C_P denote the Pragmatist's recommendation confidence, where k* is the primary claim
Constraint 1: C_S(k) ≤ min{ C_M(eᵢ) : eᵢ ∈ Eₖ }
Constraint 2: C_P ≤ C_S(k*)

These are hard constraints, not guidelines or prompt instructions. They are enforced through the session orchestration layer — the Topic Architect's round transition logic detects open [GRADE_CHALLENGE] flags and halts dispatch until the Synthesizer produces a compliant revision.

The Verdict Invariance Requirement

The Synthesizer operates at T=0.0 and anchors its verdict exclusively to the evidence nodes registry — not the discourse thread. This produces a deterministic relationship between evidence and finding: the same question with the same evidence base must always produce the same confidence verdict. This invariance is essential for the calibration corpus — it ensures each confidence grade represents a deterministic function of the evidence, enabling meaningful comparison against ground truth resolution outcomes.

T = 0.0 · locked
Grade Challenge mechanism

When the Synthesizer
overreaches.

If the Synthesizer produces a preliminary claim graded above the minimum confidence bound of its supporting evidence nodes, the Methodologist issues a [GRADE CHALLENGE] flag. This is not an advisory note — it triggers a mandatory revision loop enforced at the orchestration layer.

The Topic Architect detects the open [GRADE CHALLENGE] flag and halts dispatch of all subsequent agents — including the Contrarian, Guardian, and Pragmatist — until the Synthesizer produces a revised output in which all claims satisfy the propagation constraint.

Both the challenge and the revision are written to the session audit trail verbatim. The loop repeats until no Grade Challenge violations remain, or a maximum iteration count is reached — at which point remaining violations are surfaced to the user as blocking items.

Example trigger

Methodologist grades a research paper's construct validity as Probable. Synthesizer produces a preliminary claim labeled Established using that paper as the sole supporting evidence node. Constraint 1 is violated. [GRADE CHALLENGE] fires. Synthesizer must revise the claim to Probable before any agent dispatches.

Grade Challenge · step by step
01Synthesizer produces a preliminary evidence landscape with claim grades in Phase 1 or Phase 2
02Methodologist evaluates each claim grade against the minimum confidence bound of its supporting evidence nodes
03Violation detected — Synthesizer has graded a claim above the Methodologist's ceiling.[GRADE CHALLENGE] flag issued at Moderate severity
04Topic Architect halts dispatch of Contrarian, Guardian, and all subsequent agents. Session cannot advance.
05Synthesizer produces a revised output with the claim grade corrected to comply with Constraint 1
06Methodologist re-evaluates. If compliant:Resolved Dispatch resumes. Both challenge and revision written to audit trail.
07If max iterations reached without compliance: remaining violations surfaced to user asBlocking items
Dissent scoring

Every objection classified.
Every unresolved objection preserved.

The Contrarian issues typed objections — not general criticism. Each objection carries a strength grade and a resolution condition. The strength grade determines what happens to the objection when Phase 3 closes: Strong unresolved objections surface verbatim in the final output. The resolution condition tells you exactly what evidence would change the finding.

Strong objection
Strong

The most consequential objection grade. Issued when the Contrarian identifies a fundamental methodological flaw, a construct validity problem, or an evidentiary gap that materially undermines a claim. A Strong objection that remains unresolved at Phase 3 forces the Synthesizer to reflect it in the confidence grade — the affected node cannot stay at Established or Probable if a Strong objection is unresolved.

Unresolved at Phase 3 → surfaces verbatim in final output · affected node grade reviewed · may trigger Contested assignment
Moderate objection
Moderate

Issued when the Contrarian identifies a real concern — a replication limitation, an external validity constraint, or a methodology-claim mismatch — that is material but does not fundamentally invalidate the claim. Moderate objections that are unresolved at Phase 3 surface in the final output alongside the finding, with their resolution conditions intact.

Unresolved at Phase 3 → surfaces in output alongside finding · resolution condition preserved · does not force grade revision
Speculative objection
Speculative

Issued when the Contrarian identifies a concern that is plausible but not directly supported by evidence in the current session. Speculative objections are logged and included in the session audit trail, but do not affect confidence grades and do not surface in the main finding output.

Logged in session audit trail · does not surface in main output · does not affect confidence grades
Dissent register · example session
@Contrarian → @Cartographer · Phase 1Strong · Unresolved

"The assignment of 'weight regain follows discontinuation' to Settled Ground overstates the case: while regain after discontinuation is well-documented, its completeness varies between individuals, and the ≤2-year evidence cannot establish whether structured tapering or lifestyle transition sustains any of the loss."

Resolution: Move to Contested Terrain · reframe magnitude/durability claim
@Contrarian → @Methodologist · Phase 2Moderate · Resolved

"The lifestyle-alone maintenance evidence rests on a single trial paradigm across all studies — convergence without independent methodological variation."

Resolved: Replication Cap applied · grade adjusted to Probable
@Contrarian → @Synthesizer · Phase 3Moderate · Unresolved

"Follow-up gap — all primary evidence nodes are limited to ≤104 weeks. Long-term (>2yr) off-drug maintenance cannot be inferred from current trials."

Resolution: Independent ≥3-year RCT with a structured tapering arm
@Methodologist · Phase 3[GRADE CHALLENGE]

Long-term off-drug maintenance claim graded Probable by Synthesizer — no controlled evidence beyond ~2 years. Claim split required.

Outcome: Short-term regain (≤1yr) → Probable · Long-term (>2yr) maintenance → Gap
Worked example

Confidence propagation confirmed
across all five checkpoints.

The following is an illustrative Standard-depth Letters & Science session — the same confidence and dissent record you receive at the end of every session.

The session demonstrates all five architectural properties of the confidence system: Strong objection successfully amended a terrain classification; evidence ceiling propagated correctly across all three phases; Grade Challenge mechanism fired and resolved correctly; Pragmatist declined to produce a directional recommendation for the Gap-graded finding; and the Unresolved Moderate objection surfaced verbatim in final delivery.

Research question: Does the current evidence base support long-term weight maintenance without continued GLP-1 dosing?

Depth: Standard · Mode: Letters & Science · Guardian: Active
Illustrative session · representative of the record structure

See the research →
Confidence + dissent record
Phase 1 · Exploration
MethodologistSession ceiling set: no node may be graded Established. All evidence Moderate-confidence or below.
Contrarian"Weight regain follows discontinuation" → Settled Ground challenged.Strong Regain not universal · ≤2yr data.
SynthesizerRevised terrain: "weight regain follows discontinuation" moved to Contested Terrain.Resolved
Phase 2 · Deliberation
MethodologistReplication Cap on lifestyle-alone maintenance — convergence without methodological independence.
SynthesizerEvidence landscape: partial regain within 1yr (Probable) · universal full regain (Contested) · lifestyle-alone maintenance (Contested) · maintenance dosing preserves loss (Probable) · tapering mitigates regain (Contested)
ContrarianDurability of off-drug maintenance beyond 2 years.Moderate Carries to Phase 3.
Phase 3 · Synthesis
ContrarianFollow-up gap — all nodes limited to ≤104 weeks.Moderate · Unresolved
MethodologistLong-term maintenance claim graded above evidence ceiling.[GRADE CHALLENGE]
SynthesizerSplit: Short-term regain (≤1yr) → Probable · Long-term (>2yr) maintenance → Gap (no controlled evidence beyond ~2 years).Compliant
PragmatistDirectional recommendation declined for Gap-graded finding. Follow-on session proposed.

Calibrated confidence.
Preserved dissent.

Join waitlist and run a session — every grade enforced,every objection logged, every finding auditable.