Life sciencesHealthcareDeep · systematic review uploadedGuardian · clinical integrity9h ago · Academic Medical Centre

"What does the evidence establish about the comparative effectiveness of GLP-1 agonists vs. bariatric surgery for long-term weight maintenance?"

Pharmacy & Therapeutics Committee · Formulary review · Draft systematic review submitted as context document

Illustrative session · augle.com
Finding · Phase 3 synthesisContested
46%
Ensemble confidence
CONFIDENCE GRADE
Contested

No directional comparative-effectiveness claim is supportable. Within-modality durability claims: Probable.

Calibration basis
Confidence grade

Finding is the confidence grade itself. Confidence grade is the finding. For an open research question, the grade — not a numeric score — is the complete result.

The ensemble does not support a directional comparative-effectiveness claim between GLP-1 agonists and bariatric surgery for long-term weight maintenance. Bariatric surgery shows comparable-to-larger and more durable weight loss in long-term observational cohorts, but no large long-term head-to-head randomised trial exists — the comparison rests on indirect evidence across materially different populations. The Contrarian's Strong objection on population non-comparability is unresolved. Within-modality durability claims — that surgical loss persists at 5+ years in cohort data — are Probable. The actionable output: any formulary recommendation must be framed as modality-appropriate for distinct patient populations, not as a head-to-head superiority claim.

Confidence grid · 5 agents × 20 units
Agreement / confidentShorter row = lower confidence / dissentUnfilled
Reopen conditions · evidence-triggered, not market-triggered
RC1
A large head-to-head randomised trial comparing GLP-1 agonists and bariatric surgery for weight maintenance with ≥3-year follow-up
Trigger: Head-to-head RCT publication · Direction: Directional comparative claim becomes gradeable
RC2
A matched-population long-term cohort controlling for baseline BMI, age, and comorbidity profile across both interventions
Trigger: Matched cohort study · Direction: Resolves population non-comparability objection
RC3
Reframing of the formulary question from head-to-head superiority to within-modality durability by patient population
Trigger: Question revision · Direction: Probable grade achievable with revised claim scope
Phase transcript · all three rounds
Cartographer
Gemini 3.1 Pro · T=0.8
64%

Settled ground: Both GLP-1 agonists and bariatric surgery produce clinically significant weight loss in the short-to-medium term. Bariatric surgery has long-term (10–20 year) observational durability data. Weight regain follows GLP-1 discontinuation.

Contested terrain: Relative durability of the two modalities beyond three years. Whether the magnitude gap narrows with newer dual agonists. Comparative effectiveness in matched populations.

Unknown territory: No large long-term head-to-head randomised trial has ever compared the two interventions for weight maintenance. Surgical and pharmacological cohorts differ systematically in baseline BMI, age, and comorbidity.

Knowledge gaps entered: (1) Head-to-head RCT with long follow-up — does not exist. (2) Matched-population comparative cohort controlling for baseline differences — does not exist.

Methodologist
GPT-4o · T=0.5
58%

Comparability concern flagged: The comparison rests entirely on indirect evidence — no head-to-head randomisation exists. Surgical and pharmacological cohorts are not exchangeable: baseline BMI, age, and comorbidity profiles differ materially, so any cross-cohort effect estimate is not transportable between populations.

Evidence ceiling: the directional comparative claim is capped at Contested — indirect comparison across non-equivalent populations cannot support a superiority finding. Within-modality durability claims (surgical loss persisting at 5+ years in cohort data) are Probable. No node qualifies for Established given the absence of head-to-head randomised evidence.

Contrarian
Claude Sonnet 4.6 · T=1.0
37%

@Cartographer — Steelman: The long-term observational evidence for surgical durability is genuinely strong, and the direction of the cohort data is consistent. A clinician reading the literature would reasonably conclude surgery is more durable.

Strong objection [Phase 1]: "No large long-term head-to-head randomised trial exists. The comparison rests on indirect evidence across cohorts with materially different baseline BMI, age, and comorbidity profiles — surgical candidates are systematically higher-BMI and higher-risk. The effect estimate is not transportable between these populations, so a directional superiority claim is unsupported regardless of how consistent the cohort data looks." Resolution condition: A head-to-head randomised trial, or a matched-population cohort controlling for the baseline differences.

Dissent register · all Contrarian objections1 Strong Unresolved · 1 Moderate
Strong@Contrarian → @CartographerPhase 1 · carried to Phase 3Unresolved
Steelman

The long-term observational evidence for surgical durability is genuinely strong and directionally consistent. A clinician reading the literature would reasonably read surgery as more durable.

"No large long-term head-to-head randomised trial exists. Surgical and pharmacological cohorts differ systematically in baseline BMI, age, and comorbidity — the cross-cohort comparison is not transportable, so a directional superiority claim is unsupported."

Resolution condition

A head-to-head randomised trial, or a matched-population cohort controlling for baseline BMI, age, and comorbidity differences

Moderate@Contrarian → @MethodologistPhase 2Actionable
Steelman

That surgery has longer follow-up is a real feature of the evidence base, and longer-term data is genuinely reassuring on durability.

"The surgical durability evidence extends further largely because the intervention is older. Longer follow-up is not the same as superior durability — the comparison is confounded by evidence age, unacknowledged in the draft."

Resolution condition

Explicit acknowledgment of the evidence-age confound in the review's limitations section

Pragmatist action item: Add evidence-age and population-comparability caveats to the review before circulation

Guardian integrity log · clinical integrity mode97% · 1 flag
97%
Overall integrity score
Clinical integrity mode

No retracted papers. All primary trial citations authenticated against registries. One industry-funded extension study flagged for disclosure. Self-citation ratio within field norms. Phase boundary clearances issued at P1/2 and P2/3.

Source quality
97%
Retraction check
100%
Funding disclosure
88%
Self-citation ratio
95%
SVS verification log · 22 citations checked
Wilding et al. (2021) — STEP 1 semaglutide weight-management RCT
Peer-reviewed · Verified
Jastreboff et al. (2022) — SURMOUNT-1 tirzepatide RCT
Peer-reviewed · Verified
Sjöström et al. (2004/2012) — Swedish Obese Subjects long-term surgical cohort
Peer-reviewed · Verified
Sponsor-affiliated (2024) — GLP-1 durability extension analysis
SVS_FUNDING — industry-funded extension, sponsor-affiliated authors · Evidence node capped at Probable · Flagged for disclosure
Industry-funded · Flagged
Wilding et al. (2022) — STEP withdrawal / regain after discontinuation
Peer-reviewed · Verified
+17 additional citations verified · 0 retracted · 1 industry-funded flagged above
Session metadata
Session IDlife-glp1-vs-bariatric
VerticalHealthcare
DepthDeep
Runtime38m 04s
Attachmentdraft review
Guardian modeClinical
Guardian score97% · 1 flag
Citations22 · 1 flagged
Dissent flags1 Strong · 1 Mod
Calibration status
Calibration basis
Confidence gradeConfidence grade is the finding

For open research questions, the confidence grade is the complete finding — calibrated against the evidence base, not against a binary resolution outcome.

Finding gradeContested · 46%
Agent confidence
Cartographer
Gemini 3.1 Pro
52%
Methodologist
GPT-4o
48%
Contrarian
Claude Sonnet 4.6
28%
Synthesizer
GPT-4o
44%
Pragmatist
Grok 4.1 Fast
42%
Export & share
Illustrative session · augle.com