📄較早收集於 4h

Mirror 在內分泌委員會考試中勝過 GPT-5

Mirror 在內分泌委員會考試中勝過 GPT-5
PostLinkedIn
📄閱讀原文: ArXiv AI
#clinical-reasoning#medical-ai#benchmark#evidence-ragjanuary-mirror

💡Curated med AI beats GPT-5 on board exam w/ traceable evidence (87.5% acc)

⚡ 30-Second TL;DR

有什麼變化

87.5%準確率(105/120)對比GPT-5.2的74.6%及人類62.3%

為什麼重要

證明精選證據層可實現優於通用LLM與網路工具的亞專科推理,提升臨床應用的可審核性。突顯專門AI在醫學領域超越廣泛模型的潛力。

下一步行動

Benchmark January Mirror against your LLMs on medical reasoning datasets.

誰應關注:Researchers & Academics

關鍵要點

  • 87.5%準確率(105/120)對比GPT-5.2的74.6%及人類62.3%
  • 30最難題76.7%(人類<50%)
  • 74.2%輸出引用指南來源;引用準確率100%驗證
  • Top-2準確率92.5%對比GPT-5.2的85.25%
  • 封閉證據設定勝過有網路存取的LLM

🧠 深度解析

背景與延伸:來自公開資料,非原文內容。引用 7 個來源。

🔑 增強重點摘要

  • January Mirror achieved 87.5% accuracy on ESAP 2025 endocrinology board exam, representing a 25.2 percentage point improvement over human endocrinologists (62.3%) and 12.9 points over GPT-5.2 (74.6%)[1]
  • Mirror's architecture uses an ensemble-style clinical reasoning stack with specialized components organized around clinical question archetypes (diagnosis, testing, treatment, prognosis, mechanistic reasoning) and an arbitration layer for final output selection[1]
  • The system demonstrated particular strength on difficult questions where human accuracy fell below 50%, achieving 76.7% accuracy on the 30 hardest questions, suggesting it captures clinical reasoning patterns that challenge subspecialty-trained physicians[1]
  • Mirror operated under closed-evidence constraints without external retrieval, outperforming frontier LLMs (GPT-5, GPT-5.2, Gemini-3-Pro) that had real-time web access to guidelines and primary literature[1]
  • FDA's January 2026 Clinical Decision Support Software guidance shift enables single-recommendation outputs when presented as 'Glass Box' systems with transparent, verifiable reasoning—a regulatory framework that aligns with Mirror's evidence-linked output design[3]
📊 競品分析▸ Show
SystemAccuracyArchitectureEvidence AccessCitation Accuracy
January Mirror87.5%Ensemble reasoning with arbitration layerClosed-evidence corpus100% verified
GPT-5.274.6%General-purpose LLMReal-time web accessNot specified
GPT-5~70-72% (inferred)General-purpose LLMReal-time web accessNot specified
Gemini-3-Pro~70-72% (inferred)General-purpose LLMReal-time web accessNot specified
Human endocrinologists (reference)62.3%Clinical expertiseDomain knowledgeN/A

🛠️ 技術深入

Reasoning Architecture: Ensemble-style clinical reasoning stack with multiple specialized reasoning components generating candidate answers and supporting evidence links, followed by arbitration layer using evidence quality and internal agreement signals • Question Archetype Organization: Components structured around common clinical reasoning tasks—diagnosis, testing, treatment, prognosis, and mechanistic reasoning—reflecting how clinicians approach distinct problem types • Evidence Integration: System integrates curated endocrinology and cardiometabolic evidence corpus with structured reasoning architecture to generate evidence-linked outputs; operates under closed-evidence constraint without external retrieval • Output Design: Outputs include traceable citations from guidelines with 100% citation accuracy verification; 74.2% of outputs cited guideline sources • Performance Metrics: Top-2 accuracy of 92.5% vs GPT-5.2's 85.25%, indicating strong confidence in alternative diagnoses • Regulatory Alignment: Design aligns with FDA's January 2026 'Glass Box' CDS guidance requiring transparent, clinician-reviewable logic and verifiable source grounding rather than hallucinated citations[3]

🔮 前景展望AI analysis grounded in cited sources

Mirror's performance demonstrates that domain-specific evidence curation and structured clinical reasoning can achieve subspecialty-level performance exceeding general-purpose frontier LLMs, suggesting a strategic shift toward specialized clinical AI systems rather than relying on general-purpose models with web access. The January 2026 FDA guidance shift toward 'Glass Box' transparency requirements creates regulatory tailwinds for evidence-grounded systems like Mirror while restricting black-box approaches. This validates a design philosophy where AI augments rather than replaces clinician judgment through transparent reasoning. The closed-evidence advantage over web-enabled LLMs indicates that curated, high-quality evidence corpora may outperform real-time information access in specialized domains. Healthcare organizations may increasingly adopt domain-specific clinical reasoning systems for subspecialty applications, particularly in high-stakes diagnostic and treatment planning contexts where explainability and citation accuracy are regulatory and clinical requirements. The framework's success on difficult questions suggests potential for AI-assisted continuing medical education and quality improvement in areas where even subspecialty physicians struggle.

時間線

2022-01
FDA issued Clinical Decision Support Software guidance establishing framework for AI medical device classification
2021-01
Study site tested AI clinical decision support system in radiology with small group of radiologists (n=4) reporting positive experiences
2025-01
ESAP 2025 endocrinology board-style examination administered (120-question assessment used for Mirror evaluation)
2026-01
FDA issued updated Clinical Decision Support Software guidance superseding 2022 version, enabling single-recommendation outputs when presented as transparent 'Glass Box' systems with verifiable source grounding
2026-01
OpenAI, Anthropic, and Amazon launched enterprise healthcare products following FDA guidance shift
📰

AI 週報

閱讀本週精選 AI 大事摘要 →

👉相關動態

AI 策展新聞聚合。所有內容版權歸原始發布者所有。
原始來源: ArXiv AI

這是摘要,不是原文。去看原站,或訂閱每週簡報。

每週 AI 簡報

每週一封,可隨時退訂。