Baichuan Intelligence to Launch New Medical LLM

💡New medical LLM claims 3.3% hallucination rate, challenging general models in high-stakes healthcare.
⚡ 30-Second TL;DR
What Changed
New medical-specific LLM focuses on high accuracy.
Why It Matters
This release highlights the trend of vertical-specific models outperforming general-purpose models in high-stakes domains like healthcare.
What To Do Next
Benchmark your current RAG pipeline against this 3.3% hallucination rate to evaluate if vertical fine-tuning is necessary for your use case.
Key Points
- •New medical-specific LLM focuses on high accuracy.
- •Factual hallucination rate reduced to 3.3%.
- •Wang Xiaochuan argues general models fail to meet medical industry standards.
🧠 Deep Insight
Web-grounded analysis with 13 cited sources.
🔑 Enhanced Key Takeaways
- •The latest medical LLM from Baichuan, M4, was unveiled on May 22, 2026, alongside an Agent product named "Baixiaoyi," indicating continuous rapid development beyond the M3 Plus model.
- •Baichuan Intelligence has strategically pivoted to an "all-in on healthcare" approach since approximately August 2024, significantly downsizing its general model teams and other industry lines like finance to focus exclusively on medical AI.
- •The company's earlier open-source medical model, Baichuan-M3, demonstrated superior performance on the HealthBench and HealthBench Hard evaluations, reportedly outperforming OpenAI's GPT-5.2 and the average level of human doctors in certain tasks.
- •Baichuan's M3 Plus model incorporates an "evidence anchoring" feature that links AI-generated medical conclusions directly to specific paragraphs in original research papers, enhancing verifiability and accountability.
- •Wang Xiaochuan, the founder, envisions creating "AI doctors" capable of operating at the level of top-tier physicians, addressing the global shortage of skilled medical professionals, and views healthcare as the "crown jewel" of large models.
📊 Competitor Analysis▸ Show
| Feature/Model | Baichuan-M3 (Open-source) | Baichuan-M2 (Open-source) | OpenAI GPT-5.2 | DeepSeek (Chinese LLM) | MedGPT (Chinese LLM) |
|---|---|---|---|---|---|
| Focus | Medical LLM | Medical LLM | General LLM (with medical capabilities) | General/Medical LLM | Medical LLM |
| Hallucination Rate | "Lowest medical hallucination rate" (M3), 2.6% (M3 Plus) | N/A | N/A | N/A | N/A |
| HealthBench Score | 65.1 (Total), 44.4 (Hard) | 60.1 (Total), 34.7 (Hard) | 57.6 (gpt-oss120b), surpassed by M3 | N/A | N/A |
| NMLE Performance | N/A | N/A | Outperformed by DeepSeek | Highest among Chinese LLMs (2018-2024) | N/A |
| Deployment Cost | N/A | ~ $1,400 (single RTX 4090, quantized) | N/A | ~ 57x Baichuan-M2 (DeepSeek-R1 H20 dual-node) | N/A |
| Key Features | Clinical decision-making, active inquiry, Fact-Aware RL, open-source | Lightweight, private deployment, AI Patient Simulator | General capabilities, medical assistants (HIPAA-compliant) | Strong in Chinese medical exams | Human expert-level scores in double-blind studies |
🛠️ Technical Deep Dive
- Baichuan-M3 is trained to explicitly model the clinical decision-making process, moving beyond static question-answering to support real-world medical practice.
- It utilizes a specialized three-stage training pipeline that includes Task-Specific Reinforcement Learning and Multi-Teacher Online Policy Distillation.
- A key innovation is Segmented Pipeline Reinforcement Learning, which mimics a physician's workflow across inquiry, testing, and diagnosis stages.
- The SPAR algorithm (Step Penalized Advantage with Relative Baseline) and a hybrid Verify System are employed to ensure logical consistency and adherence to medical protocols.
- Fact-Aware Reinforcement Learning (RL) is specifically used to suppress hallucinations in the model's outputs.
- For efficient deployment, Baichuan-M3 features W4 quantization, which reduces memory usage to 26% of the original, and Gated Eagle3 speculative decoding, achieving a 96% speedup.
- The M3 Plus model introduces "evidence anchoring," a feature that provides citation sources and links each AI-generated medical conclusion to the corresponding evidence paragraph in original research papers for verifiability.
- Baichuan-M2 is designed to be lightweight, capable of running on a single RTX 4090 GPU after quantization, significantly reducing deployment costs to approximately $1,400.
🔮 Future ImplicationsAI analysis grounded in cited sources
⏳ Timeline
📎 Sources (13)
Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.
Weekly AI Recap
Read this week's curated digest of top AI events →
👉Related Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: 钛媒体 ↗


