Ox Alpha: The Mystery Model Goes Public

💡A free 1M-context multimodal model may rival leading systems—but its anonymous origin and safety limits need testing.
⚡ 30-Second TL;DR
What Changed
Ox Alpha supports a 1,048,576-token context window, up to 131,072 output tokens, text/image/video input, and tool calling.
Why It Matters
If the early results hold up, Ox Alpha could become a strong free option for long-context coding agents and multimodal workflows. However, its anonymous provenance, unclear data-retention policy, and limited independent benchmarking create meaningful production and security risks.
What To Do Next
Run Ox Alpha through OpenRouter on non-sensitive coding tasks, compare it with your current agent model on a repeatable benchmark, and keep proprietary code and secrets out of the prompts.
Key Points
- •Ox Alpha supports a 1,048,576-token context window, up to 131,072 output tokens, text/image/video input, and tool calling.
- •The model is distributed under anonymous identifiers, including stealth/ox-alpha on OpenRouter and Ox Alpha Free in OpenCode Zen.
- •A small 10-task DeepSWE sample reported an 80% pass rate, but the result is high variance and not a substitute for the full 113-task benchmark.
- •The model reportedly solved a large-bound numerical recurrence problem and was used in coding agents and visual game-control experiments.
🧠 Deep Insight
Background and context from public sources — not the original article. 9 sources cited.
🔑 Enhanced Key Takeaways
- •The model is currently operating under a limited free preview window scheduled to terminate on August 27, 2026.
- •Operators claim the infrastructure supporting Ox Alpha can handle a throughput of 100 trillion tokens per day.
- •Technical fingerprinting of tokenizer patterns and error logs has led researchers to hypothesize a connection to Zhipu AI's GLM architecture.
- •Security experts have issued formal warnings against using the model for sensitive or proprietary data due to the anonymous nature of the provider and data retention practices.
- •Community-led investigations on platforms like Reddit and X are actively attempting to differentiate the model's architecture from other potential candidates like Xiaomi's MiMo or DeepSeek.
📊 Competitor Analysis▸ Show
| Feature | Ox Alpha | Claude Fable 5 | GPT-5.6 |
|---|---|---|---|
| Context Window | 1M Tokens | 200K Tokens | 512K Tokens |
| Pricing | Free (Limited) | Paid/Subscription | Paid/Subscription |
| Coding (DeepSWE) | 80% (Sample) | Lower (Reported) | Lower (Reported) |
| Multimodal | Text/Image/Video | Text/Image | Text/Image |
🛠️ Technical Deep Dive
- Tokenizer analysis suggests structural similarities to the GLM (General Language Model) family of architectures.
- Implements a high-throughput inference stack capable of sustaining 100 trillion tokens per day.
- Supports native multimodal processing for video inputs, distinguishing it from standard text-image models.
- Utilizes a specialized error-handling protocol that has become a primary identifier for community researchers tracing the model's origin.
🔮 Future ImplicationsAI analysis grounded in cited sources
⏳ Timeline
📎 Sources (9)
Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.
Weekly AI Recap
Read this week's curated digest of top AI events →
👉Related Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: 虎嗅 ↗
This is a summary, not the original. Read the source, or get the weekly briefing.
Weekly AI briefing
One email a week. Unsubscribe anytime.


