SourceReddit r/LocalLLaMA•Stalecollected in 9h
oQ: Data-Driven Quant for Apple Silicon

#quantization#mixed-precision#apple-siliconoq-quantizationoqmlx-lmapple-siliconqwen3.5
💡2-bit quant hits 64% MMLU on Apple Silicon—beats mlx-lm defaults
⚡ 30-Second TL;DR
What Changed
Sensitivity-driven bit allocation using calibration datasets
Why It Matters
Lowers barrier for high-quality quantized models on Apple hardware, boosting local inference speed and accessibility for developers.
What To Do Next
Quantize Qwen3.5-35B with oQ at omlx.ai and load into LM Studio.
Who should care:Developers & AI Engineers
Key Points
- •Sensitivity-driven bit allocation using calibration datasets
- •2-bit oQ: 64% MMLU, 78% HumanEval on Qwen3.5-35B
- •Outperforms uniform mlx-lm quants across MMLU, TruthfulQA, etc.
- •Quantize from omlx.ai, compatible with any mlx inference
📰
Weekly AI Recap
Read this week's curated digest of top AI events →
👉Related Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: Reddit r/LocalLLaMA ↗
This is a summary, not the original. Read the source, or get the weekly briefing.
The weekly digest
One email a week. Unsubscribe anytime.