SourceStalecollected in 9h

oQ: Data-Driven Quant for Apple Silicon

oQ: Data-Driven Quant for Apple Silicon
PostLinkedIn
🦙Read original on Reddit r/LocalLLaMA
#quantization#mixed-precision#apple-siliconoq-quantizationoqmlx-lmapple-siliconqwen3.5

💡2-bit quant hits 64% MMLU on Apple Silicon—beats mlx-lm defaults

⚡ 30-Second TL;DR

What Changed

Sensitivity-driven bit allocation using calibration datasets

Why It Matters

Lowers barrier for high-quality quantized models on Apple hardware, boosting local inference speed and accessibility for developers.

What To Do Next

Quantize Qwen3.5-35B with oQ at omlx.ai and load into LM Studio.

Who should care:Developers & AI Engineers

Key Points

  • Sensitivity-driven bit allocation using calibration datasets
  • 2-bit oQ: 64% MMLU, 78% HumanEval on Qwen3.5-35B
  • Outperforms uniform mlx-lm quants across MMLU, TruthfulQA, etc.
  • Quantize from omlx.ai, compatible with any mlx inference
📰

Weekly AI Recap

Read this week's curated digest of top AI events →

👉Related Updates

AI-curated news aggregator. All content rights belong to original publishers.
Original source: Reddit r/LocalLLaMA

This is a summary, not the original. Read the source, or get the weekly briefing.

The weekly digest

One email a week. Unsubscribe anytime.