Qwen3.8-Flash-Next Runs on 48GB Mac

π‘A practical test of running a 104GB Qwen model on a 48GB Mac at usable speed.
β‘ 30-Second TL;DR
What Changed
The reported model size is 104GB.
Why It Matters
The report suggests that local developers may be able to experiment with models larger than their machineβs physical memory. Actual usability will depend on the quantization, memory-management approach, prompt length, and workload.
What To Do Next
Reproduce the setup with your preferred Mac inference runtime and record memory usage, quantization format, context length, and sustained tok/s.
Key Points
- β’The reported model size is 104GB.
- β’The host machine is a 48GB Mac.
- β’Reported generation speed is approximately 12 tokens per second.
- β’The result is a community demonstration rather than a formal benchmark.
Weekly AI Recap
Read this week's curated digest of top AI events β
πRelated Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: Reddit r/LocalLLaMA β
This is a summary, not the original. Read the source, or get the weekly briefing.
Weekly AI briefing
One email a week. Unsubscribe anytime.
