DeepSeek-V4 Flash Vision Outpaces Qwen in Coding
π‘Raw speed is not the whole story: this local test found DeepSeek finishing coding tasks twice as fast.
β‘ 30-Second TL;DR
What Changed
DeepSeek-V4-Flash-Vision was about 40% slower in raw token generation at the same quantization level.
Why It Matters
The report suggests that end-to-end task completion and reliability can matter more than raw tokens per second for local coding assistants. Results are anecdotal and hardware-dependent, so practitioners should validate them against their own workloads.
What To Do Next
Run the same coding task with Q8_K_XL builds of DeepSeek-V4-Flash-Vision and Qwen3.8-Flash-Next, measuring completion time, retries, and correctness rather than tok/s alone.
Key Points
- β’DeepSeek-V4-Flash-Vision was about 40% slower in raw token generation at the same quantization level.
- β’It completed one task in 12 minutes on medium mode versus Qwen3.8-Flash-Next's 25 minutes.
- β’Qwen3.8-Flash-Next reportedly failed a task after nearly three hours in xhigh mode.
- β’The comparison used Q8_K_XL quantization on dual StrixHalo 128GB hardware.
- β’The author observed that Qwen tends to overinterpret underspecified instructions.
Weekly AI Recap
Read this week's curated digest of top AI events β
πRelated Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: Reddit r/LocalLLaMA β
This is a summary, not the original. Read the source, or get the weekly briefing.
Weekly AI briefing
One email a week. Unsubscribe anytime.

