๐Ÿค–Stalecollected in 49h

4.5B ColQwen3.5-v1 Hits ViDoRe V1 SOTA

PostLinkedIn
๐Ÿค–Read original on Reddit r/MachineLearning

๐Ÿ’กOpen 4.5B model sets new SOTA on ViDoRe doc retrieval benchmark

โšก 30-Second TL;DR

What Changed

SOTA #1 on ViDoRe V1 (nDCG@5 0.917), competitive on V3

Why It Matters

Pushes open-source boundaries in vision-document retrieval, enabling better RAG for finance/tables.

What To Do Next

Download ColQwen3.5-v1 from Hugging Face and benchmark on ViDoRe V1.

Who should care:Researchers & Academics

Key Points

  • โ€ขSOTA #1 on ViDoRe V1 (nDCG@5 0.917), competitive on V3
  • โ€ขBuilt on Qwen3.5-4B with ColPali late-interaction approach
  • โ€ข4-phase training: hard negative mining, finance/table doc specialization
  • โ€ขApache 2.0 weights on HF: https://huggingface.co/athrael-soju/colqwen3.5-v1

๐Ÿง  Deep Insight

Background and context from public sources โ€” not the original article. 8 sources cited.

๐Ÿ”‘ Enhanced Key Takeaways

  • โ€ขColQwen3 family extends Qwen3-VL with ColBERT-style late interaction heads for per-token embeddings, enabling efficient multimodal retrieval across both text and image inputs[6].
  • โ€ขTomoro ColQwen3 achieves 13x storage cost reduction compared to previous generation models, storing 1 million images in 0.82 TB versus 10.3 TB for baseline approaches[4].
  • โ€ขQwen3.5-4B demonstrates strong multimodal reasoning performance (65.4% on MMMU-Pro), outperforming larger models like Qwen3 VL 8B (56.6%) and Ministral 3 8B (46.0%), making it an effective foundation for specialized variants[2].

๐Ÿ”ฎ Future ImplicationsAI analysis grounded in cited sources

Specialized multimodal embeddings enable cost-effective document retrieval at scale
The 13x storage reduction achieved by ColQwen3 variants suggests that domain-specific training (finance documents, tables) combined with efficient embedding architectures can democratize multimodal RAG systems for enterprise applications.
Sub-10B multimodal models are becoming competitive alternatives to larger frontier models
Qwen3.5-4B's superior MMMU-Pro performance versus 8B competitors indicates that architectural innovations (Gated Delta Networks, MoE) and specialized training can compress capability into smaller deployable models.

โณ Timeline

2026-03
Qwen3.5 Small series (0.8B-9B) released with native multimodal capabilities and Gated Delta Network architecture
2026-03
Tomoro ColQwen3 multimodal embedding models released, achieving 13x storage cost reduction on ViDoRe benchmarks
๐Ÿ“ฐ

Weekly AI Recap

Read this week's curated digest of top AI events โ†’

๐Ÿ‘‰Related Updates

AI-curated news aggregator. All content rights belong to original publishers.
Original source: Reddit r/MachineLearning โ†—

This is a summary, not the original. Read the source, or get the weekly briefing.

Weekly AI briefing

One email a week. Unsubscribe anytime.