Search

Tag: #document-ai14 results

Qwen3.5-9B Beats Frontiers on Document Benchmarks

Qwen3.5-9B Beats Frontiers on Document Benchmarks

Qwen3.5 models (0.8B to 9B) were evaluated on an open document AI benchmark with 9,000+ real documents. The 9B and 4B variants outperform frontier models like Gemini 3.1 Pro and GPT-5.4 in OCR text extraction, VQA, and KIE tasks. However, they lag significantly in table extraction and handwriting OCR.

Reddit r/LocalLLaMACommunityMar 16#document-ai#benchmarks#ocr
🔬

IDP Leaderboard Benchmarks 16 VLMs

New open IDP Leaderboard evaluates 16 VLMs on 9,000+ documents across three benchmarks: OlmOCR, OmniDoc, and IDP Core. Gemini 3.1 Pro leads narrowly; cheaper variants match flagships except on reasoning tasks. Features a Results Explorer for predictions vs. ground truth.

Reddit r/MachineLearningCommunityMar 11#vlm-benchmark#document-ai#leaderboard
Evidence-Grounded Checks for Construction PDFs

Evidence-Grounded Checks for Construction PDFs

這項研究提出一套用於建築文件審查的證據導向管線,將文字、幾何、頁面與版本資訊正規化,並以決定性規則進行約束檢查。實驗顯示,增加重疊頁面區塊的解析度在重複測試中提升準確率,但在更廣泛的資料集並未穩定勝過頁面涵蓋率,顯示需要規則感知的證據路由與人工覆核。

Page 1 of 2