Ricoh releases multimodal LLM capable of processing charts

💡New open-weight multimodal model from Ricoh excels at chart reasoning—perfect for document-heavy AI tasks.
⚡ 30-Second TL;DR
What Changed
Multimodal model capable of handling text and charts
Why It Matters
By providing a lightweight, multimodal model capable of chart analysis, Ricoh is lowering the barrier for enterprises to implement document-heavy AI workflows.
What To Do Next
Download the Ricoh LMM from Hugging Face and benchmark it against your current document processing pipeline for chart-heavy data.
Key Points
- •Multimodal model capable of handling text and charts
- •Developed under the Japanese government's GENIAC project
- •Lightweight model version released for free on Hugging Face
🧠 Deep Insight
Web-grounded analysis with 6 cited sources.
🔑 Enhanced Key Takeaways
- •Ricoh's multimodal LLM, developed under the GENIAC project, is a 'reasoning LMM' capable of multi-stage inference to achieve higher accuracy in interpreting complex documents that include charts.
- •The model is based on Alibaba Cloud's Qwen2.5-VL-32B-Instruct model, with a lightweight 8-billion parameter version (Qwen-3-VL-Ricoh-8B-20260227) released for free on Hugging Face.
- •A key focus of Ricoh's LLM development, including this multimodal model, is secure on-premise deployment to address the needs of industries like finance and healthcare that handle confidential data.
- •The model was trained using approximately 600,000 images derived from business documents, encompassing various visual elements such as characters, pie charts, bar graphs, and flow charts.
- •Ricoh has also developed a separate 70-billion parameter LLM that supports Japanese, English, and Chinese, featuring a customized tokenizer that improves Japanese language processing efficiency by 43%.
🛠️ Technical Deep Dive
- Base Model: Built upon Alibaba Cloud's Qwen2.5-VL-32B-Instruct model.
- Lightweight Version: A Qwen-3-VL-Ricoh-8B-20260227 model with 8 billion parameters has been released.
- Reasoning Capability: Features multi-stage inference for enhanced accuracy in document and chart interpretation.
- Training Data: Trained on approximately 600,000 images from business documents, including diverse chart types (pie, bar, flow charts) and characters.
- Evaluation: Performance validated using the Japanese question-answering dataset JDocQA (combining text and visual information) and proprietary benchmark tools, demonstrating superior results compared to other models.
- Deployment: Designed for on-premise environments, allowing for fine-tuning with company-specific data while maintaining security.
- Efficiency: A 4-bit quantized version is available for resource-constrained setups.
- Vision Encoder: Incorporates a vision encoder module to convert visual information into a format understandable by the language model.
🔮 Future ImplicationsAI analysis grounded in cited sources
⏳ Timeline
📎 Sources (6)
Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.
Weekly AI Recap
Read this week's curated digest of top AI events →
👉Related Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: ITmedia AI+ (日本) ↗