Ricoh releases multimodal LLM capable of processing charts

New open-weight multimodal model from Ricoh excels at chart reasoning—perfect for document-heavy AI tasks.
30-Second TL;DR
What Changed
Multimodal model capable of handling text and charts
Why It Matters
By providing a lightweight, multimodal model capable of chart analysis, Ricoh is lowering the barrier for enterprises to implement document-heavy AI workflows.
What To Do Next
Download the Ricoh LMM from Hugging Face and benchmark it against your current document processing pipeline for chart-heavy data.
Key Points
- •Multimodal model capable of handling text and charts
- •Developed under the Japanese government's GENIAC project
- •Lightweight model version released for free on Hugging Face
Deep Insight
Background and context from public sources — not the original article. 6 sources cited.
Enhanced Key Takeaways
- •Ricoh's multimodal LLM, developed under the GENIAC project, is a 'reasoning LMM' capable of multi-stage inference to achieve higher accuracy in interpreting complex documents that include charts.
- •The model is based on Alibaba Cloud's Qwen2.5-VL-32B-Instruct model, with a lightweight 8-billion parameter version (Qwen-3-VL-Ricoh-8B-20260227) released for free on Hugging Face.
- •A key focus of Ricoh's LLM development, including this multimodal model, is secure on-premise deployment to address the needs of industries like finance and healthcare that handle confidential data.
- •The model was trained using approximately 600,000 images derived from business documents, encompassing various visual elements such as characters, pie charts, bar graphs, and flow charts.
- •Ricoh has also developed a separate 70-billion parameter LLM that supports Japanese, English, and Chinese, featuring a customized tokenizer that improves Japanese language processing efficiency by 43%.
Technical Deep Dive
- Base Model: Built upon Alibaba Cloud's Qwen2.5-VL-32B-Instruct model.
- Lightweight Version: A Qwen-3-VL-Ricoh-8B-20260227 model with 8 billion parameters has been released.
- Reasoning Capability: Features multi-stage inference for enhanced accuracy in document and chart interpretation.
- Training Data: Trained on approximately 600,000 images from business documents, including diverse chart types (pie, bar, flow charts) and characters.
- Evaluation: Performance validated using the Japanese question-answering dataset JDocQA (combining text and visual information) and proprietary benchmark tools, demonstrating superior results compared to other models.
- Deployment: Designed for on-premise environments, allowing for fine-tuning with company-specific data while maintaining security.
- Efficiency: A 4-bit quantized version is available for resource-constrained setups.
- Vision Encoder: Incorporates a vision encoder module to convert visual information into a format understandable by the language model.
Future ImplicationsAI analysis grounded in cited sources
Timeline
- 1980sRicoh begins AI development, focusing on image recognition and natural language processing.
- 2020Ricoh announces its transformation into a digital service provider and starts leveraging NLP for office documents and customer feedback.
- 2022-12Ricoh develops its proprietary AI model, 'Ricoh GPT,' based on GPT-3, in three months.
- 2023-03Ricoh develops its own large language model (LLM).
- 2024-02Japan's Ministry of Economy, Trade and Industry (METI) and NEDO launch the Generative AI Accelerator Challenge (GENIAC) project.
- 2024-08Ricoh announces a 70-billion parameter LLM supporting Japanese, English, and Chinese.
- 2024-10Ricoh is selected for the GENIAC project (Phase 2) to develop multimodal LLMs.
- 2025-05Ricoh and Sompo Japan begin joint development of a private multimodal LLM for insurance operations under the GENIAC project.
- 2025-06Ricoh completes development of the basic model for its multimodal LLM under GENIAC, capable of reading Japanese documents with diagrams.
- 2026-01Ricoh announces the development of a multimodal LLM based on Qwen2.5-VL-32B for reading Japanese company documents with diagrams and charts under GENIAC Phase 2.
- 2026-02Ricoh releases the lightweight multimodal model 'Qwen-3-VL-Ricoh-8B-20260227' for free on Hugging Face.
Sources (6)
Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.
Weekly AI Recap
Read this week's curated digest of top AI events →
AI-curated news aggregator. All content rights belong to original publishers.
Original source: ITmedia AI+ (日本) ↗
This is a summary, not the original. Read the source, or get the weekly briefing.
The weekly digest
One email a week. Unsubscribe anytime.
