🗾Stalecollected in 81m

Ricoh releases multimodal LLM capable of processing charts

Ricoh releases multimodal LLM capable of processing charts
PostLinkedIn
🗾Read original on ITmedia AI+ (日本)

💡New open-weight multimodal model from Ricoh excels at chart reasoning—perfect for document-heavy AI tasks.

⚡ 30-Second TL;DR

What Changed

Multimodal model capable of handling text and charts

Why It Matters

By providing a lightweight, multimodal model capable of chart analysis, Ricoh is lowering the barrier for enterprises to implement document-heavy AI workflows.

What To Do Next

Download the Ricoh LMM from Hugging Face and benchmark it against your current document processing pipeline for chart-heavy data.

Who should care:Researchers & Academics

Key Points

  • Multimodal model capable of handling text and charts
  • Developed under the Japanese government's GENIAC project
  • Lightweight model version released for free on Hugging Face

🧠 Deep Insight

Web-grounded analysis with 6 cited sources.

🔑 Enhanced Key Takeaways

  • Ricoh's multimodal LLM, developed under the GENIAC project, is a 'reasoning LMM' capable of multi-stage inference to achieve higher accuracy in interpreting complex documents that include charts.
  • The model is based on Alibaba Cloud's Qwen2.5-VL-32B-Instruct model, with a lightweight 8-billion parameter version (Qwen-3-VL-Ricoh-8B-20260227) released for free on Hugging Face.
  • A key focus of Ricoh's LLM development, including this multimodal model, is secure on-premise deployment to address the needs of industries like finance and healthcare that handle confidential data.
  • The model was trained using approximately 600,000 images derived from business documents, encompassing various visual elements such as characters, pie charts, bar graphs, and flow charts.
  • Ricoh has also developed a separate 70-billion parameter LLM that supports Japanese, English, and Chinese, featuring a customized tokenizer that improves Japanese language processing efficiency by 43%.

🛠️ Technical Deep Dive

  • Base Model: Built upon Alibaba Cloud's Qwen2.5-VL-32B-Instruct model.
  • Lightweight Version: A Qwen-3-VL-Ricoh-8B-20260227 model with 8 billion parameters has been released.
  • Reasoning Capability: Features multi-stage inference for enhanced accuracy in document and chart interpretation.
  • Training Data: Trained on approximately 600,000 images from business documents, including diverse chart types (pie, bar, flow charts) and characters.
  • Evaluation: Performance validated using the Japanese question-answering dataset JDocQA (combining text and visual information) and proprietary benchmark tools, demonstrating superior results compared to other models.
  • Deployment: Designed for on-premise environments, allowing for fine-tuning with company-specific data while maintaining security.
  • Efficiency: A 4-bit quantized version is available for resource-constrained setups.
  • Vision Encoder: Incorporates a vision encoder module to convert visual information into a format understandable by the language model.

🔮 Future ImplicationsAI analysis grounded in cited sources

Ricoh's multimodal LLM will significantly enhance efficiency in Japanese business operations.
The model's specialization in complex Japanese documents with diagrams and its on-premise deployment capability directly addresses specific needs of Japanese enterprises, improving document processing and knowledge utilization.
The technology will drive increased AI adoption in highly regulated and sensitive sectors.
The emphasis on secure on-premise deployment makes the model suitable for industries like finance and healthcare, which handle confidential data and have strict privacy requirements.
The release of the lightweight model will democratize advanced document analysis capabilities.
Making a lightweight version freely available on Hugging Face lowers the barrier to entry for businesses and developers to integrate multimodal AI for document understanding and create new applications.

Timeline

1980s
Ricoh begins AI development, focusing on image recognition and natural language processing.
2020
Ricoh announces its transformation into a digital service provider and starts leveraging NLP for office documents and customer feedback.
2022-12
Ricoh develops its proprietary AI model, 'Ricoh GPT,' based on GPT-3, in three months.
2023-03
Ricoh develops its own large language model (LLM).
2024-02
Japan's Ministry of Economy, Trade and Industry (METI) and NEDO launch the Generative AI Accelerator Challenge (GENIAC) project.
2024-08
Ricoh announces a 70-billion parameter LLM supporting Japanese, English, and Chinese.
2024-10
Ricoh is selected for the GENIAC project (Phase 2) to develop multimodal LLMs.
2025-05
Ricoh and Sompo Japan begin joint development of a private multimodal LLM for insurance operations under the GENIAC project.
2025-06
Ricoh completes development of the basic model for its multimodal LLM under GENIAC, capable of reading Japanese documents with diagrams.
2026-01
Ricoh announces the development of a multimodal LLM based on Qwen2.5-VL-32B for reading Japanese company documents with diagrams and charts under GENIAC Phase 2.
2026-02
Ricoh releases the lightweight multimodal model 'Qwen-3-VL-Ricoh-8B-20260227' for free on Hugging Face.

📎 Sources (6)

Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.

  1. Google Search Source
  2. Google Search Source
  3. Google Search Source
  4. Google Search Source
  5. Google Search Source
  6. Google Search Source
📰

Weekly AI Recap

Read this week's curated digest of top AI events →

👉Related Updates

AI-curated news aggregator. All content rights belong to original publishers.
Original source: ITmedia AI+ (日本)