Source量子位•Stalecollected in 2h
Meituan's Token-Based Native Multimodal

Meituan's token multimodal has no ceiling—rethink your vision tokenizers now
30-Second TL;DR
What Changed
Images and speech tokenized like text
Why It Matters
Pushes boundaries for efficient multimodal LLMs, potentially lowering compute for vision-audio integration.
What To Do Next
Prototype image tokenization in your Llama-based multimodal fine-tune using Hugging Face.
Who should care:Developers & AI Engineers
Key Points
- •Images and speech tokenized like text
- •Native multimodal prediction pipeline
- •Discrete vision scalability unlimited
- •Aggressive approach by Meituan
Deep Insight
AI-generated analysis for this event — not the original article.
Enhanced Key Takeaways
- •Meituan's approach utilizes a unified tokenizer that maps diverse modalities into a shared discrete latent space, enabling the model to process visual and auditory inputs using the same transformer architecture as text.
- •The research emphasizes the elimination of modality-specific encoders, aiming to reduce computational overhead and latency by treating all inputs as a continuous stream of tokens for autoregressive prediction.
- •This architecture is specifically optimized for Meituan's local service ecosystem, targeting real-time multimodal understanding for tasks like food delivery logistics, autonomous vehicle navigation, and customer service automation.
Competitor Analysis
Architecture
- Meituan (Native Multimodal)
- Unified Discrete Tokenization
- OpenAI (GPT-4o)
- Native Multimodal (End-to-End)
- Google (Gemini 1.5)
- Multimodal Mixture-of-Experts
Primary Focus
- Meituan (Native Multimodal)
- Local Services/Logistics
- OpenAI (GPT-4o)
- General Purpose/Conversational
- Google (Gemini 1.5)
- General Purpose/Long Context
Inference Strategy
- Meituan (Native Multimodal)
- Discrete Token Prediction
- OpenAI (GPT-4o)
- End-to-End Latent Space
- Google (Gemini 1.5)
- Cross-Modal Attention
| Feature | Meituan (Native Multimodal) | OpenAI (GPT-4o) | Google (Gemini 1.5) |
|---|---|---|---|
| Architecture | Unified Discrete Tokenization | Native Multimodal (End-to-End) | Multimodal Mixture-of-Experts |
| Primary Focus | Local Services/Logistics | General Purpose/Conversational | General Purpose/Long Context |
| Inference Strategy | Discrete Token Prediction | End-to-End Latent Space | Cross-Modal Attention |
Technical Deep Dive
- Unified Tokenization: Employs a VQ-VAE (Vector Quantized Variational Autoencoder) variant to discretize image patches and audio frames into a shared vocabulary.
- Architecture: Utilizes a decoder-only Transformer backbone that operates on a unified sequence of tokens, regardless of the source modality.
- Training Objective: Implements a next-token prediction task across the combined sequence, allowing the model to learn cross-modal dependencies implicitly.
- Scalability: The discrete nature of the tokens allows for the application of standard scaling laws, theoretically avoiding the performance plateaus associated with continuous feature alignment methods.
Future ImplicationsAI analysis grounded in cited sources
Meituan will integrate this model into its autonomous delivery fleet by Q4 2026.
The native multimodal architecture's low-latency token prediction is critical for real-time obstacle detection and navigation in complex urban environments.
The model will reduce cloud-side inference costs by at least 30% compared to traditional multi-encoder architectures.
Removing modality-specific encoders simplifies the inference pipeline and increases hardware utilization efficiency on standardized GPU clusters.
Timeline
2024-05
Meituan establishes a dedicated research lab for multimodal foundation models.
2025-02
Initial internal testing of discrete tokenization for image-to-text tasks.
2025-11
Meituan publishes internal white paper on unified token modeling for local services.
2026-03
Official announcement of the native multimodal token-based framework.
- 2024-05Meituan establishes a dedicated research lab for multimodal foundation models.
- 2025-02Initial internal testing of discrete tokenization for image-to-text tasks.
- 2025-11Meituan publishes internal white paper on unified token modeling for local services.
- 2026-03Official announcement of the native multimodal token-based framework.
Weekly AI Recap
Read this week's curated digest of top AI events →
AI-curated news aggregator. All content rights belong to original publishers.
Original source: 量子位 ↗
This is a summary, not the original. Read the source, or get the weekly briefing.
The weekly digest
One email a week. Unsubscribe anytime.