McKinsey: Tech value matters more than L3/L4 labels
💡Insights on how VLM and AI are disrupting traditional autonomous driving roadmaps and consumer expectations.
⚡ 30-Second TL;DR
What Changed
69% of Chinese consumers view City NOA as a standard feature for new vehicles.
Why It Matters
Shifting focus to value-based metrics will likely accelerate the adoption of end-to-end AI models in automotive software development.
What To Do Next
If building for automotive, prioritize integrating VLM-based perception modules over legacy rule-based classification systems.
Key Points
- •69% of Chinese consumers view City NOA as a standard feature for new vehicles.
- •Industry experts suggest VLM (Vision Language Models) are bridging the gap between L2 and L4.
- •McKinsey advises focusing on user-perceived value rather than adhering to traditional SAE autonomy levels.
- •Consumer demand is shifting from price-driven to value-driven, favoring technical innovation.
🧠 Deep Insight
Web-grounded analysis with 16 cited sources.
🔑 Enhanced Key Takeaways
- •The adoption of City NOA (Navigation on Autopilot) in China has rapidly expanded, with cumulative sales reaching 3.129 million units from January to November 2025, achieving a 15.1% penetration rate in insured passenger vehicles.
- •City NOA has transitioned from a premium feature to a mainstream offering, with over 68.9% of sales in passenger vehicles priced below 300,000 yuan equipped with this functionality by late 2025.
- •Vision-Language Models (VLMs) enhance autonomous vehicles by integrating computer vision and natural language processing to interpret multimodal data, improving scene understanding, object recognition, and human-vehicle interaction, particularly for complex 'long tail problems' and ambiguous scenarios.
- •McKinsey's 2025 survey of autonomous vehicle experts indicated that adoption timelines for Level 4 (L4) and Level 5 (L5) autonomous vehicles have slipped by an average of one to two years compared to their 2023 projections, with the global rollout of L4 robotaxis now anticipated by 2030 instead of 2029.
- •High development costs are identified as the primary pain point in the Advanced Driver-Assistance Systems (ADAS) pipeline, signaling a shift in industry focus from technological development to cost-effective deployment.
🛠️ Technical Deep Dive
- Vision-Language Models (VLMs):
- Integrate computer vision (CV) and natural language processing (NLP) to process multimodal data from sensors like cameras and LiDAR, along with textual or semantic information (e.g., road signs, user commands).
- Enable enhanced scene understanding and object recognition by linking visual patterns to language-based concepts, allowing for more accurate interpretation of complex driving scenarios, including temporary detour signs with handwritten text or distinguishing pedestrian intent.
- Improve handling of 'long tail problems' and ambiguous situations that traditional rule-based systems struggle with, by leveraging pre-training on large-scale internet data for foundational world understanding.
- Facilitate natural human-vehicle interaction through language interfaces and support safety-critical decision-making by generating semantic explanations of vehicle actions.
- Example: DriveVLM, a project by Li Auto and Tsinghua University, employs a vision transformer encoder alongside a large language model (LLM) to generate detailed linguistic descriptions of the environment.
- Challenges include real-time processing of continuous, high-dimensional video streams, advanced 3D scene understanding, and inference latency (e.g., DriveVLM showed a 1.9-second processing time for a single scene).
- City NOA (Navigation on Autopilot) in China:
- Moving towards 'mapless' solutions for mass production deployment, as maintaining high-definition (HD) maps for city roads is challenging and costly.
- Utilizes Transformer-based Bird's Eye-View (BEV) in the perception layer to aggregate context from different sensor inputs in a unified space.
- Increasingly adopting learning-based, end-to-end paradigms for prediction and planning, moving away from traditional sophisticated rule-based designs that become inefficient as features roll out to more cities.
🔮 Future ImplicationsAI analysis grounded in cited sources
⏳ Timeline
📎 Sources (16)
Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.
- Google Search Source
- Google Search Source
- Google Search Source
- Google Search Source
- Google Search Source
- Google Search Source
- Google Search Source
- Google Search Source
- Google Search Source
- Google Search Source
- Google Search Source
- Google Search Source
- Google Search Source
- Google Search Source
- Google Search Source
- Google Search Source
Weekly AI Recap
Read this week's curated digest of top AI events →
👉Related Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: 虎嗅 ↗
