來源較早收集於 2h

清華院士教授領銜Qujing AI Token新紀元

清華院士教授領銜Qujing AI Token新紀元
PostLinkedIn
閱讀原文: 雷峰网
#ai-inference#token-production#hpc#prod-learn-fusionqujing-techzheng-weiminwu-yongweitsinghuagl-ventures

💡清華精英加盟超效AI推理—關鍵降本部署利器。(28字)

⚡ 30 秒速覽

有什麼變化

鄭緯民(中國工程院院士、清華教授)出任首席科學顧問。

為什麼重要

頂尖學者強化Qujing Tech於AI基礎設施的研發,提升大模型推理效率。此舉透過產學融合助中國AI競爭,利企業擴展性。

下一步行動

試用Qujing Tech推理平台,優化每GPU AI Token產出。

誰應關注:Enterprise & Security Teams

關鍵要點

  • 鄭緯民(中國工程院院士、清華教授)出任首席科學顧問。
  • 武永衛(IEEE Fellow、清華教授)任首席科學家。
  • 首創「以存換算」及異構協同技術,Token產出倍增。
  • 獲高瓴創投、清華校友基金等投資。
  • 專注統一AI推理部署中算力碎片化問題。

🧠 深度解析

本篇為 AI 生成分析,非原文內容。

🔑 增強重點摘要

  • Qujing Tech's core technology, often referred to as 'AI-native storage' or 'compute-storage synergy,' specifically targets the memory wall bottleneck in Large Language Model (LLM) inference by optimizing data movement between HBM and system memory.
  • The company's strategic focus is on the 'inference-as-a-service' market, aiming to reduce the Total Cost of Ownership (TCO) for enterprise-grade AI deployments by increasing Token-per-second (TPS) throughput on existing GPU clusters.
  • The involvement of Zheng Weimin and Wu Yongwei signals a strong alignment with China's national 'East Data, West Computing' (Dongshu Xisuan) strategy, positioning Qujing to provide infrastructure software for large-scale, distributed AI data centers.
📊 競品分析▸ Show
FeatureQujing TechvLLM (Open Source)NVIDIA TensorRT-LLM
Primary FocusHeterogeneous compute-storage synergyPagedAttention memory managementHardware-specific kernel optimization
DeploymentEnterprise-grade infrastructureResearch/General purposeNVIDIA-exclusive hardware
Key AdvantageMemory-compute bottleneck reductionHigh flexibility/community supportMaximum hardware utilization

🛠️ 技術深入

  • Heterogeneous Synergy Architecture: Utilizes a tiered memory management system that dynamically swaps model weights between GPU VRAM and high-speed system memory to accommodate models larger than available VRAM.
  • Token Production Optimization: Implements custom kernels that overlap data transfer (I/O) with compute operations, effectively hiding latency during the KV-cache generation phase of inference.
  • Fragmentation Unification: Employs a software-defined abstraction layer that aggregates disparate compute resources (CPU/GPU/NPU) into a unified inference pool, reducing the overhead of managing fragmented hardware clusters.

🔮 前景展望基於引用來源的 AI 分析

Qujing Tech will likely pursue a partnership with major Chinese cloud providers to integrate their inference engine into public cloud offerings.
The company's focus on enterprise-scale inference and the backing of GL Ventures suggests a strategy of scaling through existing cloud infrastructure providers.
The company will face significant pressure to maintain performance parity as NVIDIA releases newer generations of hardware with larger HBM capacities.
As hardware-level memory bandwidth increases, the relative advantage of software-based 'compute-storage synergy' may diminish, forcing the company to innovate further up the stack.

時間線

2024-05
Qujing Tech completes early-stage financing round led by GL Ventures.
2025-02
Zheng Weimin and Wu Yongwei officially join the company in advisory and scientific leadership roles.
2026-01
Qujing Tech achieves milestone in heterogeneous synergy performance, demonstrating multi-fold Token gains in internal benchmarks.
📰

AI 週報

閱讀本週精選 AI 大事摘要 →

👉相關動態

AI 策展新聞聚合。所有內容版權歸原始發布者所有。
原始來源: 雷峰网

這是摘要,不是原文。去看原站,或訂閱每週簡報。

每週電子報

每週一封,可隨時退訂。