Search

Tag: #moe-models14 results

Chinese Open Models Power Global AI Tools

Chinese Open Models Power Global AI Tools

Cursor's Composer 2 uses Moonshot's open-source Kimi K2.5 model via Fireworks AI, sparking discussions on China's AI supply chain. DeepSeek and Kimi models are increasingly foundational for global apps like OpenClaw agents. This shift mirrors manufacturing supply chains, with tokens as the new infrastructure.

🤖

Current Chinese LLM Landscape Overview

Detailed breakdown of China's LLM scene featuring ByteDance's Doubao as proprietary leader and Alibaba's Qwen excelling in open-weight models. Deepseek innovates with MLA and other tech, while 'Six Small Tigers' like Zhipu and Minimax release large open-weight MoE models for cheap inference. Meituan pushes aggressive open releases like 562B LongCat-Flash.

Reddit r/LocalLLaMACommunityMar 23#chinese-llms#open-weight#moe-models
China's AI Compute Independence Push

China's AI Compute Independence Push

DeepSeek announces V4 multimodal model using fully domestic chips, bypassing Nvidia. Chinese firms optimize algorithms like MoE to cut costs and advance国产 chips for training. This marks a shift from inference to full training on non-Nvidia hardware.

虎嗅MediaMar 4#china-ai#chips#moe-models
V100 Benchmarks: Power Limits & Offload Tested

V100 Benchmarks: Power Limits & Offload Tested

6-hour benchmarks on V100 32GB GPU across 20 LLM models (MoE & dense) with Ryzen 7600X, testing power limits from 300W-150W and CPU offload levels up to 32K context. Power limiting to 200W yields <2% loss for generation; MoE models handle offload far better than dense. Top performers: Nemotron-30B Mamba2 at 152 t/s.

Reddit r/LocalLLaMACommunityMar 28#v100-benchmarks#cpu-offload#moe-models
Page 1 of 2