Search

直接匹配不多,已補上最新動態。

Tag: #kernels4 results

⚙️

Triton MoE 核心擊敗 Megablocks

純 Triton 的融合 MoE 分派核心在 Mixtral-8x7B 推理批次大小下擊敗 CUDA 優化的 Megablocks(32 個 token 時快 131%)。透過融合操作減少 35% 記憶體流量,並支援 NVIDIA 與 AMD 硬體的多模型。GitHub 有程式碼與文章。

Reddit r/MachineLearningCommunityApr 5#moe#inference#kernels
🔬

SSMs 在 25M 參數訓練中掙扎

OpenAI Parameter Golf 實證顯示,SSMs 在微型模型(25M 參數、16MB、10分訓練)壓縮比 transformer 差,且大詞彙時架構優勢喪失。部落格詳述 Mamba-3 Triton kernel 實驗,揭露融合延遲、量化 bug 與精度修正。

Reddit r/MachineLearningCommunityMay 4#compression#kernels#parameter-golf
Custom Kernels for All Users

Custom Kernels for All Users

Hugging Face introduces custom kernels powered by Codex and Claude, now available to everyone. This expands access to advanced customization options on the platform. Users can integrate these models seamlessly into their workflows.

Hugging Face BlogOfficialFeb 13#new-feature#hugging-face#codex-claude