Search

Tag: #gpu-acceleration26 results

⚙️

Moore Threads Hits Nvidia FP8 Parity

Moore Threads achieves systemic FP8 breakthrough, syncing with Nvidia globally as one of few domestic GPU makers. MTT S5000 delivers 1000 TFLOPS FP8 AI compute with 80GB VRAM and full precision support. MUSA platform enables compatibility with PyTorch, vLLM, and other AI frameworks.

IT之家MediaMar 13#fp8#gpu-acceleration#ai-hardware
Bytedance's AI Agent Writes CUDA Code

Bytedance's AI Agent Writes CUDA Code

Import AI 448 highlights ongoing AI R&D, including Bytedance's new agent capable of writing CUDA code for GPU tasks. It also covers advancements in on-device AI for satellites. The issue questions when the first major AI war might occur, paralleling Ukraine's drone war.

📰

ibu-boost: GBDT with Absolute Split Rejection

ibu-boost is a new gradient-boosted tree library applying a screening transform to absolutely reject poor splits, eliminating the need for min_gain_to_split tuning. It supports non-oblivious and oblivious trees, MSE/binary loss, missing values, and Triton GPU kernels for speedups. Benchmarks on California Housing show 12% RMSE gap to LightGBM but 3x GPU acceleration.

Reddit r/MachineLearningCommunityApr 10#gbdt#gradient-boosting#screening-transform
Page 2 of 3