Search

Tag: #amd-gpu12 results

Zyphra Launches Efficient ZAYA1-8B MoE Model

Zyphra Launches Efficient ZAYA1-8B MoE Model

Zyphra released ZAYA1-8B, an open 8B-parameter mixture-of-experts reasoning model trained on AMD Instinct MI300 GPUs, achieving competitive benchmarks against GPT-5-High and DeepSeek-V3.2. Available under Apache 2.0 on Hugging Face for free download and customization. It features innovative MoE++ architecture for superior efficiency.

VentureBeatMediaMay 7#moe#amd-gpu#reasoning-model
⚙️

ZINC: Zig LLM Inference for AMD GPUs

ZINC is a new LLM inference engine written in Zig, enabling 35B models on $550 AMD GPUs via Vulkan. It loads GGUF models, achieves 7.1 tok/s on RDNA4, and addresses AMD consumer GPU gaps. Repo at github.com/zolotukhin/zinc.

Reddit r/LocalLLaMACommunityMar 29#amd-gpu#vulkan#gguf
DIY Tiled Attention for AMD GPUs

DIY Tiled Attention for AMD GPUs

A user built a PyTorch-based tiled attention mechanism as a flash-attention alternative for unsupported AMD MI50 GPUs (gfx906), enabling video generation without OOM. Inspired by llama.cpp, it uses query chunking, softmax fallbacks, and optimizations like BF16-to-FP16 conversion. Pure PyTorch, no custom kernels needed.

Reddit r/LocalLLaMACommunityMar 28#flash-attention#amd-gpu#tiling
Run Massive Qwen 397B on 8x R9700 GPUs

Run Massive Qwen 397B on 8x R9700 GPUs

Tutorial enables running Qwen3.5-397B-A13B with vLLM on 8x AMD R9700 GPUs using MXFP4 quantization. Achieves 30 t/s single request, up to 100 t/s batched. Includes Dockerfile patches and launch script for high throughput.

Reddit r/LocalLLaMACommunityApr 11#amd-gpu#vllm#quantization
Page 1 of 2