Search

Few direct matches — filled in with the latest updates.

Tag: #tensor-parallel2 results

Qwen3.5 27B at 100+ t/s on 2x3090s

Qwen3.5 27B at 100+ t/s on 2x3090s

A user optimized Qwen3.5 27B dense model on 2x RTX 3090 GPUs using vLLM, achieving 100+ t/s decode, 1500 t/s prefill, and 585 t/s throughput for 8 requests. Key tweaks include tensor parallelism, MTP with 5 tokens, specific int4 quantization, and custom vLLM compilation. Scripts and a forked repo with fixes are shared for replication.

Reddit r/LocalLLaMACommunityMar 1#local-inference#quantization#mtp
⚙️

Cybersecurity Must Protect Physical Reality

As industrial systems, infrastructure, robots, and AI agents gain the ability to change physical conditions, cybersecurity must protect actions and outcomes—not only data and access. The article argues that future defenses need to evaluate whether an authorized action is appropriate for the current environment, state, and safety boundaries.