Search

Tag: #on-device-inference14 results

LFM2.5 Runs at 17 Tok/s on OnePlus 13

LFM2.5 Runs at 17 Tok/s on OnePlus 13

A developer demonstrated LFM2.5-2.6B running at approximately 17 tokens per second on a OnePlus 13 using only the phone’s CPU. The setup uses a Q4_K_M GGUF model, a custom 450 KB inference engine, and an ADB-based device probe suite, with a target of around 30 tokens per second.

Reddit r/LocalLLaMACommunityAug 5#on-device-inference#android#gguf
Qwen 0.8B Runs on Old S10E at 12 t/s

Qwen 0.8B Runs on Old S10E at 12 t/s

Qwen released its new 0.8B model, which runs locally on a 7-year-old Samsung S10E at 12 tokens per second using llama.cpp and Termux. After resolving missing C libraries, it handles conversations and complex tasks effectively. This demo highlights powerful tiny LLMs on edge devices.

Reddit r/LocalLLaMACommunityMar 2#on-device-inference#mobile-llm#tiny-models
Page 1 of 2