SourceStalecollected in 60m

74% RAM cut for SmolLM2 on Galaxy Watch

PostLinkedIn
🦙Read original on Reddit r/LocalLLaMA
#edge-ai#memory-optimization#wearables#androidllama.cppsmol(lm2-360mllama.cppsamsung-galaxy-watch-4ggml

💡Run 360M LLM on 380MB watch RAM – game-changer for edge inference

⚡ 30-Second TL;DR

What Changed

74% RAM reduction: 524MB to 142MB on 380MB device

Why It Matters

Enables ultra-low-resource LLM inference on wearables, expanding edge AI applications. Potential upstream PR to main llama.cpp repo could benefit broader embedded deployments.

What To Do Next

Clone the axon-dev branch from Perinban/llama.cpp and test on low-RAM Android devices.

Who should care:Developers & AI Engineers

Key Points

  • 74% RAM reduction: 524MB to 142MB on 380MB device
  • Boot time improved: first 19s to 11s, second ~2.5s
  • Fix via host_ptr and mmap for CPU tensors; Vulkan tensors copied
  • Code at github.com/Perinban/llama.cpp tree axon-dev
📰

Weekly AI Recap

Read this week's curated digest of top AI events →

👉Related Updates

AI-curated news aggregator. All content rights belong to original publishers.
Original source: Reddit r/LocalLLaMA

This is a summary, not the original. Read the source, or get the weekly briefing.

The weekly digest

One email a week. Unsubscribe anytime.