Search

Tag: #distillation27 results

11x Token Cut for Agent Memory

11x Token Cut for Agent Memory

Structured distillation compresses personalized AI agent conversation histories into compact retrieval structures, achieving 11x token reduction from 371 to 38 tokens per exchange. Evaluated on 14k exchanges, it preserves 96% of verbatim recall and exceeds baselines in cross-layer search. Open-source implementation released.

PACED: Frontier LLM Distillation

PACED: Frontier LLM Distillation

PACED optimizes LLM distillation by targeting the zone of proximal development with Beta-weighted pass rates, avoiding compute waste on mastered or unreachable problems. It proves theoretical SNR optimality and minimax-robustness of the weighting. Empirical results show gains in forward KL, self-distillation, and two-stage schedules on reasoning benchmarks.

ArXiv AIResearchMar 13#distillation#model-training
🤖

Savant Commander 48B: 12-Distill MOE

Savant Commander 48B is a custom Qwen3-based 4x12B MOE merging distills from Claude, Gemini, OpenAI, DeepSeek, and more with hand-coded routing. Users control activation via prompts and test differences easily. GGUF versions include regular and Heretic uncensored, available on Hugging Face.

Reddit r/LocalLLaMACommunityMar 24#moe#distillation#uncensored
🤖

Uncensored Qwen 3.5 9B Distilled from Claude Opus

Community creator released a fully uncensored 9B Qwen model by merging uncensored tensors from HauhauCS with Claude Opus reasoning distillation from Jackrong. It eliminates refusals, boosts creativity for roleplay and image prompts, with thinking disabled via custom chat template. GGUF version optimized for RTX 3060 in LM Studio.

Reddit r/LocalLLaMACommunityMar 15#uncensored#distillation#local-llm
Page 2 of 3