Search

7 results on this page

A Smaller Qwen3.8 Built by Pruning Layers

A Smaller Qwen3.8 Built by Pruning Layers

A community developer created Qwen3.8-23B-Mini-Me by strategically removing layers from Qwen3.8-27B, reducing the model to approximately 22.7B parameters without severe reasoning degradation. The model is reported to work well for coding, agentic tasks, and multi-turn chats, but it has not yet been benchmarked and struggles more with edge cases and underspecified prompts.

Reddit r/LocalLLaMACommunity1d ago#model-pruning#model-compression#apple-silicon
🤖

Fine-Tuning Gemma for Complex Legal Reasoning

A practitioner reports that fine-tuning Gemma 4 26B A4B on 100,000 court decisions failed to outperform a prompted base model for generating legal principles. The discussion explores whether the bottleneck lies in data quality, evaluation design, task complexity, or the fine-tuning workflow.

Reddit r/MachineLearningCommunity1d ago#legal-ai#fine-tuning#evaluation
Page 2