Fine-Tuning Gemma for Complex Legal Reasoning
A practitioner reports that fine-tuning Gemma 4 26B A4B on 100,000 court decisions failed to outperform a prompted base model for generating legal principles. The discussion explores whether the bottleneck lies in data quality, evaluation design, task complexity, or the fine-tuning workflow.






