Search

10 results on this page

Transfer More Knowledge with Less Multilingual Data

Transfer More Knowledge with Less Multilingual Data

Apple Machine Learning presents a lexical-intervention approach for improving cross-lingual knowledge transfer when target-language data is scarce. The work targets downstream capabilities such as scientific reasoning, commonsense inference, and world knowledge without relying heavily on parallel data, translation systems, or auxiliary models.

Apple Machine LearningOfficial1d ago#multilingual-models
Z.ai Reframes Scaling Beyond Parameter Counts

Z.ai Reframes Scaling Beyond Parameter Counts

Z.ai argues that model scaling should account for data, compute allocation, inference cost, sparsity, effective depth, and post-training—not parameters alone. The post presents GLM-5.3 as a controlled experiment using the same total and activated parameters as GLM-5.2 while scaling long-horizon environments and reinforcement learning for one month.

Reddit r/LocalLLaMACommunity1d ago#scaling-laws#mixture-of-experts#post-training
🤖

Fine-Tuning Gemma for Complex Legal Reasoning

A practitioner reports that fine-tuning Gemma 4 26B A4B on 100,000 court decisions failed to outperform a prompted base model for generating legal principles. The discussion explores whether the bottleneck lies in data quality, evaluation design, task complexity, or the fine-tuning workflow.

Reddit r/MachineLearningCommunity1d ago#legal-ai#fine-tuning#evaluation
Page 2