Layer Duplication Tops Open LLM Leaderboard

๐กTop LLM leaderboard #1 via simple layer dupโno retraining, just 2x 4090s!
โก 30-Second TL;DR
What Changed
Duplicated exactly 7 middle layers for top leaderboard score
Why It Matters
Enables performance boosts for open models without costly retraining. Suggests pretraining creates preservable functional circuits. Democratizes leaderboard-topping capabilities for hobbyists with consumer GPUs.
What To Do Next
Duplicate 7 middle layers in Qwen2-72B and test on Open LLM Leaderboard benchmarks using dual 4090s.
Key Points
- โขDuplicated exactly 7 middle layers for top leaderboard score
- โขNo weight modifications needed; works only for ~7-layer circuit blocks
- โขDeveloped on 2x RTX 4090s; top 4 leaderboard models are descendants
- โขBlog details findings; code and Qwen3.5 RYS versions soon
๐ง Deep Insight
Background and context from public sources โ not the original article. 8 sources cited.
๐ Enhanced Key Takeaways
- โขThe RYS (layer duplication) technique is orthogonal to fine-tuning and weight optimization, enabling stacking with other training methods like ORPO and instruction tuning to achieve even higher leaderboard scores[1].
- โขLayer duplication works only for ~7-layer 'circuit blocks' discovered through orthogonal probing; single-layer duplication has no effect, suggesting pretraining carves discrete functional circuits that must be preserved whole[1][3].
- โขAs of early 2026, the top four Open LLM Leaderboard models are all descendants of the original RYS technique, with MaziyarPanahi's calme-3.2-instruct-78b achieving 52.08 points[1].
๐ ๏ธ Technical Deep Dive
- โขLayer duplication mechanism: Duplicates ~7-layer functional blocks within the model architecture without modifying weights, providing additional iterations through the model's internal reasoning space[1].
- โขOrthogonal probing approach: Two narrow, orthogonal probes were used to identify optimal layer configurations that generalized across diverse benchmarks[1].
- โขHardware efficiency: Developed on 2x RTX 4090 GPUs, demonstrating that significant architectural improvements can be achieved without enterprise-scale compute[1].
- โขStacking compatibility: The method combines with fine-tuning (calme-2.4-rys-78b), ORPO training (CalmeRys-78B-Orpo-v0.1), and instruction tuning (calme-3.1, calme-3.2) for cumulative performance gains[1].
- โขCircuit discovery: Empirical testing revealed that too few layers (single-layer) produce no effect, while too many layers degrade performance, indicating pretraining creates discrete functional units of approximately 7 layers[3].
๐ฎ Future ImplicationsAI analysis grounded in cited sources
โณ Timeline
๐ Sources (8)
Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.
- dnhkng.github.io โ Rys
- o-mega.ai โ Top 10 Open Source Llms the Deepseek Revolution 2026
- news.ycombinator.com โ Item
- onyx.app โ Open LLM Leaderboard
- magazine.sebastianraschka.com โ A Dream of Spring for Open Weight
- Hugging Face โ Open LLM Leaderboard
- contabo.com โ Open Source Llms
- stormap.ai โ Latest Open Source AI Model Releases
Weekly AI Recap
Read this week's curated digest of top AI events โ
๐Related Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: Reddit r/LocalLLaMA โ
This is a summary, not the original. Read the source, or get the weekly briefing.
Weekly AI briefing
One email a week. Unsubscribe anytime.