๐Ÿฆ™Stalecollected in 6h

Layer Duplication Tops Open LLM Leaderboard

Layer Duplication Tops Open LLM Leaderboard
PostLinkedIn
๐Ÿฆ™Read original on Reddit r/LocalLLaMA
#layer-duplication#benchmark-beat#consumer-gpuqwen2-72bqwen2-72bopen-llm-leaderboardrtx-4090gladosgh200

๐Ÿ’กTop LLM leaderboard #1 via simple layer dupโ€”no retraining, just 2x 4090s!

โšก 30-Second TL;DR

What Changed

Duplicated exactly 7 middle layers for top leaderboard score

Why It Matters

Enables performance boosts for open models without costly retraining. Suggests pretraining creates preservable functional circuits. Democratizes leaderboard-topping capabilities for hobbyists with consumer GPUs.

What To Do Next

Duplicate 7 middle layers in Qwen2-72B and test on Open LLM Leaderboard benchmarks using dual 4090s.

Who should care:Researchers & Academics

Key Points

  • โ€ขDuplicated exactly 7 middle layers for top leaderboard score
  • โ€ขNo weight modifications needed; works only for ~7-layer circuit blocks
  • โ€ขDeveloped on 2x RTX 4090s; top 4 leaderboard models are descendants
  • โ€ขBlog details findings; code and Qwen3.5 RYS versions soon

๐Ÿง  Deep Insight

Background and context from public sources โ€” not the original article. 8 sources cited.

๐Ÿ”‘ Enhanced Key Takeaways

  • โ€ขThe RYS (layer duplication) technique is orthogonal to fine-tuning and weight optimization, enabling stacking with other training methods like ORPO and instruction tuning to achieve even higher leaderboard scores[1].
  • โ€ขLayer duplication works only for ~7-layer 'circuit blocks' discovered through orthogonal probing; single-layer duplication has no effect, suggesting pretraining carves discrete functional circuits that must be preserved whole[1][3].
  • โ€ขAs of early 2026, the top four Open LLM Leaderboard models are all descendants of the original RYS technique, with MaziyarPanahi's calme-3.2-instruct-78b achieving 52.08 points[1].

๐Ÿ› ๏ธ Technical Deep Dive

  • โ€ขLayer duplication mechanism: Duplicates ~7-layer functional blocks within the model architecture without modifying weights, providing additional iterations through the model's internal reasoning space[1].
  • โ€ขOrthogonal probing approach: Two narrow, orthogonal probes were used to identify optimal layer configurations that generalized across diverse benchmarks[1].
  • โ€ขHardware efficiency: Developed on 2x RTX 4090 GPUs, demonstrating that significant architectural improvements can be achieved without enterprise-scale compute[1].
  • โ€ขStacking compatibility: The method combines with fine-tuning (calme-2.4-rys-78b), ORPO training (CalmeRys-78B-Orpo-v0.1), and instruction tuning (calme-3.1, calme-3.2) for cumulative performance gains[1].
  • โ€ขCircuit discovery: Empirical testing revealed that too few layers (single-layer) produce no effect, while too many layers degrade performance, indicating pretraining creates discrete functional units of approximately 7 layers[3].

๐Ÿ”ฎ Future ImplicationsAI analysis grounded in cited sources

Layer duplication may represent a new scaling paradigm orthogonal to model size and fine-tuning.
The technique achieves top leaderboard performance by optimizing internal reasoning iterations rather than adding parameters or knowledge, suggesting future models may balance width, depth, and reasoning iterations differently.
Discrete functional circuits in transformer architectures are a fundamental design principle exploitable for efficiency gains.
The discovery that only ~7-layer blocks work suggests pretraining naturally organizes layers into functional units, which could inform future architecture design and layer allocation strategies.

โณ Timeline

2026-02
RYS layer duplication technique achieves top position on Open LLM Leaderboard with Qwen2-72B variant
2026-02
MaziyarPanahi releases calme-2.4-rys-78b by fine-tuning RYS-XLarge, advancing leaderboard rankings
2026-02
dfurman applies ORPO training to calme-2.4-rys-78b, producing CalmeRys-78B-Orpo-v0.1 (51.23 score)
2026-02
MaziyarPanahi releases calme-3.1-instruct-78b (51.29 score) and calme-3.2-instruct-78b (52.08 score), occupying top two leaderboard positions
๐Ÿ“ฐ

Weekly AI Recap

Read this week's curated digest of top AI events โ†’

๐Ÿ‘‰Related Updates

AI-curated news aggregator. All content rights belong to original publishers.
Original source: Reddit r/LocalLLaMA โ†—

This is a summary, not the original. Read the source, or get the weekly briefing.

Weekly AI briefing

One email a week. Unsubscribe anytime.