SourceStalecollected in 26h

Predict GPT-2 Edges from Weights Alone

PostLinkedIn
🤖Read original on Reddit r/MachineLearning
#path-patching#virtual-weightscheap-anchor

💡125x faster edge importance prediction for GPT-2 circuits from weights alone – interpretability breakthrough.

⚡ 30-Second TL;DR

What Changed

ρ=0.623 Spearman correlation with path patching

Why It Matters

Enables fast prioritization of edges for investigation or pruning in transformer circuits, saving compute on causal scrutiny. Promising for scaling mechanistic interpretability.

What To Do Next

Compute Cheap Anchor scores on your transformer model's induction heads using the described spectral and path metrics.

Who should care:Researchers & Academics

Key Points

  • ρ=0.623 Spearman correlation with path patching
  • 125x speedup, computable in 2 seconds from weights
  • Uses spectral concentration and downstream path weight
  • Outperforms weight magnitude (ρ=0.070) and gradients
  • Limited to known GPT-2 induction circuit edges
📰

Weekly AI Recap

Read this week's curated digest of top AI events →

👉Related Updates

AI-curated news aggregator. All content rights belong to original publishers.
Original source: Reddit r/MachineLearning

This is a summary, not the original. Read the source, or get the weekly briefing.

The weekly digest

One email a week. Unsubscribe anytime.