SourceStalecollected in 2h

ICLR 2025 Oral Paper Flaws SQL Eval

PostLinkedIn
🤖Read original on Reddit r/MachineLearning
#sql-generation#eval-flaws#conference-reviewsql-code-generation-llmiclr-2025openreviewllm

💡Exposes eval flaw in top ICLR paper—critical for code LLM researchers

⚡ 30-Second TL;DR

What Changed

Paper uses NL metrics for SQL eval, not execution-based.

Why It Matters

Highlights risks of flawed evals in ML conferences, urging better benchmarks for code gen tasks.

What To Do Next

Review the paper at openreview.net/forum?id=GGlpykXDCa and replicate SQL eval tests.

Who should care:Researchers & Academics

Key Points

  • Paper uses NL metrics for SQL eval, not execution-based.
  • Tests show 20% false positive rate in evaluation.
  • Questions acceptance as ICLR 2025 Oral despite flaw.
📰

Weekly AI Recap

Read this week's curated digest of top AI events →

👉Related Updates

AI-curated news aggregator. All content rights belong to original publishers.
Original source: Reddit r/MachineLearning

This is a summary, not the original. Read the source, or get the weekly briefing.

The weekly digest

One email a week. Unsubscribe anytime.