SourceReddit r/MachineLearning•Stalecollected in 2h
ICLR 2025 Oral Paper Flaws SQL Eval
#sql-generation#eval-flaws#conference-reviewsql-code-generation-llmiclr-2025openreviewllm
💡Exposes eval flaw in top ICLR paper—critical for code LLM researchers
⚡ 30-Second TL;DR
What Changed
Paper uses NL metrics for SQL eval, not execution-based.
Why It Matters
Highlights risks of flawed evals in ML conferences, urging better benchmarks for code gen tasks.
What To Do Next
Review the paper at openreview.net/forum?id=GGlpykXDCa and replicate SQL eval tests.
Who should care:Researchers & Academics
Key Points
- •Paper uses NL metrics for SQL eval, not execution-based.
- •Tests show 20% false positive rate in evaluation.
- •Questions acceptance as ICLR 2025 Oral despite flaw.
📰
Weekly AI Recap
Read this week's curated digest of top AI events →
👉Related Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: Reddit r/MachineLearning ↗
This is a summary, not the original. Read the source, or get the weekly briefing.
The weekly digest
One email a week. Unsubscribe anytime.