Papers With Code Adds Support for Closed-Source Model Evals

💡Easily compare your open-source AI models against the latest closed-source benchmarks on a unified leaderboard.
⚡ 30-Second TL;DR
What Changed
Added support for closed-source model benchmarks on leaderboards.
Why It Matters
This update improves transparency in AI benchmarking by acknowledging the dominance of closed-source models while maintaining a clear distinction for open-source research.
What To Do Next
Visit paperswithcode.co and use the toggle settings to compare your open-source model's performance against the latest closed-source benchmarks.
Key Points
- •Added support for closed-source model benchmarks on leaderboards.
- •Introduced a toggle to filter out closed-source models for open-model-only views.
- •Allows submission of non-arXiv sources like blog posts for model evaluations.
- •Provides scatter plots and tables for benchmarking state-of-the-art performance.
🧠 Deep Insight
Background and context from public sources — not the original article. 7 sources cited.
🔑 Enhanced Key Takeaways
- •The current Papers With Code platform (paperswithcode.co) is a community-driven revival, launched around May 2026, following the shutdown of the original paperswithcode.com by Meta in July 2025.
- •This resurrected platform is spearheaded by Hugging Face's Niels Rogge, in collaboration with Meta AI and the original creators, aiming to restore a vital resource for the machine learning community.
- •The new Papers With Code leverages AI agents to automatically parse research papers and generate leaderboards, with human verification in place to ensure data quality and accuracy.
- •Beyond traditional arXiv sources, the platform now supports external paper submissions and integrates with Hugging Face for trending papers and model hubs, enhancing the discoverability and reproducibility of research.
📊 Competitor Analysis▸ Show
| Feature / Platform | Papers With Code (paperswithcode.co) | Hugging Face Leaderboards | CodeSOTA |
|---|---|---|---|
| SOTA Leaderboards | Comprehensive, supports open & closed models, toggle filter [article] | Limited, primarily for models on Hugging Face Hub (e.g., Open LLM Leaderboard) | Comprehensive, starting with OCR/document AI, independently verified results |
| Code Links | Yes, emphasis on working implementations | Yes, for models on Hugging Face Hub | Yes, reproducible code with verification (signed container hash) |
| Dataset Registry | Yes, restored from original PWC | Different focus, not a comprehensive registry | Per-domain, fresh and maintained |
| Closed-Source Model Evals | Yes, newly added support with special 'closed' tag [article, cite: 14] | Generally requires public models on Hub, limiting closed-source visibility | Yes, publishes dated source link, prompt template, and vendor model-card snapshot for closed APIs |
| Submission Sources | Supports non-arXiv sources like blog posts [article], external papers | Primarily focused on arXiv for trending papers, models on Hub | Focuses on traceable benchmarks, publishes source links |
| Pricing | Free and open resource | Free (for basic use) | Free (for basic use) |
🛠️ Technical Deep Dive
- The platform utilizes AI agents for parsing research papers at scale and automatically generating leaderboards.
- Human verification is incorporated into the process to maintain the quality and accuracy of the data.
- It integrates with arXiv for academic papers and Hugging Face for trending papers and open-source model hubs.
- The system tracks metrics such as GitHub star velocity for trending papers to highlight high-impact research.
- Support for multiple code repositories linked to a single paper and external, non-arXiv paper sources is implemented.
- User authentication and data storage for thumbnails, paper PDFs, and backups are handled via 'Sign in with HF' (Hugging Face) and Storage Buckets.
🔮 Future ImplicationsAI analysis grounded in cited sources
⏳ Timeline
📎 Sources (7)
Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.
Weekly AI Recap
Read this week's curated digest of top AI events →
👉Related Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: Reddit r/MachineLearning ↗
This is a summary, not the original. Read the source, or get the weekly briefing.
The weekly digest
One email a week. Unsubscribe anytime.