Papers With Code Adds Support for Closed-Source Model Evals

๐กEasily compare your open-source AI models against the latest closed-source benchmarks on a unified leaderboard.
โก 30-Second TL;DR
What Changed
Added support for closed-source model benchmarks on leaderboards.
Why It Matters
This update improves transparency in AI benchmarking by acknowledging the dominance of closed-source models while maintaining a clear distinction for open-source research.
What To Do Next
Visit paperswithcode.co and use the toggle settings to compare your open-source model's performance against the latest closed-source benchmarks.
Key Points
- โขAdded support for closed-source model benchmarks on leaderboards.
- โขIntroduced a toggle to filter out closed-source models for open-model-only views.
- โขAllows submission of non-arXiv sources like blog posts for model evaluations.
- โขProvides scatter plots and tables for benchmarking state-of-the-art performance.
๐ง Deep Insight
Web-grounded analysis with 7 cited sources.
๐ Enhanced Key Takeaways
- โขThe current Papers With Code platform (paperswithcode.co) is a community-driven revival, launched around May 2026, following the shutdown of the original paperswithcode.com by Meta in July 2025.
- โขThis resurrected platform is spearheaded by Hugging Face's Niels Rogge, in collaboration with Meta AI and the original creators, aiming to restore a vital resource for the machine learning community.
- โขThe new Papers With Code leverages AI agents to automatically parse research papers and generate leaderboards, with human verification in place to ensure data quality and accuracy.
- โขBeyond traditional arXiv sources, the platform now supports external paper submissions and integrates with Hugging Face for trending papers and model hubs, enhancing the discoverability and reproducibility of research.
๐ Competitor Analysisโธ Show
| Feature / Platform | Papers With Code (paperswithcode.co) | Hugging Face Leaderboards | CodeSOTA |
|---|---|---|---|
| SOTA Leaderboards | Comprehensive, supports open & closed models, toggle filter [article] | Limited, primarily for models on Hugging Face Hub (e.g., Open LLM Leaderboard) | Comprehensive, starting with OCR/document AI, independently verified results |
| Code Links | Yes, emphasis on working implementations | Yes, for models on Hugging Face Hub | Yes, reproducible code with verification (signed container hash) |
| Dataset Registry | Yes, restored from original PWC | Different focus, not a comprehensive registry | Per-domain, fresh and maintained |
| Closed-Source Model Evals | Yes, newly added support with special 'closed' tag [article, cite: 14] | Generally requires public models on Hub, limiting closed-source visibility | Yes, publishes dated source link, prompt template, and vendor model-card snapshot for closed APIs |
| Submission Sources | Supports non-arXiv sources like blog posts [article], external papers | Primarily focused on arXiv for trending papers, models on Hub | Focuses on traceable benchmarks, publishes source links |
| Pricing | Free and open resource | Free (for basic use) | Free (for basic use) |
๐ ๏ธ Technical Deep Dive
- The platform utilizes AI agents for parsing research papers at scale and automatically generating leaderboards.
- Human verification is incorporated into the process to maintain the quality and accuracy of the data.
- It integrates with arXiv for academic papers and Hugging Face for trending papers and open-source model hubs.
- The system tracks metrics such as GitHub star velocity for trending papers to highlight high-impact research.
- Support for multiple code repositories linked to a single paper and external, non-arXiv paper sources is implemented.
- User authentication and data storage for thumbnails, paper PDFs, and backups are handled via 'Sign in with HF' (Hugging Face) and Storage Buckets.
๐ฎ Future ImplicationsAI analysis grounded in cited sources
โณ Timeline
๐ Sources (7)
Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.
Weekly AI Recap
Read this week's curated digest of top AI events โ
๐Related Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: Reddit r/MachineLearning โ