๐Ÿค–Stalecollected in 24m

Papers With Code Adds Support for Closed-Source Model Evals

Papers With Code Adds Support for Closed-Source Model Evals
PostLinkedIn
๐Ÿค–Read original on Reddit r/MachineLearning

๐Ÿ’กEasily compare your open-source AI models against the latest closed-source benchmarks on a unified leaderboard.

โšก 30-Second TL;DR

What Changed

Added support for closed-source model benchmarks on leaderboards.

Why It Matters

This update improves transparency in AI benchmarking by acknowledging the dominance of closed-source models while maintaining a clear distinction for open-source research.

What To Do Next

Visit paperswithcode.co and use the toggle settings to compare your open-source model's performance against the latest closed-source benchmarks.

Who should care:Researchers & Academics

Key Points

  • โ€ขAdded support for closed-source model benchmarks on leaderboards.
  • โ€ขIntroduced a toggle to filter out closed-source models for open-model-only views.
  • โ€ขAllows submission of non-arXiv sources like blog posts for model evaluations.
  • โ€ขProvides scatter plots and tables for benchmarking state-of-the-art performance.

๐Ÿง  Deep Insight

Web-grounded analysis with 7 cited sources.

๐Ÿ”‘ Enhanced Key Takeaways

  • โ€ขThe current Papers With Code platform (paperswithcode.co) is a community-driven revival, launched around May 2026, following the shutdown of the original paperswithcode.com by Meta in July 2025.
  • โ€ขThis resurrected platform is spearheaded by Hugging Face's Niels Rogge, in collaboration with Meta AI and the original creators, aiming to restore a vital resource for the machine learning community.
  • โ€ขThe new Papers With Code leverages AI agents to automatically parse research papers and generate leaderboards, with human verification in place to ensure data quality and accuracy.
  • โ€ขBeyond traditional arXiv sources, the platform now supports external paper submissions and integrates with Hugging Face for trending papers and model hubs, enhancing the discoverability and reproducibility of research.
๐Ÿ“Š Competitor Analysisโ–ธ Show
Feature / PlatformPapers With Code (paperswithcode.co)Hugging Face LeaderboardsCodeSOTA
SOTA LeaderboardsComprehensive, supports open & closed models, toggle filter [article]Limited, primarily for models on Hugging Face Hub (e.g., Open LLM Leaderboard)Comprehensive, starting with OCR/document AI, independently verified results
Code LinksYes, emphasis on working implementationsYes, for models on Hugging Face HubYes, reproducible code with verification (signed container hash)
Dataset RegistryYes, restored from original PWCDifferent focus, not a comprehensive registryPer-domain, fresh and maintained
Closed-Source Model EvalsYes, newly added support with special 'closed' tag [article, cite: 14]Generally requires public models on Hub, limiting closed-source visibilityYes, publishes dated source link, prompt template, and vendor model-card snapshot for closed APIs
Submission SourcesSupports non-arXiv sources like blog posts [article], external papersPrimarily focused on arXiv for trending papers, models on HubFocuses on traceable benchmarks, publishes source links
PricingFree and open resourceFree (for basic use)Free (for basic use)

๐Ÿ› ๏ธ Technical Deep Dive

  • The platform utilizes AI agents for parsing research papers at scale and automatically generating leaderboards.
  • Human verification is incorporated into the process to maintain the quality and accuracy of the data.
  • It integrates with arXiv for academic papers and Hugging Face for trending papers and open-source model hubs.
  • The system tracks metrics such as GitHub star velocity for trending papers to highlight high-impact research.
  • Support for multiple code repositories linked to a single paper and external, non-arXiv paper sources is implemented.
  • User authentication and data storage for thumbnails, paper PDFs, and backups are handled via 'Sign in with HF' (Hugging Face) and Storage Buckets.

๐Ÿ”ฎ Future ImplicationsAI analysis grounded in cited sources

The inclusion of closed-source model evaluations will lead to more comprehensive and realistic comparisons of state-of-the-art AI performance.
By allowing both open and closed models on leaderboards, researchers and practitioners can gain a more complete understanding of the true SOTA, regardless of model accessibility, fostering a more level playing field for evaluation.
Papers With Code's revival and new features will reinforce its role as a central hub for AI research discovery and reproducibility.
The platform's renewed focus on discoverability, reproducibility, and actionability, combined with AI-driven automation and community collaboration, positions it to regain its status as an indispensable daily tool for AI professionals.
The ability to submit non-arXiv sources will broaden the scope of recognized AI research and accelerate the dissemination of findings.
Accepting blog posts and other non-traditional sources for evaluations allows for quicker sharing of results, especially for rapidly evolving closed-source models, potentially reducing the time lag between development and public benchmarking. [article]

โณ Timeline

2018-07
Original Papers With Code (paperswithcode.com) launched as an independent project.
2019-12
Papers With Code acquired by Facebook AI (now Meta), with a pledge to remain neutral and open.
2025-07
Original Papers With Code (paperswithcode.com) was shut down by Meta without prior notice, redirecting to Hugging Face.
2026-05
Papers With Code (paperswithcode.co), a community-driven revival spearheaded by Hugging Face's Niels Rogge, was launched.
2026-06
The revived Papers With Code adds support for closed-source model evaluations on its leaderboards. [article, cite: 14]

๐Ÿ“Ž Sources (7)

Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.

  1. codesota.com
  2. tib.eu
  3. youtube.com
  4. quasa.io
  5. reddit.com
  6. similarweb.com
  7. medium.com
๐Ÿ“ฐ

Weekly AI Recap

Read this week's curated digest of top AI events โ†’

๐Ÿ‘‰Related Updates

AI-curated news aggregator. All content rights belong to original publishers.
Original source: Reddit r/MachineLearning โ†—