SourceStalecollected in 53m

Papers with Code adds SOTA badges and external evals

Read original on Reddit r/MachineLearning
#benchmarking#model-evaluation#research-tools

Discover how the revamped Papers with Code tracks SOTA models using new community-driven evaluation metrics.

30-Second TL;DR

What Changed

Reintroduced SOTA badges for papers ranking in the top 3 of specific benchmarks.

Why It Matters

These features make it significantly easier for researchers to identify state-of-the-art models and verify performance claims through community-driven external benchmarks.

What To Do Next

Visit the updated tasks page on paperswithco.de to verify if your latest model benchmarks are correctly represented with external evals.

Who should care:Researchers & Academics

Key Points

  • •Reintroduced SOTA badges for papers ranking in the top 3 of specific benchmarks.
  • •Updated trending algorithm to include GitHub star velocity and Hugging Face artifact activity.
  • •Added support for external, third-party evaluation results beyond the original paper's metrics.
  • •Expanded coverage of benchmarks including ImageNet, 3D semantic segmentation, and object counting.

Deep Insight

AI-generated analysis for this event — not the original article.

Enhanced Key Takeaways

  • •Papers with Code was acquired by Meta AI in 2021, which significantly accelerated its integration into the broader open-science ecosystem.
  • •The platform utilizes a crowdsourced approach combined with automated extraction pipelines to link GitHub repositories to research papers.
  • •The integration of Hugging Face artifacts marks a shift toward 'living' benchmarks where model weights and inference code are directly executable within the browser.
  • •Third-party evaluation support addresses the 'reproducibility crisis' by allowing community-verified metrics to supersede potentially biased self-reported results.
  • •The platform's metadata is increasingly used by downstream AI agents and automated literature review tools to train and validate RAG (Retrieval-Augmented Generation) systems.

Competitor Analysis

Primary Focus
Papers with Code
Benchmark Tracking
Hugging Face Spaces
Model Hosting/Demos
arXiv Sanity
Paper Discovery
Pricing
Papers with Code
Free (Open)
Hugging Face Spaces
Free/Enterprise
arXiv Sanity
Free
Benchmarks
Papers with Code
Extensive/Automated
Hugging Face Spaces
Limited/Community
arXiv Sanity
None

Technical Deep Dive

  • The platform employs a combination of Natural Language Processing (NLP) models to parse PDF research papers and extract tables containing benchmark results.
  • SOTA (State-of-the-Art) badges are calculated using a dynamic ranking system that normalizes metrics across different evaluation protocols.
  • External evaluation support is implemented via a standardized API that allows researchers to submit JSON-formatted results verified by cryptographic hashes.
  • The trending algorithm utilizes a weighted moving average of GitHub API events (stars, forks, commits) and Hugging Face Hub activity (downloads, likes, and inference requests).

Future ImplicationsAI analysis grounded in cited sources

Papers with Code will transition to a fully automated, agentic benchmark verification system.
The integration of third-party evaluation and Hugging Face artifacts suggests a move toward real-time, agent-driven validation of model claims.
The platform will become the primary data source for training 'Scientific Reasoning' LLMs.
By structuring the world's research and benchmark data into machine-readable formats, the platform provides the highest quality training corpus for specialized scientific AI.

Timeline

2018-09
Papers with Code is launched as a community-driven project.
2021-05
Meta AI acquires Papers with Code to support open science initiatives.
2023-02
Introduction of the 'Methods' and 'Tasks' taxonomy expansion.
2025-11
Initial pilot program for third-party evaluation verification.

Weekly AI Recap

Read this week's curated digest of top AI events →

AI-curated news aggregator. All content rights belong to original publishers.
Original source: Reddit r/MachineLearning ↗

This is a summary, not the original. Read the source, or get the weekly briefing.

The weekly digest

One email a week. Unsubscribe anytime.