๐Ÿค–Stalecollected in 53m

Papers with Code adds SOTA badges and external evals

Papers with Code adds SOTA badges and external evals
PostLinkedIn
๐Ÿค–Read original on Reddit r/MachineLearning
#benchmarking#model-evaluation#research-toolspapers-with-codepapers with codehugging face

๐Ÿ’กDiscover how the revamped Papers with Code tracks SOTA models using new community-driven evaluation metrics.

โšก 30-Second TL;DR

What Changed

Reintroduced SOTA badges for papers ranking in the top 3 of specific benchmarks.

Why It Matters

These features make it significantly easier for researchers to identify state-of-the-art models and verify performance claims through community-driven external benchmarks.

What To Do Next

Visit the updated tasks page on paperswithco.de to verify if your latest model benchmarks are correctly represented with external evals.

Who should care:Researchers & Academics

Key Points

  • โ€ขReintroduced SOTA badges for papers ranking in the top 3 of specific benchmarks.
  • โ€ขUpdated trending algorithm to include GitHub star velocity and Hugging Face artifact activity.
  • โ€ขAdded support for external, third-party evaluation results beyond the original paper's metrics.
  • โ€ขExpanded coverage of benchmarks including ImageNet, 3D semantic segmentation, and object counting.

๐Ÿง  Deep Insight

AI-generated analysis for this event โ€” not the original article.

๐Ÿ”‘ Enhanced Key Takeaways

  • โ€ขPapers with Code was acquired by Meta AI in 2021, which significantly accelerated its integration into the broader open-science ecosystem.
  • โ€ขThe platform utilizes a crowdsourced approach combined with automated extraction pipelines to link GitHub repositories to research papers.
  • โ€ขThe integration of Hugging Face artifacts marks a shift toward 'living' benchmarks where model weights and inference code are directly executable within the browser.
  • โ€ขThird-party evaluation support addresses the 'reproducibility crisis' by allowing community-verified metrics to supersede potentially biased self-reported results.
  • โ€ขThe platform's metadata is increasingly used by downstream AI agents and automated literature review tools to train and validate RAG (Retrieval-Augmented Generation) systems.
๐Ÿ“Š Competitor Analysisโ–ธ Show
FeaturePapers with CodeHugging Face SpacesarXiv Sanity
Primary FocusBenchmark TrackingModel Hosting/DemosPaper Discovery
PricingFree (Open)Free/EnterpriseFree
BenchmarksExtensive/AutomatedLimited/CommunityNone

๐Ÿ› ๏ธ Technical Deep Dive

  • The platform employs a combination of Natural Language Processing (NLP) models to parse PDF research papers and extract tables containing benchmark results.
  • SOTA (State-of-the-Art) badges are calculated using a dynamic ranking system that normalizes metrics across different evaluation protocols.
  • External evaluation support is implemented via a standardized API that allows researchers to submit JSON-formatted results verified by cryptographic hashes.
  • The trending algorithm utilizes a weighted moving average of GitHub API events (stars, forks, commits) and Hugging Face Hub activity (downloads, likes, and inference requests).

๐Ÿ”ฎ Future ImplicationsAI analysis grounded in cited sources

Papers with Code will transition to a fully automated, agentic benchmark verification system.
The integration of third-party evaluation and Hugging Face artifacts suggests a move toward real-time, agent-driven validation of model claims.
The platform will become the primary data source for training 'Scientific Reasoning' LLMs.
By structuring the world's research and benchmark data into machine-readable formats, the platform provides the highest quality training corpus for specialized scientific AI.

โณ Timeline

2018-09
Papers with Code is launched as a community-driven project.
2021-05
Meta AI acquires Papers with Code to support open science initiatives.
2023-02
Introduction of the 'Methods' and 'Tasks' taxonomy expansion.
2025-11
Initial pilot program for third-party evaluation verification.
๐Ÿ“ฐ

Weekly AI Recap

Read this week's curated digest of top AI events โ†’

๐Ÿ‘‰Related Updates

AI-curated news aggregator. All content rights belong to original publishers.
Original source: Reddit r/MachineLearning โ†—

This is a summary, not the original. Read the source, or get the weekly briefing.

Weekly AI briefing

One email a week. Unsubscribe anytime.