πŸ“„Stalecollected in 11h

ThermoQA LLM Thermodynamic Benchmark Launch

ThermoQA LLM Thermodynamic Benchmark Launch
PostLinkedIn
πŸ“„Read original on ArXiv AI

πŸ’‘New benchmark proves top LLMs falter on thermo reasoningβ€”evaluate yours!

⚑ 30-Second TL;DR

What Changed

293 problems across three tiers: 110 property lookups, 101 component analysis, 82 full cycles.

Why It Matters

ThermoQA exposes LLM weaknesses in scientific reasoning, guiding targeted improvements for engineering tasks. It quantifies consistency via multi-run sigmas, advancing reliable evaluation methods.

What To Do Next

Download ThermoQA dataset from Hugging Face and benchmark your LLM on thermodynamic reasoning.

Who should care:Researchers & Academics

Key Points

  • β€’293 problems across three tiers: 110 property lookups, 101 component analysis, 82 full cycles.
  • β€’Ground truth via CoolProp 7.2.0 for water, R-134a, variable-cp air.
  • β€’Leaderboard: Claude Opus 94.1%, GPT-5.4 93.1%, Gemini 3.1 Pro 92.5%.
  • β€’Cross-tier degradation up to 32.5pp proves memorization β‰  reasoning.
  • β€’Open-source dataset/code at huggingface.co/datasets/olivenet/thermoqa
πŸ“°

Weekly AI Recap

Read this week's curated digest of top AI events β†’

πŸ‘‰Related Updates

AI-curated news aggregator. All content rights belong to original publishers.
Original source: ArXiv AI β†—