πArXiv AIβ’Stalecollected in 11h
ThermoQA LLM Thermodynamic Benchmark Launch

π‘New benchmark proves top LLMs falter on thermo reasoningβevaluate yours!
β‘ 30-Second TL;DR
What Changed
293 problems across three tiers: 110 property lookups, 101 component analysis, 82 full cycles.
Why It Matters
ThermoQA exposes LLM weaknesses in scientific reasoning, guiding targeted improvements for engineering tasks. It quantifies consistency via multi-run sigmas, advancing reliable evaluation methods.
What To Do Next
Download ThermoQA dataset from Hugging Face and benchmark your LLM on thermodynamic reasoning.
Who should care:Researchers & Academics
Key Points
- β’293 problems across three tiers: 110 property lookups, 101 component analysis, 82 full cycles.
- β’Ground truth via CoolProp 7.2.0 for water, R-134a, variable-cp air.
- β’Leaderboard: Claude Opus 94.1%, GPT-5.4 93.1%, Gemini 3.1 Pro 92.5%.
- β’Cross-tier degradation up to 32.5pp proves memorization β reasoning.
- β’Open-source dataset/code at huggingface.co/datasets/olivenet/thermoqa
π°
Weekly AI Recap
Read this week's curated digest of top AI events β
πRelated Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: ArXiv AI β