Search

Tag: #llm-benchmark24 results

GTO Wizard Poker AI Benchmark

GTO Wizard Poker AI Benchmark

GTO Wizard Benchmark launches a public API and framework for evaluating Heads-Up No-Limit Texas Hold'em agents against superhuman GTO Wizard AI, which outperforms Slumbot by 19.4 bb/100. It employs AIVAT for 10x variance reduction efficiency. Benchmarks reveal LLM progress but all models lag far behind the baseline.

ArXiv AIResearchMar 26#poker-ai#llm-benchmark#multi-agent
GPSBench Tests LLM GPS Reasoning

GPSBench Tests LLM GPS Reasoning

Researchers launch GPSBench, a 57,800-sample dataset across 17 tasks to probe LLMs' geospatial reasoning without tools. 14 SOTA LLMs show reliability in geographic knowledge but struggle with geometric computations like distance and bearing. Dataset, code, and findings reveal trade-offs in finetuning and augmentation benefits.

AI Outsmarts Humans in 40% Yield Scam Test

AI Outsmarts Humans in 40% Yield Scam Test

Nanjing University's latest research shows AI models remain more rational than humans when confronted with a scam promising 40% annualized returns. Under simulated investor pressure, seven mainstream large models demonstrated superior financial scam resistance by holding their底线 better. This real-world test highlights AI's edge in fraud detection.

LLMs Struggle in Clue Reasoning Test

LLMs Struggle in Clue Reasoning Test

Researchers built a text-based multi-agent Clue game to evaluate LLMs' multi-step deductive reasoning with GPT-4o-mini and Gemini-2.5-Flash agents. Across 18 simulated games, agents secured only four correct wins, struggling with consistent reasoning. Fine-tuning on logic puzzles did not reliably enhance performance and sometimes increased verbosity without precision gains.

Page 2 of 3