
Gemini vs ChatGPT vs Claude: Video Test Winner
Tested Gemini, ChatGPT, and Claude's video analysis on YouTube clips and local files. Explores if AIs truly understand videos or just simulate it. Reveals the top performer.
Tag: #llm-benchmark24 results

Tested Gemini, ChatGPT, and Claude's video analysis on YouTube clips and local files. Explores if AIs truly understand videos or just simulate it. Reveals the top performer.

Tested 8 LLMs as agentic tabletop GMs; a 27B model topped a 405B on narrative quality despite tool-calling challenges. Smaller models like Mistral Small 3.1 fail after 4-5 tool calls on low-end hardware. Threshold for reliable local use is 70B+ on 64GB RAM.

An analysis reveals Google's AI Overviews feature is inaccurate 10% of the time. This raises questions about whether 90% accuracy is sufficient for AI-driven search results.
arXiv founder personally tested LLMs for generating 'water papers' (filler content). Grok proved strongest in compliance and output quality. Claude was the least cooperative.