Search

Tag: #agentic16 results

Claude 4.7 Tops Coding Benchmarks

Claude 4.7 Tops Coding Benchmarks

Anthropic launches Claude Opus 4.7, its top model with 64.3% on SWE-bench Pro (vs GPT-5.4's 57.7%), multi-agent coordination for long workflows, 3x image resolution, and 14% better agentic reasoning with fewer tool errors. Pricing: $5 input/$25 output per million tokens.

The Next Web (TNW)MediaApr 16#benchmarks#agentic#multimodal
Gemma 4 Dominates Benchmarks at $0.20/Run

Gemma 4 Dominates Benchmarks at $0.20/Run

Gemma 4 31B achieved 100% survival and +1,144% median ROI on FoodTruck Bench, outperforming GPT-5.2, Gemini 3 Pro, and all tested Chinese open-source models. Only Opus 4.6 beats it, but at 180x the cost of $0.20 per run. Best cost-performance ratio among 22 tested models.

Reddit r/LocalLLaMACommunityApr 5#benchmark#agentic#cost-performance
Qwen3.6 Autonomously Builds Tower Defense Game

Qwen3.6 Autonomously Builds Tower Defense Game

A user tasked Qwen3.6-35B with building a tower defense game using screenshots from MCP, and it successfully implemented and tested features like upgrades. The model self-detected and fixed bugs in canvas rendering and wave completions. High excitement for the upcoming Qwen Coder model.

Reddit r/LocalLLaMACommunityApr 17#agentic#multimodal#gguf
Liquid AI's tiny 350M agentic model

Liquid AI's tiny 350M agentic model

Liquid AI released LFM2.5-350M, a 350M parameter model optimized for data extraction, tool use, and agent workflows. Trained on 28T tokens with RL, it beats Qwen3.5-0.8B while fitting <500MB quantized. Runs on CPU/GPU/mobile with reliable outputs.

Reddit r/LocalLLaMACommunityMar 31#agentic#quantized#small-model
Page 1 of 2