Search

Tag: #research297 results

AI Climate Claims Branded Greenwashing

AI Climate Claims Branded Greenwashing

A report dismisses tech industry claims that AI can fix climate issues as greenwashing. Most references are to traditional machine learning, not energy-intensive generative AI like chatbots and image tools. Explosive growth fuels massive datacenter energy demands.

The Guardian TechnologyMediaFeb 17#research#generative-ai#climate
VeRA: Scalable Verified Reasoning Data Augmentation

VeRA: Scalable Verified Reasoning Data Augmentation

VeRA is a framework that transforms static benchmark problems into executable specifications for generating unlimited verified variants. It features VeRA-E for equivalent rewrites to detect memorization and VeRA-H for hardened tasks at intelligence frontiers. The tool is open-sourced with code and datasets after evaluating 16 frontier models.

ArXiv AIResearchFeb 17#research#vera#ai-evaluation
SSLogic:代理式元合成擴展邏輯推理

SSLogic:代理式元合成擴展邏輯推理

SSLogic 是一個代理式元合成框架,透過迭代的生成-驗證-修復迴圈,在任務家族層級擴展邏輯推理任務,合成 Generator-Validator 程式對。採用多閘驗證協議,結合多策略一致性檢查與對抗性盲審,由獨立代理撰寫並執行程式碼以過濾問題任務。在演化資料上訓練,提升 SynLogic +5.2 分等基準表現。

ArXiv AIResearchFeb 17#research#sslogic#logical-reasoning
BotzoneBench: Scalable LLM Game Eval Benchmark

BotzoneBench: Scalable LLM Game Eval Benchmark

BotzoneBench introduces a scalable framework for evaluating LLMs' strategic reasoning in interactive games using fixed hierarchies of skill-calibrated game AIs. It assesses five flagship models across eight diverse games via 177,047 state-action pairs, revealing performance gaps and behaviors comparable to mid-tier game AIs. This enables linear-time absolute measurements with stable interpretability, unlike volatile LLM-vs-LLM rankings.

ArXiv AIResearchFeb 17#research#botzonebench#llm
Adversarial Self-Critique for Safer AI Underwriting

Adversarial Self-Critique for Safer AI Underwriting

New agentic AI system for commercial insurance underwriting uses adversarial self-critique where a critic agent challenges primary decisions before human review. It reduces hallucinations from 11.3% to 3.8% and boosts accuracy from 92% to 96% on 500 expert cases. The human-in-the-loop design ensures oversight in regulated environments.

ArXiv AIResearchFeb 17#research#agentic-ai#self-critique
SpaceX & xAI Join DoD Killer Drone Competition

SpaceX & xAI Join DoD Killer Drone Competition

SpaceX and xAI are participating in a secretive US DoD competition to develop voice-controlled autonomous drone swarms for offensive use. The six-month challenge, with up to $100M prizes, focuses on converting voice commands to digital instructions for multi-drone coordination and targeting. This involves xAI's AI tech and marks Musk's controversial pivot into military AI weapons.

IT之家MediaFeb 16#research#spacex-xai#defense-ai
Page 2 of 30