Search

10 results on this page

New Benchmark Tests Whether AI Can Draw Geometry

New Benchmark Tests Whether AI Can Draw Geometry

Researchers introduce an open-source benchmark separating olympiad geometry solving from accurate diagram construction. Across 954 problems, current foundation models achieved only a 36.14% average diagram compilation success rate, revealing a substantial gap between mathematical reasoning and faithful visual construction.

ArXiv AIResearch23h ago#geometry-reasoning#benchmark#asymptote
ByteDance Restructures Seed AI Team

ByteDance Restructures Seed AI Team

ByteDance’s Seed foundation-model division has reportedly completed another internal restructuring. Its foundation-model organization now includes four first-level departments focused on pretraining data, reinforcement learning, product post-training for work, and product post-training for chat.

Model Cards Alone Can’t Govern Open-Weight AI

Model Cards Alone Can’t Govern Open-Weight AI

A position paper analyzing 500 Hugging Face model cards argues that model cards alone do not provide enough information for governing open-weight foundation models. It proposes combining model cards with acceptable use policies and licenses to address safety, provenance, behavior, and enforcement gaps.

ArXiv AIResearch23h ago#model-cards#ai-governance
Page 1