Search

Few direct matches — filled in with the latest updates.

Tag: #evocodebench1 results

Benchmark for Self-Evolving Coding LLMs

Benchmark for Self-Evolving Coding LLMs

EvoCodeBench evaluates LLM-driven coding systems on self-evolution, efficiency, and human-comparable performance across languages. Tracks dynamics like solving time and improvements over iterations. Enables cross-language robustness analysis.

ArXiv AIResearchFeb 12#research#evocodebench#v1