Search

Tag: #self-evolution11 results

Benchmark for Self-Evolving Coding LLMs

Benchmark for Self-Evolving Coding LLMs

EvoCodeBench evaluates LLM-driven coding systems on self-evolution, efficiency, and human-comparable performance across languages. Tracks dynamics like solving time and improvements over iterations. Enables cross-language robustness analysis.

ArXiv AIResearchFeb 12#research#evocodebench#v1
Page 2 of 2