Competing LLMs Self-Train on Coding via DPO
Two same-model LLM agents compete on coding problems; better execution winner forms DPO pairs for fine-tuning, repeating cycles. Pure execution reward (pass rate), local hardware friendly with specialist temps and memory consolidation. Early Colab A100 results: HumanEval Pass@1 from 0.671 to 0.683 (+1.2pp).







