Search

Few direct matches — filled in with the latest updates.

Tag: #execution-reward1 results

📰

Competing LLMs Self-Train on Coding via DPO

Two same-model LLM agents compete on coding problems; better execution winner forms DPO pairs for fine-tuning, repeating cycles. Pure execution reward (pass rate), local hardware friendly with specialist temps and memory consolidation. Early Colab A100 results: HumanEval Pass@1 from 0.671 to 0.683 (+1.2pp).

Reddit r/LocalLLaMACommunityApr 16#self-play#fine-tuning#execution-reward