
VISTA Escapes Black-Box in Prompt Optimization
Reflective APO methods like GEPA suffer black-box issues, degrading GSM8K accuracy from 23.81% to 13.50% on defective seeds. VISTA, a multi-agent framework, decouples hypothesis generation from rewriting for interpretable, robust optimization. It recovers to 87.57% accuracy and outperforms baselines on GSM8K and AIME2025.





