Bi-Level Optimization for Multimodal LLM Judges
Introduces BLPO to optimize prompts for multimodal LLM-as-a-judge evaluating AI images. Overcomes context limits by converting images to text representations. Outperforms baselines on four datasets with three LLM judges.





