
AI Agents Need Behavioral Tests
This position paper argues that AI agents should be evaluated as behavioral systems, not only by their final performance outcomes. It proposes systematic observation, environmental perturbation, and action-sequence analysis to understand how agents make decisions and adapt.





