Search

Few direct matches — filled in with the latest updates.

Tag: #consistency4 results

Measuring LLM Agent Behavioral Consistency

Measuring LLM Agent Behavioral Consistency

Study reveals LLM agents like Llama/GPT/Claude produce 2-4 unique action paths per 10 runs on HotpotQA, with inconsistency predicting failure. Consistent runs hit 80-92% accuracy vs 25-60% for inconsistent ones. Variance traces to early decisions like first search query.

ArXiv AIResearchFeb 13#research#llama#gpt