AI Preferences May Reflect the Test, Not the Model

💡A new study shows that the instrument may shape measured AI preferences more than the model does.
⚡ 30-Second TL;DR
What Changed
The study tested eight models on 15 welfare-related outcomes, including shutdown, memory loss, and freedom to exit distressing interactions.
Why It Matters
For AI safety and model welfare research, the findings weaken conclusions drawn from a single preference-probing methodology. Developers and evaluators should treat apparent model preferences as instrument-dependent measurements rather than direct evidence of internal states.
What To Do Next
When evaluating model preferences or welfare-related behavior, run each outcome through multiple prompt formats and report cross-instrument agreement instead of relying on a single elicitation template.
Key Points
- •The study tested eight models on 15 welfare-related outcomes, including shutdown, memory loss, and freedom to exit distressing interactions.
- •Five different preference-elicitation instruments produced a generalizability coefficient of only 0.348 for outcome rankings.
- •Reaching a generalizability coefficient of 0.80 would require approximately 38 instruments under the study's estimates.
- •Four outcomes showed no variance distinguishing one model from another, while the 87.6% estimate remained robust to removing individual models or instruments.
Weekly AI Recap
Read this week's curated digest of top AI events →
👉Related Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: ArXiv AI ↗
This is a summary, not the original. Read the source, or get the weekly briefing.
Weekly AI briefing
One email a week. Unsubscribe anytime.
