⚖️Stalecollected in 52m

Refined Counterfactual Prisoner's Dilemma

PostLinkedIn
⚖️Read original on AI Alignment Forum

💡Exposes utility theory flaws via perfect predictor dilemma—vital for AI alignment.

⚡ 30-Second TL;DR

What Changed

Inspired by Scott Garrabrant's critique of utility theory and independence axiom.

Why It Matters

Highlights flaws in observation-based updating, pushing AI alignment researchers toward updateless approaches for robust agents. May influence decision theory in AGI safety.

What To Do Next

Simulate the dilemma in code to test your decision theory agent's updateless behavior.

Who should care:Researchers & Academics

Key Points

  • Inspired by Scott Garrabrant's critique of utility theory and independence axiom.
  • Omega punishes based on prediction of counterfactual payment behavior.
  • Improved from original for better clarity and spreadability.
  • Tests decision theories on perfect predictors and updatelessness.

🧠 Deep Insight

Background and context from public sources — not the original article. 9 sources cited.

🔑 Enhanced Key Takeaways

  • Scott Garrabrant's critique of expected utility theory identifies a hidden assumption in Von Neumann-Morgenstern axiomatization: the implicit requirement that agents update beliefs branch-by-branch, which formally encodes the independence axiom and creates vulnerabilities to counterfactual reasoning attacks[1].
  • Ergodicity economics, developed by Ole Peters, offers an alternative framework that derives utility functions from the dynamics of stochastic processes rather than postulating ad hoc utility functions, producing logarithmic utility for multiplicative dynamics and linear utility for additive dynamics[1].
  • The independence axiom violation debate extends beyond decision theory into AI alignment: recent work argues that expected utility theory is both unnecessary and insufficient for rational agency, and AI systems can be designed with locally coherent preferences not representable as utility functions[6].
  • Updatelessness versus updatefulness presents a genuine trade-off in decision theory: agents can strategically choose updatelessness in some decision problems to gain coherence benefits while remaining updateful in others to capture value-of-information gains[2].

🔮 Future ImplicationsAI analysis grounded in cited sources

Counterfactual punishment mechanisms may become central to AI safety frameworks if decision-theoretic vulnerabilities in expected utility maximizers are empirically demonstrated in deployed systems.
The refined thought experiment isolates a specific failure mode where standard updating leads to value destruction, suggesting that AI systems optimizing expected utility without counterfactual reasoning could be exploited by sufficiently capable predictors.
Ergodicity economics may reshape how AI utility functions are specified in domains with multiplicative or exotic stochastic dynamics.
If utility functions derived from process dynamics outperform ad hoc specifications, AI developers may shift from preference elicitation to process-analysis-based utility construction.

Timeline

2018-06
Scott Garrabrant publishes foundational critique of independence axiom and utility theory assumptions on LessWrong
2020-01
Ergodicity economics framework gains traction in decision theory discourse, with Ole Peters presenting multiplicative vs. additive dynamics distinction
2024-08
arXiv paper 'Beyond Preferences in AI Alignment' published, critiquing normative status of expected utility theory for AI systems
📰

Weekly AI Recap

Read this week's curated digest of top AI events →

👉Related Updates

AI-curated news aggregator. All content rights belong to original publishers.
Original source: AI Alignment Forum

This is a summary, not the original. Read the source, or get the weekly briefing.

Weekly AI briefing

One email a week. Unsubscribe anytime.