
Transformers are Bayesian Networks
This arXiv paper proves Transformers are Bayesian networks via five methods: formal proofs showing sigmoid transformers implement loopy belief propagation, constructive exact BP implementation, uniqueness of BP weights, AND/OR structure matching Pearl's algorithm, and experiments. It argues hallucinations arise from lacking finite grounded concepts, unverifiable without them.





