
Debiasing-DPO Cuts LLM Bias 84%
Researchers propose Debiasing-DPO to counter LLM biases from spurious social contexts like teacher demographics. Using NCTE classroom transcripts, it reduces bias by 84% and boosts accuracy 52% on Llama and Qwen models. Standard DPO fails, but this self-supervised method pairs neutral and biased reasoning effectively.


