In a recent paper, Statistical impossibility and possibility of aligning LLMs with human preferences: From Condorcet paradox to Nash equilibrium, published in the Annals of Statistics, Drs. Jiancong Xiao, Weijie Su, Qi Long, and their team show that popular AI alignment methods (e.g., RLHF) that represent human preferences with a single reward score are flawed because such a score can only exist if there is no circular disagreement among preferences (a “Condorcet cycle”), and these cycles become almost certain to arise in the presence of diverse human opinions. As an alternative, they analyze Nash Learning from Human Feedback (NLHF), a game-theoretic approach in which two AI systems compete to produce responses preferred over each other’s, aiming for a stable “Nash equilibrium” strategy rather than one fixed best answer. They find that when preferences are circular, the fair solution is not a single best response but a “mixed strategy,” whereby AI randomly generates different responses to honestly reflect the conflicting preferences.
This work has important implications for AI alignment in medicine, such as treatment recommendations and patient-facing chatbots, which often involve genuinely conflicting values among clinicians, patients, and guidelines (e.g., aggressive treatment vs. quality-of-life preservation or differing risk tolerances across patient populations). By leveraging their proposed NLHF methods, a clinical LLM can present a range of clinically reasonable options (reflecting genuine disagreement among experts or patient values) rather than collapsing to one “consensus” answer. This is particularly useful in clinical settings where context-dependent judgment matters and suppressing minority-supported treatment paths could be harmful.
This work was led by Dr. Xiao when he was a postdoctoral researcher under the joint supervision of Drs. Su and Long.