If artificial intelligence were making the call, would you trust it to decide who receives a life-saving transplant?
A new study suggests that when large language models are asked to weigh scarce medical resources, they do not think like human doctors — and in some cases, they do not think like humans at all.
Follow THE FUTURE on LinkedIn, Facebook, Instagram, X and Telegram
AI Weighs The Wrong Things, Or At Least Different Ones
Researchers at Penn State University tested large language models using hypothetical kidney-allocation scenarios drawn from prior human research. In each case, the AI had to choose between two eligible patients competing for a single available kidney.
The patients were defined by traits such as age, health and drinking habits, allowing researchers to compare how AI systems prioritized competing factors against how people had previously made the same decisions in human studies.
The results were striking. Human participants tended to place more weight on age, often favoring younger patients. By contrast, many AI models gave greater priority to lower alcohol consumption. More importantly, the models frequently narrowed complex ethical judgments to a single attribute, while humans tended to consider the broader context.
“AI chatbots often diverge from human values in how they weigh a patient’s traits,” said Hadi Hosseini, who led the study at Penn State University. “They fixate on a single factor, like drinking habits, rather than balancing multiple considerations the way people do.”
Indecision Is A Human Feature — And An AI Weakness
Another key difference was hesitation. Human respondents often recognized that there is no single objectively correct answer in a scarcity decision such as organ allocation. Their choices reflected nuance, ambiguity and moral trade-offs.
The models, by contrast, typically committed to one answer with little sign of uncertainty.
That may seem efficient, but in high-stakes settings, certainty is not always a virtue. Decisions about kidneys, jobs or other scarce resources often involve values that cannot be reduced to a clean formula. Humans often absorb that ambiguity through discussion, debate and institutional safeguards. AI systems, the researchers argue, tend to skip over it.
“When we allocate something scarce, whether it’s a kidney, a job or access to some other resource, there isn’t always a single objectively correct answer,” said John Dickerson, chief executive officer at Mozilla.ai, who collaborated on the study. “Humans recognize that ambiguity and codify it via open debate into the allocative process. AI models often don’t.”
Why This Matters For Healthcare
The study arrives at a moment when AI is moving rapidly into healthcare, where it is already being used to support diagnosis, clinical workflows, treatment planning and the allocation of scarce medical resources.
That growing role makes the question of alignment especially urgent. In healthcare, the issue is not simply whether a model can produce an answer, but whether that answer reflects the moral standards and professional judgment that society expects from life-altering decisions.
Kidney allocation is a particularly sensitive example because it sits at the intersection of ethics, medicine and resource scarcity. Choosing one patient over another is never just a technical decision; it is a judgment about fairness, need, prognosis and social values.
“The ethical stakes are high, and AI’s role in such life-altering decisions requires deep reflection,” Hosseini said. “Moral decisions in settings like organ allocation directly determine who lives and who dies, so getting AI’s role in them right isn’t optional.”
The Broader Debate Over AI And Moral Judgment
The researchers say their findings speak to a wider debate over whether AI can make moral decisions — or whether it can ever truly align with human values.
That debate is no longer theoretical. As organizations increasingly rely on AI systems for recommendations, rankings and triage decisions, understanding how those systems reason has become a practical governance issue.
The study does not argue that AI should replace professional judgment in medicine. If anything, it reinforces the opposite conclusion: the more consequential the decision, the more important it is to understand where AI diverges from human reasoning.
In healthcare, as in business and public policy, the danger is not only that AI may be wrong. It may also be confidently, efficiently and consistently wrong in ways that humans would immediately question.







