When AI explains its decision, humans may stop thinking independently
AI is known to be confidently wrong, and now it’s influencing humans to be that way, too.
In a new study, researchers tested AI’s influence on humans reviewing innovation proposals, and found that AI recommender tools were persuasive enough to convince the evaluators to reject decisions made by independent human experts, thus causing them to pass on promising innovations. Similarly, they went along with AI approval of ideas that the human experts found sub-par.
Interestingly, reviewers were also more inclined to defer to an incorrect AI decision when the model explained itself. Narrative explanations degraded human judgment, rather than enhancing it. People did better when they weren’t given a reason for the AI’s decision.
“Our findings reveal that LLM explanations do not necessarily improve decision-making,” the researchers, associated with Harvard Business School, MIT, and the University of Washington explained in their findings. “Effective human-AI collaboration requires designs that preserve rather than supplant independent human judgment.”
AI rationale can undermine human judgment
Every enterprise screens proposed projects before pursuing them, but there is always uncertainty, and the risk of trade-offs like false positives (going forward with projects that ultimately fail) or false negatives (rejecting ideas that might have succeeded). For an example of the former, the researchers point to Google Glass or Amazon’s Fire Phone; for the latter, Xerox terminating early Ethernet and PostScript projects.
Because they have limited time and only basic information to go on, decision-makers are increasingly turning to LLMs that use predictive algorithms to generate recommendations and rationales based on context.
The researchers set out to explore AI’s role in what they called “early-stage innovation screening.” They judged how human evaluators were influenced by LLM recommendations, both with and without explanations from the model on how and why it reached its decision.
Their experiment asked 228 experienced evaluators to assess nearly 50 submissions to an MIT challenge. They tested three different scenarios: human-only proposals with no AI assistance; LLM evaluations with a written rationale for the decision; and black-box AI pass-fail recommendations with no accompanying explanation.
Evaluators’ decisions were then compared to those made by four human experts. Those decisions were considered the ‘correct’ baseline. They were judged on whether they outright complied with the LLM’s recommendations, overrode them, or productively overrode them, meaning they independently verified persuasive model outputs before making a decision.
Their decisions were classified as correct (agreeing with human experts’ positive/negative decisions), false positive (supporting submissions that experts would reject), and false negative (rejecting submissions experts would move forward with).
Overall, the evaluators accepted LLM recommendations 67% of the time. They agreed with both black-box and narrative LLM decisions roughly 75% of the time, but only agreed with human decisions 54% of the time.
Seemingly counterintuitively, black-box recommendations improved the quality of decisions (aligning them with human experts) but recommendations with narratives did not. When given an LLM recommendation to reject a submission and an accompanying reason why, evaluators disproportionately agreed, which reduced false positives, but “substantially” increased false negatives.
The researchers posit that this is because narrative explanations “suppress” productive overrides; LLMs provide a convincing argument that is easy to accept, essentially discouraging independent human verification. This contradicts a common assumption that LLM explanations augment human decision-making.
The researchers pointed out that people are cognitively predisposed to weigh negative information more heavily than positive information; the phenomenon is known as ‘negativity bias.’
“Rejection is an active, eliminative decision that feels more consequential and accountable than preserving optionality,” they wrote. It also maintains the status quo, avoids risk and bias, and requires no resource commitment.
LLM explanations provide “ready-made justifications” for going along with rejection decisions without independently verifying them; humans effectively offload their thinking to AI, researchers explained. Evaluators often rely on surface cues such as fluency, coherence, and seeming credibility. LLMs are particularly well-suited to exploit this because they are linguistically fluent and expert-like, creating an “illusion of explanatory depth.”
Thus, “individuals tend to overestimate their understanding of a decision despite limited insight into its reasoning,” the researchers wrote.
Finding a balance in recommendation systems
The researchers pointed out that their findings have “clear implications” for enterprises designing AI-assisted evaluation systems.
Enterprises should be cautious with LLM explanations in high-stakes decision-making, they advised. AI recommendations should not be taken at face value; they should always be tested before any associated deployment. This helps improve accuracy and encourages human reviewers to detect errors and learn how models operate, or potentially can even increase human-AI agreement.
In decision contexts such as quality control, compliance screening, or fraud detection, LLM explanations could support conservative human decision-making, the researchers noted. On the other hand, in tasks like early-stage screening, LLM narratives could undermine performance by “discouraging independent judgment and suppressing productive human override.” In this context, simpler or more opaque recommendations may preserve human discretion and verification.
Future design of explanation systems should factor in the nature of the task and the potential cost of errors made by AI, the researchers advised. Enterprises could experiment with models that support contrasting narratives (reasons to reject an idea alongside reasons to accept it) or uncertainty disclosures based on a fixed threshold, rather than on purely binary decisions. Systems could also be structured to invite human disagreement.
The researchers also noted that there is opportunity to test whether narrative explanations have different impacts at later stages of decision-making, when evaluators have fewer options, more information, and increased incentive to verify outputs and think the problem through.
Ultimately, the researchers emphasized, “organizations should treat AI explanations not as universally beneficial transparency tools, but as behavioral interventions whose effects depend on how evaluators process information under uncertainty.”
This article originally appeared on CIO.com.ComputerworldRead More