In a Harvard Business School field experiment, 228 evaluators screened 48 real submissions with help from a large language model. When the model gave a bare recommendation โ accept or reject, no reasoning โ decision quality improved against an independent expert benchmark. When the identical recommendation arrived wrapped in a persuasive written rationale, that improvement vanished, even though evaluators followed the AI more often (Lane et al., Harvard Business School, 2025).
Higher compliance. No better outcomes. If you run an AI screener anywhere in your operation โ candidate shortlists, vendor selection, stage-gate reviews โ that pairing should stop you, because the feature most vendors sell as the safety mechanism is the one that made the humans worse.
The Experiment That Inverts the Explainability Assumption
The study โ The Narrative AI Advantage?, HBS Working Paper 25-001 โ ran three arms against real early-stage innovation submissions: human evaluation alone, human plus a black-box LLM recommendation, and human plus the same LLM recommendation accompanied by a narrative explanation. Every decision was scored against an independent expert benchmark, so "quality" here is not self-reported confidence. It is measured accuracy against a standard the evaluators could not see.
The black-box arm won. The explained arm did not beat human-only performance, despite evaluators deferring to the model more often in that condition. The researchers' reading is direct: narrative explanations suppress productive overrides by substituting persuasive text for independent verification.
That phrasing is worth sitting with. The entire justification for explainable AI in decision support is that a rationale lets a human audit the machine. What this experiment found is that the rationale replaces the audit. Fluent, well-structured prose reads as evidence rather than as a claim that still requires checking. The evaluator stops evaluating the submission and starts evaluating the argument โ which the model has already optimized to be convincing.
A note on vintage, because it matters for how much weight you put on this: the working paper itself is not new. It re-entered the operational conversation through a Harvard Business Review synthesis in August 2026 that folded it into a broader four-stage framework for where AI helps and hurts across an innovation pipeline (Harvard Business Review, 2026). The experiment is well-aged evidence newly applied, not a fresh headline. That is an argument for taking it more seriously, not less.
Asymmetric Compliance: Why the Reject Path Is Where It Breaks
The mechanism has a name in the paper โ asymmetric compliance โ and it is the operationally important finding.
Evaluators did not follow the explained AI uniformly. They followed it disproportionately when it recommended rejection. Accept recommendations were still questioned; reject recommendations were largely accepted. The result was a substantial increase in false negatives, concentrated in exactly the borderline cases where human judgment is supposed to earn its cost.
There is a plain psychological reason for the asymmetry, and it does not require the evaluators to be careless. Overriding a reject means committing your own name to a candidate the machine has argued against, in writing, with reasons. Overriding an accept costs nothing โ the pipeline simply continues and someone else looks later. So the explanation raises the personal price of one direction of disagreement and leaves the other unchanged. Compliance is not distributed evenly because risk is not distributed evenly.
Now map that onto a 200-person company's hiring funnel. Your AI screener processes 400 applications for a role. It surfaces reject recommendations with two-paragraph rationales. Your recruiter, working through the queue at volume, agrees with nearly all of them โ the reasoning is coherent, specific, and superficially checkable. The screener has not just filtered the pipeline. It has quietly transferred the accept/reject decision from your recruiter to a model, while leaving your process documentation claiming a human made the call.
False Negatives Are the One Error Your Funnel Cannot See
Here is why this specific failure is worse than the size of the effect suggests.
Screening errors come in two types, and your metrics see only one of them. A false positive โ a weak candidate advanced โ shows up eventually as a bad hire, a failed probation, a re-opened req. It is painful, visible, and attributable. A false negative โ a strong candidate rejected โ produces no artifact at all. There is no cohort to measure, no exit interview, no cost line. The person goes somewhere else and performs well there, and you never learn about it.
So an AI screener that shifts your error distribution toward false negatives will look, on every dashboard you have, like it is working. Time-to-screen falls. Recruiter throughput rises. Quality-of-hire holds steady, because the candidates you did hire are still fine. The degradation is real and structurally invisible.
That asymmetry compounds in the same direction as another well-documented risk in the same layer. Stanford research analyzing millions of applications across employers using a shared algorithmic screener found that when many organizations rely on correlated evaluation logic, a candidate rejected by one is systematically rejected by others โ a monoculture effect that turns a single model's blind spot into market-wide exclusion (Stanford HAI, 2026). Combine the two findings and the picture sharpens: correlated reject logic across employers, plus a within-employer nudge toward accepting reject recommendations, means the reject path is where both the invisible damage and the systemic damage concentrate.
The counter-case for explanations, and its actual boundary
The obvious objection: you cannot remove explanations. Regulators, auditors, and increasingly candidates themselves expect a reason. Adverse-action requirements, emerging AI transparency rules, and basic professional decency all point the same way.
That objection is correct, and it is not in conflict with the finding โ because it concerns a different audience.
Explanations serve two entirely separate functions that vendors bundle into one feature. The first is accountability: an auditable record of why a decision was made, reviewed after the fact by a compliance function, a regulator, or the affected person. The second is decision support: text shown to a human evaluator at the moment of judgment, intended to help them decide. The Harvard result indicts the second use, not the first.
You can log the model's reasoning in full, make it retrievable for audit, and attach it to any adverse-action notice โ while not displaying it to the evaluator before they form their own view. Nothing in a transparency obligation requires that the rationale be the first thing your recruiter reads. Bundling those two things is a product-design default, not a legal requirement.
There is a second, narrower defense of explanations worth naming: they help when the human can independently verify the claims. If the rationale says "no SOC 2 certification" and your reviewer can check that in ten seconds, the explanation is a pointer to a fact. If it says "limited evidence of strategic ownership," it is an interpretation dressed as a finding, and there is nothing to check. Verifiability is the dividing line โ the same distinction that separates the tasks practitioners trust agents with from the ones they do not, where confidence tracks whether an output has an objective grading metric rather than how capable the underlying model is (MIT Technology Review, 2026).
What to Change in Your AI Screener This Quarter
Four moves, in order of cost.
1. Audit where explanations surface by default. Most screening, sourcing, and vendor-evaluation tools display rationale text alongside the recommendation with no configuration option exposed. Find out, per tool, whether reasoning can be suppressed at the point of decision and retained in the log. If the answer is no, that is a procurement question for your next renewal, and a specific one.
2. Pilot verdict-only mode on the reject path. Not the whole funnel โ the reject path. Show the recommendation and the underlying artifact, withhold the narrative until the evaluator records a judgment, then reveal it. This is a sequencing change, not a removal, and it is usually configurable in the review UI even when the model output is not.
3. Instrument override rates in both directions. Track how often your evaluators overturn accept recommendations versus reject recommendations. A healthy ratio is roughly symmetric. If your team overrides 15% of accepts and 2% of rejects, you have measured asymmetric compliance in your own operation, and you now have the one number that makes this problem visible to a budget holder.
4. Sample the reject pile. Pull 20 rejected candidates or vendors per quarter and have a second evaluator assess them blind, without the AI output. It is a small, boring control that generates the only false-negative data you will ever have. Without it, you are managing an error class you have no instrument for.
The finding underneath all four is uncomfortable for the direction the market is moving. Every AI screener roadmap is adding richer, more conversational explanation. The best available field evidence says the richer the explanation, the more it substitutes for the judgment you are paying a human to apply โ and that the substitution lands hardest on the decisions that never come back to tell you they were wrong.
Before your next screening tool renewal, ask the vendor a single question: can I turn the explanation off at the point of decision and keep it in the audit log? If the answer is no, you are not buying decision support. You are buying compliance, and paying an evaluator to ratify it.