Give the same AI assistant to your strongest operator and your weakest one, and you do not get two smaller versions of the same gain. You get opposite signs. In a five-month randomized field experiment with 640 small-business owners, high performers who received a GPT-4 business adviser saw revenue and profit rise about 15%, while already-struggling owners came out nearly 10% worse than the control group (MIT Sloan Management Review, 2026). Same tool. Same access window. Same volume of advice. The AI performance gap did not narrow โ it widened, and part of that widening was outright damage.
That result should change a sequencing decision most Heads of Operations are making right now: who gets the copilot first.
Why the AI performance gap widened instead of closing
Nicholas Otis, Rowan Clarke, Solรจne Delecourt, David Holtz, and Rembrand Koning randomized access to a GPT-4 assistant, prompted to act as a business adviser, among 640 Kenyan entrepreneurs running fast-food outlets, poultry farms, cybershops, and similar small operations. Delivery was deliberately frictionless: a WhatsApp contact, on the platform roughly 90% of Kenyans already use, tracked over five months (Berkeley Haas, 2026). The study was published as The Uneven Impact of Generative Artificial Intelligence on Entrepreneurial Performance (Management Science, 2026).
The design matters more than the setting. Access was randomized, so the split is not a story about who opted in. The tool was identical across arms. Usage volume did not explain the divergence. What separated the two groups was what they brought to the assistant and what they did with what came back.
Read that as an operating condition, not a curiosity about Nairobi. The variable being tested โ does a general-purpose advisory AI help someone whose judgment is already shaky โ is exactly the variable in play when you hand a copilot to the bottom quartile of a 200-person operations org and call it enablement.
The "AI closes the skill gap" evidence is real โ and narrower than you think
The leveling story is not a myth. It is a finding about a specific class of work.
Brynjolfsson, Li, and Raymond studied 5,179 customer-support agents on a staggered rollout of a generative AI assistant. Productivity rose 14% on average, with a 34% improvement among novice and low-skilled agents and minimal effect on the most experienced ones (Quarterly Journal of Economics, 2025). The mechanism was that the model disseminated the practices of the strongest agents and moved newer ones down the experience curve faster.
Look at what that task actually is. Bounded scope. A known-good answer that exists somewhere in the firm. Immediate feedback from the customer. Quality checked by supervisors within hours. Under those conditions, an AI that encodes tacit best practice compresses the distribution โ the weakest gain the most because the gap between their behavior and the encoded standard is largest.
Now change one variable: remove the known-good answer. Ask the model an open question about pricing, hiring, or where to spend the next thousand dollars. There is no verified target for it to converge on, no supervisor closing the loop, and no immediate signal that the advice was wrong. Compression stops being the default. Amplification takes over.
The leveling result is a property of verifiable tasks, not of the technology.
The moderator is judgment, not access or prompt skill
The Kenya study's authors are explicit about the mechanism: what determined the sign of the effect was "whether an entrepreneur had the judgment to distinguish good AI advice from bad." Weaker performers followed generic or misleading advice because they lacked the filter to reject it (MIT Sloan Management Review, 2026).
The behavioral detail underneath is sharper. High performers brought the assistant tractable problems โ questions with structure, where the model's output could be checked against something they already knew. Struggling owners brought it the hardest, most ambiguous problems they had: a competitor undercutting them, a drought, no working capital. Those are questions neither an AI nor a consultant can answer well from outside the business (Berkeley Haas, 2026). The model answered anyway. It always does. Generic advice โ cut prices, spend more on ads โ is fluent, plausible, and margin-destroying when applied without situational judgment.
So the causal chain runs: weaker judgment โ harder and vaguer questions โ more generic answers โ higher implementation rate of bad advice โ measurable damage.
Note what is not in that chain. Not access. Not usage volume. Not prompt engineering skill. The intervention that most enablement programs buy โ more licenses, more prompt training โ operates on none of the steps that produced the loss.
This is the same shape as Dell'Acqua and colleagues' finding with 758 BCG consultants: across 18 tasks inside the AI's capability frontier, consultants using AI completed 12.2% more tasks, 25.1% faster, at significantly higher quality; on one complex managerial task deliberately placed outside that frontier, AI-assisted consultants were 19% less likely to reach a correct solution than consultants working without AI at all (Harvard Business School, 2023). Two studies, different populations, same structure: the harm arrives when confident output meets a user who cannot evaluate it.
Why the remediation instinct is backwards
Here is the decision this reframes.
The intuitive move โ the one that shows up in most enablement plans โ is to give AI tools first to the people who are struggling. It reads as fair, it reads as efficient, and it borrows credibility from the call-center result. Lift the bottom, close the gap, bank the average.
On open-ended advisory work, that sequencing is inverted. You are handing an unfiltered advice engine to precisely the population least equipped to filter it, and the field evidence says the expected value is negative, not merely small.
The cost lands twice. First in the direct damage โ decisions made worse than they would have been unassisted. Second in attribution. When the pilot posts a flat or negative aggregate, the conclusion drawn is usually "the tool does not work here," and the program dies. The tool worked fine for the cohort that could use it. The rollout design was what failed, and nobody measured at the level that would have shown it.
Averages hide sign reversals. If your pilot reports one blended number, it cannot tell you whether you helped anyone.
What to change in your rollout this quarter
Scope open-ended AI use to demonstrated domain judgment
Open advisory use โ strategy questions, pricing calls, prioritization, anything where the model's answer cannot be checked against a known target โ should be scoped to people who already have the domain judgment to reject a bad answer. That is a capability gate, not a seniority gate: the ten-year employee who has never owned a P&L does not automatically pass it, and the two-year analyst who has been calling forecasts accurately does.
Give the struggling quartile narrow, verifiable work behind a review gate
The same people who are hurt by open advisory use are the ones who gain most from bounded tasks with checkable outputs โ the conditions that produced the 34% novice gain in the support-agent study. Drafting against a template, summarizing a known document, first-pass classification, structured data extraction. Add a human review gate that has authority to reject, not just to rubber-stamp. That is the configuration in which AI genuinely levels.
Measure the effect by performance quartile, not in aggregate
Instrument your pilot so the outcome metric is split by prior performance quartile before rollout begins. If you only look at the blended average, a +15% top quartile and a โ10% bottom quartile net out to a number close enough to zero that you will conclude nothing happened โ and you will have paid for the damage without seeing it.
Change what you train
Prompt training teaches people to get more fluent output. It does not teach them to reject fluent output that is wrong. What the evidence calls for is evaluation training: what a good answer looks like in this domain, what the model's known failure modes are, and โ most useful and least taught โ which questions to bring it and which to bring a person.
The counterargument worth taking seriously
The honest objection: Kenyan micro-entrepreneurs are not mid-market operations staff, and a WhatsApp advice bot is not an enterprise deployment with retrieval over your own documents, defined workflows, and audit trails. Both are true. The effect size will not transfer.
But the mechanism is not about the setting. It is about the interaction of open-ended questions, confident output, and an evaluator who cannot verify. Every enterprise deployment that puts a general assistant in front of an employee and invites them to ask it anything reproduces those three conditions. Grounding the model in your own data narrows the failure surface; it does not close it, and it makes the wrong answers sound more institutionally authoritative when they come.
The defensible position is not "this will happen to us at the same magnitude." It is "we have not measured whether it is happening to us at all" โ which, for most mid-market pilots reporting a single blended productivity number, is simply accurate.
The decision on your desk
The AI performance gap is not a fairness problem to be solved later. It is a design parameter you are setting right now, by default, every time you decide who gets access next.
Before your next tranche of licenses goes out, split the deployment: open-ended advisory use for people whose judgment you would already trust in a room without the tool, narrow verifiable tasks behind a review gate for everyone else โ and one metric, reported by performance quartile, so you find out which half you were right about.