Fifty-six percent of U.S. workers who used AI on the job in the previous week saved two hours or less. Thirteen percent saved nothing at all, or needed extra time to finish the work (U.S. Census Bureau, 2026; Nextgov/FCW, 2026).
That is the first federal, vendor-independent measurement of AI time savings at work. It is not a licensing pitch, a customer-success telemetry export, or a survey run by a company that sells seats. And for a Head of Operations building a 2026 AI business case, it settles a question that vendor data structurally cannot: what the median user actually gets back.
The answer is one to two hours. The problem is not that the number is small. It is that one to two scattered hours is not the unit your business case is denominated in.
What the Census Actually Measured About AI Time Savings at Work
The March 2026 Household Trends and Outlook Pulse Survey, released August 11, asked working respondents whether they had used AI for any of 11 specific job tasks. About 55% said yes (Census Bureau, 2026).
The task list is worth reading closely, because it describes the shape of real adoption rather than the shape of a roadmap:
- 37% โ searching for information or technical help
- 32% โ writing communications, documentation, or instructions
- 32% โ generating ideas
- 31% โ interpreting, translating, or summarizing information
- 27% โ administrative tasks
None of that is agentic. It is retrieval, drafting, and compression โ the substitutable middle of knowledge work, not the end-to-end process ownership most 2026 deployment plans are written around.
Then the time question. Among workers who used AI in the previous week, the distribution ran: 25% saved less than an hour, 31% saved one to two hours, roughly 30% saved three or more (split about evenly between three hours and four-plus), 10% saved no time, and 3% said the work took longer with AI than without it (Census Bureau, 2026; CBS News, 2026).
One methodological note that matters for how you cite this internally: the Census question is anchored to tasks completed in the prior week, not to a clean weekly total per person. If anything, that framing is generous. Usage is also thinner than adoption headlines imply โ only 24% of AI users touched it every day in the reference week, while 46% used it on at least one day.
Two Independent Federal Reads, One Number: About Two Hours
The most useful feature of this dataset is that it is the second federal-adjacent measurement to land near the same place.
The Bick, Blandin and Deming nationally representative survey โ the one behind the St. Louis Fed's adoption work โ asked generative AI users how many additional hours they would have needed to finish the prior week's work without it. The answer came to 5.4% of work hours, or roughly 2.2 hours in a 40-hour week (St. Louis Fed, 2024).
Two different instruments, two different years, two different question designs. Both land at about two hours.
Here is the part most people get backwards. The vendors are not far off on this number either. Forrester's Total Economic Impact study of Microsoft 365 Copilot builds its value engine on roughly 9 hours saved per user per month โ about 2.25 hours a week โ and still arrives at a 116% ROI for its composite organization, with $36.8M in benefits against $17.1M in costs over three years (Forrester, 2025).
So the disagreement between the federal data and the vendor model is not about hours. It is about what an hour is worth. And that is a modeling assumption you own, not one the vendor can validate for you.
Your Business Case Isn't Wrong About the Hours. It's Wrong About What They Are.
The standard arithmetic is hours saved ร loaded hourly rate ร headcount. It is clean, it survives a finance review, and it is wrong in a specific way: it assumes recovered time is fungible with purchased time.
It isn't. Two hours a week, distributed across a dozen interrupted tasks, is slack. It absorbs into the working day โ into slightly earlier finishes, slightly longer reviews, slightly less end-of-week compression. It does not consolidate into a reallocatable increment of capacity, because nothing in the operating model is designed to collect it. There is no mechanism in a 200-FTE company that sweeps forty people's scattered ninety minutes into one funded project.
This is the mid-market version of a macro result. Acemoglu's task-level accounting of AI's aggregate effect โ impacted task share multiplied by average task-level cost saving โ produces no more than a 0.66% increase in total factor productivity over ten years, precisely because meaningful per-task savings on a modest slice of work does not aggregate the way intuition insists it should (Acemoglu, 2024). Your P&L runs the same arithmetic at smaller scale, with the same result.
The operational test is binary and unforgiving: did a whole task leave someone's plate?
If a monthly reconciliation, a first-draft cycle, a tier-one ticket class, or a report build is now fully absorbed โ that is capacity, and it is bankable. It shows up as a vacancy you do not backfill, a contractor you do not renew, a queue that clears without overtime. If instead every task still requires a person and each one is somewhat faster, you have bought comfort. Comfort is a legitimate purchase. It is not a cost line, and it should never have been modeled as one.
The 13% Nobody Prices
Ten percent of AI users saved no time. Three percent needed more time than they would have without it (Nextgov/FCW, 2026).
I have not seen a mid-market business case that carries a negative tail. Every model I review treats savings as a floor of zero and a distribution above it. The federal data says roughly one in eight users is at or below zero โ and those users still consume a licence, a training slot, and manager attention.
That reframes deployment breadth. Universal rollout at a fixed per-seat price buys you the full distribution, including its bottom eighth. Targeted rollout to task profiles that match the top of the distribution โ high-volume drafting, summarization, retrieval-heavy roles โ buys you a truncated one. At 500 seats, the difference between those two curves is the whole ROI argument, and it is decided by procurement scope, not by the technology.
The Objection: "Our Pilot Showed More Than That"
It probably did. Two structural reasons why, both of which should temper how far you extrapolate.
Pilots select on enthusiasm. Volunteers for an AI pilot are drawn from the right tail of the very distribution the Census just published โ the 30% getting three or more hours. Rolling out to the full population moves you toward the median, not toward the pilot mean. A pilot that returned four hours a week per participant does not forecast four hours a week per employee; it forecasts the ceiling.
Self-reported time savings are estimates of a counterfactual. Both the Census and the St. Louis Fed instruments ask people to imagine how long the work would have taken without AI. That is a hard cognitive task, and it is not obviously biased downward. Nothing here corrects for the possibility that self-reports overstate.
Neither point means the technology underperforms. Both mean the same thing operationally: your business case should be built on the median of a full population, and stress-tested at the bottom decile โ not built on the pilot cohort and hoped forward.
Three Decisions Before the Quarter Closes
- Re-denominate the model in tasks, not hours. Take your current AI business case and identify every line where hours saved are converted to dollars. For each, name the specific task that has left a specific person's plate. Lines that survive stay. Lines that cannot name a task get reclassified from cost savings to quality or experience benefits โ still real, still worth funding, no longer load-bearing for the ROI.
- Price the bottom eighth explicitly. Add a line for the roughly 13% of users who net zero or negative. Multiply by fully loaded seat cost plus enablement. If that line changes your payback period materially, the answer is narrower deployment, not a better prompt library.
- Instrument one process end to end before renewal. Pick a single high-volume, well-bounded workflow. Measure cycle time and headcount touches before and after โ at the process level, not the individual level. One clean process-level measurement is worth more than a year of self-reported hours, and it is the only evidence that distinguishes slack from capacity.
The Question Worth Asking Before the Renewal
The Census Bureau's contribution is not that AI time savings are small. Two hours a week for more than half of a 55%-adopting workforce is a genuine and broad-based gain, and it converges with the independent federal estimate.
The contribution is that the number is now public, non-commercial, and stable enough to plan against โ which removes the last excuse for a business case that multiplies scattered minutes by a loaded rate and calls the product capacity.
Before you sign the renewal, ask your team one question and require a name, not a number: which task no longer requires a person on it? If nobody can answer, you have not bought capacity this year. You have bought a slightly easier week โ and it is worth knowing that before finance finds out in the variance report.