Workslop: When AI Output Looks Finished and Isn't
Artificial Intelligence
12 mins

Workslop: When AI Output Looks Finished and Isn't

Workslop is AI-generated work that looks polished but has no substance, so the person who receives it does the thinking the sender skipped. This article looks at the research behind the term, what it costs, why it happens, and what it turns into when it reaches a court or a regulator.

A word for something everyone had already noticed

In September 2025 a group of researchers from BetterUp Labs and the Stanford Social Media Lab published an article in Harvard Business Review with a title that did most of the work: "AI-Generated 'Workslop' Is Destroying Productivity". The authors, Kate Niederhoffer, Gabriella Rosen Kellerman, Angela Lee, Alex Liebscher, Kristina Rapuano and Jeffrey T. Hancock, defined workslop as "AI generated work content that masquerades as good work, but lacks the substance to meaningfully advance a given task".

The term travelled quickly, for the reason good coinages usually do. It named something people had already been experiencing without having a word for it: the document that is well formatted, correctly structured, appropriately long, and useless. The deck that covers all the headings and says nothing. The analysis that reads fluently until you check a number.

What makes it more than a complaint is the second half of the authors' description, which identifies where the cost lands: "The insidious effect of workslop is that it shifts the burden of the work downstream, requiring the receiver to interpret, correct, or redo the work."

That is the whole mechanism. Workslop does not destroy value at the point of creation, where it looks like a productivity gain. It destroys value at the point of receipt, where nobody is measuring.

The numbers, and their limits

The underlying research was an online survey of 1,150 full-time US desk workers, conducted in September 2025 by BetterUp in partnership with the Stanford Social Media Lab.

Around 40% of respondents reported receiving workslop in the previous month. Those who had encountered it estimated that roughly 15% of the content they receive at work qualifies. Each instance cost the recipient just under two hours to deal with. Translated into money, the authors put the invisible tax at $186 per employee per month, and calculated that for an organisation of 10,000 people this comes to over $9m a year.

The finding that ought to give pause to anyone sending it is not about time at all. Roughly half of recipients said they viewed the colleague who sent it as less creative, less capable and less reliable. About 42% said less trustworthy. A third reported being less likely to want to work with that person again, and a third told a teammate or a manager about it. Sending workslop is not a neutral act of efficiency. It is a reputational transaction, and the sender is on the wrong side of it.

The direction of flow is worth noting too: roughly 40% of it moves between peers, 18% travels upwards from direct reports to managers, and 16% travels down from managers to teams.

Two honest caveats. The figures come from magazine-published practitioner research, not from a peer-reviewed paper: there is no published questionnaire, dataset or methods appendix. And some numbers appear inconsistently across the original article and its authors' own later write-ups, with prevalence given as both 40% and 41%, and the time per incident as both one hour 51 minutes and one hour 56 minutes. The differences are small and the direction is not in doubt, but a reader in a regulated industry should know that this is a well-designed survey rather than a replicated study.

Corroboration from elsewhere

The phenomenon has independent support, arriving from different directions.

Workday published research in January 2026, based on 3,200 respondents surveyed by Hanover Research at organisations with revenue over $100m, all of them active AI users. Its central finding is the same shape as workslop measured from the other end: "Nearly 40% of AI time savings are lost to rework, including correcting errors, rewriting content, and verifying outputs from one-size-fits-all AI tools." Only 14% of employees consistently got clear, positive net outcomes from AI.

Gartner adopted the term outright in its January 2026 release on future of work trends for HR leaders, describing an "overwhelming focus on AI adoption and improving individual employee productivity" as having "led to 'workslop' - an abundance of fast but poor-quality work produced by or with AI". In the same release it reported that 88% of HR leaders said their organisations had not realised significant business value from AI tools.

Deloitte used the word once, in a single sentence in its Tech Trends 2026 chapter on agentic AI, noting that "poorly designed agentic applications can actually add work to a process, with some enterprises finding agentic 'workslop' can make processes even less efficient". A single sentence is not an endorsement of a framework, but it is a signal about which vocabulary is being adopted at that level.

The wider cultural marker: Merriam-Webster made "slop" its word of the year for 2025, defining it as "digital content of low quality that is produced usually in quantity by means of artificial intelligence".

Why it happens

The mechanism is an asymmetry. A preprint from a group of researchers including authors at the Alan Turing Institute names it directly, identifying "asymmetry effort" as a defining feature of AI slop: it takes vastly less effort to generate than it would without AI, while the effort required to check it is unchanged. Generation got a hundred times cheaper. Verification did not get cheaper at all.

Christian Catalini, a research scientist at MIT Sloan, framed the consequence as a balance sheet item: "If we do not invest in verification, we're accumulating hidden risk. It is technical debt accumulating behind the scenes, and, at some point, it'll come due."

The second cause is that fluency reads as competence. A language model produces text that has the surface properties of careful work, because those properties are what it was trained to reproduce. There is no correlation, at the level of the prose, between how finished something looks and whether it is right. Every heuristic a professional has developed over a career for judging whether a document was taken seriously, formatting, structure, tone, length, completeness, has been decoupled from the thing it used to indicate.

The third cause is that nobody is measuring the receiving end. If the sender's productivity is counted and the recipient's rework is not, the system will produce more workslop indefinitely, and the numbers will look good throughout.

There is a fourth, and it is the most uncomfortable, because it applies to people who are trying hard rather than to people cutting corners. In 2025 METR ran a randomised controlled trial with 16 experienced open-source developers across 246 real issues. Developers using AI tools took 19% longer to complete their work. Before starting, they expected to be about 24% faster. Afterwards, having actually been slower, they still believed they had been roughly 20% faster.

That is the finding to take seriously. These were skilled people, measured, on their own code, and they could not detect a substantial slowdown in themselves. The idea that someone will spontaneously notice their own output has degraded is not supported by the evidence. METR themselves have since said the result should be read carefully: in February 2026 they announced they were redesigning the experiment, having hit severe selection effects in a follow-up study, and they have always cautioned that the finding does not generalise beyond experienced developers working in large mature codebases. But the perception gap is the part that generalises, and it is the part that matters here.

What workslop becomes in a regulated setting

In most organisations workslop is expensive and annoying. In a setting where documents are relied upon by someone else, it becomes something else.

The clearest public record is in the courts, because courts write down what happened. Damien Charlotin maintains a database of decisions in which a court or tribunal has found that a party relied on hallucinated material. As of early September 2026 it listed just over 2,000 such cases, the largest numbers in the United States, Canada and Australia.

The founding case is Mata v Avianca in 2023, in which two New York attorneys submitted a brief citing six judicial opinions that did not exist, and were sanctioned $5,000. Judge P. Kevin Castel's formulation has held up: "a fake opinion is not existing law", and "an attempt to persuade a court by relying on fake opinions is an abuse of the adversary system."

The leading English authority is Ayinde v London Borough of Haringey, heard with Al-Haroun v Qatar National Bank, decided by the Divisional Court in June 2025. One case involved five non-existent authorities; the other involved 45 problematic citations, 18 of them entirely fictitious. The court's statement of the duty is the sentence to circulate internally: "Those who use artificial intelligence to conduct legal research notwithstanding these risks have a professional duty therefore to check the accuracy of such research by reference to authoritative sources." The court also endorsed an earlier formulation, from Mrs Justice Dias, of what putting fake authorities before a court amounts to: it "is prima facie only explicable as either a conscious attempt to mislead or an unacceptable failure to exercise reasonable diligence to verify the material relied upon".

Professional services is not exempt. In October 2025 it was reported that Deloitte Australia had refunded more than A$97,000 on a report for the Australian Department of Employment and Workplace Relations after it was found to contain a fabricated quotation from a Federal Court judgment and references to academic papers that did not exist.

And then there is the case that says the most about how this happens. In Kohls v Ellison, decided in Minnesota in January 2025, a court excluded an entire expert declaration because it contained three citation errors including two non-existent academic articles. The expert was Professor Jeff Hancock of Stanford, who is a co-author of the original workslop article. Judge Laura Provinzino's line has been widely quoted since: "Professor Hancock, a credentialed expert on the dangers of AI and misinformation, has fallen victim to the siren call of relying too heavily on AI."

That is not a point against the research. It is the strongest possible evidence for it. If one of the people who documented the phenomenon can be caught by it, in a sworn declaration, in his own field of expertise, then nobody should be confident that awareness alone is a control.

What actually works

FINRA's 2026 regulatory oversight report offers the most usable framework, because it starts from the position that nothing has changed. Existing rules "continue to apply when firms use GenAI or similar technologies in the course of their businesses", covering communications, recordkeeping and fair dealing. On supervision it is specific: "If a firm is relying on Gen AI tools as part of its supervisory system, its policies and procedures may consider the integrity, reliability and accuracy of the AI model", with "validation and human-in-the-loop review of model outputs, including performing regular checks for errors or bias" and "ongoing monitoring of prompts, responses and outputs".

Beyond that, four things follow from the research rather than from anyone's policy.

Measure the receiving end. Any productivity metric that counts output and ignores rework will reward workslop. If you are going to measure AI adoption at all, measure what happens to the work downstream of it.

Make provenance ordinary. The reputational damage in the survey attaches to being caught, not to using AI. Teams where saying "I drafted this with AI, I have checked the figures and not the framing" is a normal sentence do not generate this problem, because the recipient knows what they are receiving.

The follow-up HBR piece in January 2026 identified the two causes as leaders "issuing vague directives for employees to start using extremely powerful tools", and employees being "overburdened, psychologically depleted, and operating in environments where it doesn't feel safe to admit uncertainty or ask for help". Both of those are management conditions rather than technology problems, which points at the next two.

Do not mandate usage without defining quality. Gartner's reading is that workslop follows from pressure to adopt AI without the autonomy to judge whether the output is fit for purpose. An adoption target with no quality standard attached is an instruction to produce volume.

Put verification where the risk is. Verification capacity is the scarce resource. Spending it evenly across everything is the same as spending it on nothing. The document going to a regulator, a court or a client needs a named human who has checked the citations. The internal summary does not.

The uncomfortable conclusion of the research is that the tools are working as intended. They generate plausible, well-formed, complete-looking work at very low cost, which is precisely what was asked of them. What has not kept pace is the organisational habit of checking, and the habit of noticing when something that looks finished is not.

References

The original research

Kate Niederhoffer, Gabriella Rosen Kellerman, Angela Lee, Alex Liebscher, Kristina Rapuano and Jeffrey T. Hancock - AI-Generated 'Workslop' Is Destroying Productivity, Harvard Business Review (22 September 2025) https://hbr.org/2025/09/ai-generated-workslop-is-destroying-productivity

BetterUp Labs - Workslop: the hidden cost of AI-generated busywork https://www.betterup.com/workslop

Emily Peck - AI 'workslop' sabotages productivity, study finds, Axios (24 September 2025) https://www.axios.com/2025/09/24/ai-workslop-workplace-efficiency-study

Kate Niederhoffer, Alexi Robichaux and Jeffrey T. Hancock - Why People Create AI 'Workslop' and How to Stop It, Harvard Business Review (16 January 2026) https://hbr.org/2026/01/why-people-create-ai-workslop-and-how-to-stop-it

Corroborating research

Workday - New Workday Research: Companies Are Leaving AI Gains on the Table (14 January 2026) https://newsroom.workday.com/2026-01-14-New-Workday-Research-Companies-Are-Leaving-AI-Gains-on-the-Table

Gartner - Gartner Identifies the Top Future of Work Trends for CHROs in 2026 (12 January 2026) https://www.gartner.com/en/newsroom/press-releases/2026-01-12-gartner-identifies-the-top-future-of-work-trends-for-chros-in-2026

Deloitte - The agentic reality check, Tech Trends 2026 (10 December 2025) https://www.deloitte.com/us/en/insights/topics/technology-management/tech-trends/2026/agentic-ai-strategy.html

Merriam-Webster word of the year 2025, reported by Anna Furman for the Associated Press via PBS NewsHour (15 December 2025) https://www.pbs.org/newshour/nation/merriam-websters-word-of-the-year-for-2025-is-ais-slop

Why it happens

METR - Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity (10 July 2025) https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/

METR - We are Changing our Developer Productivity Experiment Design (24 February 2026) https://metr.org/blog/2026-02-24-uplift-update/

Cody Kommers et al. - Why Slop Matters, arXiv:2601.06060 preprint (23 December 2025) https://arxiv.org/abs/2601.06060

Seb Murray - Seeing real value from AI depends on being able to verify its outputs, MIT Sloan (8 June 2026) https://mitsloan.mit.edu/ideas-made-to-matter/seeing-real-value-ai-depends-being-able-to-verify-its-outputs

Consequences in regulated settings

Damien Charlotin - AI Hallucination Cases database https://www.damiencharlotin.com/hallucinations/

Mata v Avianca, Inc., 678 F.Supp.3d 443 (S.D.N.Y., 22 June 2023) https://www.law.berkeley.edu/wp-content/uploads/archive/2025/12/Mata-v-Avianca-Inc.pdf

Ayinde v London Borough of Haringey; Al-Haroun v Qatar National Bank [2025] EWHC 1383 (Admin) (6 June 2025) https://www.judiciary.uk/wp-content/uploads/2025/06/Ayinde-v-London-Borough-of-Haringey-and-Al-Haroun-v-Qatar-National-Bank.pdf

Kohls v Ellison, No. 0:24-cv-03754 (D. Minn., 10 January 2025) https://law.justia.com/cases/federal/district-courts/minnesota/mndce/0:2024cv03754/220348/46/

Alexei Alexis - Deloitte refunds over $60K for report with AI errors, Australian government says, CFO Dive (21 October 2025) https://www.cfodive.com/news/deloitte-refunds-60k-report-ai-errors-australian-government-accounting/803321/

FINRA - GenAI: continuing and emerging trends, 2026 FINRA Annual Regulatory Oversight Report https://www.finra.org/rules-guidance/guidance/reports/2026-finra-annual-regulatory-oversight-report/gen-ai

Image

Image by shogun on Pixabay (Pixabay image ID 9000198), used under the Pixabay Content License, which permits free use without attribution. Credit given as a courtesy.

July 15, 2026

Read our latest

Blog posts