Preventing Enterprise AI Slop Grenades: An Audit Guide to Unvetted AI Output
What a slop grenade is
A slop grenade is unvetted AI output handed to a colleague or a client as if it were finished work. Shopify’s chief executive, Tobi Lütke, used the phrase on The Knowledge Project in September 2026 to describe what he was already seeing inside his own company. The AI bill does not move; someone else’s afternoon does.
Farnam Street’s Shane Parrish states the definition in Brain Food No. 699: a slop grenade “is when you let AI produce the work and pass it on without adding any value (including checking it). Someone else has to wade through it, catch the mistakes, and clean up the mess. You save time and look productive but someone else pays for it.”
Lütke’s own wording was plainer: “So we call those Slop grenades that people toss at each other.” Nothing in that description is metaphorical. The sender converts their own review work into someone else’s queue, and the queue is usually read by the more senior, more expensive person. The rest of this page turns the phrase into a twelve-point gate you can attach a name to.
Where the term came from — and who actually coined it
Press coverage in mid-September 2026 credited Tobi Lütke with coining “slop grenades”. The public record says otherwise, and every step of it is dated.
On 2026-09-15 Lütke posted: “Full credit to @harrybrundage for coining the term Slop Grenade. Let’s make it a thing.” The credited account, @harrybrundage, replied the same afternoon: “@tobi I stole it!” — with a shortened link that expands to noslopgrenade.com. Asked about the exchange, Lütke answered: “More like slop hot potato then”.
That site is not new. The Internet Archive holds a capture from 2026-01-31, the earliest recommendation of its etiquette page in our record is a Hacker News comment from 2026-01-28, and a May 2026 Hacker News thread pointing at it reached 717 points and 417 comments.
Lütke’s podcast words are narrower than the coverage. He said the phrase was already in internal use at Shopify — “internally we have come to call these things that people are lobbing slop grenades at each other which i think is a really fun term that we should push into industry” — and separately: “So we call those Slop grenades that people toss at each other.” Popularise, yes. Coin, no.
If you are writing the history of this term, the credit post and the reply are one click apart and both carry timestamps. Anyone naming a single founder as its author has skipped them. That gap between the primary record and the coverage is the argument for this page: a claim this easy to check travelled this far unchecked.
Why this is an audit problem, not an etiquette problem
The etiquette framing is real, but it stops at the message. The page the term came from puts it sharply: “Even when your answer is technically correct, the format is hostile to how humans communicate.” That is a description of a social problem. An audit begins when you stop treating it as one.
The reason is that the cost is transferred, not created. The AI-industry glossary that dates the coinage frames it as “the slop grenade is verification debt handed to someone else: the sender skips the review, the recipient pays it” A debt has a counterparty, and a counterparty can be named — which is the whole of what an audit needs.
Scale is what turns a habit into a queue. Lütke put the volume at “probably up to about 50% of pull requests in Shopify” And there is no machine to bill when it goes wrong: “Machines can't take responsibility, and I think this is actually probably the most overlooked thing in the entire, um, uh, stack”
Someone will object that the behaviour is self-limiting — that anyone who made a habit of it would be removed from a team. Read carefully, that is an argument for detection, not for leniency: make the unread hand-off visible. A label is a signal; a gate is a mechanism.
What unvetted AI output actually costs
Two bodies of work measure this, and they have to be kept apart.
The primary is the workslop research. Its definition is “Workslop is AI-generated content that looks good, but lacks substance.” Its method statement reads: “Insights are based on an online survey of 1,150 full-time U.S. desk workers conducted in September 2025 by BetterUp in partnership with the Stanford Social Media Lab.” BetterUp’s own summary of that survey: “40% Percentage of U.S. desk workers who received workslop last month. 2 hrs Average time it takes to resolve each incident. $186 Monthly cost per employee caused by these incidents. $9M Annual cost for a 10,000-person company.”
The secondary is the 2026 wave, which at the time of writing exists only as Fortune’s report: “BetterUp Labs and Stanford’s Social Media Lab surveyed 962 American full-time desk workers this year and found over half (52.7%) reported sending workslop to colleagues and it was more common in people whose organizations encouraged AI use. Over a third (38%) reported receiving workslop and estimated it cost them 3.4 hours per month on average to revise it.” Treat those figures as that outlet’s reporting rather than as a page the vendor has published, and do not sum the two waves — they are separate surveys with separate samples and separate years.
For scale on the sending side, Fortune also quotes Duolingo’s chief executive: “We may need to write 1,000 different stories for people to learn a language, then you’ll find that 20% of the things were just pure slop”
All of those are labour hours, which is a different cost line from token spend and infrastructure. For that one, see the cost of rework.
The 12-point AI output verification checklist
Run every AI-assisted artefact through these twelve checks before it leaves the hands of the person who produced it. Each row names the check, the artefact it leaves behind, and the failure it prevents. Keep the artefact: that is what makes a gate auditable rather than a matter of taste.
- 1. Named owner — artefact: The name of the human authoring the output. Failure prevented: the “the AI wrote it” defence, which is not a defence.
- 2. Source list — artefact: Every claim and figure traced to a named, dated source. Failure prevented: confident fabrication that reads exactly like knowledge.
- 3. Numeric re-derivation — artefact: Each number recomputed or re-read at the source, not copied. Failure prevented: stale rates and plausible-looking arithmetic.
- 4. Link check — artefact: Every citation fetched and confirmed to say what you claim. Failure prevented: sources that 404, redirect, or argue the opposite.
- 5. Claim-by-claim red-team — artefact: A second person tries to falsify each load-bearing claim. Failure prevented: the one wrong sentence that discredits the whole document.
- 6. Client-data scan — artefact: A check for customer names, credentials and internal figures. Failure prevented: a data exposure that outlives the document.
- 7. Tone pass — artefact: A human rewrites the machine cadence, hedging and restated headings. Failure prevented: the format hostility that makes a correct answer unusable.
- 8. Length bound — artefact: The artefact sits inside an agreed limit for its type. Failure prevented: offloading the reading to the recipient.
- 9. Label — artefact: The disclosure line telling the reader what was AI-assisted. Failure prevented: a reader assuming a verification that never happened.
- 10. Sign-off — artefact: A named approver accepts ownership before it is sent. Failure prevented: work leaving the team with nobody accountable for it.
- 11. Retention — artefact: Prompt, output and checklist stored somewhere findable. Failure prevented: an incident nobody can reconstruct three months later.
- 12. Post-gate review — artefact: A quarterly review of what the gate missed, and a dated change. Failure prevented: a checklist that ossifies into ritual.
Points 1, 5, 10 and 12 fail first. The others can be delegated to tooling; those four cannot. For the agent-side equivalent — what an autonomous system may do without a human in the loop — start from the AI agent risk checklist. This gate governs the hand-off between people, not the run itself.
Where the gates run: pre-send, client-deliverable, production
Those twelve checks are one list, but they run in three places, and the standard tightens at each one.
Gate 1 — pre-send (internal). Anything going to a colleague: a summary, a draft, a status note. Owner and source list are mandatory; the red-team is optional for low-stakes internal work. The floor is that the sender has read every word they are sending.
Gate 2 — client-deliverable. Everything in gate 1, plus numeric re-derivation, link check, client-data scan, tone pass and the label. This is the point where “the model produced it” stops being context and becomes a disclosure obligation. Our existing human checkpoints cover where a person must intervene inside an agent run; this gate covers the hand-off between two people.
Gate 3 — production. Anything reaching a customer system, a repository or a payment: sign-off, retention and a recorded rollback path are required, and the agent’s own permissions are part of the review — see agent permissions.
The common failure is applying gate 3 discipline to nothing at all and calling everything else “just internal”. Internal is where the habit forms. The test is not which gate a piece of work falls under; it is that somebody decided, in writing, and that the decision is findable later.
The enterprise AI quality control audit (5 artefacts)
If you have to show a board, a client or an auditor that output control exists, these are the five artefacts that prove it. Each is a file, not a practice.
- Register — the AI systems, tools and agent runs in use, each with a named owner.
- Sample — a dated random pull of AI-assisted outputs from the last quarter, including the ones that shipped.
- Findings — for each sampled item, which of the twelve checks it passed and which it missed.
- Owner map — one accountable human per output class, plus a deputy.
- Sign-off log — dated approvals, refusals and rollbacks in a single place.
The output of that audit is a table of gaps, not a score. A score invites an argument; a gap invites a fix.
This is the operational half of the work. For the regulatory and vendor-oversight half, see AI safety and compliance audit.
Policy, ownership, and the sign-off rule
Two sentences do most of the work in a written AI policy.
One named human is accountable for every output that leaves the team. Not a team, not a tool, not “the AI”. If a policy cannot answer who signed this, it is not a policy.
No unlabelled AI output leaves the team. The label is the cheapest control on the list and the one people resist hardest, because it feels like a confession. Framed as provenance rather than confession it is the same move your disclosure process already makes; AI content disclosure covers what to mark and where.
Attach the consequence to the gate, not to the use of AI. A missed sign-off is an incident; drafting with an assistant is not. Punish the skipped step and people will stop confessing it.
And keep the policy short enough to be read. A twelve-page AI policy nobody has opened protects nobody. A one-page gate that is actually run does.
Rolling this out without banning AI
Start with one team and one output class. Client deliverables are usually the right first target, because the pain is already visible and the standard is already agreed.
Run the gate for four weeks and collect two numbers: how many artefacts were caught by checks 2, 4 and 6, and how many minutes the gate adds per artefact. A zero catch rate for a month means the gate either is not being run or is aimed at the wrong class. An unmanageable time cost means cutting checks — never the sign-off.
Expect the label to be the contentious part. The label is the signal; the gate is the mechanism. Keep both, and be explicit that the policy governs what leaves the team, not which tools people may open. For the wider automation programme this sits inside, see the AI automation checklist.
Frequently asked questions
What is a slop grenade?
A slop grenade is unvetted AI output passed to a colleague or a client as if it were finished work. The sender skips the checking; the recipient has to read it, find the errors, and rebuild it. The phrase reached mainstream business coverage in September 2026 when Shopify's chief executive, Tobi Lütke, used it on The Knowledge Project. The cost is not the AI bill — it is the working hours the recipient spends re-deriving work nobody verified.
Who coined the term slop grenade?
Not Tobi Lütke, and not Shane Parrish. Lütke's own post credits another account: “Full credit to @harrybrundage for coining the term Slop Grenade. Let’s make it a thing.” That account replied “@tobi I stole it!” with a link to noslopgrenade.com — a site the Internet Archive captured on 2026-01-31 and that Hacker News discussed on 2026-01-28, months before the podcast. Several September 2026 write-ups attribute the coinage to Lütke; the two dated posts above are the primary record.
Is a slop grenade the same as workslop?
No. Workslop names the artefact — “Workslop is AI-generated content that looks good, but lacks substance.” It comes from research by BetterUp Labs with the Stanford Social Media Lab, reported through Harvard Business Review in September 2025. A slop grenade names the hand-off: the act of passing an unchecked artefact along. A team can have one without the other, and the twelve-point gate below targets the hand-off, not the artefact.
How many hours does unvetted AI output cost a team?
The primary survey of 1,150 U.S. desk workers, run in September 2025 by BetterUp with the Stanford Social Media Lab, reports that 40% received workslop in a month and that resolving each incident takes about 2 hours — roughly $186 per employee per month. A separately reported 2026 wave puts revision time at 3.4 hours a month; treat that figure as the reporting outlet's, not the vendor's published page, and do not add the two surveys together.
What should an AI output verification checklist contain?
Twelve checks: a named owner, a source list, numeric re-derivation, a link check, a claim-by-claim red-team, a client-data scan, a tone pass, a length bound, a label, a sign-off, retention, and a quarterly review of what the gate missed. The gate is only real if each check leaves an artefact — a name, a calculation, a log, a dated approval — because an artefact is what an auditor can read.
Who is accountable when AI-generated work is wrong?
The person who sends it. A model cannot hold a duty of care, and Lütke put the gap plainly: “Machines can't take responsibility, and I think this is actually probably the most overlooked thing in the entire, um, uh, stack” Internal accountability therefore has to be assigned before the work moves: one named human per output class, with a deputy. This is an operational answer about ownership and sign-off, not legal advice about liability.
Does this mean we should ban AI at work?
No — that is our recommendation, not a research finding. The failure mode is unverified output leaving a team, not the use of an assistant. The fix is a gate plus a label: a twelve-point check that produces artefacts, and a disclosure line the reader can see. Teams that ban the tools lose the speed and keep the risk, because the work moves to personal accounts where nothing is visible.
Sources
Sources: The Knowledge Project episode transcript with Tobi Lütke, 2026-09-15 (podscripts mirror); the credit post @tobi, 2026-09-15; the sourcing reply @harrybrundage, 2026-09-15; the definition post @shaneparrish, 2026-09-15; noslopgrenade.com and its Internet Archive capture of 2026-01-31; Hacker News item 48219992, 2026-05-21; BetterUp / Stanford Social Media Lab workslop survey, September 2025; Fortune, 2026-09-17; Business Insider, 2026-09-17; Farnam Street and Brain Food No. 699; ADI Pod glossary. All links resolved live on 2026-09-18.