You’ve seen it even if you never gave it a name. A landing page with a purple gradient, three rounded cards and a heading that says “Supercharge your workflow”. A pull request that adds four hundred lines where forty would do, with a try/catch around everything. A chatbot answer that’s fluent, confident, nicely formatted, and wrong.
That’s slop. The shape of good work w/o the substance.
People made slop long before AI. What changed is the cost: it’s nearly free now, and it shows up faster than anyone can review it.
I use AI every day btw. I’ve trained more than twenty engineers on AI coding assistants and I build AI systems that run in prod, so this isn’t an anti-tools post. It’s about where a human has to stay in teh loop.
1. The average is nobody’s taste
A model gives you the most likely continuation, and “most likely” drags everything toward the average (the average landing page has a gradient and three cards). The average has nobody’s taste and nobody’s context in it.
Good design is mostly subtraction. AI is bad at subtraction. Ask it for a dashboard and you get twelve charts, because it has no way of knowing the operator on the production floor is wearing gloves and has thirty seconds between tasks.
When we moved Genba follow-up into WhatsApp, the big decision was to not build another app. People on the floor already had WhatsApp open all day, so the interface became a message they reply to with #CA, #PA and #STATUS. no model would’ve suggested dropping the app. That took someone standing on the floor, watching how people actually work.
Code too. Generated code is cheap to write and expensive to own. If the AI produced 800 lines, the task was too big, and “the AI wrote it” is never an answer in code review.
2. Answers you can check
For Q&A systems slop is worse, because a wrong answer looks exactly like a right one.
The ebook assistant I built answers production questions from hundreds of documented factory cases, and people act on those answers. So it assumes the model will sometimes be wrong:
- every answer cites the exact case or SOP, one click to check
- links are checked against the catalogue on the server, because the model once invented plausible-looking ones
- which cases a user may see is a permission check on every search, not a line in the prompt (a prompt can be talked around)
- a fixed set of real questions runs as a test suite w/ an LLM judge, so “it seems better” isn’t a result
- operators can flag answers; in one review cycle 43 actionable reports came in and 33 were fixed and deployed
That’s the harness imo. Everything around the model that catches its mistakes before a person acts on them.
So the human stays on deciding what to build (and what not to), reviewing what ships, and owning it when it breaks.
AI made producing things nearly free. It didn’t make judging them free.
Keep using it, tbh. Just make sure someone who cares is still deciding what ships.