Originally published on Product Tribe

You can hand off the doing, but not the watching. AI made execution cheap - and turned the leftover job into review work nobody budgeted for.

Somewhere around the fourth AI tool in your stack, you stopped designing and started supervising, and nobody sent a memo about it. One week you were sketching flows. A few months later you were mostly reading things other systems made, deciding if they were good enough, fixing what wasn’t. Somebody in your all-hands called this 10x productivity. Somebody selling a $497 “AI Workflows for Designers” course called it something even better, with a bigger number attached, because bigger numbers sell more seats. What happened is you got pushed into a job nobody gave you a title for, and you didn’t exactly say no: the reviewer.

You can hand off the doing, but not the watching, and that second part was never the free half of the job - it just used to come bundled in, so nobody had to name it, let alone pay for it.

The delegation illusion

Boston Consulting Group asked 1,488 workers, in a study published this March, whether AI leaves them mentally fried, and 14% said yes - worst in marketing, at roughly one in four. Sure, BCG sells AI transformation for a living, so read the exact numbers the way you’d read a headline from someone with a horse in the race. But the shape of the finding doesn’t need BCG to be true: people whose job is mostly overseeing AI reported meaningfully more mental effort, more fatigue, more information overload than people whose job isn’t - the tool got faster, but the watching never caught up.

None of this is new, which is the part that should worry you. In 1983, an ergonomics researcher named Lisanne Bainbridge wrote a paper about chemical plants and other automated control rooms that’s still one of the most-cited things in her field, called “Ironies of Automation.” Her point: better automation makes the human’s leftover job harder, not easier, because you stop practicing the skill it replaced. Your judgment about it goes soft. And the one job left for you - deciding whether the machine is right - is the job you’re now worst prepared for.

There’s a study I trust more than the BCG one, because nobody’s selling anything in it. Researchers put sixteen drivers on ~150 kilometers of real road, part of it on autopilot, and measured their mental load two ways: what the drivers said they felt, and what an actual reaction-time test picked up. What the drivers said didn’t match what the test found: they said it felt easier, while their measured load went up, especially in traffic. They were wrong about their own effort, and wrong in the specific direction that matters - they underestimated the cost, not overestimated it. Vigilance has been known to be hard mental work since at least 2008. Watching just doesn’t feel like work, which is exactly why it gets skipped first.

Where it actually gets complicated

I’d be lying if I stopped there, though. Microsoft ran a survey of 319 knowledge workers across 936 real tasks last year, and the honest finding cuts the other way half the time: on routine, low-stakes work, people genuinely think less with AI in the loop, and that’s fine - that’s the delegation working as advertised. It’s only on the tasks where being wrong costs something that people reported putting in more effort with AI than without it.

That’s the annoying part - those are exactly the tasks you can’t afford to phone in, and the ones your calendar keeps scheduling back to back with the low-stakes ones, as if your brain can tell the difference on autopilot. It can’t, and that’s rather the point. Checking your third AI-generated onboarding flow before lunch is lighter than building it from scratch. Checking your ninth one, once you’ve stopped actually reading and started pattern-matching for “looks about right,” is a worse kind of tired, and it’s the kind that produces the bug you find in QA three weeks later.

Nobody puts any of this on a roadmap, because nobody’s measuring it. Your company tracks how many AI tools shipped this quarter. It doesn’t track the hour you spent unpicking a beautifully wrong flow at nine at night, or the review you rushed because you’d stopped believing your own tiredness counted as information.

I’ve watched this exact bet get made. A client’s leadership pushed teams onto early-stage AI tools - Make, whatever the pitch deck was selling that quarter - while cutting headcount at the same time, because the tools were supposed to cover the gap. The tools weren’t ready for anything production-grade. The designers who stayed ended up rebuilding most of that output by hand, after hours, because what shipped simply didn’t work. Upstairs called it a success. It had, after all, cut the budget.

Do the math yourself too, since no one up there will: BCG’s own data shows fried reviewers make 39% more major errors and report 33% more decision fatigue than reviewers who aren’t. A major error caught in QA costs an afternoon. The same error caught after ship costs a hotfix, a postmortem, and someone else’s afternoon too. Multiply either one by every reviewer on your team, every sprint, and you get a number that’s never once made it into a slide - not because it’s small, but because putting it there would mean admitting where it came from.

If anyone upstairs wanted the real cost of the rollout, they’d ask how many of those bugs trace back to a reviewer too fried to catch them. Nobody’s asking, because the honest answer would mean admitting the tool didn’t do what the slide deck said it would - and that’s a worse quarter than just letting you absorb it quietly.

I’ve done this too, more than I’d like to admit. Some days I’m four or five rounds into iterating something with AI, tweaking, re-prompting, switching between two or three tools that are all half-handling different pieces of the same problem, and somewhere around round six my standards quietly drop. I stop caring whether it’s good and start caring whether it’s done. Context-switching between tasks was already hard. Context-switching between the tasks and the tools supposedly doing them for me is a different order of tired, and most days that end like that, I close the laptop more drained than I ever felt doing the same work by hand - which was supposed to be the whole point of not doing it by hand.

So the honest version of this letter’s whole argument is narrower than “supervising AI always costs more than doing it yourself.” It costs more when the stakes are high, the output is unfamiliar, the tool is unreliable, or you’re four hours past when you should’ve stopped - and you’ll underestimate that cost, the same way those drivers did, because attention spent catching nothing feels like attention you never spent.

It’s also why your tool count matters more than your tool choice. Open Product Hunt on any given morning and you’ll find dozens of new AI tools launched that week alone, each one promising to be the thing that finally saves you time. Designers surveyed this year reported their kit more than doubling, from three tools to seven, in twelve months. (That number comes from a VC-backed report, and even their own methodology note admits the sample skews toward people who’ve gone all-in on AI.)

But it rhymes with what BCG found independently: productivity climbs through your second and third tool, then falls off a cliff. You didn’t decide seven was the number. It happened one plugin, one “just try this too” at a time, the same way scope creep happens to a project, except this scope creep happens behind your own eyes, and there’s no PM around to flag it.

What actually changes in practice

I’m not telling you to use less AI. If that’s the letter you were bracing for, close the tab - I have no interest in writing you a eulogy for doing everything by hand. What I’m telling you is narrower: budget the cost of watching the way you already budget the cost of doing, because right now you don’t, and the costs nobody budgets for are the ones that eventually eat the project.

Cap how many AI tools you’re actively supervising at once to two or three, roughly where BCG’s own curve turns over; past that point you’re not reviewing more, you’re just missing more. Batch your review passes into two or three fixed windows instead of checking as things trickle in. Research on notification batching, not AI-specific but close enough, found batched checking cut daily stress, while both constant checking and going fully dark made things worse.

Protect one real block of your day too, ninety minutes if you can get it, where you’re doing the work yourself instead of grading someone else’s attempt. Call it practice, not nostalgia. It keeps your judgment sharp enough to catch the machine when it’s confidently, fluently wrong, which, quietly, has become the actual job.

Be honest, too, about which of your reviews are the ones that don’t matter much and which are the ones where a miss costs something, and treat them differently on purpose - skim the first kind, and slow down for the second, on purpose, before you’re tired enough that slowing down stops being a decision and turns into something that just happens to you late in the evening.

You can delegate the sketch. You can delegate the first draft, the variants, the boilerplate. What you can’t delegate is the moment you decide whether any of it was right. That moment was always yours. All that’s changed is how easy it’s gotten to skip it - and call the skipping trust.