If you're a founder or a technical leader right now, you've probably had this thought: AI can build software fast, so why pay an agency at all? Worth taking seriously, the instinct isn't crazy. AI didn't kill agency work, though. What it did was expose which agencies were only ever doing the part AI can now do by itself.
Start with the part nobody should argue about. Today's coding agents index an entire repository, follow your conventions, and run for hours without supervision. An idea becomes a working prototype in days instead of months. In the hands of someone who knows what they're doing, the output is good.
That's real. The gap starts right after, in the part that never makes it into the demo.
AI ships what it understands from your prompt, nothing more. Ask for a feature and you get that feature, scoped exactly to what you described. Andrej Karpathy called this "vibe coding" back in 2025: hand over a prompt, don't look too closely at what comes back. The model isn't guessing wrong when it does this. It's just answering the question you actually asked instead of the one you meant to.
This isn't a flaw in the model, it's the model working as designed. These systems are trained on an enormous amount of excellent work, so the capability to produce something great is genuinely there. The judgment isn't there, knowing when "good enough for the prompt" isn't good enough for the product. That call still belongs to whoever's asking.
AI isn't writing bad code. It's writing exactly what it was told, and most teams building solo haven't learned yet what that gap actually costs them.
Addy Osmani, who spent over a decade in engineering leadership at Google before recently joining Anthropic, calls it the 70% problem. AI gets you roughly 70% of the way there, fast. The remaining 30% is where most AI-built products actually fail.
It fails predictably. A database query that returns instantly against fifty rows crawls at fifty thousand. A request that works fine for one user starts colliding with itself the moment ten people trigger it at once. An input nobody thought to test takes something down at 2 a.m. None of it surfaces while you're building, because nobody's using the product yet. The prototype worked well enough to attract the traffic that eventually broke it.
This isn't only a knowledge gap, either. An experienced engineer building solo knows to guard against most of it, that part of the 30% does shrink with skill. But some of it is a timing gap, not a knowledge gap: nobody can fully verify how a system behaves under fifty thousand real users before fifty thousand real users actually show up. That data doesn't exist yet, no matter how good the person building is. Experience closes maybe half of this gap. Real traffic is the only thing that closes the other half, which is exactly why it has to get caught on the way there, not guessed at in advance.
That same 30% shows up in the budget too. Uber burned through its entire 2026 AI coding budget in four months, after an internal leaderboard pushed teams to maximize AI tool usage. By May, Uber's own COO couldn't trace a clear connection between what the company spent on AI coding tools and the features users actually received. "That link is not there yet," he said. When even Uber can't connect the spend to shipped value, that's not a budgeting problem. It's the same gap, showing up on the invoice instead of in the code.
When the models got good, the assumption was that AI would replace headcount outright, and plenty of companies acted on it. The correction never came.
Big tech is still cutting. Meta and Microsoft announced roughly 20,000 combined job cuts in April 2026, with Meta reducing about 10% of its workforce and closing 6,000 open roles on top of that. Nobody is quietly rehiring the engineers they let go. Adidas cut roughly half of its India Tech Hub in Gurugram this week too, an estimated 350 or more technology roles. Fresh enough that the exact scope is still being confirmed across outlets, but the shape matches everything above: a technology team specifically, with no claim anywhere that AI is covering the gap on its own.
The companies with the deepest engineering benches are thinning them at the exact moment AI is generating more code than ever, and the review burden hasn't gone anywhere. The demand has shifted toward engineers who can read what AI produces and spot what it missed. That's a different skill from writing the code yourself, and it's getting scarcer at the same rate the gap is getting wider.
AI can help you build the prototype. It cannot hand you the product.
It isn't one thing, it's a set of them, and none are visible from the outside. Load behavior under real traffic is one: query plans, indexing, and caching that hold at 50,000 rows rather than 50, since most prototypes are never tested past the developer's own clicking. Concurrency is another, what happens when two requests hit the same endpoint at the same moment, one of the most common categories of bug AI-generated code introduces. Then there are the error paths and input validation nobody wrote a prompt for, because nobody had hit them yet. Security review, authentication flows and credential handling checked by someone who's seen these fail before. And observability, logging and monitoring, so when something breaks you find out from a dashboard instead of a customer.
An experienced engineer checks these by reflex, not because they distrust the AI, but because they've watched each one fail in production before. Someone without that background reads the same code and sees working software. Which it is, right up until it isn't.
A few signals tend to arrive right before an AI-built prototype needs real engineering behind it. Real users show up, and the traffic that validates the idea is the same traffic that exposes what wasn't built to handle it. Sensitive data enters the picture, payments, customer records, anything regulated, and "it mostly works" stops being an acceptable standard. Nobody left on the team can explain why the code works the way it does, which means nobody can maintain it either. Or the AI spend starts outpacing user growth, the Uber pattern, just smaller.
Any one of those means the gap has arrived, not that anything failed. It means the thing worked well enough to need what comes next.
No. It relocated the job. The work used to be writing the code. Now it's knowing what the code is missing, and closing that before someone else finds it the hard way.
That's not a smaller job than it used to be. For most teams right now, it's a job nobody's actually staffed.