Engineering

Half of designers are shipping AI-generated code to production. Here's where it breaks.

Half of designers are shipping AI-generated code to production. Here's where it breaks.

A 2026 industry survey put a number on something every engineering lead already felt in their gut: 91% of designers now use AI daily, and half of them say they've shipped AI-generated code straight to production. Not prototypes. Not throwaway experiments. Real frontend polish, component fixes, motion systems, accessibility patches — the actual product, written by a model and merged with a thumbs-up.

We're not here to tell you to stop. We use these tools every day, and we're not going back either. But we write production code for a living — React, Next.js, TypeScript, the stuff that has to survive past the handoff — and we've watched enough codebases age to know that "it works" and "it's done" are two very different claims. The gap between them is where the bill comes due.

The trap isn't bad code. It's code that looks finished.

The popular story is that AI writes sloppy code. That's not quite right, and getting it wrong leads teams to the wrong fix.

The more precise problem is that AI writes plausible code — code that compiles, passes the tests you thought to write, and reads cleanly enough that a tired reviewer approves it at 6pm. The output looks finished before it's actually ready. You get the time savings up front, for free. Then you pay it back later, with interest, in the form of testing, accessibility work, performance tuning, and security review that nobody scheduled because the feature already "shipped."

The data backs this up. A GitClear analysis of over 100 million lines of changed code found that code churn — lines reverted or rewritten within two weeks of being committed — jumped 39% in projects leaning heavily on AI tools. A separate study of 806 repositories that adopted an AI coding assistant measured a roughly 41% increase in code complexity and a 30% increase in static-analysis warnings, alongside a velocity boost that turned out to be transient. The speed is a sugar high. The complexity stays.

That's the shape of the problem: you move faster this quarter and slower every quarter after, until refactoring quietly becomes the job.

The real liability has a name: cognitive debt

Traditional technical debt is messy code you can see — the tangled function, the duplicated helper, the TODO that's been there since 2023. You can point at it. You can grep for it.

The debt AI introduces is different, and worse, because it's invisible. When a human writes a function, something happens alongside the typing: they build a mental model. They know why the retry logic is set to three attempts and not five. They remember the edge case that forced the weird conditional. That understanding lives in the team, not just the file.

AI-generated code skips that step. The function works, but nobody owns the reasoning behind it. Every approved suggestion that you didn't genuinely understand adds to a balance that no linter can detect. The industry has started calling it cognitive debt, and the failure mode is brutal in its honesty: in 2022, when something broke, the developer said "I'll fix it — I wrote it, I know where to look." In 2026, the developer says "let me try regenerating it with a different prompt." That second sentence should worry you. It means the team is debugging code nobody wrote and nobody comprehends.

For an interface, this is especially dangerous, because the parts AI gets superficially right are exactly the parts that hide depth. A dropdown that renders is not a dropdown that handles focus trapping, keyboard navigation, screen-reader announcements, and the seventeen states it occupies between empty and error. The happy path is cheap. The edge cases are the actual work — and they're the first thing a model under-delivers.

Where it actually breaks, specifically

In our experience, AI-generated frontend code fails in predictable places. Knowing them turns a vague anxiety into a checklist.

State and edge cases. Models design for the demo: the populated, logged-in, everything-loaded view. They quietly skip empty states, loading skeletons, partial failures, race conditions, and the offline case. These aren't polish — they're most of the real surface area of a component.

Accessibility. Generated markup tends to be visually correct and semantically hollow. Divs doing the job of buttons. Missing ARIA. Focus order that breaks the moment you unplug your mouse. With WCAG 3.0 raising the bar, "looks accessible" is not a defense.

Duplication. Ask a model the same thing in three files and you'll get three slightly different implementations of the same idea. The Don't-Repeat-Yourself principle erodes one convenient suggestion at a time, and the duplication doesn't announce itself until you need to change the behavior in all three places.

Silent decisions. This is the dangerous one. Every generated feature embeds choices you never explicitly made — a data shape, an error-handling strategy, a dependency — and those choices compound with every change built on top of them. You don't discover them when they're cheap to fix. You discover them six months later, in a debugging session that touches fifteen files.

The fix is a standard, not a tool

None of this is an argument against AI. It's an argument against the workflow most teams accidentally adopted: accept the suggestion, watch the tests pass, move on. That phase is over. The teams that treat AI output as a first draft — reviewed, understood, and refactored like any other code — keep their velocity. The teams that rubber-stamp it into production are scheduling the most expensive rewrite of their careers for sometime next year.

Here's the standard we hold, and the one we'd suggest you adopt regardless of who's writing your code:

Specify before you generate. The model is only as good as the contract you give it. Vague prompt, vague output, silent decisions. A clear spec — states, edge cases, accessibility requirements, the shape of the data — turns the AI from an author into a very fast implementer of your design. The thinking still has to happen first. It always did.

Review for comprehension, not just correctness. The question at the PR is not "do the tests pass?" It's "can someone on this team explain why this works?" If the answer is no, it's not ready, no matter how green the checks are. Code nobody understands is a liability disguised as a feature.

Own the architecture; let AI own the material. The most useful framing we've seen has engineers moving from bricklayers to architects: you define the blueprint, the constraints, the quality bar — the model lays the bricks. As AI gets more capable, human oversight gets more valuable, not less. The team that can orchestrate AI beats the team that merely consumes it.

Make the edge cases the deliverable. Design every state before a line of code is written — empty, loading, error, success, the awkward in-between. When the spec already contains the hard parts, the AI has far less room to skip them. This is just good interface practice, accelerated.

The bottom line

AI didn't change what good engineering is. It changed how fast you can produce code that merely looks like it. The velocity is genuine and worth having. But velocity without comprehension isn't progress — it's debt accumulating at machine speed, in a form your tools can't see and your future self will inherit.

The studios and teams that win the next two years won't be the ones who adopted AI fastest. Everyone adopted it. They'll be the ones who kept their standard while doing it — who questioned the output the same way they question a brief, and refused to ship anything they couldn't explain.

Built for performance. Built for teams. Built to last past the handoff. That part hasn't changed, and a model can't do it for you.

Related Articles

More on
this topic.

All Articles →

Where We Work

One studio, five markets.

Northeast India, Munich, Dubai and Denver each have their own page — the work we've done there, the people, and how to reach us.