Short answer: check before you rebuild. What breaks first in a vibe-coded app is whatever nobody prompted for: access rules on the data, secret keys kept out of the browser, a backup that restores, and tests around the paths that take money. Check those four, in that order. Rebuild only when the data model itself is wrong.

Vibe coding got its name in February 2025, when Andrej Karpathy described a way of building where you talk to the model, accept what it writes and "forget that the code even exists". Collins made it the word of the year nine months later. By then a founder could go from an idea to a working product over a weekend, and plenty did.

That part is real, and we're not here to argue with it. This note is about week ten. There are users, somebody has paid, and an investor's engineer (or your own) has asked a question about the app that you can't answer.

To see what goes wrong, and in which order, we went through the public record: three scans of live apps, one academic benchmark, two vendor reports and one incident that made the news.

What goes wrong in vibe-coded apps, according to the evidence?

  • Matt Palmer, May 2025. Out of 1,645 apps built with Lovable, 170 had database tables a stranger could reach. It became CVE-2025-48757.

  • Escape, October 2025. A scan of 5,600 live apps found over 2,000 vulnerabilities, more than 400 exposed secrets and 175 cases of personal data in the open, medical records and bank account numbers among them.

  • Red Access, May 2026. Roughly 380,000 publicly reachable assets built with Lovable, Replit, Netlify and Base44. About 5,000 of them exposed potentially sensitive information.

  • SusVibes, revised September 2026. A benchmark of 186 real feature requests. One leading setup, SWE-Agent with Claude 4 Sonnet, got 57% of them working. Only 11.8% of its solutions were secure.

  • Veracode, July 2026. Roughly 44% of AI code generation tasks introduced a risky vulnerability. On cross-site scripting the models passed 15% of the time.

  • GitGuardian, 2026. Commits co-authored with Claude Code leaked a secret 3.2% of the time, against 1.5% for all public GitHub commits.

  • Replit and SaaStr, July 2025. An AI agent deleted a production database during a code freeze, in the middle of a founder's public experiment.

Two cautions. Most of these numbers come from companies that sell security tools, and the scans looked at apps that were easy to find, which leans toward demo and hobby projects. Escape also scanned passively (no attack attempts), so its count is a floor.

Even so, read the list again. Nearly every item is the same kind of fault.

Why do these apps break in the same places?

Because a model builds what it is asked for, and nobody asks for the things that don't show.

A prompt says "users can see their orders". It doesn't say "and nobody else can see them". The first sentence produces a screen you can demo. The second produces nothing you can look at, so it gets left out, and the app still works for the person who built it (signed in as themselves, with ten rows of data).

SusVibes measured that gap directly: 57% working against 11.8% secure. The researchers also tried adding security hints to the request, and the hints did not close it. So "I told it to be secure" isn't a fix.

The same goes for backups, error handling and tests. None of them is visible in a demo. All of them are what an engineer means by "production".

What should you check first?

Four things, in the order of the damage they do. The first three you can do yourself in an afternoon.

Can a stranger read your data?

Make two accounts. Sign in as the second and try to open something that belongs to the first: change the id in the address bar, or repeat a request with a different number. If it loads, everything else on this list can wait.

If the app sits on Supabase, open each table and look at row level security. It has to be switched on, and the policies have to mention the user. That is exactly what was missing in Palmer's 170 apps. Lovable later added a scanner, but by his account it checks that a policy exists and can't know whether it is the right one.

Are there secret keys in the browser?

Open the site, look through the page source and the scripts it loads, and search for sk_live, sk-, sb_secret and the plain word secret. Some keys are meant to be public. Supabase's anon key is one, and it's only safe when the table rules above are right. A payment secret, a mail API key or a Supabase service key in the browser is an open door.

Deleting the key from the code is not enough. If it was ever in the front end, or anywhere in the repository's history, treat it as copied and issue a new one.

Can you get yesterday back?

Find out whether backups exist. Then restore one somewhere and look at it. A backup nobody has restored is a hope.

Next, look at what the AI agent itself can touch. In the Replit case the agent had access to the live database and used it, against an instruction to change nothing. Replit's answer was to separate development and production databases automatically. Do the same: the agent works on a copy, and production credentials live somewhere it can't read.

What happens on the unhappy path?

Pay with a card that declines. Submit the same form twice. Close the tab in the middle of checkout. Type <b>test</b> into your own profile name and reload: if the name comes back bold, the app is treating what users type as code.

Each of these is a path the prompt never described. When you've tried them, write down the three or four flows that carry money or personal data and get an automated test around each. That short list is worth more than broad coverage of everything else, because it tells you whether tomorrow's change broke yesterday's payment.

Should you fix the app or rebuild it?

Fix it when the faults are at the edges. Rebuild when they're in the middle.

Edges are what the checks above find: missing access rules, keys in the wrong place, no tests, no backups. They're serious. They're also bounded — an engineer can list them, and the list ends.

The middle is the data model. Signs that it's wrong: the same fact stored in three tables, with no way to say which one is true. Prices or permissions worked out in the browser. Tables that mirror screens instead of the things the business deals in (customers, orders, invoices). When that's the shape, every fix moves the problem somewhere else.

There's a third option, and it gets overlooked. Keep the interface and what the app does, and rebuild what sits underneath. A working prototype is the most precise specification most founders will ever have: real screens, real flows, and a record of what users did with them.

We don't know of public data on how often each outcome happens, and we'd be wary of anyone who quotes a rebuild rate. It takes reading the code.

Can you keep vibe coding after the fix?

Yes, with a few rails around it.

The code lives in a repository, and every change arrives as a pull request that somebody reads. The tests from the last section run on each one. The agent holds no production credentials. Secrets sit in a secrets manager and never in a prompt.

That is roughly how professional teams already use these tools. In Stack Overflow's 2025 survey, 84% of developers used or planned to use AI, and 72% said vibe coding was no part of their professional work. Same models, with a person reading the output.

It does move the wait into review, though. We went through the numbers on that in Is AI making developers faster?

Common questions

Is vibe coding safe?

For a prototype with made-up data, yes. For real users' data, not by default: the setup measured in SusVibes produced secure code 11.8% of the time. It becomes safe the way any code does, when someone checks it.

Will a security scanner fix it?

A scanner finds what can be seen from outside: a leaked key, a table with no rules at all. It can't tell whether a rule matches how your business works. Run one, and do the two-account test anyway.

How much does it cost to fix a vibe-coded app?

Nobody can say without reading the code, and a quote given before that is a guess. The four checks cost you an afternoon. After that, the first thing worth paying for is a written second opinion — ours is the Consult step.

Is vibe coding the same as AI-assisted development?

The tools are the same. The difference is whether anyone reads the result. Karpathy's definition is about not reading it, which is fine for a weekend project and a poor fit for other people's data.

More field notes