Why Your AI Says "It's Secure" (and Why That Is Not Evidence)
Your AI says "it's secure." Why a model's self-report is not evidence, and what an AI-generated code security review should actually check.
Why does your AI say "it's secure"?
Language models produce confident security claims because confidence is a property of the text, not of the code. The model that wrote your checkout flow never ran it, never loaded your environment variables, and never queried your database. "It's secure" summarizes the paragraph the model just finished — nothing more.
SecurityReview.ai names the trap directly: AI coding tools make it easy to generate features quickly, the code looks clean, passes checks, and moves toward production. Clean formatting and a passing build answer a different question. Linters catch style, compilers catch syntax, and neither one knows whether an anon key can read your users table.
Regarding AI-generated code security review, the first job is separating "the build passed" from "the known failure patterns are absent." I ship with Cursor and Claude Code daily, and the reassurance arrives in the same confident register as the code. Register is not evidence. A scan result with a check ID is.
Fluency reads as assurance, and nothing in the training objective corrects for that. (Source: SecurityReview.ai)
What does a model actually know about your code?
The model knows the text it generated in the current session. The model does not know whether Row Level Security is enforced on your Supabase tables, whether the service-role key leaked into the client bundle, or whether a .env file got committed four commits ago. Session memory ends at the token window; production reality does not.
Regarding Supabase RLS, the classic failure starts with scaffolding: the AI creates the table, skips RLS enablement, or writes a permissive policy, and every row becomes readable with the anon key. The model will describe the policy it intended. The database enforces the policy that exists. Those are different artifacts, and only one of them runs.
Snyk puts a number on the base rate: 65–70% of production code is now AI-generated, and nearly half of it contains vulnerabilities. Veracode's 2025 GenAI Code Security Report found 45% of AI-generated code samples introduced an OWASP Top 10 vulnerability. The failures land in territory the OWASP Top 10 and the CWE catalog have mapped for years: broken access control, injection, misconfiguration.
The model describes its intention; the runtime enforces the file. (Source: Mindgard)
Is a second AI reviewing the code real evidence?
A second model reviewing the first model's pull request adds opinion, not proof. MindStudio makes the argument bluntly: the obvious fix — add another agent whose job is to review the pull request — loses to deterministic gates, because probabilistic code reviewed probabilistically compounds uncertainty instead of reducing it. Two models agree or disagree based on weights and prompts, and neither one runs your app.
AI review tools still earn a place in the stack. CodeRabbit provides context-aware, line-by-line PR review, and Endor Labs ships AI security review that catches issues rule-based scanners miss. Both work best on top of deterministic checks, not instead of them. SonarSource's answer from the quality-gate side: for projects containing AI-generated code, you apply an AI-qualified quality gate — a deterministic gate, defined in advance, that the code passes or fails.
Regarding model-reviewed code, one question sorts any review output: does the result change if the model feels different today? Deterministic checks say no. Model verdicts say yes.
Two opinions are not a measurement. (Source: MindStudio)
What does evidence look like in an AI-generated code security review?
Evidence is a deterministic check with a reproducible result: a rule that either matches the pattern or does not, aimed at the known failure modes of AI-written code. DevMeth's catalog encodes the 48 known AI-code failure patterns (C1-C48) plus 12 code-health checks (R1-R12), with engines named for the job — gitleaks for secrets, Semgrep for patterns, OSV for dependency CVEs.
| Your AI tells you | Evidence replaces it with |
|---|---|
| "Row Level Security is on" | A scan finding that shows the table without RLS, or a clean result on the same check |
| "No secrets in the client" | A gitleaks run over the repo and the built bundle |
| "Dependencies are fine" | An OSV query against your lockfile |
| "The API routes check the session" | A pattern check for authless API routes, plus a live URL probe |
Table: Model claims versus the deterministic evidence that replaces them.
The named patterns read like a highlight reel of AI-code failures: RLS disabled, service-role key exposure, secrets in client bundles, committed .env files, open admin routes, authless API routes, weak JWT secrets, IDOR, client-side role checks. The free tier runs the 10 critical checks (C1, C2, C3, C5, C6, C7, C8, C9, C13, C25). Regarding secrets in client bundles, the check inspects what actually shipped, because the bundle — not the source — is what the browser downloads.
A check either fires or it does not; that property is the whole point. (Source: SonarSource)
How do you replace "it's secure" with a verification loop?
The loop has four steps: run deterministic checks, read plain-English findings, paste the fix prompt back into your AI tool, and re-scan to verify the fix landed. Verification closes the loop because the same check that flagged the pattern confirms its absence. A green re-scan is a statement about the code on a specific date — evidence with a timestamp.
Regarding the verification loop, the paste-ready fix prompt does the heavy lifting: your AI wrote the bug, and the same tool fixes it fastest when the finding names the file, the pattern, and the check ID. Vague findings produce vague fixes. Specific findings produce diffs.
DevMeth is a pre-flight check, not a pentest — it checks the 48 known AI-code failure patterns and reports which ones stay open, then re-scans after your AI applies the fixes. The R1-R12 ruleset covers the tech-debt side models leave behind: hallucinated imports, dead code, drift. Regarding AI-generated code security review as a practice, the shift is permanent: stop accepting adjectives, start accepting check results. Your AI's confidence is where the review starts. Deterministic evidence is where it ends.
Shift-left now happens at generation time, not after the code exists. (Source: Arnica.io)
About the Author: Nick Thorp is the founder of DevMeth, the pre-flight security check for AI-built apps. He also built Proven Duty, file-review tooling for UK advice firms.
Frequently asked questions
Why does your AI say "it's secure"?
Language models produce confident security claims because confidence is a property of the text, not the code. The model never ran your app, loaded your environment, or queried your database. "It's secure" summarizes the paragraph it just wrote — a prediction, not a measurement.
What does a model actually know about your code?
The model knows the text it generated in the current session. The model does not know whether Row Level Security is enforced on your Supabase tables, whether the service-role key leaked into the client bundle, or whether a .env file got committed four commits ago.
Is a second AI reviewing the code real evidence?
A second model reviewing the first model's pull request adds opinion, not proof. Probabilistic code reviewed probabilistically compounds uncertainty. Deterministic gates — rules defined in advance that the code passes or fails — beat agent review, which is why SonarSource prescribes AI-qualified quality gates.
What does evidence look like in an AI-generated code security review?
Evidence is a deterministic check with a reproducible result: gitleaks over the repo and bundle for secrets, Semgrep for patterns like authless API routes, OSV against your lockfile for CVEs. DevMeth encodes the 48 known AI-code failure patterns (C1-C48) plus 12 code-health checks (R1-R12).
How do you replace "it's secure" with a verification loop?
Run deterministic checks, read plain-English findings, paste the fix prompt back into your AI tool, and re-scan to verify the fix landed. The same check that flagged the pattern confirms its absence, so a green re-scan is evidence with a timestamp.