How do you catch a security flaw in AI-generated code that still runs correctly?

Run the code, don’t just read it: execute the generated code against the task’s test suite, and separately scan the diff for known vulnerability patterns, because code that works can still be insecure. Veracode’s 2025 GenAI Code Security Report tested more than 100 models against 80 curated coding tasks and found a security flaw in 45% of them, even as functional correctness kept improving across the same models. A grader that scores fluency or checks that the code compiles agrees with all of those 45%: nothing about a working answer says it defended against the input that breaks it.

A code-generation agent needs at least two separate verdicts on the same output: did it pass the test suite, and did the diff introduce a known vulnerability class. Neither verdict implies the other, the same way a tool call can return a result without erroring and still be wrong. A model can earn one of these verdicts routinely while failing the other, and a grader built only for correctness never finds out.

sources

keep reading

More on this.

Two ways to run Tessary.

Tessary is an open-source agent reliability platform. Cloud and self-hosted run the same workflow on the OpenTelemetry traces your agent already emits.

Tessary Cloud

We host it for you. Send your first trace with nothing to deploy and no model key.

what's includedper organization
traces
10,000 per calendar month
stored trace data
1 GB
retention
30 days
model credit
$10, one-time, for triage and root-cause analysis
credit card
not required

Self-hosted Tessary

Run the open-source code on your own infrastructure with one command. Add your own model key for triage and root-cause analysis.

Self-host Tessary for me by following https://github.com/tessaryai/tessary/blob/main/setup.md

docker compose -f oci://docker.io/tessaryai/tessary:compose up -d -y