How do you catch a security flaw in AI-generated code that still runs correctly?
Run the code, don’t just read it: execute the generated code against the task’s test suite, and separately scan the diff for known vulnerability patterns, because code that works can still be insecure. Veracode’s 2025 GenAI Code Security Report tested more than 100 models against 80 curated coding tasks and found a security flaw in 45% of them, even as functional correctness kept improving across the same models. A grader that scores fluency or checks that the code compiles agrees with all of those 45%: nothing about a working answer says it defended against the input that breaks it.
A code-generation agent needs at least two separate verdicts on the same output: did it pass the test suite, and did the diff introduce a known vulnerability class. Neither verdict implies the other, the same way a tool call can return a result without erroring and still be wrong. A model can earn one of these verdicts routinely while failing the other, and a grader built only for correctness never finds out.