All articles AI

AI Fix Proposals in AppSec: What Works and What Does Not

Sophie Laurent

AI Fix Proposals in AppSec

Attaching a fix proposal to a vulnerability finding is a natural extension of what AppSec tools should do. If the tool knows enough to call something a vulnerability, it should know enough to suggest how to address it. In practice, the quality range between "this is useful and safe to review" and "this is worse than the original code" is wide, and understanding where that range comes from is important before relying on fix proposals in production workflows.

What Makes a Fix Proposal Hard to Get Right

The naive version of an AI fix proposal looks at the flagged line and substitutes a known-safe pattern. Swap the string-concatenated query for a parameterized one. Replace the MD5 call with SHA-256. Add the missing authorization check. For isolated, well-defined vulnerability classes, this works reliably. The difficulty starts when the fix requires understanding context the tool cannot see from the affected line alone.

Consider a stored XSS finding in a rendering function. A simple fix might add HTML encoding to the output. But if the same value is used elsewhere in the call stack, encoding it in one place may not close the vulnerability. If the value is legitimately trusted markup in one rendering path and user-supplied in another, encoding it universally may break functionality. The correct fix depends on where the data originates, how it is routed, and what downstream components expect. A model that looks at the flagged line in isolation will miss all of this.

Call Graph Awareness

Fix quality correlates closely with how much call graph context the proposal generation has access to. A fix that was produced after tracing the data flow from user input through multiple service calls to the vulnerable output is meaningfully different from one generated by pattern matching on the flagged function.

This is technically expensive to do well. Full interprocedural analysis across a large codebase requires the same infrastructure that powers high-quality SAST scanning, extended to cover multiple files and services. The upside is that proposals generated with this context tend to be accurate at the scope they target and do not introduce regressions in adjacent paths. Our approach builds fix proposals on top of the same call graph analysis that produces the finding, rather than treating proposal generation as a separate downstream step.

Scoping Is a Separate Problem From Accuracy

A fix can be technically correct and still be the wrong thing to apply. Scoping errors are common in generated fix proposals and they take a different form than accuracy errors. An accuracy error means the fix does not actually remediate the vulnerability. A scoping error means the fix addresses the vulnerability but is too broad, too narrow, or touches code outside the intended change boundary.

An example of a scoping error: the tool proposes adding input validation to a function, and the proposed validation is correct, but it also modifies the function signature in a way that breaks callers. The proposal is technically sound at the line level and incorrect at the interface level. Another: the proposal adds a security check but duplicates logic already present in a middleware layer the tool did not trace through.

Catching these requires understanding the module's interface contract, not just its internal logic. For teams adopting fix proposals in CI/CD pipelines, this means the review step is not optional. The goal is not to auto-apply fixes without human review. The goal is to make the review fast enough that it gets done, rather than deferred.

Language and Framework Specificity

Fix quality degrades noticeably when a model trained primarily on one language ecosystem is applied to another. A parameterized query proposal in Python with SQLAlchemy looks different from one in Java with JDBC, which looks different from one in Go with database/sql. Getting the syntax right is the easy part. Getting the idiomatic pattern right, using the project's existing abstractions rather than introducing new dependencies, requires the model to have seen enough of that ecosystem's conventions to follow them.

This is an argument for language-specific fine-tuning and for proposals that read the project's existing code patterns before suggesting changes. A team using a custom ORM wrapper should receive fix proposals that use that wrapper, not proposals that introduce a new raw query abstraction.

What We Have Seen Work

From our early-access deployments, a few patterns stand out. Fix proposals for high-signal vulnerability classes with predictable remediation patterns, specifically secrets exposure, insecure direct object reference, and missing output encoding, tend to be accurate and ready to review with minimal friction. These classes have well-defined fix shapes and limited inter-function dependencies.

Fix proposals for logic-heavy authorization bugs or complex deserialization vulnerabilities tend to require more review time. The proposals are often directionally correct but need adjustment to fit the application's specific permission model or data handling conventions. For these classes, the value is still there: the developer has a starting point and a concrete articulation of what the fix should accomplish, rather than an abstract CWE description.

We are not claiming that AI fix proposals eliminate review. They do not, and a workflow that applies them without review will introduce regressions. The claim is narrower: a well-scoped, call-graph-aware fix proposal reduces the research burden on the developer reviewing the finding, which reduces the time between finding and committed fix, which is the metric that actually matters for AppSec program health.

The Review-to-Apply Ratio

A reasonable internal benchmark for evaluating fix proposal quality is the ratio of proposals accepted with minor or no edits versus proposals that required significant rework. A high ratio of minor-edit accepts indicates the model is well-calibrated to the codebase's patterns. A high ratio of significant reworks suggests either scoping problems, framework unfamiliarity, or false positives in the underlying findings driving proposals for vulnerabilities that are not real.

Tracking this ratio per vulnerability class rather than in aggregate reveals which classes are performing well and which are not, and guides where to focus improvement effort on the proposal generation side.

Tenzai

See findings and fixes together, not just findings.

Free plan available. Connect your first repository in under five minutes.

Start Free Trial

More from the Tenzai blog