A linter is a pattern matcher with good intentions. It parses one file, looks for shapes it recognises, and reports them. That model is genuinely good at a large class of problems: unparameterised queries in the function you are looking at, a hardcoded credential two lines below, a `dangerouslySetInnerHTML` in a component.
It breaks down in one specific and very common situation: when the vulnerability is not in any single file, but in the relationship between files.
The shape of the problem
Consider a missing tenant filter. The bug is not in the query. The query is perfectly correct SQL. The bug is the *absence* of a `WHERE organization_id = ?` clause that was supposed to be added by a caller three directories away.
No single-file analyser can see this. Not because the implementation is weak, but because the information required to detect the bug is not present in the file being analysed. This is a category error, not a coverage gap.
// repositories/order.ts — looks fine in isolation
export async function findOrders(db, tenantId) {
return db.query("SELECT * FROM orders WHERE status = $1", [status]);
}
// services/checkout.ts — the bug lives HERE, and is invisible from above
const order = await findOrders(db, req.user.tenantId); // tenantId never applied
return order.items.map((i) => chargeCard(i));Why chunk-level AI review inherits the same limit
The natural response was to add an LLM. Models are excellent at reasoning about intent, and reviewers assumed that reasoning would paper over the missing context.
It does not, because of how context is delivered. A model reviewing a diff is shown a window of the change. If the sanitising call happens in a file that is not in the window, the model is not able to reason about its absence. It will confidently review what it can see, which is exactly the failure mode that makes teams disable tools.
The fix is not a bigger window. It is compiling the whole codebase to a structure the reviewer can query, then asking questions against that structure instead of against a text slice.
- Parse every file in the dependency cone into an AST, not just the diff.
- Build call and data-flow edges across file boundaries, including through interfaces.
- Ask the model whether a required edge exists, rather than whether a pattern is present.
- Report the path when it is missing, so the developer can verify the reasoning.
The measurable difference
This is the axis our evaluation isolates. On single-file categories such as secret detection, pattern matching and AST analysis both do well. On cross-file taint tracking the gap is structural: a pattern tool can only see the file in front of it, while a whole-codebase index can follow the flow.
The takeaway is not that linters are bad. They are fast, deterministic, and excellent at what they cover. They are simply the wrong tool for architectural questions, and treating them as a security control creates a false sense of coverage.