Every organisation eventually writes a rule that is technically correct and practically hated. The rule is right, the matches are real, and within a week someone adds a suppression comment to a file they were supposed to be improving.
The failure is almost never the rule's logic. It is shipping the rule without measuring it against the codebase that has to live with it.
Rehearse on history, not on the future
You cannot evaluate a rule's precision by reading it. You can only evaluate it against real commits. That is what a dry run is for: replay the rule over historical pull requests and report what it would have flagged, before it is allowed to comment on anything new.
# Replay a rule against a diff before enabling it
scandrix rules test --diff pr.diff --rule ./drixy/no-tenant-filter.yaml
# Reports: flags raised, files touched, precision estimate, suggested suppressionsRead three numbers
Precision tells you whether to keep the rule. Recall tells you whether it is worth anything. Blast radius tells you whether you can ship it today.
- Precision — of everything it flagged, how much was genuinely wrong? Below ~70% and you are generating noise.
- Recall — of the seeded issues it should have caught, how many did it? Below ~80% and it is giving false confidence.
- Blast radius — how many files in the last 90 days would it have touched? Above a few hundred, ship it as warn-only first.
Ship it in stages
The rollout that works in practice is three-stage and takes about two weeks. Report-only for one week, so the team sees findings without a CI failure. Then warn on changed lines only, so pre-existing debt does not block anyone. Then enforce on new code, leaving grandfathered lines alone.
Teams that follow this sequence keep their rules enabled. Teams that enable on day one do not — and the difference is almost entirely down to whether the team got to see the rule's output before it started gating their merges.