Scandrix
Evaluation methodology · work in progress

AI code review,
measured on real PRs.

How we test the ScanDrix engine: an 11-domain eval gate with behavioral checks, a weighted composite score, and a fail-closed gate. No competitor runs, no leaderboard — yet.

Eval domains
11
Score pillars
4 · 0.75 bar
Location tolerance
±2 lines
Gate profiles
harness · local · CI

Evaluation Design

01
Step

Large diffs break shallow reviewers

Once a PR grows past a few hundred lines, chunk-based tools lose cross-file context and miss taint flows that span handlers, services, and DB layers. ScanDrix compiles the full AST before reasoning.

Explore AST compilation ↗
02
Step

Static rules plateau early

Pattern linters struggle because they cannot reason about architecture or cross-file data flow. ScanDrix combines tree-sitter AST queries with multi-model contextual reasoning over call graphs.

See Drixy rule engine ↗
03
Step

Deterministic evaluation gate

11 engine eval domains enforce contracts, schemas, and behavioral checks. A fatal failure blocks — it never warns and continues, guaranteeing sub-minute reviews with ±2 lines ground-truth tolerance.

Read benchmark methodology ↗

02 — Category breakdown

What the harness checks.

Six capability areas, each exercised by the evaluation harness. We publish the methodology — not leaderboard scores, and no competitor numbers we did not measure.

Eval domains

11 engine areas

Composite score

4 pillars · 0.75 bar

Leaderboard scores

Not published yet

OWASP Top 10 security

Injection, XSS, auth & access flaws

Our approach

AST parsing plus cross-file taint tracking

Evaluation

Severity + secondary evals

Logic bug detection

Broken invariants & edge cases

Our approach

Multi-model reasoning over call graphs

Evaluation

Scorer + investigation evals

Secret & credential leaks

Tokens, keys & passwords in diffs

Our approach

Deterministic secret patterns on every diff

Evaluation

Promotion gate

Cross-file taint tracking

Source-to-sink flows across packages

Our approach

Dataflow analysis across file boundaries

Evaluation

Parser + anchoring evals

Drixy rules enforcement

Custom org & repo policies

Our approach

Declarative rules with Dry Run previews

Evaluation

Behavioral scorer (±2 lines)

PR turnaround

Time from webhook to posted review

Our approach

Async worker queue, streaming comments

Evaluation

Design goal: sub-minute

Evaluation status

Internal evaluation in progress

No public leaderboard yet — and no competitor scores we did not measure. Scores stay with the harness until independent runs exist.

Method: an 11-domain engine eval gate — anchoring, dedup, severity, format, investigation, rules, parser, summary, promotion, scorer, secondary — with a weighted four-pillar composite score and a fail-closed gate. No competitor runs, no leaderboard — yet.

03 — How to read these results

Three things we actually track.

Marketing benchmarks cherry-pick one metric. A review tool lives or dies on all three — miss bugs, slow the team, or cry wolf, and it gets disabled. So these are the three axes of the harness, not three scores.

Location-true scoring

A finding counts when it names the right file within ±2 lines — recall and precision against ground truth, plus F1.

Machine-checkable output

Format and anchor location are scored pillars, so malformed or unplaceable findings fail the same gate as wrong ones.

One bar, fail-closed

Four pillars weighted into a composite with a 0.75 pass bar. Below it the gate blocks — in harness, local, and CI profiles.

04 — Methodology

Reproducible by design.

No tuned demos. One fixed gate, one composite bar — and no rankings we did not earn. The bar is exact file and line, checked automatically.

Read the full evaluation spec
01

Eleven engine domains

Anchoring, dedup, severity, format, investigation, rules, parser, summary, promotion, scorer, secondary — each with behavioral checks.

02

Ground truth with tolerance

Rule violations scored against ground-truth file + line sites within ±2 lines: recall, precision, F1.

03

One composite bar

Recall, precision, format and anchor location weighted into a single score with a 0.75 pass bar.

04

Fail-closed gate

Three profiles — harness, local, CI. A fatal failure blocks the gate instead of warning and continuing.

Stop trusting. Start measuring.

Test ScanDrix on your own repositories.

Connect GitHub or GitLab in minutes. Read your first line-precise review before believing any number on this site — free trial, no credit card.

Methodology, not a leaderboard · live service state on /status