02Our testing, plainly
Where our testing numbers come from.
These describe the evaluation harness, not customer telemetry. The harness scores findings against ground truth automatically — and publishes no leaderboard.
We do not present a modelled saving as a result. The ROI calculator applies a stated 65% review-time assumption. That is an estimate to test against your own numbers, not a measurement, and it is labelled that way there.
- Engine eval gate
- 11 domainsAnchoring, dedup, severity, format, investigation, rules, parser, summary, promotion, scorer, secondary — each with behavioral checks.
- Composite score
- 4 pillars · 0.75 barRecall, precision, format and anchor location weighted into one number per run. Below the bar, the gate blocks.
- Ground-truth tolerance
- ±2 linesRule violations score against ground-truth file + line sites: recall, precision, F1.
- Public scores
- Not publishedNo leaderboard and no competitor numbers until independent runs exist. Methodology first.