Engine calibration record
How our engine scores against known answers.
We run our assessment engine against applications we built with known answers — planted weaknesses and clean controls whose correct verdict we already have — and publish the results. This is that record: single and repeated runs, so you can see the accuracy to expect, and check the figures yourself. It is a measurement, not a marketing claim.
Method
The measurement is the same product a customer gets.
Each run drives the exact production engine over the entire published check set — not a reduced subset. The model reads only the code and the framework requirements; the answer key is never in its context, and is applied afterward, in scoring. We measure in both directions — what it catches and what it wrongly flags — and any flag on a check we had not pre-labelled is reviewed and locked by a person before it counts. Repeated runs measure the same code several times, so a single lucky pass cannot inflate the result and instability cannot hide.
To be plain about it: this is us testing ourselves. We authored these applications and we authored their answer key — it is an internal calibration, not an independently-verified benchmark. We publish the counts, the model, the index version and the date so the arithmetic is yours to check, but the ground truth is ours.
Results
No calibration run has been published yet.
This page records measured accuracy — the real engine against known answers — as each run is reviewed and published. There is nothing to show yet; please check back.
How to read these figures
Scope: these are controlled measurements of the engine against applications we built and labelled — not a statement about any customer's application, not a third-party audit, and not a regulatory certification. Every figure is tied to a named model, a published index version, and a date, and is derived from the counts shown beside it.