01Repeatable PL/EN evals
Controlled scenarios test evidence, factuality, prompt injection resistance, report language and the distinction between a gap and a missing skill.
02Versions, not silent changes
We record the model, prompt and rubric version so that a result can be tied to a specific configuration and compared after a change.
03Product quality signals
We monitor error categories, analysis success, cost and voluntary feedback — without sending CV or report content to analytics.
04Human review for material changes
A new model or material rubric change must pass the same tests and a review of representative outputs.