The question
A vendor tells you their agent scored 87 and the bar is 85. Three follow-ups usually end the conversation: Why 85? Who set it? What would have happened if a different panel had set it? ISO/IEC 17024 requires a documented standard-setting study for exactly this reason — a threshold with no procedure behind it is an opinion wearing a number’s clothes.
Where EvalQA currently stands
EvalQA does not today publish a formally derived cut score, because one has not been produced. planning The design below is what we have committed to run, and the results will be published here with the panel composition and the raw ratings. Until that happens we report severity-classified findings and coverage rather than a pass/fail verdict — which is also why there is no badge.
Why we report findings instead of a verdict, for now
A single pass/fail number is the most commercially attractive output an assurance vendor can produce and the easiest to abuse. Without a defensible cut score, a verdict transfers our judgement into your governance process while hiding the fact that the boundary was chosen rather than derived. Severity-classified findings put the decision where the Independence Charter says it belongs: with your executive sponsor, who owns the residual risk.
The study we have committed to run
- Method: modified Angoff with a borderline-group check. Panellists estimate, item by item, the probability that a minimally acceptable agent would handle it correctly. The cut score is the aggregate of those estimates, then cross-checked against the observed score distribution of agents the panel independently judges borderline.
- Panel composition. Analytics engineers who own semantic layers, a data governance lead, a risk or audit representative, and at least one practitioner with no commercial relationship to EvalQA. Panel size and names published with the result.
- Two rounds with feedback. Panellists rate, see the group distribution and impact data, then rate again. Round-one and round-two ratings are both published, because the movement between them is itself evidence about how stable the standard is.
- Consequence weighting. Separate cut scores by consequence level. A Sev-1 miss on a regulatory report is not interchangeable with a Sev-4 inefficiency, and a single threshold across both is a category error.
- Published standard error. The cut score is an estimate with uncertainty. A result within the standard error of the boundary is reported as indeterminate, not rounded into a pass.
What we will publish
- The panel: roles, sector, and whether each member has any commercial relationship with us.
- Both rounds of item-level ratings, in full.
- The derived cut score per consequence level, with its standard error.
- The borderline-group distribution used as the cross-check.
- The date, and the rubric version the study was run against — see the rubric changelog.
What would invalidate it
A standard-setting study is tied to the instrument it was run on. A material change to the rubric, to the failure taxonomy, or to the judge configuration invalidates the cut score and requires a re-run. We will say so on this page rather than quietly carrying an old number forward, which is the most common failure mode in credentialing and the easiest one to check for.
Further reading
The underlying method is covered at length in the curriculum: borderline-group standard setting, item analysis, and NCCA and ISO 17024.