Glossary
The vocabulary of AI assurance is used loosely, and the difference between an audit, an evaluation, and a certification decides what a reader is entitled to conclude. These are the definitions we work to.
Each entry says what the term means when we use it. Where a definition is set by a standard rather than by us, the standard governs.
AI assurance
The activities by which confidence in an AI system is established and communicated to somebody who has to rely on it, covering evaluation, audit, red teaming, and conformity assessment.
AI safety
The reduction of harm caused by AI systems, covering both the harm a system does when it works as designed and the harm it does when it fails.
Evaluation
Measurement of how a system performs against a specified task, on a specified population, under specified conditions. It establishes what happened rather than whether a rule was followed.
Audit
Examination of a system against a stated criterion, by a party independent of the one being examined. An audit asks whether a requirement is met. An evaluation asks what the system does.
Validation
Confirmation that a system is fit for its intended use, as distinct from verification, which confirms it was built to specification. A system can be correctly built and still be unfit.
Access tier
The level of access an evaluator was given, which bounds what may honestly be claimed. Black box, grey box, and white box are set out on Methods.
Red teaming
Adversarial testing in which the evaluator tries to make the system fail. It establishes that a failure is reachable, not how often it occurs in service.
Conformity assessment
A determination that a product meets the requirements of a standard or a regulation, carried out under rules the standard sets, by bodies accredited for the purpose.
Independent, or third party
Carried out by a party that neither supplies nor operates the thing being examined. Testing produced by the operator is first party, however rigorous, and does not meet the independence expectation.
Model risk
The risk that a model is wrong, is used wrongly, or is relied on beyond the conditions its evidence covers. The term comes from financial supervision and transfers cleanly to AI.
Released under CC BY 4.0. May be quoted with attribution.