Quantile Labs Contact

AI Evaluation & Assurance

We help operators, regulators, and the public understand the performance and risk of deployed AI systems.

Areas of work

Six domains where AI already carries weight in decisions at scale.

Financial services and credit

Scorecards and fraud models decide who is lent to. We test them on thin-file applicants and on households paid in cash.

Health

Triage and screening decide who is seen first. We measure them against the presentations of the clinic where they will run.

Public administration and identity

Identity and eligibility systems establish who the state recognises. We measure how often recognition holds, and for whom.

Climate, agriculture, and disaster risk

Earth observation sits behind early warnings and insurance payouts. We establish how much of the population at risk it covers.

Security and defence

Detection and targeting systems mark who is treated as a threat. We test how often that mark falls on the wrong person.

General-purpose and frontier models

Models are consulted as advisors before anyone certifies them. We measure how often they are confident and wrong, language by language.

Publications

  1. The 94% Problem

    Kayode Adeniyi

    Everyone can tell a restaurant with five reviews from one with five thousand, and then loses that intuition the moment the same structure arrives with a percent sign attached. Two evaluations reporting 94% can differ by fourteen points in what they actually support, and the convention for printing results hides the difference from the people deciding.

All publications

Enquiries

Describe the system, the decision it takes part in, and the access you can give us. We reply within three working days.