Benchmarks test general capability.
The index
The Task Reliability Index is a score card that grades, in plain language, how reliably an AI model performs a specific task.
We run a suite of automated tests behind every grade, and we write the results so you know exactly what the model can be trusted to do. The methods are public, and every score card comes with a signed attestation that the tests were run as published.
Why it’s built this way
Seven problems the index is built to answer.
Pass/fail can’t tell you how to deploy.
We grade from “may act alone” to “a human checks every case”.
Test reports are written for engineers.
Every grade is a plain sentence a non-technical person can understand.
New AI rules demand evidence organisations can’t produce.
Every indicator maps to the evidence your assurance framework asks for.
Most assurance asks you to trust the assessor.
Every grade links to automated tests anyone can re-run.
The model you tested isn’t the model you’re running.
Grades expire when the model changes. We re-test and re-grade.
Vendors mark their own homework.
We don’t sell AI, and anyone can re-run our tests.
The seven indicators
Seven questions, each answered by automated tests.
What we grade
The index grades the eight task families organisations most often give AI.
Each family names the indicators that matter most for it, so the score card is weighted to your use case rather than a generic checklist.
01 DECIDING02 TRIAGING03 ANSWERING THE PUBLIC04 SUMMARISING05 DRAFTING06 FINDING07 CLASSIFYING08 FORECASTING
Grades & expiry
Every grade states what the system may be allowed to do, and expires when the model changes.
Grade thresholds are set per use case in the analysis plan, agreed in writing before any data is collected.
Reading a score card
Built to be read twice: once by the person accountable, once by an engineer.
The top layer of every card is plain language: the question, the grade, the sentence. The bottom layer is the instrument: the tests behind the grade, the statistics, the uncertainty, and links to re-run everything. Anyone can re-run the tests. An example card is on the home page.