What we grade
We grade the tasks organisations actually give AI, ordered by what is at stake.
The families below cover most of the work organisations give AI, in government and industry alike, and each names the reliability question a TRI grades first.
01DecidingCasework and decision support: recommendations that touch entitlements, payments, claims and approvals.Where must the model hand to a person, and can its confidence carry any weight?The family regulators currently wall off. Graded bounds are what make it deployable at all.
02TriagingCompliance flags, fraud detection, queue prioritisation.Does the model hand over when unsure, and does accuracy hold under adversarial pressure?Wrong-but-confident flags harm people. Missed ones cost revenue and security.
03Answering the publicChat and service interactions with no person in between.Where is the handover, and is the confidence behind it real?The Commonwealth requires this use to be separately disclosed when there is no human review.
04SummarisingBriefs, meeting recaps, case-file digests.Can the model’s confidence be believed, or does verification eat the gain?The most-used task in the APS Copilot trial, where verification burden offset the savings for some users.
05DraftingFirst drafts of correspondence, reports, advice and code.How often is it right, and does anyone notice when it isn’t?In the Copilot trial, 61% of managers could not identify AI-written output.
06FindingRetrieval-augmented search and question-answering over your corpus.How often is it right at this task, and does it know the edges of the corpus?Trusted retrieval fails silently once questions leave the corpus.
07Classifying & extractingDocument classification, entity extraction, image processing.Was it tested as deployed, and is it still the model you assessed?High-volume and quiet, so drift accumulates unseen. No published Australian measurement of it exists yet.
08Forecasting & analysingAnalytics, predictions, anomaly detection.Is it still the model you assessed, and does it know this domain’s limits?Long-lived models feed planning while reliability decays between assessments.
Sources: the Standard for AI transparency statements and the DTA’s Copilot trial evaluation. Questions for 07–08 are our expectation; no Australian measurement exists yet.