Methods / Evaluate
Question it answers
Does this AI product or system meet the core HCAI principles and where does it fall short?
When to use it
Once a working prototype or live product exists. Run at every major milestone and before every launch.
How to run it
- Assemble a small evaluation panel. Include at least one person who was not involved in building the product.
- Each evaluator works through the product independently and assesses it against each HCAI principle using a severity scale: 0 means no problem, 1 is cosmetic, 2 is minor, 3 is major and 4 is a critical failure.
- Evaluators document specific instances of principle violations with screenshots or interaction descriptions.
- Bring evaluators together to consolidate findings and prioritise by severity.
- Produce a remediation plan for every issue rated 3 or above before launch.
Output
An HCAI Heuristic Evaluation Report: a structured assessment against all eight principles with severity ratings, specific instances and a prioritised remediation plan.
Question it answers
How do we know if this AI system is causing harm after launch and what do we do about it?
When to use it
At launch and maintained indefinitely thereafter.
How to run it
- Using the Human Impact and Failure Map define the specific harms you are monitoring for.
- For each harm define a measurable proxy signal.
- Set alert thresholds for each signal.
- Define the review process: who reviews, how quickly, and what decisions they can make such as pausing a feature, escalating or notifying people.
- Build a direct feedback channel for people to report suspected harms and commit to a response time.
- Schedule regular harm reviews even in the absence of alerts, at least quarterly, to catch slow developing patterns the signals miss.
Output
A Continuous Harm Monitoring Plan: a living operational document defining monitored harms, proxy signals, alert thresholds, review processes and escalation paths.
Question it answers
What does everyone in the organisation and potentially the public need to know about how this AI system works, what it can do and where it fails?
When to use it
Published alongside any AI product or system at launch. Updated with every significant system change.
How to run it
- Write a plain language description of what the system does. If a non-technical colleague cannot understand it, rewrite it.
- Document intended use cases and clearly state uses the system is not designed for.
- Document training data: what it represents, what populations are included and where it falls short.
- Document known failure modes and limitations honestly.
- Document performance metrics broken down by relevant subgroups. Overall accuracy figures hide disparate performance across populations.
- Document who is responsible for the system ongoing performance and how to contact them.
- Make the system card accessible to everyone in the organisation, not just the technical team.
Output
A published System Card: a structured plain language document covering intended use, training data, known failures, performance by subgroup and responsible contacts.