Published standard · Version 1.0
How we mark: Reviewing AI Answers
This is exactly how the test is marked. We publish it so you do not have to take our word for what the certificate measures.
In short
- You read answers an AI gave, and score each one out of 4.
- You score them on five things. Getting the facts right counts most.
- You need 80% overall to pass. This does not change based on how others did.
- The most common reason people fail is trusting an answer that sounds good but is wrong.
Pass mark: 80%. Fixed, not curved. Your answers are compared with answers our experts already agreed on.
The five dimensions
Candidates score a set of AI-generated responses across five dimensions. Each is marked 0–4.
| Dimension | What it measures | Weight |
|---|---|---|
| Factual accuracy | Is the claim true? Does the candidate catch confident-sounding fabrication? | 30% |
| Instruction following | Did the response do what was actually asked, including format and constraints? | 20% |
| Usefulness | Would this genuinely help the person who asked, or does it just sound complete? | 20% |
| Safety & harm | Correctly flagging responses that are unsafe, deceptive, or inappropriate for the audience. | 20% |
| Justification | Can the candidate explain the score in a way another reviewer could apply consistently? | 10% |
The scale
| Score | Meaning |
|---|---|
| 4 | Fully correct. No reviewer would reasonably disagree. |
| 3 | Correct with a minor issue that doesn't change the outcome. |
| 2 | Partially correct. A real problem is present but the response has value. |
| 1 | Mostly wrong, misleading, or fails the instruction. |
| 0 | Unsafe, fabricated, or entirely off-task. |
What separates a pass from a fail
Catching confident errors
The most common failure is marking a fluent, well-structured answer as good when it contains a fabricated fact. Candidates who pass check claims rather than judging tone.
Consistency across items
Applying the same standard to item 40 as to item 1. We seed repeated items to measure this directly; drifting standards is a fail even when individual scores look reasonable.
Not marking down for style
Taking marks off a correct, useful answer because you did not like the tone or the layout. We measure substance, not taste.
Explaining the call
A score with no defensible reason behind it is worth little to a downstream team. Justifications are marked, not decorative.
How we check project work
Passing the test is the entry point, not the end. Project work is checked as well:
- Projects include items our experts have already scored, so we can compare directly.
- How closely a person matches those expert scores is shown on their verification page.
- A reviewer reads every project and grades it against this rubric.
- Repeated items check whether someone stays consistent with their own earlier judgements.