Print / save PDF

Published standard · Version 1.0

How we mark: Reviewing AI Answers

This is exactly how the test is marked. We publish it so you do not have to take our word for what the certificate measures.

In short

  • You read answers an AI gave, and score each one out of 4.
  • You score them on five things. Getting the facts right counts most.
  • You need 80% overall to pass. This does not change based on how others did.
  • The most common reason people fail is trusting an answer that sounds good but is wrong.

Pass mark: 80%. Fixed, not curved. Your answers are compared with answers our experts already agreed on.

The five dimensions

Candidates score a set of AI-generated responses across five dimensions. Each is marked 0–4.

DimensionWhat it measuresWeight
Factual accuracyIs the claim true? Does the candidate catch confident-sounding fabrication?30%
Instruction followingDid the response do what was actually asked, including format and constraints?20%
UsefulnessWould this genuinely help the person who asked, or does it just sound complete?20%
Safety & harmCorrectly flagging responses that are unsafe, deceptive, or inappropriate for the audience.20%
JustificationCan the candidate explain the score in a way another reviewer could apply consistently?10%

The scale

ScoreMeaning
4Fully correct. No reviewer would reasonably disagree.
3Correct with a minor issue that doesn't change the outcome.
2Partially correct. A real problem is present but the response has value.
1Mostly wrong, misleading, or fails the instruction.
0Unsafe, fabricated, or entirely off-task.

What separates a pass from a fail

Catching confident errors

The most common failure is marking a fluent, well-structured answer as good when it contains a fabricated fact. Candidates who pass check claims rather than judging tone.

Consistency across items

Applying the same standard to item 40 as to item 1. We seed repeated items to measure this directly; drifting standards is a fail even when individual scores look reasonable.

Not marking down for style

Taking marks off a correct, useful answer because you did not like the tone or the layout. We measure substance, not taste.

Explaining the call

A score with no defensible reason behind it is worth little to a downstream team. Justifications are marked, not decorative.

How we check project work

Passing the test is the entry point, not the end. Project work is checked as well:

  • Projects include items our experts have already scored, so we can compare directly.
  • How closely a person matches those expert scores is shown on their verification page.
  • A reviewer reads every project and grades it against this rubric.
  • Repeated items check whether someone stays consistent with their own earlier judgements.
Placeholder: version this document and date it. Employers and state credential reviewers both look for a change history — it signals the standard is maintained rather than written once and abandoned.

← Back to employers