Use this pack to review whether marked work has been judged consistently against the same rubric and the same annotated standards. It is for teachers, heads of department and assessment leads who need an auditable moderation record, rather than a general summary of marking.
Remove names, email addresses and any other unnecessary personal information before pasting student work. Keep the anonymous student ID consistent across every prompt so that a proposed change can be traced back to the correct record.
Stop
Protect student information
Use anonymous IDs and paste only the work, marks and annotations needed for the moderation decision.
Run the prompts in order
- Start with Prepare the moderation evidence register. It exposes missing annotations, unclear marks and rubric labels that do not match. Do this before asking for a judgement. A neat-looking moderation record built on incomplete source material is still incomplete.
- Use Compare one marked response with benchmarks for cases flagged High priority, borderline responses, and a small sample from each marker. Keep the original annotations in the input. They show whether the written justification supports the awarded mark.
- Use Find inconsistent judgement across markers when you have a set of comparable responses. This separates a genuine difference in student performance from a different interpretation of the criterion.
- Finish with Produce the subject lead moderation record. Paste only findings that you have checked. The record distinguishes a proposed amendment from a decision that has actually been agreed.
Key point
Moderate the evidence, not the marker
Ask whether the response meets the descriptor and matches the standard, rather than whether one teacher is generally strict or lenient.
Set up the source material
Put the task brief, rubric and benchmarks beside the marked responses before starting. The task brief matters because a strong response to a different question is not evidence for this assessment. Benchmark annotations matter because the mark alone does not explain why the sample represents a particular standard.
Use the rubric exactly as issued. Do not simplify criterion names halfway through the review. If the rubric says Analysis and the teacher annotation says Evaluation, record that mismatch rather than silently treating them as the same thing.
For a large set, moderate a purposeful sample first:
- Include work near each grade or mark boundary.
- Include work from every marker.
- Include responses with sparse annotations.
- Include work already queried by a teacher or student.
- Include a small number of apparently straightforward marks.
The model you are using may handle document formats and input sizes differently. Check the relevant handling guidance in the xAI documentation overview before using a large evidence set. Split the work by assessment and criterion if needed. Do not combine different tasks merely because they share a subject.
Check the output before changing any mark
The useful test is traceability. For every proposed change, you should be able to follow a short chain: student evidence, rubric descriptor, benchmark comparison, then proposed mark. If one link is absent, the finding should be Unresolved, not a confident recommendation.
Check
Test one proposed change
Read the quoted student evidence yourself, then locate the descriptor and benchmark reference. If either does not support the proposed mark, refer it for second moderation.
Use this table to deal with common results.
| If you see | Treat it as | What to do |
|---|---|---|
| A changed mark with no quoted evidence | Unsupported judgement | Return to the individual comparison prompt with the full response. |
| Similar responses receiving different marks | Possible inconsistency | Check whether the same criterion and descriptor were applied. |
| A mark that fits the rubric but not the benchmark | Standard-setting question | Put it in Decisions for subject lead. |
| Different annotations and marks | Audit weakness | Keep the mark under review and request a clearer marker rationale. |
Watch for false precision. A response may contain evidence for part of a descriptor but not all of it. The prompts should identify that gap. They should not fill it using assumptions about what the student meant, their prior attainment, or their effort.
Watch out
Benchmarks are evidence, not replacement rubrics
Do not award a mark because a response resembles a sample in one feature while missing a required criterion.
When the result does not work
Do not ask for a cleaner answer when the source evidence is unclear. Fix the input. Add the missing page, identify the mark scale, paste the relevant benchmark annotation, or ask the original marker to explain the criterion-level judgement.
If two reasonable readings of the rubric remain, record both in the subject lead decisions section. The subject lead should decide the interpretation, document it, and apply it consistently to affected work. Then rerun the final record prompt with that decision in Decisions already agreed.