What actually happens when software marks a paper.
And, more usefully, the places where it should not be believed. If you are deciding whether to let any of this near your students’ marks, the second half matters more than the first.
The four steps
Every system doing this, ours included, is really doing four separate jobs. They fail in different ways, and lumping them together as “AI checking” is how people end up surprised.
1. Reading the handwriting
A photograph of a page becomes text. This is the part that has improved most, and it is now genuinely good on ordinary handwriting at an angle, in bad light, on ruled paper. It is still the step most likely to go wrong, and it goes wrong quietly — a misread digit is not obviously a misread.
2. Working out what the question wanted
Marks per question, and what a full answer has to contain. Some systems need you to type a marking scheme in. Ours builds one from the question paper itself, which is the difference between fifteen minutes of setup and an afternoon of it — and the reason a centre actually uses the thing on the second test rather than only the first. A whole paper we marked shows the output of all of this on six real pages. What a scheme actually has to contain is narrower than most people picture, which is why this works at all.
3. Comparing the answer to the scheme
Step marks for a derivation. Error carried forward where the method holds and the arithmetic slips. Credit for a labelled diagram. This is the part that most resembles what a teacher does, and the part people most expect to be wrong. On structured subjects it is more consistent than a tired human at 11pm — not because it is cleverer, but because it does not get tired.
4. Deciding what to do when it is unsure
This is the step that separates a useful tool from a dangerous one, and it is the one nobody asks about.
Where it must not be trusted
A system that always produces a mark is telling you something false. Some pages are genuinely unreadable. Some answers are genuinely ambiguous. A tool that returns a confident number for those is not more accurate than one that admits it — it is the same accuracy with the uncertainty hidden from you.
Concretely, these are the cases where no marking software should be believed:
- A page that did not scan properly. Blurred, half-cut, upside down, or simply missing. The right behaviour is to say which page and stop, not to mark what it can see and present a total.
- A free-hand diagram with no labels. There is nothing to read. Judging it is guessing, and we hand it back rather than guess.
- An answer that argues something unusual but correct. A student who takes an unexpected route to the right answer is exactly the student you least want mis-marked.
- Anything where the marks matter formally. A board result, a scholarship cut-off, a promotion decision. Internal tests are the right place for this; a decision a family will act on needs a human who has looked.
Four questions worth asking any vendor
Including us. If a tool cannot answer these plainly, that is the answer.
“Show me what happens when it cannot read a page.”
Ask for a demonstration on a deliberately bad photograph. You are looking for it to name the page and decline, not to produce a mark anyway.
“Can I see why it gave that mark?”
A mark with no reason behind it cannot be checked, defended to a parent, or corrected with any confidence. Every mark should point at the words on the sheet that earned it.
“What does your accuracy figure actually measure?”
You will see numbers like 95% and 98% quoted in this market. Ask what was measured, on whose handwriting, against whose marking. A percentage with no test behind it is a marketing number. We do not publish one, because the only figure that would mean anything to you is the one from your own papers — which is why the first one is free.
“Who releases the mark to the student?”
If the answer is anything other than the teacher, walk away. Software should prepare the marking. A person should decide it.
What we do about all this
- Uncertain answers are flagged and handed back, with the reason said first. A typical paper comes back with a handful of questions to look at rather than all thirty-three.
- Every mark quotes the text that earned it, so checking a paper is reading, not re-marking.
- The teacher approves before anything is released. There is no setting that turns this off.
- Answers are transcribed, so you read typed text at a legible size rather than pinch-zooming a photograph of handwriting on a phone.
None of that makes it infallible. It makes it checkable, which is the property that actually matters when a parent is on the phone asking why their child lost a mark.
Test it on a paper you have already marked.
Your question paper, your students’ handwriting, your marks to compare against. It is the only evaluation that tells you anything, and the first one is free.
Message us and we will set up a time. We reply the same day.