Back to AI Research

AI Research

HANS tests handwriting recognition on corrected, mixed-format mathematics answers

Key Takeaways

  • A new answer-sheet dataset and noise-aware model distinguish surviving work from strikethroughs and erasures.
  • Xiazhen Wu and colleagues introduce HANS, a dataset for reading handwritten mathematics answer sheets containing corrections, formulas and prose.
  • Their [paper](https://arxiv.org/abs/2610.12363) also presents NA-GOT, a recognition model that tries to suppress correction noise before and during text generation.
  • The target is the student's written problem-solving process.
  • Recognizing that process is an input to possible grading systems, but this study evaluates transcription rather than whether an automated grader awards correct marks.

Xiazhen Wu and colleagues introduce HANS, a dataset for reading handwritten mathematics answer sheets containing corrections, formulas and prose. Their paper also presents NA-GOT, a recognition model that tries to suppress correction noise before and during text generation.
The target is the student's written problem-solving process. Recognizing that process is an input to possible grading systems, but this study evaluates transcription rather than whether an automated grader awards correct marks.

Preserve the work that remains after corrections

HANS contains 5,213 answer-region images from 745 students across eighteen problem types. The sources are eighth-grade mathematics assessments conducted over three semesters in a municipal school district in China.
The authors say they cropped personally identifiable information before release and use. The captured paper describes public dataset availability as planned upon publication, so it does not establish that readers can already download the complete release.
Annotations cover mathematical expressions, natural-language explanations and hand-drawn tables. Noise annotations identify erasures, strikethroughs and overwriting. Deleted content is excluded from the target transcription, while annotators aim to preserve the final legible writing.
The annotation process combines model-generated drafts with character-level human correction and expert inspection. That matters because the intended transcription depends on distinguishing discarded work from a student's surviving answer, not merely detecting ink.

Suppress noise at two stages

NA-GOT extends the GOT recognition architecture. A lightweight module predicts the probability that each visual patch contains noise and reduces the contribution of likely noise patches.
A second mechanism uses those probabilities to bias decoder attention away from noisy regions. The model then generates Markdown containing prose, formulas and derivation steps.
Noise-region annotations supply supervision during training. At inference, the model predicts noise distributions without requiring a person to mark corrections first. Its predictions can still be wrong, especially where a deletion stroke overlaps valid content.
The authors test different suppression strengths. Their discussion reports that excessive suppression can also impair valid features, making the setting a tradeoff between reducing interference and preserving useful handwriting.

Domain adaptation accounts for much of the gain

The evaluation uses an 80/20 split of HANS and transcription metrics including BLEU, edit distance and F1. These scores describe agreement with reference text; they do not measure educational judgement.
The paper reports BLEU of 0.594 and F1 of 0.810 for NA-GOT, compared with 0.577 and 0.801 for GOT fine-tuned on the same dataset. Untuned GOT performs much worse, with BLEU of 0.098.
That comparison separates the benefit of adapting to this domain from the additional benefit of the proposed noise mechanisms. Both matter, but attributing the entire difference from untuned GOT to noise-aware attention would overstate the experiment.
HANS tests a defined educational setting with particular writing and correction patterns. The results support further work on noisy answer-sheet transcription, while leaving performance on other grades, languages and document distributions to separate evaluation.

Comments