A pottery drawing is only part of an archaeological record. Someone still has to isolate it from a scanned sheet, transcribe its inventory information, ink the lines, produce editable vectors and arrange a publication page. Errors linking the drawing to its metadata can matter as much as an imperfect outline.
PyPottery brings those post-production tasks into a modular, human-in-the-loop suite. It is designed to process existing archaeological documentation, not replace the interpretive work of drawing an artifact or infer its history automatically.
Four modules follow the publication workflow
PyPotteryScan pairs images with text from scanned sheets. The user first marks drawing and text boxes in a graphical interface. This remains assisted annotation because sheet layouts vary. The software then extracts images, recognizes handwriting and parses the resulting text into structured fields.
Recognized text is not assumed correct. A manual correction interface lets users inspect and fix transcription errors. Parsed records are exported in XLSX format, linked to the drawing by an identifier, so they can be checked in spreadsheet or database workflows.
PyPotteryInk produces inked raster drawings. PyPotteryTrace converts those outputs into component-aware vectors, while PyPotteryLayout handles publication arrangements. Layout options include a user-defined grid and a packing algorithm intended to use page space efficiently, with sorting, scale calibration and metadata-based captions.
Vectorization preserves editable components
PyPotteryTrace is not simply a generic outline tracer. The user selects archaeological components such as a profile, handle or decoration, assisted by SAM 2 segmentation. Each component is processed separately so it can be edited or exported independently.
Its pipeline reduces lines to skeletons, traces paths, simplifies them and constructs smooth Bézier curves. A specialized step extracts ceramic profiles. The module also provides manual vectorization when automatic processing fails, curve editing and standard SVG or raster exports.
These intervention points are part of the method rather than exceptions to an otherwise autonomous pipeline. The software organizes repeatable processing around expert selection, correction and validation.
Separate measured time from perceived speedup
The case study uses 50 hand-drawn sheets containing 240 pottery drawings from Terramara di Montale in Italy. The drawings follow consistent conventions, while cursive Italian metadata comes from multiple people and field seasons. The paper says the full archaeological dataset cannot yet be released because the material is unpublished.
The author reports approximately 103 minutes for one complete run over that dataset. That timing is a single measured run, not an average across independent laboratories. Its breakdown includes manual filing, cleaning, OCR correction, parsing correction and component segmentation.
The abstract also reports a median perceived speedup of 40 times from usability participants. That is a perception measure, separate from the timed pipeline run. The usability study involves five domain experts on a standardized subset of tasks, so the two results should not be treated as the same experiment.
Recognition and parsing still need checking
The reported handwriting results average 6.2% character error and 18.6% word error. Those measurements support the need for correction rather than a claim of flawless transcription.
Parsing uses examples supplied by the user. With ten examples, the table reports 93.13% overall accuracy but 67.60% exact match across all fields in a record. Overall accuracy includes empty-field matches; non-empty-field F1 is 59.07%. A high aggregate score therefore should not be read as proof that nearly every completed record is entirely correct.
The paper describes local execution without external APIs or proprietary models and reports tests on several consumer-hardware configurations. It also keeps human oversight throughout. For an archaeological team, PyPottery offers a concrete processing workflow to evaluate against its own handwriting, drawing conventions and metadata requirements—not a demonstrated universal speedup or an unattended replacement for scholarly judgment.
Comments (0)
to join the discussion
No comments yet
Be the first to share your thoughts!