We have scanned and captured around 140,000 pages of land registry data for Valueguard — some 6 million rows of data in total. The project was completed in a month.

The real challenge lay not in the volume but in the UUID codes. A UUID is a 36-character random string, and unlike running text it offers no dictionary, no context and no pattern that could reveal a misread character. A single wrong character makes the whole code useless — and the error is invisible in the data itself, surfacing only when the record no longer matches.

For that reason the entire material was captured twice, with two independent OCR programs, and the results were compared against each other. When both readings agree on a random 36-character string, the chance that they made exactly the same error is very small. Where they differ, you know precisely where the problem lies — and every such difference was reviewed visually against the scanned image.

On top of that, every page was checked structurally: the number of rows and columns per page was compared against the expected value, so that pages with a deviating structure were flagged automatically. In one batch reviewed this way, just over one row in a thousand was singled out for inspection.

Read more about our service for OCR conversion of registers and tables — we capture address registers, property data and tabular material into Excel, a database or searchable PDF.

Staplade och batchmärkta kartonger med lantmäterihandlingar efter avslutad skanning
The material once scanning was complete — ready for destruction