Text Data Collection
Books. Journals. Archives.
Digitized in partnership with the institutions that hold them.
High-fidelity scanning and OCR of print collections. Books, journals, archives — converted into clean, structured text.
Deduplication, normalization, metadata. Delivered in the format your pipeline expects, at whatever scale you need.
We negotiate and manage partnerships with universities and libraries. Clear provenance, documented rights, one point of contact.
Our text comes from agreements with the institutions that hold it. Documented origin, on every page.
Native-language teams across Europe. Collections in dozens of languages, captured correctly — from blackletter type to the footnotes.
We don't publish client names or project details. Engagements stay confidential, on both sides.
If your work depends on text at scale — text that isn't already everywhere — we should talk.