Scanned real-world documents and structured forms in authentic layouts, suited to document AI, OCR-adjacent NLP and information-extraction training.

Serbian Document Corpus (Validation Samples)

Volume
7,200 documents
Scale
Validation sample set
Language
Serbian
Region
Serbia
Format
PDF
Licensing
Non-exclusive

Seventy-two hundred real-world Serbian documents delivered as PDFs, assembled as a validation corpus. Suited to document AI, OCR-adjacent natural language processing and information extraction, particularly for teams needing coverage of Cyrillic and Latin Serbian script in authentic document layouts rather than synthetic renders.

US Scanned Tax Filing Forms

Volume
5,000 form packages
Scale
Structured forms
Language
English
Region
United States
Format
PDF
Licensing
Non-exclusive

Five thousand scanned United States tax filing packages, delivered as PDFs. Structured forms with consistent field layouts and real-world scan quality, which makes them well suited to training extraction models on tabular and key-value document understanding where the layout carries as much meaning as the text.