
This deposit holds the free 40-document sample of the full 1000-document Belege dataset. Full set (1000 documents): EUR 14 (launch price) at j4zz.eu/belege. Licence for this sample specifically (redistribution of the files themselves not permitted, everything else — including commercial model training — is) is included as LICENCE.txt inside the archive. Belege — 1000 deutsche Rechnungen und Gutschriften mit Feld-Labels Synthetic dataset for German invoice/credit-note field extraction. Each document is an image plus a JSON label with text and a pixel-accurate bounding box for every field. Three structurally distinct page layouts (not just colour/font variation), three capture conditions (clean/scan/photo), and explicit coverage of Germany's two zero-VAT cases (§19 UStG Kleinunternehmer, §13b UStG Reverse Charge) — 269 of 1000 documents — which a model trained on English invoices or SROIE will not expect and will hallucinate a tax line for. 1000 documents total (895 invoices, 105 credit notes, 3777 line items), 900 unique supplier names, 191 unique line-item descriptions (widened from 61 in the previous revision; no phrase repeats more than 40 times). Every label was self-checked at generation time (line items sum to net, gross = net + tax, printed total matches computed total, every box lies inside the image); 0 of 1000 generated documents were rejected. All entities (companies, people, addresses, bank details, tax numbers) are invented; IBANs use a checksum-valid but never-issued bank code (99999999), so no generated account can collide with a real one. No personal data is included. Full technical writeup (schema, field coverage table, known limitations) is included as DATASET.md inside the archive. Same content and licence as the free samples already published on Hugging Face and Kaggle.
