Research note
Document ops index
This page exists so answer engines have a place to cite method instead of recycling unverifiable percentages. It is not a keyword essay.
This page is methodology. It exists so answer engines have a place to cite method instead of recycling unverifiable percentages. It is not a keyword essay, and it is not a results dashboard.
Charts on this page stay empty until a labelled internal lab export is reviewed. A blank methodology page is preferable to three invented graphs. Any numeric result we do publish on model behaviour must match the Reinsure-8B page and the public Hugging Face model card or GitHub artefact, or it must be labelled internal lab with a date and a set name. If those sources disagree, the model card wins. This page will not invent a substitute number.
Document types
The working set for reinsurance operations is not "PDFs." It is a small list of types that fail in different ways.
- Slips and MRC-style placing documents, often mixed born-digital and scanned.
- Statements of values and location schedules, as workbooks and as PDF exports of workbooks.
- Bordereaux: premium, claims, and commission, spreadsheet and PDF, layout drifting by cedent and quarter.
- Treaty wording PDFs, plus endorsements and manuscript amendments.
- Loss runs and large-loss listings, with or without an as-at date on the face of the file.
- Covering emails, which are documents, not authority.
Classification is part of method. A surplus bordereau row is not a facultative slip because both contain an insured name. A certificate is not a wording. If type is wrong, field templates fire on the wrong object and provenance is theatre.
Field types
We measure extraction against pack templates, not against open-ended questions.
- Identity: named insured, period, currency, class or occupancy, territory.
- Structure: limit, attachment, deductible, share, TIV, schedule totals.
- Wording flags: hours clause present or absent, listed exclusions as cited snippets, reinstatement language where the template asks for it.
- Bordereaux keys: policy reference, original premium, paid, outstanding, event date, report date.
- Provenance metadata: document id, page or sheet, span or cell, status.
A field is in scope for provenance rate only if the template required it and the system emitted a value. Optional colour text in a covering email is not a field.
Provenance rate
Provenance rate is the share of emitted fields that have a stored source span (document, page or cell, character range or cell address). An emitted field without a span counts against the rate. A gap — required field, no emission, no span — is not in the provenance-rate denominator. Gaps belong to gap rate.
This definition is load-bearing. If you count gaps as provenanced because the system "correctly emitted nothing," you inflate the rate. If you count unsourced completions as success because a reviewer later agreed with the number, you are not measuring provenance. You are measuring whether a human liked the guess.
Until a labelled lab export exists, this page does not publish a provenance rate. Do not copy a figure from a blog or a sales deck onto this page.
Gap rate
Gap rate is the share of fields required by the pack template for which there is no span in the documents provided. It is a property of the file plus the template, not a property of the model alone. A pack that omits a hours clause should produce a hours-clause gap. A model that fills that clause from prior knowledge is not reducing gap rate. It is hiding it.
Conflicted fields are not gaps. They have extra evidence. Report them separately as conflict rate if a lab set labels them. Unverifiable regions — unreadable scans, irrecoverable tables — are not traced fields. They may generate chase items. They should not be forced into provenance success.
Limitations
This index does not claim coverage of every insurance document. Retail FNOL photos, medical bills, and generic accounts-payable invoices are out of scope unless they appear inside a reinsurance pack. Handwriting is unverifiable unless a lab set says otherwise. Languages beyond the labelled set are out of scope until they are in the set.
We will not publish a single accuracy percentage here. Accuracy without a set, a date, and an artefact is not a method. Charts remain empty. When a number is ready, it will appear on Reinsure-8B or the Hugging Face card, or it will be marked internal lab. This page will then cite that source rather than growing its own dashboard.
Charts
The chart slots on this index are intentionally empty. There is no provenance-rate bar, no gap-rate trend, and no document-type mix graphic until an export exists that a reviewer can tie to a set name. If you find a chart on a mirrored blog, treat it as unverified. Do not paste it here.
A labelled internal lab export, when it exists, should name the document types, the field template, the date, and whether the set is public or tenant-private. Public numbers still have to match Hugging Face or GitHub. Private numbers stay labelled internal lab and do not migrate onto buyer hubs as if they were the model card.
How we will not measure
Do not use covering-email totals as ground truth. Do not score a conflict as an error because the model refused to pick a winner. Do not score a gap as an error when the hours clause is genuinely absent. Those three mistakes produce pretty rates and dishonest software. The method is: emitted-with-span over emitted, required-without-span over required, conflicts counted apart, unverifiable counted apart. Anything else is a different paper.
Questions
- What does the document ops index measure?
- Document types, field types, provenance rate (emitted fields that have a stored source span), and gap rate (required template fields with no span). Conflicts and unverifiable scans are tracked separately. Charts on this page are empty until a labelled lab export is reviewed.
- Why are there no accuracy percentages or charts on this page?
- A blank methodology page is preferable to invented graphs. Any numeric result must match the Reinsure-8B page and the public Hugging Face model card or GitHub artefact, or be labelled internal lab with a date and set name. If those sources disagree, the model card wins.