Answers / Provenance

Does AI hallucinate treaty terms?

Ungrounded models will invent clauses, limits, and named insureds when the PDF is unclear. A production system should refuse to fill those fields and list them as gaps instead of sounding sure.

Ungrounded models will invent clauses, limits, and named insureds when the PDF is unclear. A production system should refuse to fill those fields and list them as gaps instead of sounding sure.

Hallucination in this domain is not a fun fact. It is a wrong attachment point, a missing hours clause written in as 72 hours, a reinstatement formula collapsed into the word reinstatement, an exclusion dropped from a summary. Binders then argue about what was said, not about what was written.

Why fluency is the hazard

A treaty wording is definitions, attachments, exclusions, hours, reinstatements, and claims procedures. A general model asked to summarise it will produce a fluent page. Summaries drop the clause that bites at the loss. They restate an attachment from a heading and miss the manuscript carve-out on a later page.

The same failure shows up on facultative slips. Ask a public model for the limit on a pack that contains two figures and it will pick one, or blend them, and speak in a complete sentence. Ask it for a hours clause that is not in the file and it will often supply a market-standard figure. That is completion. It is not administration.

Source-grounded extraction is the architectural mitigation: extract only with spans, leave gaps empty, keep conflicts as two spans, keep files in the tenant. Dual-control on bind stays with people. Mystery-shop a vendor by asking it to quote a limit from a redacted slip with a torn schedule. If it answers with a number and no span, it has hallucinated a term.

Treaty admin is where invented clauses land

Treaty reinsurance is standing capacity administered against a wording, not against a chat transcript. Attachment, limit, reinstatement, hours, event definition, class, territory: each term in the administration system needs a page. "Two reinstatements" in a summary is not a formula. The premium for a paid reinstatement is in the wording or it is a gap.

A domain model does not get a pass because the jargon is better. If it emits an unsourced attachment, the attachment is still unsourced. Pair the model with stored spans or you are back to summaries.

How this site talks about measuring extraction — provenance rate on emitted fields that have a span, gap rate on required fields with no span, conflicts counted apart, no single accuracy percentage — is the document ops index. That page is method. It is not a results dashboard, and it does not substitute a marketing figure for a labelled lab export.

Worked example: ACME Construction Ltd

Hand the product the fictional ACME Construction Ltd pack, acme.example. Slip.pdf page 2 states USD 10,000,000 any one occurrence and TIV USD 42,000,000. SOV.xlsx totals USD 47,100,000. Hours clause: not found.

If the product writes 72 hours because that is common on construction, it has hallucinated a treaty term. If it emits one TIV, it has hallucinated a schedule. If it leaves hours clause empty, stores both TIV spans, and traces the occurrence limit to page 2, it has refused to hallucinate. That refusal is the feature.

Do not paste a live wording into a public chatbot to see whether the model "knows" your treaty. You will get fluency, and you may have exported the file. Production ingest is tenant-scoped. The marketing site uses a sample pack only. Humans still bind the layer the wording actually contains.

Written by Shen Pandi · Updated 2026-08-25 · Definitional page, not a product claim sheet