Vai al contenuto principale

Questa pagina è disponibile solo in inglese

All Articles
Technical Architecture

Where the Language Model Does Not Belong

Designing Defensible AI in Claims Operations

Reuben John, CEO and Co-Founder of True Aim
7 min read

Executive Summary

A claim decision is not a single step. It is an end-to-end chain of smaller actions categorised into three distinct types: Reading, Computing and Judging.

  • Reading belongs to AI models: Turning unstructured PDFs, photographs, emails and handwritten notes into structured facts.
  • Computing belongs strictly to rules: Calculating limits, excesses, deductibles and co-insurance splits has exactly one right answer. A model can only offer a likely one.
  • Judging can use a model, provided it is marked as such: Where wording leaves an item open, the record must show that the line was resolved by a model and not by a rule.

The design choice that matters is not "rules or models". It is keeping the boundary between them visible on every step of the record.

The Three Core Operations Inside a Claim Decision

AI in claims is usually framed as a choice between rules and models. But a claim decision is not one single act. It is a chain of steps and each requires the right tool:

  1. Reading (Evidence to Facts)

    Turning raw evidence into structured facts. Claims arrive as photographs, PDFs, handwriting, emails and free text. Reading pulls out what they say (the policy number, date of loss, repaired part or invoiced amount) and links each value back to its source document.

  2. Computing (Applying Rules & Numbers)

    Applying written rules and calculations to those facts. Is the part covered under the wording? What excess applies? Does the amount breach a limit? How is liability split? Once the facts are fixed, each of these has one correct answer set by the policy, catalogue or rulebook.

  3. Judging (Resolving Open Items)

    Resolving what written wording leaves open. An invoice line describes a part in words no catalogue entry uses. Or a clause could reasonably be read two ways. No rule settles it, so someone (or something) has to make a call.

Most claims contain all three. What matters is whether anyone can see, afterwards, which kind each step was.

Why Computing Belongs to Rules, Without Exception

Arithmetic on a written limit has one right answer. A co-insurance split of a known amount at a known percentage does not have a likely outcome. It has an outcome.

  • A deterministic step is one where the same inputs always produce the same output. A rule written as a decision tree behaves this way: given the facts, the outcome is computed, not predicted.
  • A probabilistic step produces the most likely answer based on what it has seen. Because a language model is probabilistic by construction, the same inputs can produce a different answer on another run or model version.

That is a strength when reading messy evidence. On a calculation, it is simply a new way to be wrong, with no upside. A model that applies an excess correctly on every test claim has still not become a rule. The next run is not bound by the previous ones.

Excess, deductibles, thresholds, caps, limits, liability splits, co-insurance and betterment are all computing. None of them needs a model.

Where a Model Earns Its Place

A model earns its place by solving jobs where fixed rules cannot work:

  1. Document & Evidence Extraction (Reading)

    Rules cannot read a photograph or a handwritten note. Agents built on language models can. Turning unstructured evidence into structured facts is the work that suits them best. The discipline is to keep each agent narrow. An agent that only extracts is far easier to check than one that reads, extracts, sorts and explains in one pass.

  2. Marked Ambiguity Resolution (Judging)

    When an item is genuinely ambiguous and no rule resolves it, a model can do the matching. What makes that acceptable is the label. The record marks that line as probabilistic rather than deterministic, so a reviewer can see exactly which parts of the decision involved judgement and where to look first.

What Happens When You Cannot Separate the Two

If extraction, judgement and arithmetic happen inside one single model call, the decision can only be described as a whole. The output is a single answer with nothing inside to inspect.

Any explanation is written afterwards, as a narration of what probably happened. It may sound plausible, but it is not a record.

What makes a decision hard to defend is rarely that a model was involved. It is that no one can separate the steps where it was involved from the steps where it was not. A decision that is not a black box is one where every step can be opened on its own, with its kind written on it.

That property must be designed into the system from day one. A record that never stored which steps were computed cannot be split into steps later.

What the Boundary Looks Like on One Claim

Take a glass line on a motor claim:

  1. Reading

    An agent extracts the facts the rules need from repair documents: what was repaired and what caused the damage. Each value points back to its source document.

  2. Computing

    The engine applies policy wording and business rules to those facts. The line comes back covered under clause 2.2A. The explanation is rooted in the wording, not composed after the fact. Any excess, limit or cap is then computed, with the calculation shown line by line.

  3. Judging

    Where an invoice line matches no rule cleanly, a model does the matching and that line alone carries a probabilistic label. The lines around it stay deterministic.

  4. Human Review

    The whole audit trail can be exported at any point. A handler reviews the claim. Any change they make requires a reason logged directly in the record.

Where Rules Stop: The Three Handover Points

A decision tree is exactly as complete as the facts it is given and the wording it was built from. A well-designed engine treats rule boundaries as defined handovers, not gaps:

  • Handover 1: Missing Evidence. If a document is missing or a value cannot be extracted, the correct response is to ask for it. The claim waits for the fact rather than proceeding on an approximation.
  • Handover 2: Genuine Ambiguity. Where an invoice line matches no catalogue entry cleanly, a model matches the item and the record marks it as probabilistic. The judgement sits beside the arithmetic, clearly labelled.
  • Handover 3: Insurer-Defined Thresholds. Thresholds on claim value, deviation from expected outcomes or confidence levels are set by the insurer. A claim that crosses a threshold is routed to a handler, who sees the facts, the applied clause and every assumption made. If an assumption does not hold, the handler overrides it and the outcome is recomputed.

How True Aim Draws the Line

True Aim's decision engine is a hybrid of two technologies. The boundary between them is the core design choice:

  • Deterministic Decision Trees: Policy wording, catalogues and rulebooks are converted into decision trees. Given the facts, the outcome is computed rather than predicted. Same inputs, same decision, every time.
  • Task-Specific Agents: Agents read photographs, PDFs, handwriting, emails and free text to extract structured facts. Each agent does one narrow job.
  • Labelled Model Matching: Where an item is genuinely ambiguous, a model matches it and marks that line as probabilistic on the record.
  • Human-in-the-Loop Control: Coverage outcomes trace back to the clause behind them. Human review is triggered by thresholds the insurer sets. Every change a person makes carries a reason.

What a Claim Record Should Be Able to Show

Every closed claim is read again: by the handler before sign-off, by the customer when an outcome is questioned or by an auditor a year later. Each asks the same question: how was this reached?

A record that keeps the three kinds of steps apart answers directly:

  • Each fact carries the document page it was taken from.
  • Each figure carries the clause or rule that produced it.
  • Each judgement is marked as judgement, so the reviewer knows where to test first.

A record that holds only the final outcome cannot answer at all. It can only offer a description written after the event. A description is not evidence.

Key Takeaways

  • Tool Separation: Reading, computing and judging are different kinds of steps and need different tools.
  • Rules for Computation: Computing belongs to rules. A likely answer to arithmetic is a new failure mode with no upside.
  • Labelled Model Usage: A model earns its place in reading and, marked as probabilistic, in judging.
  • Defined Handover Boundaries: Rules stop at missing evidence, genuine ambiguity and insurer thresholds. Each is a handover the record shows, not a gap it hides.

Frequently Asked Questions

Should a language model calculate a claim payout?

No. Excess, limits and liability splits have one correct answer once facts are fixed. They belong to deterministic rules.

What does "probabilistic" mean on a claim record?

That the step was resolved by a model's most likely answer rather than by a fixed rule, so it may not repeat exactly. Marking it shows a reviewer where judgement was involved.

Can rules handle every claim?

Not alone. They are not designed to. Where evidence is missing, the system asks for it. Where wording is ambiguous, a model matches the item under a probabilistic label. Where a claim crosses a threshold, a handler reviews it with all facts, clauses and assumptions in view.

How do you check whether a system keeps the boundary?

Take one closed claim and ask the system to show, per step, whether it was read, computed or judged. If it cannot, the boundary is not in the record.

Zero Risk Does Not Exist. Explainable Decisions Do.

Discuss how this boundary applies to your claims operation.