The short answer
- Auditability is decided by architecture. A decision record written without provenance cannot be given it afterwards.
- The EU AI Act's high-risk deadlines moved to December 2027. Claims handling was never on that list, and GDPR has governed automated decisions since 2018.
- Twelve questions in four groups: provenance, reconstruction, human judgement and continuity.
- The decisive test is reconstruction. Take one claim closed last year and rebuild the decision from the record alone.
- Only 24% of insurance leaders are very confident they could pass an independent review of AI governance and controls, with centralised evidence, in the next 90 days. Confidence is not evidence.
A claim is a series of decisions. Which documents to trust. Which clause applies. What the policy owes. Whether a line item is fair. Automating those decisions is now within the reach of any claims operation. Automating them so that each one can be explained and defended, to a customer, an auditor or a supervisor, years after the file was closed, is a different discipline. That discipline is the subject of this article.
The regulatory calendar has recently offered an excuse to postpone it. In July 2026 the European Union deferred most of the AI Act's high-risk obligations, to December 2027 for standalone systems. Claims handling was never among the uses the Act lists as high-risk, and the rights of a person subject to a solely automated decision, including the right to meaningful information about the logic involved, have applied under GDPR since 2018. The deferral changes the calendar for other systems. It does not change what a claims operation owes the people whose claims it decides. We have set out the detail separately in our article on the AI Act and claims handling.
The stronger reason not to wait is practical. Evidence is decided by architecture. A system that records where every value came from, which rule or model produced it and who confirmed it can answer any question put to it later. A system that does not cannot be repaired after the fact, because whatever the record failed to capture on the day is gone. Compliance is either built in from the first claim or added afterwards as a layer of explanation over decisions nobody can reconstruct.
Grant Thornton's 2026 AI Impact Survey, fielded across 950 executives including 100 in insurance, found that only 24% are very confident they could pass an independent review of AI governance and controls, with centralised evidence, in the next 90 days. At that sample size the figure is indicative, but the direction is clear, and the gap is a records problem before it is a modelling problem.
The twelve questions below are the ones we believe every claims leader should be able to answer about any AI operating in their name, whether it was built in house or brought in. Under each we set out what good looks like and how True Aim approaches it. In most cases those are the same answer, because True Aim was designed around these questions.
Can you show where each answer came from?
Provenance is the foundation of everything that follows, and the easiest property to test. Ask to see one real claim.
For any closed claim, can you trace every value the decision used back to the document and the page it came from?
The standard is a record in which every value the decision relied on points to its source: the claimed amount to the invoice line, the date of loss to the report, the cover position to the clause. A statement that the system had access to all the documents describes an input. It does not describe a record.
True Aim draws a visual line from each data point to the exact place in the source document it came from, so an assessor validates instead of re-investigating. Extracted values carry that provenance down to the page and the text region, and we measure how much of each claim carries it instead of assuming the coverage is complete.
Which parts of the decision were computed by rules and which were produced by a model, and does the record show the difference?
The larger the share of a decision that runs on deterministic logic, the less there is to explain afterwards, because the rule is the explanation. The standard is a record that shows, line by line, which parts were computed from the wording and which were produced by a model, so that judgement is never mistaken for arithmetic.
True Aim converts a policy, catalogue or rulebook into a decision tree. Where the rules cover the facts, the outcome is computed, not predicted, and the same inputs always produce the same decision. Agents read the unstructured evidence, the photographs, PDFs, handwriting, emails and free text, and turn it into the structured facts the engine needs. Each agent performs one narrow task, so when something is wrong it is immediately clear where, and it is corrected at that point. Where an item is genuinely ambiguous and no rule resolves it, a model performs the matching and the record marks that line as probabilistic.
What happens when the evidence is incomplete or ambiguous, and who decided that behaviour?
Uncertainty is not a failure state. It is a routing decision, and it belongs to the claims operation, not to the software. The standard is a defined behaviour the operation controls: route to a handler, hold the claim, request the missing document. A confidence score with no consequence attached to it is a number, not a control.
At True Aim, thresholds, routing rules, straight-through limits and tolerances are configured by the client. A claim is decided automatically only where value, deviation and confidence all pass the tests the client has set. What the evidence confirms is separated from what only a person can confirm, so assessor time is spent where it changes the answer.
Can you rebuild a decision you closed last year?
Reconstruction is what separates being ready for an audit from preparing for one. Preparation is an exercise that begins when the examination is announced. Readiness is a property the system has while it runs, or does not have at all. Only the second can answer a question about a claim closed eighteen months ago.
Given the same claim and the same documents a year later, would the decision be the same, and could you prove it?
The standard is yes for every deterministic part of the decision, with the re-run to prove it, and an exact account of which parts were probabilistic and are therefore reproducible only under the model version and settings that produced them. A system that cannot reproduce its own output cannot defend it.
True Aim can re-validate or re-execute a decision chain on demand. Rule-based outcomes reproduce by construction. For every model call the record holds the model, the prompt template, the engine, the temperature, the seed, the determinism setting and the confidence, so a probabilistic step can be examined under the conditions that produced it.
Are the policy wording, the rules and the model that produced a decision recorded on the decision itself, at the version in force?
On the record itself, not in a release note. Version history kept elsewhere is a different document about a different day. The standard is a decision that names the clause, catalogue entry or business rule applied, at the version in force at that moment, together with the model and prompt that took part.
True Aim writes every action, by a person or a machine, to the record with the model, prompt, rule version and confidence attached, and records which version of the wording, catalogue or business rule was in force when the decision was made. Decisions are judged against the rules that actually applied.
Who can alter a decision record, how would you know, and how long is the record kept?
A record that can be altered without trace is a draft. The standard is a tamper-evident record and a retention period set deliberately to match the operation's own obligations, which usually outlast the claim by years. The claim you are asked about is often older than the log you kept.
True Aim hashes every recorded item, so any alteration is detectable and the record is built to withstand challenge in dispute, audit or court. Retention is configured per client to match its own record-keeping obligations, and the retention status of every claim sits in the same record as its documents, events and decisions.
Is the human in the process genuinely deciding?
Human judgement is the control most operations rely on and the one least often tested. A review step is a genuine control only when the reviewer can see the evidence and has the authority to change the outcome. Under GDPR, a decision that is solely automated gives the individual specific rights, and a sign-off that nobody could realistically refuse does not alter that. Designed well, the review step is also where a claims professional's expertise is used best: on the judgement, not the arithmetic.
At the moment of sign-off, what does the reviewer actually see, and can they change the outcome?
A review step protects the customer only when the reviewer sees the recommendation, the evidence behind it and the assumptions beneath it, and holds the authority to change any of them. A reviewer who sees only a verdict is approving a conclusion, not reviewing a claim.
True Aim produces a recommendation with its full reasoning for a person to confirm. Every assumption behind the determination is stated on a coverage assumptions panel and can be overridden, every line carries the clause or rule it rests on, and every total recalculates live as cover is toggled, so the handler sees the financial effect of a change before committing to it.
When a handler overrides the system, is the reason captured, and does anyone learn from it?
Overrides are the most valuable quality signal a claims operation owns. Each one is a case in which an experienced person disagreed with the system, and the reason is either captured and studied or lost. The standard is a mandatory reason on every override and a routine that reads them.
True Aim requires a reason for every human change and flags goodwill separately, so that a payment made outside cover never trains the model. Human and automated decisions are counted side by side in the claim record, which gives the operation a clean audit trail and clean training data at the same time.
Could you explain any single decision to the claimant, in plain language, from the record alone?
The person on the other side of a claim is entitled to understand why it was decided as it was. The standard is an explanation assembled from the record itself: the facts relied on, the clause applied and the calculation performed, in language the claimant can follow. If the explanation has to be reconstructed by hand, the record was incomplete.
True Aim decides cover line by line, each line justified against a specific clause or configured business rule, quoted and explained, with the arithmetic for excess, limits, caps and liability set out in full. Rejected lines are struck through on the source document itself and exported as a PDF, so the customer sees the reasoning against their own invoice.
Does the evidence survive new claims, time and change?
The last group looks beyond a single decision to the conditions around it: how the system was proven, what happens to the records when circumstances change, and where accountability sits throughout.
How does the system perform on claims it has never seen, and what changed after the last such test?
An accuracy figure measured on the claims a system was tuned against measures familiarity, not competence. The standard is a test on claims the system has never seen, labelled by an experienced adjuster, with the failure modes written down and the remedy recorded.
At True Aim, thirty unseen claims surfaced failure modes the development set had not covered, which is precisely what such a test is for. Closing the gap meant more ground truth and, for the payout formula, sitting with an adjuster and walking the calculations by hand. No amount of prompt engineering fixes a business-logic error. That work is what decides whether a pilot reaches production.
If you change systems, do the decision records leave with you in a form you can read without the software?
Decision records belong to the claims operation, not to the software that produced them. The standard is an export a person can read without the system: the documents, the decisions, the reasons and the versions, in an open format. An audit trail that only opens inside the system that produced it is not an audit trail.
True Aim exports the complete record of a claim, every action by a person or a machine with its model, prompt, rule version and confidence attached, as a single audit pack in one click.
Where is claim data processed, by whom, and who is accountable for each decision made in your name?
A claims operation remains accountable for every decision made in its name, whichever software made it. Swiss supervision is explicit on the point: FINMA expects institutions that use AI to align their governance, risk management and control systems accordingly, and an insurer that outsources a function remains accountable to FINMA as if it performed that function itself. The standard is a published list of processors, a known processing location and a record that shows, for every decision, whether a person or a system made it.
True Aim stores and processes customer data on Microsoft Azure in the Switzerland North region and does not transfer it outside Switzerland for routine processing. Our list of sub-processors is available on request, with the service, the data categories and the processing location for each. The claim record counts human and automated decisions separately and lists every action chronologically by actor, so accountability is visible for every decision, not asserted for the whole.
The principle behind True Aim
Everything True Aim does either produces a decision, feeds a decision or proves a decision. A claim is brought into a single, complete record. The evidence is tested for integrity before anything is decided, and anything that fails that gate never reaches the decision stage. Each decision is computed against the wording in force, with its reason attached. Every action, by a person or a machine, is written to a tamper-evident record that can be reconstructed, exported and explained for as long as it is kept. The architecture is described in full in our article on evidence chains.
That is what it means for compliance to be built in from the first claim. The judgment stays with you. The arithmetic, the wording and the evidence no longer need to.
How to use the scorecard
Choose one claim you closed last year, ideally a declined one, and work through the twelve questions against that claim, from the record alone. If the decision can be rebuilt field by field in minutes, the operation can answer a customer, an auditor or a supervisor under any regime, and the 2027 date is a formality. If it takes a quarter, that is the real deadline, and it was never 2027.
The test we would put to any board is simple. If a regulator asked tomorrow to walk through one automated claims decision, line by line, could you?
The scorecard applies to any system, including ours. Our Trust Center sets out how True Aim answers each of these questions, and we welcome the conversation.
Key takeaways
- Choose auditability at the point of design. Provenance recorded from the first claim is the only kind there is.
- Compute what can be computed. Rules where the wording covers the facts, models where the evidence is unstructured, and a record that shows which is which.
- Make human judgement a real control. The reviewer sees the evidence, holds the authority to change the outcome, and every override is captured with its reason.
- Prove performance on unseen claims first. An accuracy figure measured on the claims a system was tuned against measures familiarity, not competence.
- Keep the records yours. Tamper-evident, exportable in one click and retained to match your own obligations.
Frequently asked questions
What questions should you ask before adopting AI in claims handling?
Twelve, in four groups. Provenance: where each value came from, whether a rule or a model produced it and what the system does when evidence is incomplete. Reconstruction: whether a decision can be rebuilt a year later, whether the versions that produced it are on the record and who can alter the record and for how long it is kept. Human judgement: what the reviewer sees, whether overrides are captured and learned from and whether any decision can be explained to the claimant from the record alone. Continuity: performance on claims the system has never seen, portability of the records at exit and where data processing and accountability sit.
Does the EU AI Act's December 2027 deferral change anything for claims handling?
No. Regulation (EU) 2026/1744 moved the date on which most high-risk obligations begin to apply, to 2 December 2027 for standalone systems listed in Annex III. Claims handling is not listed in Annex III, and the rights of a person subject to a solely automated decision have applied under GDPR since 25 May 2018. The deferral changes the calendar for other systems. It does not change what a claims operation owes the people whose claims it decides.
What makes an AI claims decision auditable?
Four properties. Every value the decision used traces to its source document and page. The decision can be reconstructed later, because the rules, the policy version and the model that produced it are recorded on the decision itself. A person with the authority to change the outcome reviewed it with the evidence in front of them. And the record is tamper-evident, exportable and retained for as long as the operation's obligations require.
Is a human sign-off enough to make an automated claims decision defensible?
Only if the review is real. The reviewer must see the evidence behind the recommendation and hold the authority and competence to change it. A sign-off that nobody could realistically refuse leaves the decision solely automated in substance, and under GDPR the person affected retains the rights that follow from that.
Can AI claims decisions be deterministic?
Largely, yes. Where a policy, catalogue or rulebook is converted into a decision tree, the outcome is computed from the facts, and the same inputs always produce the same decision. Models are needed to read unstructured evidence and to resolve genuinely ambiguous items, and a well-designed record marks those lines as probabilistic so the two are never confused.
What should a Swiss insurer expect of an AI claims system?
FINMA's Guidance 08/2024 of 18 December 2024 sets out what supervision found on AI governance, inventories and risk classification, data quality, testing and ongoing monitoring, documentation, explainability and independent review, and states that FINMA expects supervised institutions using AI to align their governance, risk management and control systems accordingly. Under FINMA Circular 2018/3 an insurer that outsources a function remains accountable to FINMA as if it performed the function itself. A system should therefore be able to evidence each of those points from its own records.
Related reading
The architecture behind provenance and reconstruction.
Why testing on unseen claims decides whether a pilot reaches production.
Why the deferral changes nothing for a claims operation.
Sources
- Regulation (EU) 2026/1744 (Digital Omnibus on AI), Official Journal 24 July 2026, in force 27 July 2026. eur-lex.europa.eu
- Regulation (EU) 2024/1689 (AI Act), Annex III and Article 113. eur-lex.europa.eu
- Regulation (EU) 2016/679 (GDPR), Articles 15(1)(h) and 22. eur-lex.europa.eu
- FINMA, Guidance 08/2024: Governance and risk management when using artificial intelligence, 18 December 2024. finma.ch
- FINMA, Circular 2018/3: Outsourcing – banks, insurance companies and selected financial institutions under FinIA, last amended 4 November 2020, in force 1 January 2021, margin no. 24. finma.ch
- Grant Thornton, 2026 AI Impact Survey Report, insurance findings; fielded 23 February to 18 March 2026, 950 executives, 100 in insurance. grantthornton.com
Regulatory position stated as at 10 September 2026. This is not legal advice.