
Let’s start with the obvious: AI is worth adopting. It’s proven useful because it removes so much friction from everyday work: finding information or summarizing long documents, drafting an email or (ugh) putting together PowerPoint slides.
Productivity has become effortless. Anyone can generate a sophisticated five-page document in five minutes.
What’s missing is a layer of healthy discrimination: Just because you’ve produced a lot doesn’t mean you produced the right thing, or produced the thing in satisfactory quality (as opposed to quality that looks good on the surface but is threadbare underneath, as is the case with much AI slop).
But where do we start with evaluating whether AI is doing the job right or not? Let’s start where quality becomes immediately quantifiable: With the numbers.

The thing about accountability
There are legal precedents that hold companies accountable for whatever their AI tells their customers. In a legal precedent in 2024, Air Canada was sentenced to reimburse a customer for fares that the chatbot said were applicable to him and actually were not. In other words, if your chatbot hallucinates, you pay for it, not your customer and not the AI provider.
This also means that any output you productively generate at your job remains your company’s responsibility — and could cost you your job if you were the one messing things up.
Which is dangerous because AI is great at giving nice-sounding, but in actuality confidently bogus output. What we need is AI that people, and companies, can actually stand behind. And that’s not another model, it’s a separate architectural layer.
When the wrong numbers get you into jail
At my company, we wanted to solve this with a use case where errors could get you into jail but where verification is fairly straightforward. We found that in insurance reporting: It’s mandatory; if the numbers are wrong then people can end up in jail; and to verify them you compare the generated with the real numbers with no textual understanding required.
The real concern here is neither whether an AI can calculate a number — it can but pre-existing tools do so much more reliably already, nor whether the AI can write a number in a sentence (it can). The tricky part is whether an AI-generated sentence in an insurance report contains the right number on the right basis for the right period from the right source with the right units.
Higher standards for AI
To accomplish this, an AI chatbot should tie the source data — whichever tool or database the numbers come from — to totals, hierarchies, and business rules, not to mention public and external datasets.
The hard part here is not building a retrieval-augmented generative AI (a so-called RAG); that problem has been solved. The hard part is not about reading tables and letting AI pick the right numbers; it’s about understanding what those numbers mean and then put them all together in ways that really make sense and that the firm can stand behind.
In an insurance company, an AI may give a plausible reason for some quarterly loss ratio movement. But before this sentence becomes a management conclusion, somebody — the AI and also, crucially, a human as a final instance — needs to check whether the premium basis and the reserve development treatment, portfolio mapping, and all the other connected things are coherent with this claim.
Confabulation is a thing
I didn’t invent this — this tendency for confidently stating junk is what the NIST poetically called confabulation. So confidently, in fact, that users often believe it until it’s too late.
To mitigate confabulation, the NIST says one should compare outputs with known ground truth, use more than one method of evaluation, document data suitability, and fact-check any general information.
This sounds like duplicate effort, but it isn’t: because it catches things before they go bad. The temptation is to just give AI a nice database, so-called RAG technology, and then let it query that database and call it a day.
But the problem with it is that retrieval doesn’t solve any semantics or any calculation issues or reconciliation issues. It can’t explain why the number in a table is the way it is. And so we need a bigger architectural solution to make this kind of AI “sense making” layer happen.

High stakes, high risk?
Regulators are keenly aware of this. The EIOPA, an important insurance governance body, highlighted in 2025 that data governance, record keeping, explainability, and human oversight are vital for responsible AI use. The NAIC model bulletin, in turn, identifies potential AI-generated inaccuracies, data vulnerabilities and bias, and calls for governance, risk management, validation, and documentation.
So the regulators are not just blanketing a ban on AI, which we all know would be harmful and would just make people use AI under the radar.
What’s needed now is a layer that implements what regulators are asking for: a layer that sits on top of the numbers that are being calculated anyway, and which involves the humans that carry responsibility for its output in intelligent ways. (These are the same humans who so far have been painstakingly using their own brilliant minds on boring reporting rather than more important tasks.)
This can be fully automated with AI if, and only if, AI is actually made reliable. We thus need AI systems that not only determine what the right number is, but also what it does and doesn’t imply, what went into the calculation of the number, and under which hypothesis this calculation is sound.
That’s a set of capabilities that we can, in fact, implement architecturally by using AI agents. It’s not just another dashboard or a chatbot. It’s real infrastructure which resolves semantics and grounds any claims in authoritative data.
This kind of AI should be repeatable and calculate things logically, and it should signal uncertainty. So where there is no evidence or where it just isn’t clear, it should say, “I’m not sure.”
Which is something that AI by default is very bad at — but we can indeed make it say “I’m not sure” when it isn’t. Such a layer should produce an audit trail that’s reviewer-friendly, citing its sources and assumptions and transformations.
The technical implementation
How to build this layer without building a more powerful foundational model?
The secret, to my mind, is causal and agentic intelligence. Causal intelligence (also called causal inference in data science circles) is the discipline to tell not only what happened, but why, in a mathematically provable way.
Causal intelligence has remained small so far, however, because it used to require a lot of manual work of deeply skilled people. Agentic AI supercharges that, and because the methods themselves are inherently verifiable, this doesn’t create an additional AI reliability risk.
Causal intelligence, deployed and orchestrated by agents, makes causal relationships and underlying assumptions more explicit, which means that AI can treat an explanation as something to test or to qualify, and not just as a nice story that it can autocomplete.
How this ties to insurance reporting
In insurance reporting, we have plenty of numbers that need verifying. We also have governed data and time-sensitive metrics and multiple ambiguous definitions and reconciliation that needs to happen all across various data silos. And then there’s regulatory scrutiny and expert reviews. In short, everything is very complex.
That same pattern appears in finance, in risk, in capital planning, in operational reporting, in sustainability reporting, — the list doesn’t end.
Insurance reporting is thus our starting point, in order to test and to build out our architecture around this problem. Ultimately, we want to become the go-to provider of this AI sense-making layer for any enterprise data, beyond reporting and for all industries.
Towards sense-making enterprise AI
The future is not less AI. It’s AI that creates speed without dissolving accountability even a bit.
In order to respect that accountability, we will need to distinguish between enterprises that just use a nice interface and enterprises that build the context and the controls that are required to stand behind the answer that AI gives them.
And it’s the latter group that will save itself a lot of manual, boring work in the long run, which allows it to become much more competitive also over time.
Numbers, here, are not the forcing element — it’s just that numbers are very checkable, which is again a wonderful stress test for AI.
Once an organization can make numerical answers traceable, it has a foundation for more dependable AI everywhere else.
Meanwhile, at Wangari
On September 17, I’ll be speaking about AI in Solvency II reporting in Leipzig!
Insurance reporting is going through some key changes, including a reform on the narrative reporting from early 2027. This gives rise to some key risks and opportunities for AI solutions sold to or being developed within insurances.
If you’d like to meet me in Leipzig, registration for the event is still open (note that it’s in German). I’m looking forward to many inspiring discussions with German insurers and solution providers.
Reads of the Week
Sheri Oz dives into AI sycophancy. In a very cleanly set up experiment, she tests how agreeable AI is with the user prompt (ever heard AI tell you that your idea was brilliant?), and whether the content changes. The verdict: AI is highly sycophantic, but the factual evidence remains stable, whether one asks for compliments or criticism — which is good news, really.
I’m not usually one to hype a particular venture capitalist, but Ruben Dominguez’ piece on VC Sarah Guo’s contrarian AI bets is actually fantastic and thought-provoking. She bet on legal AI firm Harvey when they didn’t have an investor deck, and her fund has backed 6 of the 21 AI-native companies whose revenue runs over $100 million — all running on the thesis that massive AI labs can’t build the products that actually bring value themselves. She’s the genius identifying the business that comes after the AI.
A short and rather philosophical piece on Fernando’s Substack reminds us that AI is but a tool — it’s humans who give it a purpose. What should we train AI for, really? Our values may be well-encoded into language, but language is tiny compared to the “infinite richness of life.” How do we encode that into AI? Not an answer, but the questioning itself is worth reading.


