
Let me start with a simple question. Would you sign an AI assistant sentence in a regulatory report? And I don’t mean, would you publish it on an internal website or use it as a first idea in a workshop? I mean, would you put your name on it in an SFCR, an RSR, or an ORSAR related?
Narrative. And that’s a more useful test than asking whether AI can write well. Of course it can write well. In many cases, it can write much more smoothly than humans do when humans are tired or under pressure or looking at the fifth version of the same paragraph.
However, fluent language isn’t the same as evidence. A sentence can be perfectly grammatical, logically structured and useful, and it might still be impossible to defend if somebody asks where the claim of the sentence comes from.
The standard I want to suggest today is not, can AI write a good sentence? The standard is: Can a named subject matter expert trace, test, and stand behind that sentence?
The distinction may sound small, but in spaces like regulated reporting, it changes the entire process design.
AI isn’t production-ready
This isn’t just a theoretical question. A market study conducted by EIOPA covering 347 insurance undertakings across 25 EU and EEA countries showed that 65% were already using generative AI and another 23% plan to implement it within the following three years. This data is from July 2025, so these numbers likely have only increased.
Across insurance, the question of whether generative AI will enter day-to-day work has largely been answered. However, EIOPA also found that 64% of reported use cases were still in proof of concept or experimentation, and only 32% were in production.
That’s the big gap. Adoption is real, but dependable production processes are still being built, and that’s exactly where reporting functions should focus their attention.
The interesting piece of this is not to choose a tool that produces impressive language in a demonstration. It’s in building a process that still works on the evening before a filing when a model run has changed and three people are waiting for an answer.
Regulated reporting is also particularly pertinent for this type of automation. The same EIOPA survey notes that 64% of reported use cases were in internal back office applications, just like, you guessed it, regulated reporting. Internal reporting is of course not risk-free, but it does make a sensible place for a bounded and human-reviewed use case for AI.
One caveat to add here to these figures is that they are sector-wide, not specific to reporting. Nevertheless, I think they show the trend.

The real problem starts after the numbers
When I speak with reporting, finance, actuarial and risk teams, I constantly hear a version of the same story. The calculation itself is fairly straightforward. Those models run and they don’t need AI.
The real pressure begins after that number is available. There are four jobs to consider here:
1. We have to assemble the narrative. Tables, explanations, comments, and context often live in different systems and belong to different teams.
2. We have to explain the movements. What changed, which drivers matter, and which drivers are material enough to mention.
3. We have to keep the evidence recoverable. Sources, version, methods, exceptions, and the owner of the answer.
4. We have to preview, review questions, edits, challenges and approvals, move across functions, legal entities, and sometimes several rounds of review.
These four jobs are actually tightly coupled, and that becomes visible usually at the least convenient moment. So imagine it’s 20 to 6 on a Tuesday evening before a filing. Suddenly, a number shifts by 0.4% and now you have to update the whole thing. Updating the numbers may take 10 minutes, but then the real work starts.
Where else does the number appear in the report? What other numbers were calculated from it? Which explanations about that number are no longer true and which sentences now suddenly contradict other sentences further on in the report? And finally, who needs to review all of this again and sign their name?
And that’s not just a writing problem. It’s not just a data problem. It’s a linkage problem between a fact, a statement, and the evidence behind the statement. If we use AI well, then that linkage problem is where the first practical value should appear.
The four questions I want every reporting sentence to answer
For me, a sentence in a reporting document that’s defensible has to answer to four questions.
Definition. What exactly does the metric measure? A solvency ratio, an SCR component, and a movement attribution may all be related, but they’re not interchangeable.
Reference basis. Are we talking about percent or percentage points? Year end to year end or quarter on quarter, actual to plan or something else? Many reporting errors are not grammar mistakes. Their errors are unit, comparison, basis or scope.
Evidence. Which source, which version and which method supports this statement? If I can’t trace a claim back to a stable fact base, I can’t. Review it. I can only decide whether I find it plausible, and that’s what’s currently happening. It’s just vibe checks.
Responsibility. Who can amend this statement? Who reviews it and who approves it? Automation does not remove human responsibilities. If anything, it should make the responsibilities more visible.
Put together, this is by far and large not a special standard for AI. It’s a standard that should also improve human written reporting if implemented correctly.
Facts into text, never the other way around
The operating principle that I use is very simple. Facts, pre-calculated, flow into text. Text never flows back into facts.
AI may at some point be adept enough to apply changing legislation and update pre-existing calculation tools or even invent new ones. However, these days, I’m not confident enough in AI’s math capabilities, to say it bluntly, to let it loose on mission-critical calculations. A controlled package of calculated facts should at least contain the value of the unit, the period, the source, the version, and the relevant methodology.
The AI drafting layer can then take this package and turn it into a clearer and more consistent language that humans can actually digest rather than staring on a complicated table. AI can prepare a first draft. It can make recurring wording more consistent. It can highlight a missing time period or a missing source. And it can prepare useful questions for a reviewer.
It should just never recalculate a number. It should not insert a figure because the figure sounds plausible. It should not turn a correlation or a sequence of events into some causal explanation when causality isn’t established. And finally, it shouldn’t decide on its own whether a driver is material, immaterial, the largest or the most relevant.
These boundaries aren’t limitations of value, they’re conditions for value. And even if it sounds like we’re restraining AI to dummy jobs, that’s because AI at the moment is just too confident and we need to re-rail it back to what it can actually reliably do.
The goal here is not to replace any expert judgment, just to give expert judgment a better starting point with the help of AI, meaning a draft that already carries its facts, its evidence, and its open questions with it.

Are insurances AI-ready?
One objection that constantly comes up is our data landscape isn’t ready. And that may be true if we mean an enterprise-wide AI program with access to every system, every historical record, and every type of report. If that were the precondition, then most AI programs would never see the light of day.
But for a first use case, this is really not the right threshold. A perfect central data platform is not required at all, and no access to historical sources that are harmonized or having solved every remaining data issues. There is one basic requirement, and that is having an accountable fact base.
For one reporting component or one sharply defined analytical question, ideally one that carries a decision after it that can be tested. So for that one component or decision, I need to know the exact sources, the metric definition, the period, the access rate, the version and the known exceptions. Most importantly, I need to see the gaps.
I don’t want a polished AI sentence to hide a missing definition, an unresolved data quality issue, or an incomplete attribution. A visible exception is inconvenient. An invisible exception is dangerous.
So I wouldn’t ask, is our organization ready for AI everywhere? I would ask, can we trace the facts and own the exceptions for this one paragraph or for this one reporting module? This isn’t avoiding the data problem, it’s actually facing it head-on on a specific and narrowly defined case.
The result of this is that you get to tackle it piece by piece where it matters most first.
How to make an AI draft reviewable
If we accept that facts and texts need a controlled relationship between one another, then we get four practical controls.
A fact boundary. The drafting workflow should receive only approved data and approved calculation outputs. This reduces the risk that a model uses general world knowledge, an old document, or an unauthorized file to create a plausible but unapproved statement.
A claim check. Each material statement should be checked against its value unit, period, source, and where relevant method. For a movement of 12 percentage points, the test is concrete. Is the figure in the approved package? Is the unit correct? Is the period correct? Does the source match?
An evidence trail. The version, method, exceptions, and status of the draft need to remain recoverable. This is what makes the process understandable three months later to an internal reviewer, an auditor, or a supervisor.
A human decision. A reporting owner must be able to accept, edit, reject, or accelerate the draft. A polished output is not an accountable statement until that decision has been made.
This is consistent with EIOPA’s 2025 opinion, which takes a risk-based and proportionate approach and highlights data governance, recordkeeping, fairness, cybersecurity, explainability, and human oversight under existing insurance sector expectations.
The practical lesson is this: Good AI governance should be visible in the workflow itself.
Going beyond hallucinations
One word about hallucinations. When people discuss generative AI, they often focus on this, and the risk is certainly real, but in reporting, some of the more dangerous errors are not invented sentences or facts. They’re just plausible sentences with the wrong meaning.
For example, a unit error can turn 12 percentage points into 12%. That’s not the same movement. An aggregation error can present a component risk as if it were the total risk without the relevant diversification or scope. A report-level consistency error can leave one paragraph correct on its own, but inconsistent with the statement 40 pages later.
A vague quantitative phrase such as low single digits can sound careful while hiding the fact that nobody agreed on the factual basis. A causal claim can say because of interest movements, even though the documented analysis only shows that two things moved at the same time without any cause or effect proven in the data. A materiality claim can call something the largest risk or an immaterial effect without prior establishing a defined threshold or comparison logic.
Five out of six of these aforementioned errors can be grammatically flawless, but they’re not. And that’s why language quality is not the control quality here.
The correct question isn’t: “did the model make up a number?” It is: “did the statement preserve the meaning, scope, and evidence of the approved fact?”
A realistic first step
Given all these pitfalls and perils, how do we actually start? My recommendation is not to begin with an AI transformation program. That would be a huge undertaking and a costly one at that.
My recommendation is much smaller: to begin with one accountable decision. I would choose a recurring reporting component or as defined analytical question where ownership, sources and review are already reasonably clear. As a practical rule, if I can’t describe the first use case in two sentences, I have chosen a too big use case.
Before building anything, I’d define the control boundary in writing. So: which sources are in scope, which factual fields are mandatory, what may the system do, what may it explicitly not do, which tasks remain with the subject matter function under all circumstances.
Then I’d run the existing process and the assisted process in parallel for one reporting cycle. I wouldn’t judge the pilot by how impressive the AI draft looks on screen. I’d judge it by whether reviewers can’t find relevant exceptions faster, trace the evidence more easily, and spend more time on judgment rather than searching across files with the help of an AI system.
I’d also measure missed issues and false alarms. Missed issues matter because a problematic statement may pass silently into the report. False alarms matter because if a system flags every second correct statement, people will stop trusting it and the review becomes a second manual production process.
Only when the first use case proves both utility and control would I send the approach to further reporting modules, common clusters, or scenario preparation.
The bottom line: better questions, not automated decisions
I’ll close where I started. Would you sign an AI-assisted reporting draft?
The useful answer is not a simple yes or no. The useful answer is a list of things that would need to be visible: a bounded fact base, a clear comparison basis, a claim check, a recoverable evidence trail, visible exceptions, and named human approval.
Those elements are in place. The opportunity is larger than faster writing: the same structured facts, comments, and exceptions can help a team identify recurring patterns, compare explanations across periods or legal entities, and ultimately ask better questions earlier.
A pattern is of course not a diagnosis, and a correlation is not causality. That’s why an AI workflow should just prepare questions and suggestions for experts, and not make management decisions on their behalf.
The goal is not a report written by AI. The goal is a report that is easier for people to understand, challenge, and stand behind.
Meanwhile, at Wangari
The Leipzig event was a blast! I had a wonderful time connecting with key reporting and risk staff at leading insurers in the DACH region, and learning how and where they’ve started using AI.
What I enjoyed most was seeing how differently each insurance company was treating AI, and where they were at in the balancing act between caution and enthusiasm. The above post is an adaptation of the 45-minute talk I gave there.
Reads of the Week
Cass Sunstein argues that an AI Regulatory Commission is needed in the US. Indeed, voices are getting louder asking the state to step into this hot mess. Cass lays out what that would look like in practice. I follow this from Europe; still worth my eyeballs, and probably yours, too.
Another chime into that same direction: Jim Amos writes that frontier AI needs regulatory capture. For those of you who are not law experts (that includes myself — I’m a physicist by training), this piece is very useful: It explains the distinction between regulation and regulatory capture, and, interestingly enough, the latter is “what all the most evil corporations aim for.” Despite the strong language, a very informative read.
In yet another take on AI regulation (did I mention it was trending?), Jim Messina writes that much of US regulation won’t come from Washington but from the states. This is interesting… As an EU AI company (though not one of the evil ones, I think — we’re too small for that), this frankly makes the US less interesting for us, not more.


