"Where did that come from?"
The CEO is looking at slide six. Net revenue retention, 121%. It's a good number — better than last quarter, better than plan. She points at it and asks a simple question: "How is that calculated?"
The room pauses. The analyst who built the slide left in April. The figure came out of a tool that assembled it from a few systems. Someone says they can look into it. Someone else says it's probably right, it matches roughly what they'd expect. Nobody can walk the calculation.
Nothing is necessarily wrong here. NRR of 121% may be entirely accurate. But the number has become unexplainable, and an unexplainable number is a liability regardless of whether it's correct. It can't be defended to a board, reconciled to the ledger, or handed to an auditor without a research project first.
Now add AI to that picture. If the tool that produced 121% also wrote the sentence explaining it — and generated that number itself rather than reading it from a governed calculation — the problem compounds. You have a confident figure, a fluent explanation, and no way to verify either.
This is the central question finance leaders should be asking about AI right now. Not "can AI help?" — it can. The question is where AI belongs in the chain that produces a financial number, and where it emphatically does not.
Finance holds AI to a different standard
AI has landed easily in marketing, sales, and support. A drafted email that's 85% right is a productivity win — a human edits it and moves on. A support response that's slightly off gets corrected in the next message. The cost of an imperfect output is low, and the feedback loop is fast.
Financial information doesn't work that way. It carries four requirements that most AI use cases simply don't.
Accurate. Not approximately right. A revenue figure that's 3% off isn't a draft — it's a misstatement, and it may propagate into a board decision, a covenant calculation, or an investor update.
Traceable. Every reported number must be followable back to source. Not "the system produced it" but an actual path to the underlying transactions.
Repeatable. The same period must produce the same figure every time it's calculated. Period-over-period comparison is meaningless if the method drifts between runs.
Defensible. Someone has to be able to explain and stand behind the number — to a board, an auditor, an acquirer. "The AI generated it" is not a defense.
Generative AI, by design, satisfies none of these natively. It's probabilistic, it doesn't carry lineage, and the same prompt can yield different phrasing or emphasis on different runs. That's not a flaw in the technology — variability is exactly what makes it useful for language. It's a mismatch between what the tool does and what a financial figure requires.
Which means the question isn't whether AI is good enough for finance. It's whether it's being pointed at the right job.
Why fabricated numbers are a category of risk on their own
Language models generate plausible output. That's the whole capability. Ask for a revenue figure and a model will produce a number that reads like a revenue figure — formatted correctly, in a sensible range, delivered with the same fluency as a real one. Whether it corresponds to your actual data is a separate question the model has no inherent way to answer.
This is what makes hallucinated financial numbers uniquely dangerous. A fabricated fact in a marketing blog is embarrassing and correctable. A fabricated number in a board pack propagates.
Consider the paths it travels. A hallucinated churn rate informs a retention investment. A confabulated forecast assumption shapes a hiring plan. An invented cash figure influences fundraising timing. By the time anyone notices — if anyone notices — decisions have been made on it. And unlike a spreadsheet error, there's no formula to inspect. There's nothing to audit, because nothing was computed. The number simply appeared.
The detection problem is the worst part. A wrong number produced by a broken formula usually looks wrong: it breaks a total, fails a check, sits oddly against the prior period. A hallucinated number is generated to be plausible. It's designed to pass the eyeball test. Fluency is not accuracy, and in finance the more polished the output, the more scrutiny the number behind it deserves.
None of this is an argument against AI in finance. It's an argument about which part of the work AI should touch.
Deterministic calculation vs. generative AI
The distinction that resolves most of this is between two fundamentally different kinds of computation.
A deterministic calculation takes defined inputs, applies fixed logic, and produces the same output every time. Beginning ARR plus new plus expansion minus contraction minus churn equals ending ARR. Run it a thousand times on the same inputs, get the same answer a thousand times. You can inspect the logic, test it, and reproduce it. This is how financial calculation has always worked, and it's what makes numbers auditable.
Generative AI produces likely output based on patterns in training data and context. It doesn't compute — it predicts what text should come next. Run it twice and you may get different phrasing, different emphasis, occasionally different substance. That variability is a feature when you're drafting language and a defect when you're producing a figure that has to reconcile to a ledger.
These are not competing approaches to the same problem. They're tools for different problems.
Financial numbers should come from deterministic engines operating on validated source data under governed business logic — full stop. That's where accuracy, repeatability, and auditability come from.
Language about those numbers is a genuinely different task, and it's one where generative AI is well suited.
Calculating vs. explaining
Worth making the split explicit, because "AI in FP&A" gets sold as though these are one activity.
Calculating a number means determining what NRR is for the period. This requires: validated data from source systems, agreed metric definitions, reconciliation across billing, CRM, and ledger, correct period cut-offs, and deterministic logic applied consistently. The output must tie to source and reproduce exactly. This is engine work. It should never be probabilistic.
Explaining a number means articulating what 121% NRR means — what drove it, how it compares to prior periods, which movements mattered, what a board should take from it. This is language work: synthesis, context, clear communication. The inputs are already-computed figures; the output is prose.
The first job has a right answer that must be reproducible. The second has a range of good answers, judged on clarity and usefulness. Different jobs, different tools.
Where AI earns its place is squarely in the second. Given validated financial outputs, AI can genuinely help finance teams:
- Interpret results — turning a reconciled ARR waterfall into a clear account of what moved.
- Identify meaningful trends — surfacing that gross churn has climbed three quarters running, when the quarterly view alone doesn't make it obvious.
- Summarize variances — drafting the actual-vs-plan commentary that eats hours every close.
- Answer executive questions — letting a CFO ask "what drove the expansion beat?" and get an answer built from computed figures.
- Communicate more effectively — converting technically correct detail into something a board can absorb in the time they have.
That's substantial value. It's also entirely downstream of the numbers. In every case, the figures already exist, already reconcile, and already trace to source. AI is describing them, not deciding them.
And the humans stay in charge. AI drafts; finance reviews, corrects, and signs. Accountability for what reaches the board never transfers to a tool, and no AI should be making financial decisions on its own. The value is in speed and clarity of communication, not in delegating judgment.
The four requirements for responsible AI in finance
Four concepts, worth defining plainly.
Traceability. Every number in an AI-generated narrative connects back to a computed figure, which connects back to source data. When the commentary says expansion drove the quarter, that number is the same one in the waterfall, from the same reconciled base — not a figure the model produced independently.
Explainability. Anyone can articulate how a number was derived. Not the model's internals — the financial logic. What went in, what rule applied, what came out. If the answer is "the AI determined it," you don't have explainability.
Financial governance. Metric definitions are agreed and applied consistently. Data is validated before use. Sources are reconciled. Figures are finalized through a controlled process and don't shift after distribution. Governance is what makes numbers dependable by design rather than by whoever happened to build the file.
Auditability. An external party can follow the trail from a reported number to its supporting records. This is a hard requirement — for auditors, for diligence, for a board that wants proof rather than assurance. AI-generated numbers with no computational trail fail it outright.
Together these produce a clear principle: every AI-generated financial narrative should be grounded in validated financial outputs. The AI reads computed figures and writes about them. It does not independently generate metrics, invent assumptions, or estimate a number it couldn't find. When the data doesn't support a claim, the right behavior is silence, not a plausible sentence.
Where SMPL.ai fits
SMPL.ai is built on exactly that division of labor.
SMPL reads and reconciles information from your source systems — billing, CRM, and the general ledger — into one governed operating model, then computes SaaS metrics from that reconciled base through deterministic calculation: the ARR waterfall, NRR and GRR, deferred revenue, recognized revenue, cash, headcount as a driver. Same inputs, same outputs, every time. Every figure carries lineage back to source, so a headline number can be walked to the contracts underneath.
SMPL never writes back to your ERP or general ledger. It reads from your systems of record; it does not post transactions to them. Your books stay yours, owned by your team and your auditors.
And the AI narratives are grounded in validated engine outputs. They explain the movements the engine actually computed — helping finance teams understand what happened and communicate it clearly — rather than generating metrics of their own. The numbers come from deterministic calculation on reconciled data. AI puts them into language a board can read. That boundary is the design, not a limitation someone will remove later.
Questions to ask before trusting AI with your numbers
Practical diligence for any AI tool touching financial reporting:
- Where do the numbers come from? If the answer involves the model producing figures rather than reading computed ones, stop there.
- Are calculations deterministic? Ask directly: does the same period produce identical results on repeated runs? Ask for a demonstration, not an assurance.
- Can I trace any figure to source? Take a headline number in the demo and ask to see the transactions behind it. The path should be short and visible.
- What happens when the data doesn't support an answer? A good system says it can't determine something. A concerning one produces a confident number anyway. Test this deliberately.
- Does the narrative use the same numbers as the reports? Commentary and tables must draw from one computed set. If the narrative is generated separately, they can silently disagree.
- Does it write to my systems of record? Know exactly what a tool can modify. A reporting layer should read from your GL, not post to it.
- Can an auditor follow this? If you can't hand a reported number and its supporting trail to a third party, that's a diligence and audit problem waiting to surface.
- Who signs off? The tool should produce drafts a human reviews and approves. Any workflow where AI output reaches a board unreviewed is a governance failure regardless of the technology.
Ask these of any vendor, including us. A tool that answers them concretely is one worth evaluating; a tool that answers with adjectives isn't.
See it on your own numbers
The test that matters is whether the AI commentary you're shown cites figures you can trace — and whether the same period reproduces the same numbers twice.
Book a demo and we'll run that on data that looks like yours, so you can judge the grounding for yourself.