Skip to content

AI in finance

An answer that cannot show its evidence is not an answer.

Finance teams are careful with AI for a good reason, and it is not conservatism. A confident answer with nothing underneath it costs more than no answer, because somebody acts on it. The line worth drawing is between a system that explains from records and one that produces the shape of an explanation.

Published

Why is finance careful about AI?

Because the cost of a wrong answer is asymmetric. In most places an approximate answer is a small inconvenience. In treasury and accounting it becomes a figure somebody reports, a payment somebody releases, or an exception nobody escalated.

Fluency makes that worse rather than better. A well-formed explanation reads as a considered one, so the more natural an answer sounds the less it gets challenged, which is the opposite of what you want from a system that is sometimes wrong. Caution here is a correct response to the incentive, not a personality trait of the profession.

So the question a finance team asks is never whether the thing is clever. It is what the answer is built on, and whether they can go and look.

What does it mean for an answer to be grounded?

It means every claim in it points at a record you can open. Not a citation of a document: a pointer to the statement line, the journal entry, the payment run or the rate that produced the sentence.

Grounding is also about what a system refuses to say. A ledger that is missing three days is a fact, and an assistant that fills the gap with a plausible continuation has done the most damaging thing available to it. The right behavior is to name the gap, say what it leaves out of the figure, and stop.

The test is cheap to run. Ask the same question twice, a week apart, over data that has not changed. A grounded answer is the same answer, because it is a retrieval. If it drifts, it was composed, and a composed answer about money is a liability with a friendly tone.

Where does a language model belong, and where does it not?

It belongs on top. Tresora AI is an assistant layered on top of a deterministic engine. The reconciliation, the positions and the reporting do not run through a language model, and that boundary is the design rather than a limit on it.

The reason is that matching is arithmetic and evidence, not language. A movement is explained by weighing candidate entries against dates, amounts, references and counterparties, scoring them, and keeping the ones that lost. That process has to be reproducible byte for byte, because somebody will ask for it a year later, and reproducible is exactly what a probabilistic model is not.

What the assistant does is the part that genuinely is language. It explains what is on the screen, points at what needs attention, and runs routines inside the same approval path and the same permissions the person asking already has. It answers about the engine’s work rather than doing the engine’s work.

How do you test an assistant before you trust it?

Give it a question you already know the answer to, and check the working rather than the conclusion. Anybody can be right once. What you are buying is the ability to be checked.

Three probes separate the two kinds quickly. Ask it something the data cannot support and see whether it declines or improvises. Ask it for the records behind a number and see whether they open. Ask it the same thing on Monday and on Thursday and see whether the answer moves.

Then look at what happens when it is wrong. A system that revises a conclusion when new evidence arrives, retires the previous verdict rather than overwriting it, and shows you both, is a system you can live with. One that quietly changes its mind has removed the only thing you were going to review.

What to do about it

Buy the evidence, not the fluency.

The useful question about financial AI is not how advanced the model is. It is what happens between the question and the answer, and whether any of it survives for somebody else to inspect.

Ask a vendor to show you a conclusion the system changed its mind about. The ones worth having keep the first verdict, the evidence that overturned it and the scores of the explanations that lost, because that trail is what a controller needs in March about a decision made in November.

Where this is worked

The deterministic half, where the answers come from

Everything the assistant explains was decided somewhere else. These are the places where the deciding happens and where the evidence is kept.

Bring one month and one question.

One month of your own statements and the ledger extract that should agree with them. You bring the question this piece did not fully answer for your group, and we answer it on your own figures.

One session, whichever piece brought you here. No slides before the data.