Why do treasury evaluations take so long?
Because the buyer inherits the output. A marketing team that picks the wrong tool changes tool. A treasurer who picks the wrong one spends two years explaining figures they cannot reproduce, to people entitled to ask.
That asymmetry shows up as a specific behavior: the buyer keeps asking about the unhappy path. What happens when a statement does not arrive, when two names are the same company, when no rate exists for a pair on a date, when somebody has to be told no. Those questions read as obstruction from the outside and they are the actual evaluation.
The other reason is that the decision is rarely one person’s. Treasury, accounting, IT and often internal audit each check a different thing, and they check in sequence because each question depends on the previous answer. A short evaluation usually means somebody was skipped.
What is a treasury buyer actually checking?
Whether the output survives being questioned. The screens, the coverage and the plan are all proxies for that one property, and experienced buyers go at it directly.
Concretely they are checking three things. Reproducibility: can the same period be run twice and produce the same answer, from records that still exist. Traceability: does every figure resolve to the statement line or entry it came from, without anybody exporting anything. And refusal: does the system decline when the data cannot support an answer, or does it produce one anyway.
The third is the one most vendors have not thought about and the one that predicts the next two years. A missing rate, a statement that never arrived, a counterparty that resolves two ways: every group meets all three. What the system does in those moments is the difference between an exception you review and a number you find out about later.
Which demo questions separate one system from another?
The ones that ask to see something fail. A demonstration of a clean month proves the screens exist; a demonstration of a bad month proves the system does.
Five worth writing down before the call. Show me a line that did not reconcile, and the records behind why. Show me the explanations that lost, with their scores. Show me a figure whose currency had no published rate that day, and what the total says about it. Show me a case the system reopened because new evidence arrived, and where the old verdict went. Show me the same period run twice.
Ask them on your own data if you can, and on one month rather than a year. One real month with real differences tells you more than a full-year sample, because the sample was chosen by the vendor and your month was not.
What should a vendor show you on the first call?
One movement, end to end. The statement line, the entries it was matched to, the evidence for the match, the alternatives that were rejected with their scores, and the trail that exports. If that cannot be shown in the first hour, the rest of the evaluation is going to be about screenshots.
And one refusal. Ask what the system does with a figure it cannot support, and expect to be shown a decline on the surface with the reason beside it rather than a smoothed-over total. A vendor who has never had to build that has never had a customer ask for it.