Nobody would tolerate a calculator that is usually right. If your pocket calculator returned 6.99 for seven from time to time, fluently and with apparent confidence, you would not keep it for the days it gets things right — you would throw it out, because you can no longer trust any answer it gives you, including the correct ones. We hold deterministic tools to an absolute standard precisely because they are the part of thinking we have agreed to stop checking. And yet, in the one place where the tolerance for error is genuinely zero — the arithmetic of a financial model — there is now a fashion for handing the work to a tool that is, by design, only usually right.
That a generative model should not be building your financial statements is an argument worth making once and moving past. The more useful question sits underneath it: if not the calculation, then what should an AI touch in a model, and what should it never go near? The answer is not “none of it”, which throws away something genuinely useful, and it is not “all of it”, which is how you arrive at a confident, fluent, wrong cash flow. There is a line, and it is worth drawing precisely.
What a chess experiment settled
The cleanest map of where that line falls comes from chess. After Deep Blue beat Garry Kasparov in 1997, Kasparov did something more interesting than sulk: he proposed a format, later called advanced or centaur chess, in which a human plays alongside a computer rather than against one. The experiment reached its sharpest result in a 2005 freestyle tournament, open to any combination of people and machines. Grandmasters with powerful computers entered and, at first, dominated. The eventual winners were two American amateurs, Steven Cramton and Zackary Stephen, running three ordinary PCs.
They did not out-calculate the grandmasters or out-compute the supercomputers. Their advantage was process: they knew what each chess engine was good and bad at, asked it the right questions, and decided which of its answers to trust. Kasparov’s conclusion is the part that matters here. A weak player with a machine and a good process beat both the strongest computer working alone and a strong player whose process was worse. The edge lived not in the human and not in the machine, but in how the two were joined.
A model has exactly this shape
A financial model is the same kind of system, and it wants the same division of labour. The engine — the deterministic calculation — should do what engines do without error: the arithmetic, the statements, the linkages, the identical answer every time. That part should be closed to a probabilistic model for the same reason you would not let one overrule your calculator. The objection is not that the AI is unintelligent; it is that “plausible” is the wrong target in a place where only “correct” will do, and a system trained to produce the likely next token is built to hit the wrong target beautifully.
The human supplies the other end: which business this actually is, which assumptions carry the strategy, whether the result is one they can live with. The AI’s proper seat is between the two. It should not calculate. It should interrogate, explain and challenge what the engine has computed — which is precisely what the chess amateurs were extraordinary at, reading the machine’s output and pressing it with the right questions.
The middle seat is genuinely valuable
That middle role is worth having, and it is where generative AI earns its place in finance. Given a model whose numbers can be trusted, it can read the statements back in plain language, name the two or three assumptions the outcome truly hinges on, point out that the downside case turns cash-negative in the second year, notice a margin drifting in a way nobody had flagged, and compare this version with the last to say what changed and why it matters. None of that requires inventing a figure. All of it requires a figure that is already sound to reason about. AI is at its best here exactly because it is reasoning rather than computing.
This is also why asking a general chatbot about “your model” is a stranger act than it appears. A generic language model does not hold your structured assumptions, your calculation logic, your version history or your scenario results. Asked to reason about your model, it has little choice but to imagine a plausible one — useful for loosening up your thinking, hazardous for making a decision on. The difference between that and a model-aware analysis is not that one AI is cleverer. It is that one is reasoning about an actual computed model and the other about a confident guess. Context, not horsepower, is what makes the centaur work.
Drawing the line on purpose
FinModeler is built as that centaur, with the line drawn deliberately rather than left to chance. The deterministic engine owns the calculation: structured assumptions in, the same auditable statements out, every time. The AI layer sits on top and reads the result — its analysis names the strengths, the risks and the drivers behind them, generated from your own numbers rather than a generic prior, and aimed at interrogation rather than arithmetic. The randomness stays out of the formulas and the reasoning stays out of the sums. The way those two halves are joined is the part that decides whether the whole thing clarifies or misleads, which was Kasparov’s point all along.
The discipline, then, is to give each half the work it is actually good at. The engine calculates, because it never tires of being exact. The AI interrogates, because it never tires of asking. And you decide, because the decision was always going to be yours, and no tool on either side of the line was ever going to take it from you.

Build your model and ask it what matters — see how FinModeler reads your numbers on FinModeler.
Have questions? Talk to us.
FAQs
Should I use AI to build a financial model?
Not to perform the calculation. Financial arithmetic has a zero tolerance for error, and a generative model is designed to produce plausible output rather than guaranteed-correct output. Let a deterministic engine compute the statements, and use AI for the reasoning around them.
So what can AI safely do with a model?
Read and explain the results, identify the assumptions the outcome depends on, flag risks and inconsistencies, compare versions, and suggest scenarios worth testing. In short, everything that involves interpreting or challenging numbers — provided it never has to invent them.
Why is asking a general chatbot about my model risky?
Because it does not have your model. It lacks your structured assumptions, your calculation logic, your version history and your scenario data, so it reasons about a plausible-looking stand-in rather than the real thing. That is fine for brainstorming and unsafe as the basis for a decision.
Sources
- Advanced/centaur chess and the 2005 freestyle result (Steven Cramton and Zackary Stephen, two amateurs with ordinary computers, defeating grandmaster and supercomputer entrants on the strength of process): Garry Kasparov, “The Chess Master and the Computer”, The New York Review of Books (2010); accounts of the PAL/CSS Freestyle Tournament.
