When AI Gets Your Numbers Wrong: A Founder’s Guide to Fact-Checking Financial Projections
AI financial projections can look completely trustworthy and still be wrong — sometimes by a lot. Large language models are prone to calculation errors and confidently invented numbers, so any revenue forecast, growth rate, or market-size figure an AI hands you needs to be checked against real math and real sources before it ever reaches a spreadsheet, a lender, or an investor.
If you’ve used ChatGPT, Gemini, or Copilot to draft your startup’s financial model, you’re not doing anything wrong. Founders don’t have time to build a five-year forecast from scratch, and AI genuinely speeds up the boring parts. The problem is what happens next — when a polished, professional-looking model gets treated as if it’s already correct.
Why AI Struggles With Numbers in the First Place
Language models are built to predict plausible-sounding text, not to run arithmetic reliably. That distinction matters a lot when the “text” is your burn rate. A study on ConvFinQA financial reasoning tasks found ChatGPT miscalculating “$753 million + $785 million + $1,134 million” to be $3,672 million instead of $2,672 million, even though all the intermediate results were correct. That’s not a rare glitch — the researchers noted such mistakes can be critical in the financial domain, particularly in high-stakes setups.
It gets more specific when you look at CFA-style exam testing. Researchers evaluating ChatGPT and GPT-4 on mock exam questions found CoT prompting negatively affected both models particularly in Quantitative Methods, which could be due to hallucinations in mathematical formula and calculations. And when the error rate was broken down by type, calculation errors accounted for 17.2% of ChatGPT’s mistakes and 28.6% of GPT-4’s mistakes on questions the models had actually gotten right without step-by-step reasoning.
A more recent side-by-side test comparing chatbots on finance and economics tasks found the gap is still wide open. The largest performance gaps appear in finance and economics, where Grok and Gemini both reach accuracy levels of 76.7%, while ChatGPT, Claude, and DeepSeek fall below 50%. Most of those errors weren’t conceptual misunderstandings — researchers grouped them into categories and found “sloppy math” errors made up 68% of all mistakes, where the AI understood the question and the formula but failed in the actual computation. In other words, the model knew what to do and still botched the arithmetic.
Practitioners are seeing the same thing outside the lab. One financial modeling professional who ran identical inputs through different AI tools found Copilot calculated figures correctly until Depreciation, which then created knock-on effects in the EBIT and Net Income totals along with the dependent Taxes calculation. The takeaway from that test was blunt: you will need to review even the simplest calculations, because there is little logic in where an error might occur.
It’s Not Just Bad Math — It’s Invented Numbers
The scarier failure mode isn’t a wrong sum. It’s a number that never existed anywhere, stated with total confidence. This is the same behavior that got Deloitte in hot water. In a widely reported case, a Deloitte Australia report submitted to a government client contained fabricated references and a made-up quote — 19 hallucinations were identified in a single report, first flagged by a University of Sydney academic, and Deloitte acknowledged the AI usage and offered a partial refund. If a global consulting firm can ship a client deliverable full of invented citations without catching it, a solo founder rushing to finish a pitch deck at midnight absolutely can too.
Financial projections are especially vulnerable to this because they’re full of the kind of specific, checkable claims AI likes to guess at: market size figures, industry growth benchmarks, competitor revenue, customer acquisition costs. People who build AI-assisted pitch decks for a living warn about exactly this pattern — AI can confidently invent figures like ARR, growth rates, or market size, so every metric on an investor deck must come from your own real data and be fact-checked before the meeting.
Where the Errors Actually Sneak Into Your Model
Based on how these tools fail, there are a few specific spots to watch:
- Month-to-month flow. AI-generated models frequently break at the seams between periods — Month 11 does not flow into Month 12, and revenue minus costs does not equal profit even though each individual number looks reasonable on its own.
- Cited benchmarks with no source. If your AI tool tells you “the average SaaS churn rate is X%” or “your industry grows at Y% annually,” treat it as a placeholder, not a fact. Every statistic an AI cites needs a traceable source — if you cannot find the original report or dataset, don’t use the number.
- Multiplying “AI runs” instead of building real scenarios. Running the same prompt three times and calling the outputs base, optimistic, and pessimistic isn’t scenario planning. Real scenarios require deliberate, documented changes to specific assumptions, not random variation in language model output.
- Charts that hide bad extraction. If you or an AI tool are pulling numbers back out of a pitch deck, remember most financial projections live in charts, not text — tools that read only text miss chart data entirely, which is the most common and consequential gap.
A Practical Checklist Before You Send Your Model Anywhere
You don’t need a finance degree to catch most of this. You need a habit.
- Recalculate the totals by hand or in a spreadsheet formula — never trust the AI’s arithmetic, even on simple addition. That’s where errors like the $3,672 million miscalculation slip through.
- Trace every benchmark back to its original source. If the AI can’t tell you exactly where a growth rate or market-size figure came from, assume it made it up until proven otherwise.
- Check that every month connects to the next. Revenue, costs, and cash balance should reconcile cleanly from one period to the next with no unexplained jumps.
- Use AI for structure and research, not final numbers. Let it draft the layout, suggest what categories you’re missing, and speed up formatting — then run the actual math through a deterministic tool or your own formulas.
- Get a second human to read it. Not another AI tool — a person who understands your business and will ask “wait, how did you get this number?”
None of this means you should ditch AI for building your financial projections. It’s genuinely useful for getting a rough structure fast, brainstorming assumptions, and cleaning up formatting. But the moment those numbers are going in front of an investor, a lender, or a co-founder deciding whether to quit their job for you, they need to have gone through a human who checked the math and the sources. Fast and polished isn’t the same as correct — and in a fundraising conversation, that difference is the whole ballgame.
Hi! I use AI to help research and write posts on this site. I do my best to keep things accurate, but please double-check anything important — and nothing here replaces advice from a licensed or certified professional.