Why Every Retirement Calculator Gives You a Different Number
Your friend's app says 95%. This one says 82%. Nobody's lying — they're answering different questions, with different assumptions, and calling both answers a "score." Here's how to read any retirement number like someone who knows what's behind it.
- Retirement scores from different apps aren't measuring the same thing. A "90" can mean a probability, a coverage ratio in a deliberately bad market, or a count of historical survivals — three different quantities wearing the same costume.
- Five quiet assumptions move the number more than your inputs do: which average return the math uses, whether the model promises a market rescue, how long it thinks you'll live, whether it subtracts real taxes, and what it assumes you earn after you retire.
- No calculator can be "right" — nobody predicts the future. The trustworthy tool is the one that tells you what its number means, shows every assumption, and lets you change them.
First: you can't actually compare the scores
Here's the thing almost nobody tells you. When two retirement apps give you two different numbers, your instinct is to ask which one is right. But the numbers usually aren't even answers to the same question. Comparing them is like comparing a thermometer reading to a barometer reading and asking which one is the real weather.
| When this tool says "90"… | …it actually means |
|---|---|
| RetirementScenario, Boldin, ProjectionLab | "In 90% of the simulated futures we ran, your money outlasted the plan." A probability — but built on each tool's own assumptions, so the same 90% still isn't interchangeable between them. |
| Fidelity Retirement Score | Not a probability at all. It's the share of your estimated retirement expenses your plan could cover in a significantly below-average market — their deliberately pessimistic test case. A 90 here is a coverage ratio in a bad storm, not a 90% chance of anything. |
| FIRECalc / cFIREsim | "Your plan survived 90% of every actual historical stretch of markets since 1871." A backward-looking survival count — history's report card, not a forecast. |
All three framings are legitimate. But if your friend quotes their 95 against your 82, step one isn't panic — it's asking what each number is a measurement of.
The five quiet assumptions that move the number
Once two tools are at least measuring the same kind of thing — say, both running Monte Carlo simulations — the remaining gap comes almost entirely from assumptions most people never see. There are five that matter most. None of them require math to understand. All of them require honesty to disclose.
1. Which average? (Averages lie.)
Try this: your portfolio loses 50% one year and gains 50% the next. The "average" of those two years is zero — sounds like you broke even. You didn't. You're down 25%. ($100 drops to $50, then grows 50% to $75.)
That gap between the flattering average and the average you actually live gets bigger the bumpier your investments are, and over a 30-year retirement it compounds into a very large difference. At typical market bumpiness — a 12% standard deviation — the flattering average overstates compounding by roughly 0.7 percentage points a year, every year; identical plans score materially higher on that math by construction. Some tools build their simulations on the flattering average; some use the lived one. Our engine uses the lived one — the math term is volatility drag, and the exact formula is published in our methodology. Boldin, to its credit, discloses its approach too: its 2025 methodology update moved to the flattering (arithmetic) average. That's a legitimate, documented modeling choice — but it means the identical plan with identical inputs will score higher there than here, by construction.
2. Does the model promise a rescue?
Tools that replay actual market history — FIRECalc, cFIREsim, ProjectionLab's historical mode — inherit a comforting fact: American markets have always, eventually, recovered. Every crash in their data is followed by the rebound that really happened. That recovery is quietly baked into their scores.
Our simulations don't promise the rescue. Each simulated year's return is drawn fresh, so a bad stretch isn't automatically followed by the recovery that bailed out every previous bad stretch. Which approach is right? Genuinely unknowable — that's a debate professionals haven't settled. But you should know which bet your score is making. A plan that's only safe if history repeats is a different plan than one that's safe even if it doesn't.
3. How long does it think you'll live?
Most tools quietly plan to a life expectancy that is really a coin flip — half of people live longer than it. And for couples the math is stronger than intuition suggests: for a healthy couple at 65, there's roughly a 44% chance at least one of you reaches 95. Running out of money at 91 is not a rounding error; it's the single most expensive way to have been optimistic.
We plan to the longer-lived spouse's horizon and model what actually changes when one spouse dies first — the smaller Social Security check disappears, taxes shift to single filing, one household on one income. Tools that plan each person to a median age, separately, will produce a happier number for the same couple. Happier, and thinner.
4. Who's subtracting the bills?
Two tools can agree on your savings, your spending, and the market — and still disagree by a lot, because one of them is subtracting bills the other skips. Taxes on your Social Security. State income tax, which varies from zero to double digits depending on where you retire. Medicare surcharges that switch on above certain income levels. Health-insurance costs in the gap years before Medicare, and healthcare inflation that has historically outrun regular inflation.
Our engine models all of those, per year, per state. Boldin also does serious tax modeling — genuine credit to them there. Simpler tools skip some or all of it, and skipped bills always make the score look better. A smaller number that already paid the bills beats a bigger number that forgot them.
5. What do you earn after you retire?
Most people de-risk their investments as they age — more bonds, less drama. A tool that carries your working-years growth rate through your whole retirement is assuming 85-year-old you invests like 45-year-old you. Our defaults assume you'll invest more carefully after you retire, because you almost certainly will. If a friend's tool carries one optimistic rate for life — often whatever they typed in — that alone can explain most of the gap between their score and yours.
A smaller number you can see into beats a bigger number you have to take on faith.
So who's right?
Nobody — and that's the honest answer, not a dodge. No calculator predicts the future. Every one of them, ours included, is a projection: a disciplined way of asking "if these assumptions hold, how does this plan do?" The future gets a vote no model can count.
But here's what should make you comfortable with the conservative end of the range: it's where the serious institutions already live. Fidelity deliberately scores your plan against a significantly below-average market. Empower trims its own return assumptions by a full percentage point just to be careful. J.P. Morgan's 2026 retirement research shows a typical balanced portfolio at a 5% withdrawal rate succeeding only about two-thirds of the time over a long retirement — and flatly calls the famous 4% rule "good in theory, poor in practice." When our numbers read more conservative than a cheerful app, we're not the outlier. We're in the boring, careful company we want to be in.
And the reason to prefer the careful end is the asymmetry of the two ways to be wrong. Too conservative, and you work a little longer or leave a little more behind — a real cost, but a recoverable one. Too optimistic, and you find out at 85, when there's nothing left to adjust. Those aren't symmetric mistakes, so we don't treat them symmetrically.
The five questions to ask any retirement tool
Here's the whole article as a checklist. Any tool — ours included — should be able to answer all five. Marks reflect each product's own public documentation as of July 2026.
| RetirementScenario | Boldin | ProjectionLab | Empower | Fidelity Score | FIRECalc / cFIREsim | |
|---|---|---|---|---|---|---|
| Can you see and change every assumption? | ✓ Yes | ✓ Yes | ✓ Yes | ◐ Some | ✗ Nofixed scenario | ✓ Yes |
| Uses the lived average, not the flattering one? | ✓ Yes | ✗ Noarithmetic, disclosed | ◐ Dependsmode & inputs | ◐ Partial−1% haircut | ✓ Yesbelow-average by design | ✓ Yesreal history |
| Subtracts real taxes — federal, state, Social Security, Medicare? | ✓ Yes | ✓ Yes | ✓ Yes | ◐ Approximate | ◐ Simplified | ✗ Noby design |
| Plans past median life expectancy, with survivor economics? | ✓ Yessurvivor economics | ◐ Partial | ◐ User-set | ◐ User-set | ◐ Conservativedefault horizon | ✗ Nosingle horizon |
| Can you check the math itself? | ✓ Yespublished formulas + testing regimen | ◐ Docs onlyclosed engine | ◐ Docs onlyclosed engine | ✗ Minimal | ◐ Detailed PDFclosed engine | ✓ Yesopen source |
Based on each product's public documentation, verified July 2026. Methodologies change — Boldin updated its return methodology in 2025, and others will evolve too. If we've mischaracterized anything here, tell us and we'll correct it.
What each tool is genuinely good at
This isn't a takedown — these are good tools, and the comparison only means something if we say so plainly. Boldin is a deep, full-featured planner with serious tax modeling. ProjectionLab gives you unusual control over assumptions and a beautiful historical mode. Empower's planner is free and connects to your real accounts. Fidelity's score is a deliberately conservative gut-check you can get in two minutes. FIRECalc and cFIREsim are free, open-source, and the most rigorous pure-history backtests available. Plenty of people are well served by each — and if one of them fits how you think, use it with our blessing. What we built is for people who want the careful math and every assumption out on the table.
What this doesn't mean
None of this means a conservative score is automatically a better score, or that a cheerful tool is lying to you. It means a score without its assumptions is just a vibe with a percent sign. The tools above disagree because they make different, mostly defensible choices — the problem is only that most people never find out which choices produced their number. Once you can see the assumptions, the disagreement stops being alarming and starts being useful: run the cheerful case and the careful case, and plan your life somewhere that survives both.
Common questions
Every assumption in this article is a dial you can turn in the app — returns, volatility, lifespan, taxes, all of it — and the methodology behind each one is published where you can read it. See your real number, and everything behind it.
Run your plan free →