

In the Second World War, Allied statisticians studied bombers returning from raids over Europe, mapped where the bullet holes clustered, and proposed reinforcing those areas. Abraham Wald pointed out the flaw. The holes showed where a bomber could be hit and still come home. The armour belonged where the returning planes had no holes at all, because the planes hit there did not return to be studied.
Every customer health score has the same problem, and almost nobody names it.
The model is built from the accounts sitting in your customer base. Those accounts stayed. The ones that would teach you most about churn stopped generating data at the moment they became instructive. So the score learns what surviving customers look like and quietly assumes the inverse is what churn looks like, which is not the same claim at all.
That is one of four structural reasons a SaaS customer health score lies to you. This piece covers all four, what an honest score looks like instead, and where AI agents for SaaS renewal forecasting actually help rather than adding a confident number to a bad model.
A customer success health score is a composite metric combining product usage, support activity, relationship strength and commercial signals into one number meant to predict whether an account renews, expands or leaves. In practice it is a weighted average, and the weights are usually set by someone in a room rather than derived from outcomes.
The standard customer health score formula looks like this: assign each signal a normalised score, multiply by a weight, sum, and band into green, amber and red.
That mechanism is fine. The failure is in what goes into it, because most teams populate the model with whatever data is easiest to pull. Logins. Ticket counts. NPS. None of those measures whether the customer is getting value, which is the thing the renewal actually turns on.
The gap between what the score measures and what it claims to predict is where every problem below comes from.
This is the missing-planes problem. A score fitted to your current base captures what healthy accounts look like today, which is not the same as knowing which signals preceded departures.
The correction is to build the training set from churn events rather than from customer snapshots. For each account that left, reconstruct its signal history in the 90 and 180 days before it gave notice, and ask which values distinguished it from accounts that renewed in the same window. That is the difference between real churn prediction and pattern-matching against survivors, and it is what separates customer health score prediction from description.
Product usage is generated by the people who like your product. That is a self-selecting population, and it can look completely healthy while the renewal is dying.
The failure mode is specific. A power user logs in daily, adoption trends look strong, feature depth is fine, and the score sits green. Meanwhile the VP who signed the contract has changed, the new one has a preferred vendor, and nobody in your product data has any idea. Usage predicts activity. Renewal is a commercial decision made by someone who may never log in.
For enterprise accounts, executive engagement is consistently reported as the strongest single signal, and it is the one most customer health score software handles worst, because it lives in calendars and email rather than in product telemetry.
Most scoring models only move when something actively pushes them. Nothing pushes an account that has simply gone quiet, so it stays wherever it last landed.
The result is a slow accumulation of false greens. An account scored healthy in January, with no login since March, still reads green in April because no negative event has occurred. Silence is not neutral. In retention it is one of the most reliable warnings you have, and a score without time decay is structurally unable to express it.
The fix is decay: every signal loses weight as it ages, so a score drifts toward neutral when nothing new arrives. That single mechanism converts silence from invisible into visible.
Logins are trivial to pull, so logins get weight. Executive relationship strength is hard to quantify, so it gets left out. The model ends up optimised for data availability rather than for predictive power, which is how you end up with a 27-factor weighted average that CSMs override on instinct.
If product engagement depth correlates with churn several times more strongly than NPS in your data, giving them equal weight is a decision to be wrong on purpose. Weights should come from outcomes, and they should be re-derived as the product changes, because what predicted churn two years ago will not predict it now.
This is the question that exposes whether a customer health score model was thought about or assembled, and the honest answer is that tickets are not a one-directional signal at all.
High ticket volume is usually treated as unhealthy. In reality, deeply invested customers file more tickets precisely because they are using the product hard and want problems fixed. Meanwhile a disengaging account often files nothing, because nobody cares enough to complain. Weighting tickets as simply negative will misclassify your most committed customers as at-risk while missing the ones quietly heading for the door.
The practical weighting principle: product usage carries the most weight as a leading indicator, but read it as velocity rather than volume, and read tickets by type and sentiment rather than by count.
A useful split for a first version, treated as a hypothesis to be recalibrated rather than a rule: product engagement depth and trend carries the largest share, buyer and executive relationship next, support signal read directionally rather than by volume, and commercial signals such as contract value trajectory and payment behaviour making up the remainder. Then re-derive all of it against your own churn events within two quarters.
Four properties, and each one addresses a specific failure above.
Global thresholds are the largest source of false positives. A login threshold sensible for your enterprise tier is nonsense for self-serve. There is a real difference between "this account logged in 30% less this week" and "this account's primary admin has not logged in for two weeks, breaking an eighteen-month pattern." The first is noise. The second is a warning, and only a per-account baseline can tell them apart.
An account sliding from 80 to 70 to 60 over three weeks is in trouble. An account holding steady at 60 is not. An account that spikes to 75 during a campaign then falls back to 55 has not improved, despite briefly clearing a healthy bar. Rate of change is the signal; absolute level is context.
Every input loses weight with age. Accounts drift toward neutral when nothing new happens, so going quiet costs an account points instead of preserving its last good state.
The step almost nobody does. Every quarter, take the accounts that churned and check what your score said 90 days before. Take the accounts that renewed and do the same. If green accounts churn at a similar rate to amber ones, the bands are decorative. This closing of the loop is what separates retention analytics from a dashboard.
By treating the score as an input to a revenue number rather than as the output. Health scoring and revenue forecasting are different jobs, and conflating them is why so many renewal forecasts are a colour-coded list rather than a number a CFO will sign.
A renewal rate forecasting model takes each account's renewal date, contract value, expansion potential and risk-adjusted probability, and produces a forecast range rather than a point estimate. The health score supplies the probability. The contract data supplies the weight. Neither is useful alone.
Net revenue retention is the headline metric and it is the one most likely to flatter you. Expansion inside healthy accounts masks churn inside unhealthy ones, so NRR can look strong while the base is eroding. Teams have run 110% NRR with gross retention below 85% and treated it as a success until expansion stopped compounding.
For context on where you sit, 2026 benchmarks put median B2B SaaS NRR at roughly 101% to 106%, with clear separation by contract size: enterprise above $100K ACV around 118%, mid-market around 108%, and SMB under $25K ACV around 97%. Median annual logo churn sits near 3.5%. Gross retention benchmarks cluster around 92% at scale, with strong performers near 98%.
Read those by segment rather than blended. An SMB-focused business at 97% NRR is at benchmark, not failing.
Three places, and they are unglamorous. Assembling the signal history per account from product telemetry, CRM, support and billing systems, which is integration work rather than intelligence. Flagging accounts whose trajectory changed rather than whose level is low. And drafting the risk narrative for the renewal review, so the CSM starts from an evidenced position rather than a blank page.
What an agent should not do is set the probability without showing its reasoning. A renewal forecast that cannot be interrogated is a forecast nobody will defend in a board meeting.
Here is what cuts against building any of this, including against our own commercial interest.
A learned health score needs churn events to learn from. Published guidance for AI-based scoring puts the practical floor at roughly two months of history and around 40 historical churns before the model produces something better than guesswork. Below that threshold you are fitting noise, and the model will produce confident scores with no predictive content, which is worse than no score because people act on it.
If you have fewer than 40 churn events, the correct answer is a simple rules-based score with honest weights, reviewed monthly by someone who talks to customers. A spreadsheet with good weights beats a platform with bad ones, and a CSM with fifteen accounts and a calendar beats both.
There is a second uncomfortable truth. A health score is only worth building if a change in it triggers something. Scores that move from green to amber with no playbook attached and no action taken are a reporting exercise. The score is the cheap part. The operating routine around it is the expensive part, and it is the part that actually retains revenue.
Codiste builds retention and forecasting agents for B2B SaaS teams, starting with your churn history rather than a scoring template. If your health score has never been checked against what actually happened, our AI agent development services begin there.




Every great partnership begins with a conversation. Whether you're exploring possibilities or ready to scale, our team of specialists will help you navigate the journey.