Ardent Health's ambient AI documented 20% more HCCs per visit. Compliance reviews said the codes were supported. That can be entirely true and still leave the interesting question open. Here are 2,985 US counties — risk score on one axis, hospitalizations nobody can document into existence on the other.
A hierarchical condition category is a payment object. It exists because someone wrote a diagnosis in a note, and the note got coded, and the code rolled up into a risk score that tells Medicare how much this patient should cost. The condition existed either way. The code is the artifact.
So when Ardent reports a 20% lift in HCCs documented per visit after turning on ambient AI, CMIO Brad Hoyt gets out ahead of the obvious read:
That is the right answer. It is also, word for word, exactly what the wrong answer would sound like — which is why it's worth knowing how big the gap between sick and documented already is, before anyone's ambient tool touches it.
Each dot below is one US county. Horizontal position is the average HCC risk score of its fee-for-service Medicare population in 2023 — CMS normalizes this to 1.00 nationally, so it is pure cross-sectional signal about how richly a population is coded relative to the country. Vertical position is a hard event: inpatient hospital stays per 1,000 beneficiaries. Dot area is the size of the county's FFS population.
Hospitalizations are not a documentation artifact. A person is admitted or they aren't. If risk scores were a clean read on sickness, these two would move together almost perfectly.
Start it at zero and the cloud is a mess — a long tail of tiny rural counties flung to the edges. The instinct is to read that spray as a real signal: look how many places break the pattern.
They don't. Drag the floor up and the correlation gets stronger, monotonically:
| Minimum county size | Counties | r (risk vs. IP stays) | r (risk vs. std. spend) |
|---|---|---|---|
| no floor | 3,140 | 0.683 | 0.533 |
| 250 benes | 3,078 | 0.684 | 0.530 |
| 1,000 benes | 2,613 | 0.707 | 0.602 |
| 5,000 benes | 1,092 | 0.751 | 0.723 |
| 20,000 benes | 313 | 0.778 | 0.742 |
Computed on the full county file before any of this page's filtering — CMS Geographic Variation PUF, 2023, all-beneficiary stratum.
This is the direction people get backwards. Small samples usually don't manufacture a fake trend — they dilute a real one, by burying it in noise. The counties at the edges of the unfiltered cloud are mostly places with 400 beneficiaries where nine extra admissions moved the rate by fifty points.
Both failure modes are live. Small n can invent a pattern that isn't there, and it can hide one that is. The only way to know which you're looking at is to move the floor and watch which direction the number goes.
Even at the tightest filter — the 313 counties with 20,000+ FFS beneficiaries, where the sampling noise is essentially gone — the correlation tops out around 0.78. Square it and the risk score explains roughly 60% of the variation in how often those populations are actually hospitalized.
Turn on Color by residual and you can see the other 40%. Red counties are documented richer than their hospitalization rate predicts. Navy counties are documented leaner. Two counties can carry the same risk score and differ by 60 admissions per 1,000 people.
That residual is not fraud, and it is not one thing. It is coding practice, EHR templates, MA penetration reshaping who is left in fee-for-service, dual-eligible mix, how many chronic conditions get re-documented each January, and yes, real unmeasured illness. It is the space every documentation tool operates in — and the space where “we closed the gap between care delivered and care documented” and “we moved the risk score” are the same sentence viewed from two directions.
Hoyt's argument in the piece isn't that the HCC lift is meaningless. It's that it's the wrong thing to be proud of, because a vendor can help you produce it. The metric he points at is the one nobody can manufacture:
The chart above is the argument for why he's right. Every quantity on it — risk score, spend per capita, coded conditions — is downstream of a documentation decision, and therefore movable. Hospitalizations aren't, which is precisely why they only track the risk score 60% of the way.
If you are about to report an HCC lift from a documentation tool, report the residual alongside it. Pull the same population's admissions, ED visits and 30-day readmissions for the twelve months before and after. If the risk score moved and none of those did, you have not found sicker patients. You have found better stenography — which may be entirely legitimate, and is still a different claim.