Nine landmark ambient-scribe studies. UCLA, Mass General Brigham, Emory, UCSF, Yale, UC Davis, Kaiser Northern California, Penn, Stanford. Here is every federally funded health center in the country, plotted at once — and not one of them is in that list.
Dr. Gigi Magan, a family physician in a safety-net exam room, spent a month building a lecture on ambient AI scribes. Forty-one references, each checked against the primary source. Then she lined up the settings of every major study and found the same nine names: academic medical centers and one large integrated system.
Published trials or large studies conducted in a federally qualified health center: zero. Peer-reviewed studies reporting scribe accuracy stratified by patient race, language, or accent in real clinical use: also zero.
“A gap in evidence is not a verdict. It is a to-do list.”
So here is the to-do list, as data. Every HRSA Health Center Program grantee in the 2024 Uniform Data System — 1,352 organizations, 32.3 million patients. Each dot is one grantee. Horizontal: how many patients it serves. Vertical: what share of those patients HRSA records as best served in a language other than English.
The vertical axis is the axis nobody has published on. Across all 1,352 grantees, 8.98 million patients — 27.8% of the total panel — are recorded as best served in a language other than English. In 190 grantees, covering 4.9 million patients, that share is above 50%. The upper-right of this chart is where documentation burden per visit is highest and where speech recognition is weakest, and it is empty of evidence.
Add one line to your next ambient-documentation RFP: stratified accuracy data by preferred language and demographic group, from a deployment setting comparable to ours, before contract. You will probably not get it. The answer you get back — and how fast — tells you more about the vendor than the demo did.
Look at the scatter with no filter and there is a tilt to it — bigger panels look like they carry more non-English-preferred patients. Across all 1,352 grantees the correlation between log panel size and other-language share is r = 0.21. That is the kind of number a slide deck rounds up to “scale correlates with linguistic diversity.”
Now drag Min panel size to 20,000. The tilt mostly goes away: r falls to 0.12, and it keeps falling. The apparent relationship was largely produced by the long tail of very small rural grantees clustered at near-zero language share — 552 rural grantees have a median other-language share of 3.4% against 25.8% for the 800 urban ones. It is a geography artifact wearing a scale costume.
| Min panel | Grantees | r (log size × language) | Weighted language share |
|---|---|---|---|
| none | 1,352 | 0.205 | 27.8% |
| 5,000 | 1,137 | 0.204 | 28.0% |
| 10,000 | 881 | 0.136 | 28.8% |
| 20,000 | 484 | 0.115 | 30.0% |
| 50,000 | 156 | 0.111 | 31.4% |
The demographic gradient in the underlying technology is not speculative. The 2020 PNAS test of five commercial speech recognition systems found an average word error rate of 0.35 for Black speakers against 0.19 for white speakers — nearly double — tracing to the acoustic model and thin training data. Vendors have invested heavily since. Nobody has published the check in a deployment setting like the ones on this chart.
“Best served in another language” is a HRSA reporting field, not an audio measurement. It is a self-reported UDS count of patients whose care is best delivered in a non-English language. It says nothing about accent, dialect, code-switching, or the acoustic conditions of the room — which is precisely what a scribe pipeline fails on. A center at 5% could still have most of its visits conducted in accented English.
The grain is the grantee, not the clinic. One bhcmis_id can run dozens of sites across a state. A 295,386-patient grantee is an aggregate; the language share is a weighted average that can hide a single heavily multilingual site inside a mostly English network, in either direction.
There is no ambient-scribe field in this dataset. The zero on this page is not a column HRSA reports. It is the count of published studies Magan located after checking 41 references — an absence assembled by hand, and the honest characterization is “none found,” not “none exist.”
The vintage is the file date. UDS tables here carry no performance-year column; 2024-12-31 is the publication vintage used as the year proxy. Seven grantees report no language field and are excluded, which is why 1,352 and not 1,359.
Cardiology already built the machinery. The AHA AI Assessment Lab ran EchoGo Heart Failure against ~90,000 real-world echocardiograms and published subgroup findings by race and age alongside the accuracy numbers. Ambient documentation is deployed far more widely than that algorithm and has no equivalent body doing that work.
None of that needs a grant. Decline rates by language. Note quality by population. Edit burden on interpreter-mediated visits. It needs someone to decide on day one that equity gets measured instead of assumed — and the entry requirement is a QI dashboard, not an R01.