Seventy percent of the $942 million increase in Blue Cross Blue Shield's new analysis came from secondary diagnoses, the extra conditions recorded beside the main reason a patient entered hospital. That is the figure AI builders should watch. A coding tool can change the price of a stay without changing a treatment order. BCBSA says this pattern appeared in member claims from 2023 through 2025 as hospitals adopted AI coding, although its public findings do not prove that software caused every additional code.
The association's central comparison is between what hospitals wrote down and what they did. More stays were billed as medically complex, yet BCBSA found no matching rise in the care recorded on those claims. In one example, hospitals documented more anemia after major bowel surgery without a comparable increase in transfusions. The insurer interprets that gap as evidence that AI is finding more billable conditions rather than sicker patients, according to its September analysis.
That conclusion matters beyond one dispute over US hospital bills. AI coding products search records for information people miss, and the BCBSA findings show how a recovered condition can also become a price input. Better extraction and higher spending can arrive in the same deployment, even when the model has made no clinical decision.
How an extra diagnosis changes the bill
Hospitals do not bill every inpatient stay as a simple sum of individual actions. Payment often turns on a diagnosis-related group, or DRG, that places the stay into a category based on the principal diagnosis, procedures and patient complexity. A qualifying secondary condition can move the claim into a group with higher reimbursement. Fierce Healthcare's account of the BCBSA briefing says the share of Blue plan inpatient claims classed as medically complex rose from 37% at the start of 2023 to 40% by the end of 2025.
BCBSA attributed about $653 million of the increase to more than 55,000 cases above its 2023 baseline in which a secondary diagnosis moved the claim into a higher-paying DRG. The association says the full $942 million estimate excludes cases where its data showed additional care. That narrower definition is important: the reported total is meant to capture higher coding intensity without an observable treatment change, according to the briefing coverage.
AI fits this workflow because medical records contain far more detail than a human coder can inspect quickly. The BCBSA release says more than 60% of hospital systems now use AI-enabled tools that can scan notes, lab results and electronic records for secondary conditions. Finding an overlooked condition may make a claim more complete. When that condition also changes the DRG, the same retrieval feature produces more revenue.
There is still a clinician between a lab value and a valid diagnosis. The official US coding guidelines for fiscal 2026 say an abnormal lab, imaging or pathology result should not be reported as a diagnosis unless a provider documents its clinical significance. The guidelines also say a secondary diagnosis should meet reporting criteria such as requiring evaluation, treatment, monitoring or a longer stay. Software can surface evidence and propose a code, but it cannot turn an unexplained number into a clinically supported condition by itself.
What Blue Cross found
The current analysis used de-identified claims from Blue Cross and Blue Shield companies, which collectively cover one in three people in the United States. BCBSA has published the headline figures and an example involving bowel surgery, but the release does not provide row-level data, a hospital list or a full statistical appendix. Readers can inspect the claimed effect, not reproduce it from the public material.
The finding also comes from the payer in the transaction. Blue plans owe more money when a claim moves into a higher DRG, so BCBSA has a direct financial reason to challenge greater coding intensity. That does not invalidate the claims analysis. It does mean its causal language deserves the same scrutiny that a hospital vendor's savings estimate would receive. TechCrunch accurately framed the result as an insurer claim rather than a settled measure of AI's total effect on health costs.
BCBSA has seen a similar pattern before. Its March report on coding intensity examined maternity admissions and reported more acute posthemorrhagic anemia diagnoses without the expected rise in transfusions. That paper linked the change to the spread of ambient documentation and autonomous coding tools. The September analysis uses major bowel procedures as another test case, which makes the repeated diagnosis-treatment gap harder to dismiss as a quirk of a single specialty.
Repetition does not settle cause. Patients may be more complex in ways that claims treatment markers fail to capture, and hospitals may be correcting years of under-documentation. A diagnosis can be real without triggering the one treatment chosen as a comparison. The BCBSA release shows what was billed and paid; it does not supply the full clinical record needed to judge every bedside decision.
The association describes AI adoption at the hospital-system level, while its cost estimate comes from changes in claims. The public analysis does not disclose a hospital-by-hospital adoption timeline or a control group that would separate AI from new coding policies, staff training or financial pressure. BCBSA found a large association with a plausible mechanism. It has not published proof that the software invented diagnoses.
The model is doing the job it was given
Medical coding tools are built to improve code capture and documentation completeness, functions that the BCBSA report says vendors connect to higher reimbursement. Those measures reward a system for finding more supported conditions. They do not ask whether the payment schedule turns each recovered condition into a proportionate charge. The model can meet its product metric while the surrounding system spends more for the same visible care.
A classifier may be accurate at detecting a possible secondary condition, yet the workflow can still fail if clinicians rubber-stamp suggestions or auditors cannot trace the evidence. The 2026 coding guidelines require provider documentation and clinical significance, so the relevant unit of evaluation is the completed claim: source note, suggested diagnosis, clinician confirmation, DRG change and associated treatment. Accuracy on isolated code suggestions leaves the financial consequence out of the test.
The CMS rules provide a practical boundary. A hospital should be able to point from a secondary code to the responsible provider's documentation and explain how the condition met reporting criteria. An insurer should be able to contest that chain without using an opaque denial model of its own. CMS's coding guidance already treats accurate diagnosis assignment as a shared responsibility among providers, coders and billing agents. AI changes the speed and volume of suggestions, not that obligation.
For developers, an audit log matters more here than another percentage point on a benchmark. Because CMS requires provider support for a diagnosis, each suggestion should retain its source passage, model version and final human decision. Hospitals could then measure how often an AI-added diagnosis changes payment and how often a later audit reverses it. Payers could identify which additions lacked supporting care instead of treating every coding increase as waste.
The next evidence to demand
BCBSA's $942 million estimate is large enough to justify a closer audit and incomplete enough to resist a simple verdict. Its public release does not include the analysis design needed for replication. A useful follow-up would show results for hospitals before and after adoption, test several treatment markers for each diagnosis and give independent researchers access to suitably protected data.
Regulators and buyers should watch what vendors optimize. The March BCBSA report describes vendor claims about revenue capture, an incentive that rewards more payable codes. Evaluations should also track unsupported suggestions, payment changes reversed on review and diagnoses that appear without care expected in the clinical context. Those outcomes can be measured even when hospitals and insurers disagree about the thresholds.
The next BCBSA release should be judged by what it makes reproducible. If the diagnosis-treatment gap reported in its September analysis remains after hospital-level controls and independent review, the case against revenue-focused coding AI will be much stronger. If added diagnoses survive chart review as clinically supported, the problem sits partly in payment rules that attach thousands of dollars to better documentation. The $942 million figure opens the audit; it does not finish it.