Tuesday, April 7, 2026

Chat GPT 20-page report on Value in Cancer Care, Today through 2030

18 pages as PDF

4500 words

Cut-paste below

Chat GPT "Deep Research" mode (45 minutes)

##

Value-Based Cancer Care Measurement to 2030

Executive summary

Purpose. This report is designed to help plan a fall conference panel on measurement as the “missing operating system” for value-based cancer care (VBCC): why progress has been limited, what measures would actually enable VBCC, what is missing today, what a tangible future state looks like, and how to get there by 2030. The analysis prioritizes primary/official sources (CMS/CMMI, ONC/ASTP, NCI, NCQA, HL7/mCODE, ICHOM) and peer‑reviewed evidence (notably Basch ePRO trials). [1–15]

Central finding. VBCC has not failed for lack of payment experimentation; it has stalled because measurement has remained too claims-centric, too process-heavy, too misaligned across programs, and too weakly connected to patient-centered outcomes and real oncology clinical context—while also imposing high reporting burden and inviting gaming. The Oncology Care Model (OCM) is the clearest signal: practices reported substantial care redesign, yet accountability quality measures showed no significant improvement versus comparison groups, and practice-reported process gains did not translate into patient-reported or claims-based outcome gains. [2–3] At the same time, patient experience scores were already high (a ceiling effect), and the COVID public health emergency complicated several patient-reported domains. [2–3]

Why this moment is different. The plausible “springboard” is the convergence of (a) CMS’s next-generation oncology model requirements (EOM’s eight participant redesign activities include ePROs, HRSN screening, CEHRT use, and CQI data use), [4–6] (b) the federal shift toward digital quality measures (dQMs) and alignment (Meaningful Measures 2.0; Universal Foundation), [10–12] and (c) oncology-specific interoperability infrastructure—USCDI+ Cancer (ONC + NCI, with CMS/CDC/FDA input) and HL7 mCODE as a minimal structured oncology dataset with FHIR-based exchange. [13–15] Together, these can reduce abstraction burden and make patient-centered outcomes measurable at scale.

What VBCC measurement should become. A credible VBCC measurement portfolio should be small (≈8–12 measures), outcome-forward, equity‑stratified, and digitally computable from interoperable data. This report proposes 8 candidate measures spanning: symptom/toxicity control via ePROs, functional status, avoidable acute care, evidence-based regimen appropriateness, timeliness, end-of-life (EOL) goal-concordant care, financial toxicity, and equity/whole-person supports. Each is defined with required data elements and risk-adjustment needs, and compared in a single table (below). [5–9], [13–18]

2030 future state in one sentence. By 2030, a “VBCC-ready” system can (1) capture core oncology facts (diagnosis, stage, key biomarkers, therapy intent) in standardized fields (mCODE/USCDI+ Cancer), [13–16] (2) routinely capture ePROs and respond clinically, [5–9] and (3) compute a core set of dQMs that are aligned across payers, auditable, equity‑stratified, and used for rapid-cycle improvement—not merely reporting. [10–12]

Assumptions. The exact conference date, panel duration, and confirmed panel format are not specified; this report assumes a 60–90 minute panel with a policy/technical audience and the term “VBCC website” refers to the trade publication Value-Based Cancer Care (valuebasedcancer.com). [19]

Why decades of work produced limited progress

VBCC is often defined as improving outcomes that matter to patients per dollar spent. That concept is clear; what has been unclear is which outcomes to measure, how to compute them reliably, and how to avoid drowning clinicians in reporting. Porter’s foundational framing underscores that value requires outcomes measurement, not just costing or utilization tracking. [1]

Claims-first measurement shaped incentives toward what is easiest to count. Early oncology “value” programs leaned heavily on claims-based utilization and narrow process metrics (e.g., ED visits, hospitalizations, hospice timing) because these are available and standardizable. In OCM, CMS held practices accountable on several such measures, yet the final evaluation found no significant improvement versus comparison groups on the accountable quality measures, even as practices reported substantial redesign work. [2–3] This fuels the “emperor has no clothes” reaction: lots of effort, weak signal of better patient outcomes.

Misalignment and measure proliferation diluted focus and comparability. CMS acknowledges the burden and fragmentation problem through Meaningful Measures 2.0 and the Cascade of Meaningful Measures, explicitly aiming to reduce burden, align measures, and prioritize what matters. [11–12] The Universal Foundation is CMS’s attempt to streamline high-priority measures across programs; a parallel NEJM analysis describes the need for cross-program alignment because CMS historically used hundreds of measures that were not always aligned. [12], [20]

Oncology’s clinical reality is not “measurement-ready.” Valid comparisons require clinical context—stage, biomarkers, line of therapy, performance status, recurrence/progression—and longitudinal follow-up. These elements are inconsistently structured in EHRs, often buried in narrative notes, and are not reliably inferable from claims. This is precisely why USCDI+ Cancer and mCODE exist: to standardize the oncology dataset and enable interoperable exchange. [13–16]

Attribution and episode definitions are structurally hard in cancer care. Cancer patients traverse surgeons, medical oncologists, radiation oncologists, hospitals, and ancillary services; attributing outcomes or costs to a single entity can be unstable. Empirical work on newly diagnosed cancer patients highlights attribution challenges that must be addressed for accurate quality measurement and payment design. [21] Episode-based oncology payment model design also faces definitional challenges because episodes must be observable from claims and clinically meaningful—often a tension. [22]

Burden and workflow misfit undermined “measure-to-improve.” Measurement has too often been “measure-to-report.” Meanwhile, oncology EHR workload has increased: a national analysis of oncology specialists’ EHR inbox work reported a 19% increase in weekly inbox messages from 2019 to 2022, with high burdens in medical oncology/hematology—making additional manual measurement particularly fragile. [23]

Measures that would support VBCC

A VBCC measure set should satisfy four design tests:

1) Patient-centered outcome linkage: directly reflects symptom burden, function, survival proxies, goal-concordant care, or financial well-being. [1], [5–9], [18]
2) Clinical actionability: results can trigger care redesign (navigation, toxicity management, regimen selection, end-of-life conversations). [4–6]
3) Computability at scale: feasible from claims + standardized EHR data + ePROs, using dQM specifications where possible. [10–12], [13–16]
4) Fairness and auditability: explicit risk adjustment and guardrails against selection, coding inflation, and exception misuse. [10–12], [24]

Candidate measure comparison table

Candidate measure (domain)

Precise definition (example specification)

Key data elements required

Primary data source(s)

Computability today

Major barriers / risk-adjustment needs

ePRO symptomatic toxicity control (patient-centered outcomes)

Among patients initiating systemic therapy in the measurement period, % completing standardized ePRO symptom assessments at defined intervals and % of severe symptom alerts with documented clinical follow-up within 48 hours

Patient identifier; therapy start date; standardized symptom instrument (e.g., PRO‑CTCAE/PROMIS domains); timestamped alert; follow-up action

ePRO platform + EHR; partially claims (therapy trigger)

Medium (increasing in EOM)

Workflow integration; standard alert logic; missing PRO completion in underserved groups; risk adjust by cancer type, regimen intensity, baseline symptom burden, language access. [5–6], [9], [11]

Physical function preservation (function)

Mean change (or % with clinically meaningful decline) in PROMIS physical function score from baseline to 3 months after therapy start

Baseline and follow-up PROMIS PF; therapy start; demographics

ePRO + EHR

Low–Medium

Baseline capture; instrument licensing/workflow; case-mix adjustment; ensure accessibility for older/disabled patients. [11], [5–6]

Avoidable acute care utilization during episodes (utilization/outcomes proxy)

Risk-adjusted rate of ED visits not leading to admission per 6‑month episode; optionally paired with symptom-triggered preventability review

Episode attribution; ED visit claims; admission linkage

Claims (CMS/payer)

High

Attribution; preventability not captured; risk adjust by cancer type/stage, comorbidity, social risk. [2–4], [21–22]

Evidence-based regimen appropriateness / pathway concordance (appropriateness)

% of new regimens consistent with specified evidence-based guidelines/pathways given stage/biomarkers (with documented exceptions)

Diagnosis; TNM stage; key tumor markers (e.g., ER); regimen and dosing; exception reason

EHR orders; pathway system; mCODE-aligned oncology data

Medium

Proprietary pathway definitions; incomplete structured stage/biomarkers; gaming via exception overuse; risk adjust by disease subtype and treatment intent. [16], [15], [25]

Timeliness from diagnosis/staging to treatment initiation (timeliness)

Median days from “initial diagnosis date” to initiating cancer therapy; stratify by cancer type and stage (and by referral source)

Date of diagnosis; date of therapy; stage; referral/consult timestamps

EHR + registry + claims

Low–Medium

Cross‑org data; ambiguous definitions; staging completion dates; risk adjust by complexity, access constraints. USCDI+ Cancer includes timeliness-related use cases. [13–14], [16]

Goal-concordant end-of-life care (EOL quality)

% receiving systemic therapy in last 14 days of life; % hospice enrollment ≥7 days before death; paired with goals-of-care documentation rate

Date of death; therapy claims; hospice claims; (optional) structured goals-of-care

Claims + EHR

High (claims) / Low (goals)

Death date availability; clinical nuance (appropriate late therapy in select cases); gaming by shifting care settings; adjust for cancer trajectory and patient preference. [2–3], [26]

Financial toxicity screening and navigation response (financial well-being)

% screened with validated instrument (e.g., COST) within 60 days of therapy start; among “high toxicity,” % receiving financial navigation within 30 days

COST score; therapy start date; navigation referral and completion

ePRO + EHR + revenue-cycle systems

Low–Medium

Many systems lack standardized workflows; risk adjust by baseline socioeconomic status; avoid penalizing safety-net providers; ensure language/cultural validity. [18], [6], [11]

Equity and whole-person supports (equity/HRSN)

% screened for HRSN domains and % with closed-loop resource connection; report key outcome measures stratified by race/ethnicity, dual-eligibility, and neighborhood disadvantage

Demographics; HRSN screening results; referral; completion; stratification variables

EHR + community resource platforms + claims

Medium (in EOM)

Data completeness; standard capture of race/ethnicity/language; accountability for community resource availability; risk adjust for structural barriers, not just clinical factors. [4–6], [13], [27]

What is lacking today

Data and standardization gaps

Structured oncology facts remain inconsistent across EHRs. CMS’s EOM Clinical Data Elements (CDE) Guide illustrates the minimum clinical detail needed even for a limited set of models: diagnosis date, death date, recurrence/relapse status, metastatic history, TNM staging, and tumor markers (e.g., ER for breast cancer), mapped to mCODE/FHIR elements for high-tech submission. [16] The existence of such guidance is progress, but it also highlights the current reality: many practices cannot reliably compute nuanced measures because these fields are either missing, inconsistently modeled, or unstructured.

Interoperability is improving, but unevenly adopted. ONC’s Cures Act Final Rule pushes standardized APIs and addresses information blocking to enable growth of interoperable apps and data use. [28–29] Yet, interoperability alone does not guarantee semantic consistency (same meaning, same code sets, same timestamps), which is required for measure validity.

Workflows, burden, and the “last mile” problem

ePROs have strong evidence but weak operational penetration. Randomized trials by Basch and colleagues show ePRO symptom monitoring can improve quality of life, reduce symptom burden, reduce acute care use, and in some studies improve survival. [7–9] However, scaling requires (a) patient enrollment and sustained completion, (b) alert triage protocols, (c) EHR integration, and (d) clinical response capacity—each a failure point.

Reporting burden competes with improvement. Rising oncology EHR message volume and “work outside work” time heighten the risk that new measures become administrative tasks rather than improvement tools. [23] VBCC measurement that is not digitally computable risks worsening burnout and undermining adoption.

Equity, patient-reported outcomes, and financial toxicity remain under-measured

Equity stratification is more policy requirement than measurement norm. EOM requires screening for HRSNs and embeds health equity-related redesign activities. [4–6] USCDI+ Cancer explicitly aims to define core real-world data elements to support care, research, and interoperability with cross-HHS involvement (NCI + ONC, with CMS/CDC/FDA input). [13] Yet, in practice, demographic and social risk data are incomplete or inconsistent, and closed-loop referral outcomes are rarely captured in standard fields.

Financial toxicity is measurable but not systematized. COST (FACIT-COST) is validated as a patient-reported measure of financial toxicity in cancer. [18] Despite this, most measurement programs still do not treat financial toxicity as a core VBCC outcome with defined numerator/denominator logic and accountability for navigation response.

Attribution and gaming risks

Attribution is an “engineering constraint,” not an afterthought. If quality measures are tied to episodes or entities that cannot be attributed consistently, the measurement signal becomes noise. Evidence on patient attribution in newly diagnosed cancer underscores these challenges for accurate quality measurement and payment. [21] Episode-based payment design literature similarly identifies the need for observable, well-defined intervals and compatible quality measurement. [22]

Gaming risks are real and predictable. When measures are process-based or loosely specified, organizations can optimize documentation, coding, and exception pathways rather than outcomes. Digital measures help only if paired with transparent specifications, version control, auditing, and validation.

Future-state VBCC measurement vision to 2030

A tangible future state is best described as three synchronized workflows: clinic, payer/program, and patient.

Clinic workflow

In a “VBCC-ready” clinic, care teams do not “report measures”; they run care workflows that inherently generate computable data:

·       Core oncology clinical facts (diagnosis, stage, key markers, treatment intent) are captured in consistent structured fields and exchanged via FHIR/mCODE-aligned profiles. [15–16]

·       ePROs are routine during systemic therapy, with standardized instruments and triage protocols, and ePRO data are visible in the EHR and used in daily operations. [5–6], [7–9]

·       Navigation, HRSN screening, and health equity plans are integrated into the episode pathway and monitored as operational KPIs. [4–6]

Payer/program workflow

·       A small aligned measure set is computed as dQMs (standards-based specifications, code packages, interoperable data) from multiple sources (claims + EHR/FHIR + ePRO + registries). [10–12]

·       Equity stratification is standard in reporting dashboards, and incentives are structured to avoid penalizing safety‑net practices for patient risk and structural barriers. [4–6], [13]

·       Measures are paired with learning: quarterly feedback cycles, targeted supports, and measurement updates with governance.

Patient workflow

·       Patients know what “value” means operationally: symptom control, function, goal-concordant care, financial well-being, and fairness.

·       Patients can contribute data via portals/apps without friction, and they see feedback loops (“you reported severe nausea; nurse called; antiemetic adjusted”). [5–6], [7–9]

·       Patients can access and share oncology data across systems due to standardized APIs and reduced information blocking. [28–29]

Milestones and timelines to 2030

CMS’s EOM runs through June 30, 2030, providing a real-world runway for maturing measurement and digital reporting pathways. [4], [30] The federal quality ecosystem also faces a widely cited goal of transitioning to all digital measures by 2030. [31] A pragmatic milestones framework:

·       Near term (through 2027): “Make patient-centered data routine.”

·       ePRO adoption reaches operational reliability in participating oncology practices (enrollment, completion, triage, documented responses). [5–6]

·       HRSN screening and referral workflows reach stable capture with stratified reporting. [4–6]

  • Core oncology structured data capture expands using EOM CDE-like elements and mCODE patterns. [15–16]
  • Mid term (2028–2029): “Make measures digitally computable and aligned.”

·       Multi-payer pilots compute a core VBCC measure set as dQMs using claims + FHIR + ePRO feeds. [10–12]

  • Governance matures: common measure specs, semantic validation, audit pathways, and versioning. [10], [24]
  • Long term (2030): “Benchmark outcomes credibly.”

·       Outcome benchmarking includes more direct measures of symptom burden, function, and goal-concordant care, equity-stratified and risk-adjusted, with substantially reduced abstraction burden. [10–13], [15–16], [31]

timeline
  title Staged VBCC measurement timeline to 2030
  2026 : Establish "minimum viable VBCC" measure set
       : Expand ePRO + HRSN workflows in oncology episodes
  2027 : Standardize core oncology data elements (mCODE/USCDI+ Cancer alignment)
       : Routine equity stratification for core measures
  2028 : Multi-source dQM pilots (claims + FHIR + ePRO)
       : Governance: semantic validation + audit models
  2029 : Cross-payer alignment (Universal Foundation-style streamlining)
       : Reduced manual abstraction through automation
  2030 : EOM ends; mature digital measurement ecosystem
       : Credible benchmarking of patient-centered outcomes

Springboard technologies and policies, plus governance needs

Digital quality measures and alignment infrastructure

CMS defines dQMs as quality measures using standardized digital data captured and exchanged via interoperable systems, with standards-based specifications/code packages, computable in an integrated environment. [10] Meaningful Measures 2.0 explicitly emphasizes simplifying PRO-PMs and embedding them into EHR workflow via APIs and use of standardized tools (including NIH PROMIS instruments). [11] The Universal Foundation is designed to streamline high-priority measures across CMS programs—addressing the “too many unaligned measures” problem. [12]

Interoperability rails: ONC’s Cures Act Final Rule, USCDI+ Cancer, and mCODE

ONC’s Cures Act Final Rule supports secure access/exchange/use of EHI, calls for standardized APIs, and targets information blocking—critical prerequisites for multi-source measurement and patient-centered data flows. [28–29]

USCDI+ Cancer is collaboratively managed by NCI and ONC with multi-agency input, aiming to define real-world data elements to support prevention, diagnosis, treatment, research, and care, with explicit use cases. [13]

HL7’s mCODE is a core structured oncology dataset intended to increase interoperability; an early peer-reviewed overview describes its organization into domains (patient, lab/vitals, disease, genomics, treatment, outcome). [15], [32] CMS’s EOM CDE guide demonstrates an applied approach: CDEs can be reported via templates or via HL7 FHIR API mapped to mCODE elements—illustrating a bridge from low-tech reporting to high-tech computability. [16]

ePRO operationalization in EOM

CMS provides stepwise guidance for ePRO implementation in EOM and encourages valid/reliable instruments suitable for diverse populations; EOM requires increasing uptake over time and expects integration into information system workflow with EMR visualization and eligibility identification. [6] This is the most concrete federal lever currently driving routine capture of patient-reported symptom and function domains in oncology episodes. [4–6]

AI-assisted extraction: promise and constraints

AI/NLP can reduce abstraction burden by extracting stage, biomarkers, progression/recurrence events, and toxicity from unstructured notes—but only if it is validated, monitored for bias, and anchored to standardized data definitions. Evidence from mCODE implementation pilots suggests promise but also highlights limitations of current FHIR APIs and structured data availability for complex oncology analysis—implying AI will be needed as a bridge while structured capture matures. [33]

Governance and validation requirements

Digitizing measures can digitize errors if governance lags. Minimum governance requirements:

·       Specification governance: open, versioned specs; code sets; change control; alignment across payers. [10–12]

·       Semantic validation and testing: eCQI defines semantic validation as comparing formal criteria to manual computation from the same test database—still essential as measures go digital. [24]

·       Equity safeguards: stratification requirements, fairness monitoring, and avoidance of perverse incentives that penalize providers serving higher-risk populations. [4–6], [13]

flowchart LR
  subgraph Data_Sources[Data sources]
    EHR[EHR clinical data\n(stage, markers, meds)]
    PRO[ePRO platform\n(symptoms, function, distress)]
    Claims[Claims\n(utilization, cost, hospice)]
    Registry[Cancer registries\n(dx, stage, survival)]
    SDOH[HRSN/Community resource data\n(screening, referrals)]
  end

  subgraph Interop[Interoperability + standards]
    FHIR[FHIR APIs]
    mCODE[mCODE profiles]
    USCDI[USCDI+ Cancer elements]
  end

  subgraph Measure_Engine[Measure computation]
    dQM[dQM specifications\n(CQL/code packages)]
    Risk[Risk adjustment + stratification\n(case-mix, equity)]
    Audit[Validation + audit\n(semantic testing)]
  end

  subgraph Use_Cases[Use cases]
    CQI[Practice CQI + workflow triggers]
    Payment[Payment incentives + benchmarking]
    Public[Transparency/reporting\n(patient-facing summaries)]
    Research[Learning system + research]
  end

  EHR --> FHIR --> dQM
  PRO --> FHIR --> dQM
  Claims --> dQM
  Registry --> dQM
  SDOH --> FHIR --> dQM
  mCODE --> FHIR
  USCDI --> FHIR
  dQM --> Risk --> Use_Cases
  dQM --> Audit --> Use_Cases
  Risk --> CQI
  Risk --> Payment
  Risk --> Public
  Risk --> Research

Pragmatic staged roadmap and panel planning aids

Staged roadmap with win-conditions and stakeholders

Stage A (now–2027): Minimum viable VBCC measurement set becomes operational.
Win-conditions: (1) ePRO completion and triage response is reliable; (2) HRSN screening/referrals are captured; (3) core oncology facts are structured enough to compute at least two clinical-contextual measures (timeliness; regimen appropriateness). [4–6], [16]
Key stakeholders: oncology practices (especially community), EHR vendors, ePRO vendors, CMS/CMMI model teams.

Stage B (2028–2029): Digital computability and cross-payer alignment.
Win-conditions: (1) core VBCC measures are computed as dQMs from multi-source data feeds; (2) measure specs are aligned for at least a “core 8–12” across multiple payers; (3) semantic validation and audits are routine; (4) equity stratification is standard. [10–12], [13], [24]
Key stakeholders: CMS + commercial payers, NCQA/measure developers, ONC/ASTP, NCI/USCDI+ Cancer, HL7.

Stage C (2030): Credible outcome benchmarking with lower burden.
Win-conditions: (1) validated benchmarking on symptom burden/function and EOL quality is feasible; (2) manual abstraction is the exception, not the norm; (3) improvement cycles show measurable gains. The EOM endpoint (June 2030) is a natural forcing function to assess whether these win-conditions have been met. [4], [30–31]

Provocative decision-point questions for the panel

1.       What is the “minimum viable VBCC” measure set (8–12 measures) we will commit to—and which legacy measures should we retire? [11–12]

2.       Should ePRO-based measures be “process-accountable” (completion and response) first, then “outcome-accountable” (improved symptom burden/function) later—or should we jump directly to outcome accountability? [6–9]

3.       Which oncology clinical facts must be standardized first (stage, key biomarkers, line of therapy, recurrence), and who will pay for the workflow change—the payer, the EHR vendor, or the practice? [13–16]

4.       How do we prevent pathway concordance measures from becoming proprietary “black boxes” and exception-driven gaming? [25], [10–12]

5.       What is the right fairness model: do we adjust for social risk, stratify without adjustment, or use both with guardrails—and how do we avoid penalizing safety-net practices? [4–6], [13]

6.       Is timeliness a VBCC quality signal, an access signal, or both—and what data standard is required so it’s not just an EHR timestamp artifact? [13–16]

7.       What should be the “audit trigger” set for gaming or selection (e.g., abrupt shifts in exception rates or patient mix), and who should run audits? [24]

8.       If CMS/NCQA are driving a 2030 digital measurement horizon, what must happen by 2027 to avoid a ‘digital facade’ that is computable but not meaningful? [10–12], [31]

Suggested panelist types

A high-yield panel typically needs: a CMS/CMMI model leader (OCM→EOM lessons), a community oncology practice leader implementing ePRO + navigation, an ONC/ASTP interoperability or USCDI+ Cancer representative, an EHR vendor or FHIR/mCODE implementer, a payer quality lead familiar with dQMs and contracting, a measurement scientist (NCQA/NQF), and a patient advocate focused on symptoms/function/financial toxicity.

Evidence-backed talking points

·       OCM demonstrates the “care redesign without measurable outcome gain” dilemma: substantial practice effort did not translate to significant improvements on accountable quality measures versus comparison groups. [2–3]

·       EOM is a policy pivot toward patient-centered measurement: required redesign activities explicitly include ePRO collection/monitoring, HRSN screening, CEHRT, and CQI data use. [4–6]

·       ePRO symptom monitoring has unusually strong clinical evidence for a measurement domain: multiple trials show improvements in symptom burden, quality of life, acute care utilization, and in some studies survival—supporting symptom/function measurement as a VBCC cornerstone. [7–9]

·       Digital measurement is not optional; it is the burden-reduction strategy: CMS’s dQM definition and Meaningful Measures 2.0 explicitly frame digital, interoperable data and workflow-embedded PRO-PMs as the path away from manual abstraction. [10–11], [24]

·       Oncology needs a common data substrate: USCDI+ Cancer and mCODE are the clearest pathway to standardizing core oncology facts necessary for risk adjustment and meaningful comparisons. [13–16], [32]

·       Equity must be built into measurement design, not appended: EOM requires HRSN screening and equity-oriented activities; without stratification and fairness guardrails, VBCC incentives risk deepening disparities. [4–6], [13]

·       Financial toxicity is a real outcome domain with validated instruments: COST/FACIT‑COST is validated and can be operationalized as a VBCC measure paired with navigation response and stratification. [18]

Numbered bibliography

1.       Porter ME. What Is Value in Health Care? (journal article). N Engl J Med. 2010. [1]

2.       CMS / Abt Global. Evaluation of the Oncology Care Model: Final Report—Executive Summary (report, PDF). May 2024. [2]

3.       CMS / Abt Global. Oncology Care Model—Final Evaluation At-a-Glance (report, PDF). 2024. [3]

4.       CMS (CMMI). Enhancing Oncology Model (EOM) overview and model timeline (web page). Accessed 2026. [4]

5.       CMS. Update: Enhancing Oncology Model Factsheet (web page). Jun 27, 2023. [5]

6.       CMS. EOM Electronic Patient-Reported Outcomes Guide (report, PDF). Jun 2023. [6]

7.       Basch E, et al. Symptom Monitoring With Patient-Reported Outcomes During Routine Cancer Treatment: A Randomized Controlled Trial (journal article). J Clin Oncol. 2016. [7]

8.       Basch E, et al. Overall Survival Results of a Trial Assessing Patient-Reported Outcomes for Symptom Monitoring During Routine Cancer Treatment (journal article). JAMA. 2017. [8]

9.       Basch E, et al. Effect of Electronic Symptom Monitoring on Patient-Reported Outcomes Among Patients With Metastatic Cancer: A Randomized Clinical Trial (journal article). 2022. [9]

10.  CMS / eCQI Resource Center. Digital Quality Measures—About dQMs (CMS definition) (web page). 2026. [10]

11.  CMS. Meaningful Measures 2.0 (web page). Updated 2026. [11]

12.  CMS. The Universal Foundation (web page). 2025. [12]

13.  ONC / ASTP. USCDI+ (including USCDI+ Cancer program description) (web page). Dec 2023. [13]

14.  NCI (CBIIT). Real-World Data program and USCDI+ Cancer partnership (web page). Sep 2025. [14]

15.  HL7 International. mCODE (Minimal Common Oncology Data Elements) Implementation Guide (web page). Accessed 2026. [15]

16.  CMS. EOM Clinical Data Elements Guide (and mapping to HL7 FHIR API/mCODE) (report, PDF). Nov 2025. [16]

17.  ICHOM. Colorectal Cancer Standard Set and Reference Guide (web page + PDF; outcomes + case-mix variables) (guideline/resource). 2017. [17]

18.  de Souza JA, et al. Measuring financial toxicity as a clinically relevant patient-reported outcome: Validation of the COST measure (journal article). Cancer. 2017. [18]

19.  Value-Based Cancer Care. Metrics to Keep in Mind for Value-Based Cancer Care (web page). Aug 2024. [19]

20.  Jacobs DB, et al. Aligning Quality Measures across CMS—The Universal Foundation (journal article). N Engl J Med. 2023. [20]

21.  Gondi S, et al. Assessment of Patient Attribution to Care From Medical Oncologists (journal article). JAMA Network Open. 2021. [21]

22.  Kline RM, et al. Design Challenges of an Episode-Based Payment Model in Oncology (journal article). 2017. [22]

23.  Holmgren AJ, et al. National trends in oncology specialists’ EHR inbox work, 2019–2022 (journal article). J Natl Cancer Inst. 2025. [23]

24.  CMS / eCQI Resource Center. Semantic validation (definition and testing concept for eCQMs/dQMs) (web page). Updated 2025. [24]

25.  NCCN. Development and Update of Guidelines (evidence-based guideline process) (web page). Accessed 2026. [25]

26.  CMS. Oncology Care Model Fact Sheet (web page). 2016. [26]

27.  Balch A, et al. Patient perspectives on cost and quality measures in value-based cancer care models (journal article). 2026. [27]

28.  ONC / ASTP. ONC’s Cures Act Final Rule overview (web page). Accessed 2026. [28]

29.  Federal Register. 21st Century Cures Act: Interoperability, Information Blocking, and the ONC Health IT Certification Program (web page). May 1, 2020. [29]

30.  CMS. EOM Quality, Health Equity, and Clinical Data Strategy (timeline through 2030) (report, PDF). Aug 2024. [30]

31.  NCQA. Why Digital Quality (CMS goal of transitioning to all digital measures by 2030; Universal Foundation alignment) (web page). Accessed 2026. [31]

32.  Osterman TJ, et al. Improving Cancer Data Interoperability: The Promise of mCODE (journal article). 2020. [32]

33.  Li Y, et al. mCODE Genomics pilot / proof-of-concept and limitations (journal article). 2024. [33]


[1] https://www.nejm.org/doi/full/10.1056/NEJMp1011024

https://www.nejm.org/doi/full/10.1056/NEJMp1011024

[2] Evaluation of the Oncology Care Model: Final Report

https://www.cms.gov/priorities/innovation/data-and-reports/2024/ocm-final-eval-report-2024-exec-sum?utm_source=chatgpt.com

[3] Oncology Care Model (OCM) - Final Evaluation

https://www.cms.gov/priorities/innovation/data-and-reports/2024/ocm-final-eval-report-2024-aag?utm_source=chatgpt.com

[4] EOM (Enhancing Oncology Model)

https://www.cms.gov/priorities/innovation/innovation-models/eom?utm_source=chatgpt.com

[5] Update: Enhancing Oncology Model Factsheet

https://www.cms.gov/newsroom/fact-sheets/update-enhancing-oncology-model-factsheet?utm_source=chatgpt.com

[6] EOM Electronic Patient-Reported Outcomes Guide

https://www.cms.gov/priorities/innovation/media/document/eom-elec-pat-rpt-outcomes?utm_source=chatgpt.com

[7] Symptom Monitoring With Patient-Reported Outcomes During ...

https://pubmed.ncbi.nlm.nih.gov/26644527/?utm_source=chatgpt.com

[8] Overall Survival Results of a Trial Assessing Patient ...

https://pubmed.ncbi.nlm.nih.gov/28586821/?utm_source=chatgpt.com

[9] Effect of Electronic Symptom Monitoring on Patient-Reported ...

https://pmc.ncbi.nlm.nih.gov/articles/PMC9168923/?utm_source=chatgpt.com

[10] Digital Quality Measures - About dQMs | eCQI Resource Center

https://ecqi.healthit.gov/dqm?utm_source=chatgpt.com

[11] https://www.cms.gov/medicare/quality/cms-national-quality-strategy/meaningful-measures-2-0

https://www.cms.gov/medicare/quality/cms-national-quality-strategy/meaningful-measures-2-0

[12] https://www.cms.gov/medicare/quality/cms-national-quality-strategy/universal-foundation

https://www.cms.gov/medicare/quality/cms-national-quality-strategy/universal-foundation

[13] https://healthit.gov/standards-and-technology/uscdi-plus/

https://healthit.gov/standards-and-technology/uscdi-plus/

[14] https://www.cancer.gov/about-nci/organization/cbiit/projects/real-world-data

https://www.cancer.gov/about-nci/organization/cbiit/projects/real-world-data

[15] https://build.fhir.org/ig/HL7/fhir-mCODE-ig/

https://build.fhir.org/ig/HL7/fhir-mCODE-ig/

[16] https://www.cms.gov/priorities/innovation/media/document/eom-clinical-data-elements-guide

https://www.cms.gov/priorities/innovation/media/document/eom-clinical-data-elements-guide

[17] https://www.ichom.org/patient-centered-outcome-measure/colorectal-cancer/

https://www.ichom.org/patient-centered-outcome-measure/colorectal-cancer/

[18] https://pmc.ncbi.nlm.nih.gov/articles/PMC5298039/

https://pmc.ncbi.nlm.nih.gov/articles/PMC5298039/

[19] https://www.valuebasedcancer.com/issues/2024/august-2024-vol-15-no-1/metrics-to-keep-in-mind-for-value-based-cancer-care

https://www.valuebasedcancer.com/issues/2024/august-2024-vol-15-no-1/metrics-to-keep-in-mind-for-value-based-cancer-care

[20] https://www.nejm.org/doi/full/10.1056/NEJMp2215539

https://www.nejm.org/doi/full/10.1056/NEJMp2215539

[21] https://jamanetwork.com/journals/jamanetworkopen/fullarticle/2779755

https://jamanetwork.com/journals/jamanetworkopen/fullarticle/2779755

[22] https://pmc.ncbi.nlm.nih.gov/articles/PMC5508445/

https://pmc.ncbi.nlm.nih.gov/articles/PMC5508445/

[23] https://academic.oup.com/jnci/article/117/6/1253/8051594

https://academic.oup.com/jnci/article/117/6/1253/8051594

[24] https://ecqi.healthit.gov/glossary/semantic-validation

https://ecqi.healthit.gov/glossary/semantic-validation

[25] https://www.nccn.org/guidelines/guidelines-process/development-and-update-of-guidelines

https://www.nccn.org/guidelines/guidelines-process/development-and-update-of-guidelines

[26] https://www.cms.gov/newsroom/fact-sheets/oncology-care-model

https://www.cms.gov/newsroom/fact-sheets/oncology-care-model

[27] https://pmc.ncbi.nlm.nih.gov/articles/PMC12967066/

https://pmc.ncbi.nlm.nih.gov/articles/PMC12967066/

[28] ONC's Cures Act Final Rule

https://healthit.gov/regulations/cures-act-final-rule/?utm_source=chatgpt.com

[29] 21st Century Cures Act: Interoperability, Information ...

https://www.federalregister.gov/documents/2020/05/01/2020-07419/21st-century-cures-act-interoperability-information-blocking-and-the-onc-health-it-certification?utm_source=chatgpt.com

[30] https://www.cms.gov/files/document/eom-qual-health-equity-clin-data-strat.pdf

https://www.cms.gov/files/document/eom-qual-health-equity-clin-data-strat.pdf

[31] https://www.ncqa.org/digital-quality-transition/why-digital-quality/

https://www.ncqa.org/digital-quality-transition/why-digital-quality/

[32] https://pmc.ncbi.nlm.nih.gov/articles/PMC7713551/

https://pmc.ncbi.nlm.nih.gov/articles/PMC7713551/

[33] https://pubmed.ncbi.nlm.nih.gov/38935887/

https://pubmed.ncbi.nlm.nih.gov/38935887/


Tuesday, March 31, 2026

Appendix S: Two Comparisons (2025/2026 and Feb to May 2026)

 

2025 to 0326 summary (about 150 words):
The March 26, 2026 revision of Appendix S is not just an edit of the 2025 version; it is a substantial effort to turn Appendix S from a simple taxonomy into a more operational CPT policy framework for software-intensive services. The 2025 version mainly defined assistive, augmentative, and autonomous services at a high level. The 0326 version keeps those categories but adds much more about software outputs, reference services in current clinical practice, and the types of evidence needed to justify each category. It narrows assistive by warning that terms like “risk for” or “suggestive of” may require clinical validation. It raises the threshold for augmentative by demanding outputs that are not merely statistical but clinically meaningful, clinically important, and pertinent to the CPT descriptor. It also tightens autonomous claims by emphasizing transparency, guidelines, and clinical utility, suggesting the drafters want stricter boundaries and stronger evidentiary discipline.

0204 to 0326 summary (about 150 words):
The March 26, 2026 version is best seen as a tightening and sharpening of the February 4, 2026 draft rather than a wholesale rewrite. By February, Appendix S had already begun evolving beyond a simple AI taxonomy toward a framework about software outputs, evidence, and coding boundaries. The March draft pushes this further. It drops more of the device-oriented/FDA-style language and speaks more clearly in CPT terms, focusing on software outputs and their relationship to a reference service in current clinical practice. It more carefully restricts assistive status, especially for outputs using predictive language like “likelihood of” or “risk for.” It makes augmentative more demanding by tying clinical meaningfulness directly to the CPT code characteristics. It also removes February language that gave Category III applicants a more permissive developmental pathway. Overall, March appears more conservative, more evidence-calibrated, and more focused on preventing applicants from overclaiming sophistication or autonomy.



PUBLISHED VERSION VERSUS MARCH 2026 BALLOTT

What jumps out first is that the March 26 ballot draft is not a light cleanup of the 2025 Appendix S. It is an attempt to re-found the appendix on a more operational theory of software services. The 2025 text was short, elegant, and high-level: assistive detects data, augmentative analyzes/quantifies to yield clinically meaningful output, and autonomous interprets and independently generates clinically meaningful conclusions, with three escalating levels of action. The new draft keeps that skeleton, but wraps it in a much thicker framework about software outputs, evidentiary standards, clinical meaningfulness, equivalence to reference services, and CPT code criteria. In other words, the authors seem to be trying to turn Appendix S from a taxonomy into something closer to a gatekeeping policy for future code applications.

The single biggest conceptual change is the move away from talking mainly about “AI” and “work done by machines” toward talking about software, software outputs, and the service performed by the software to produce the desired outputs. The 2025 text framed the issue as classification of “AI medical services and procedures” based on the work performed by the machine on behalf of the physician or QHP. The ballot draft repeatedly substitutes the language of software outputs and adds the statement that in CPT, software outputs are recognized as useful in diagnosis, cure, mitigation, treatment, or prevention, and that their clinical meaningfulness is established through clinical evidence. That change looks intentional and strategic. It deemphasizes the fashionable word AI and shifts the debate toward the thing CPT actually codes: a clinical service with a defined output and evidentiary basis. It also makes Appendix S more future-proof, because the policy can govern software-intensive services whether or not applicants call them AI.

A second major change is that the draft inserts an explicit reference-service concept. The new text says classification is based on equivalence to a primary or usual reference service in current clinical practice and that use of an Appendix S term should be supported by evidence of analytical validity, clinical validity, or clinical utility, as appropriate to the choice of term in the code descriptor and consistent with CPT code criteria. That is a very important move. It suggests the authors are trying to prevent Appendix S from becoming a free-floating vocabulary for novel products. Instead, they want it anchored to something familiar in CPT logic: what is the usual clinical service, what is this software doing relative to it, and what level of evidence matches that claim? Apparent intention: make it harder for applicants to jump too quickly from technical performance to broad claims of augmentative or autonomous status.

The assistive section is probably the clearest example of tightening. In 2025, assistive was simply software that detects clinically relevant data without analysis or generated conclusions, and it required physician interpretation and report. The draft keeps that bottom-line concept but elaborates it substantially. It now says assistive software may draw attention to clinically relevant data without deriving a new parameter, generating an interpretation, or providing conclusions; it clarifies that parameters include indexes, scores, or classifications; and it adds an evidence discussion saying that assistive outputs are clinically supportive if they improve physician/QHP performance, with benefit to the patient substantiated by technical or analytical validation. But then it introduces an important caveat: if the output uses language such as “likelihood of,” “suggestive of,” or “risk for,” then those terms should be substantiated by clinical validation rather than mere analytical validation. That is a significant policy move. It looks like the drafters are worried that applicants have been trying to package quasi-interpretive risk language as if it were only assistive triage. The message seems to be: you may stay assistive only if you truly remain non-interpretive; once you start implying clinical inference, your evidentiary burden rises.

The augmentative rewrite is even more consequential. The 2025 definition was compact: the machine analyzes and/or quantifies data to yield clinically meaningful output; physician interpretation/report remains required. The new draft expands this into a mini-doctrine. It says the output must be a quantitative or categorical parameter qualitatively different from the input, and not merely descriptive statistics such as adding or averaging. It says augmentative output does not include a definitive interpretation or conclusion, which is reserved for autonomous. It says augmentative outputs may be reported as clinical scales, indexes, categorical classifications, or other metrics in common clinical use, or may be novel predictive/prognostic indices validated for impact on patient care. It then defines “clinically meaningful” through a three-part test: the output must be clinically validated, clinically important, and directly pertinent to the CPT code characteristics. This is much more than clarification. It is a deliberate elevation of the threshold for what counts as augmentative. The apparent intention is to stop applicants from saying, in effect, “our software produces a score, therefore it is augmentative.” Under this markup, a score alone is not enough; it must be meaningful in clinical practice, not just mathematically nontrivial.

That augmentative section also contains an underappreciated policy signal: it says software may require physician or QHP interaction during the process between input and output, for example adjusting settings based on clinical context, and then notes in a footnote that physician work related to augmentative services may already be captured by existing codes, such as E/M or presurgical planning. That looks like a direct attempt to separate two questions that otherwise get tangled: what category is the software output, and where is any physician work paid. The drafters seem to be saying that Appendix S should classify the software piece, but not automatically create new physician-work value around it. For future CPT applicants, this is potentially quite important: even if a service is accepted as augmentative, that does not mean CPT will agree there is separately payable physician work embedded in the same code.

The autonomous section is also being disciplined, though less radically than augmentative. The 2025 text already defined autonomous as software that automatically interprets data and independently generates clinically meaningful conclusions without concurrent physician/QHP involvement, and then divided autonomy into three levels. The new draft keeps the levels, but it tightens the entrance criteria. It now says autonomous software derives parameters similarly to augmentative outputs and independently generates clinically meaningful interpretations or conclusions in accordance with clinical scales/metrics in common use, clinical practice guidelines, or direct demonstration of impact on patient care. It then adds that reporting of derived parameters is essential for transparency and explicability, and states that recommendations for definitive diagnosis, specific management, or interventions should be validated for clinical utility. This is an unmistakable attempt to constrain bold autonomous claims. It implies that if software is going to generate conclusions or management recommendations, the bar is not merely analytical performance or even clinical association; it trends toward utility, guideline anchoring, and explainable transparency.

The draft’s treatment of the three autonomous levels is revealing. The levels themselves are broadly familiar: Level I recommends and requires physician action to implement; Level II initiates action with alert/opportunity to negate; Level III continues unless the physician intervenes. But the draft rewrites the prefatory language to emphasize outputs that include recommendations of definitive diagnosis and/or specific management, or medical management actions, or automatically initiated management actions. That phrasing makes the levels feel less like abstract AI maturity and more like a graded ladder of clinical authority and workflow control. The likely intention is to tie autonomy to what the software actually gets to do in the clinical workflow, not just to the sophistication of the model. That should make debates at CPT more practical: not “how smart is it?” but “what exactly does it conclude, recommend, initiate, and who has to stop it?”

The new ballot also adds a new summary table on primary objective, required evidence of clinical meaningfulness, and whether physician/QHP interpretation/report is required. This table is especially important because it exposes the authors’ true architecture. In the 2025 appendix, the summary table was descriptive. In the new draft, the new table becomes almost quasi-regulatory. It explicitly says assistive does not require evidence of clinical meaningfulness in the same way augmentative and autonomous do, although it still requires evidence of benefit to patient care; by contrast, augmentative and autonomous do require evidence of clinical meaningfulness. That is a new and consequential distinction. It formalizes a step-up in evidentiary burden across the taxonomy. I suspect the authors added this because earlier versions of Appendix S did not give enough practical help when panelists asked, “What kind of evidence is enough for each label?”

Stepping back, the draft seems to pursue at least five apparent intentions.

First, to make Appendix S more usable for actual CPT decision-making by connecting taxonomy terms to evidence standards. The old text classified. The new text classifies and tells you what sort of validation must back the classification.

Second, to shift focus from the hype term AI to the more durable concept of software outputs in clinical services. That makes the appendix harder to game with branding and easier to apply across AI, algorithms, rules engines, and software-intensive services generally.

Third, to draw a firmer line between assistive and augmentative by blocking the quiet smuggling of predictive or inferential language into assistive territory. The new “likelihood of / risk for / suggestive of” language is almost certainly there for that reason.

Fourth, to prevent weakly justified scoring systems from claiming augmentative status unless they are truly clinically meaningful, clinically important, and pertinent to the coded service. That feels aimed at the many software products that can generate a score but have a shakier claim to real-world medical relevance.

Fifth, to cabin autonomous claims by demanding more explicit linkage to guidelines, patient-care impact, transparency, and in some cases clinical utility. That makes autonomous feel less like a prestige label and more like a serious claim that must be earned.

As for impact, I think the draft, if adopted in something like this form, would make life harder for applicants seeking ambitious software codes, but easier for the CPT Panel and staff who need a principled vocabulary for saying yes, no, or not yet. It will likely favor applicants whose services resemble existing reference services, whose outputs are well-specified, and whose evidence packages are aligned to the exact claim made in the descriptor. It will be less friendly to applicants who rely on broad claims of “AI-enabled” improvement, opaque risk outputs, or arguments that a score is self-evidently meaningful because it is statistically significant. It may also reduce category inflation: fewer things called autonomous, some things pushed down from augmentative to assistive, and many proposed descriptors forced to become more precise.

There are, however, a few tensions in the markup. One is that the document is trying to do two jobs at once: define taxonomy and articulate evidence policy. That may be exactly what the drafters want, but it also risks making Appendix S more contestable, because parties may disagree not only about category assignment but about the embedded evidentiary doctrines. Another is that phrases like clinical equivalence to current clinical practice, clinically important, and directly pertinent are sensible but still somewhat elastic. Those phrases invite interpretation, which means future debates will move from “what category is this?” to “what counts as clinically important enough?” That may be unavoidable, but it means the revised Appendix S could generate more argument, not less, even as it becomes more rigorous.

My overall read is that the markup authors are trying to make Appendix S do what many coding policies eventually do after a few years in the wild: move from a clean conceptual taxonomy to a workable adjudication framework. The 2025 version was elegant and easy to summarize. The March 26 version is denser because it is trying to answer the hard questions the 2025 version left open: What exactly is the software output? How different is it from the input? Is it only drawing attention, or deriving a parameter, or making a conclusion? What evidence is needed for that level of claim? How does it relate to existing clinical practice? Who still interprets it? Who acts? Who can override it? That is why the markup feels so heavily “lawyered” and policy-laden. The authors appear to be building a document meant not just to describe AI/software services, but to control the boundary-setting battles that are now arriving at CPT.

If you want, I can turn this next into a blog-ready essay in your usual paragraph style, or a side-by-side change matrix with “old text / new text / probable purpose / coding impact.”



##

##

Feb 2026 versus March 2026 Ballott

###

Yes. Comparing the February 4 draft with the March 26 draft, the March text again looks like a document that is trying to move Appendix S away from being merely a taxonomy of AI types and toward being a more operational CPT policy for software outputs, evidence, and coding boundaries. But this time the key point is slightly different from the 2025-to-0326 comparison: the February draft had already moved in that direction. So the March draft is not a revolution from February. It is better understood as a further tightening, sharpening, and legalizing of ideas that were already present on February 4.

The broadest change is at the level of framing and ontology. The February 4 version still spoke in the language of a “software device” and explicitly defined software in FDA-like terms as software intended for diagnosis, cure, mitigation, treatment, or prevention of disease. It also said that software integral to a total service or functioning only as support was not separately codifiable unless distinct as a service with beginning, middle, and end. That language felt like an effort to define what kind of thing CPT was even talking about: a software device, perhaps regulatable, perhaps separately codable only under certain structural conditions. By March 26, that device-oriented language has been stripped back and replaced with a cleaner emphasis on software output(s) and on the service performed by the software to produce desired outputs on behalf of the physician or QHP. The March draft also adds the idea that classification is based on equivalence to a primary or usual reference service in current clinical practice. So the apparent intention was to shift from a somewhat product-centered formulation to a service-and-output-centered formulation more native to CPT. That is an important conceptual refinement. February still sounded partly like a hybrid of CPT and device-regulatory language; March sounds much more like CPT trying to speak in its own voice.

Related to that, March is markedly more explicit about evidence standards tied to terminology choice. The February version said clinical evidence should demonstrate that the output of the software device benefits patient care, and then, in its “clinically meaningful output” section, said such output must be clinically validated beyond technical or analytical validation, clinically important beyond statistical significance, and directly relevant to intended use of the CPT code. March preserves that structure but tightens it by saying that, to use a term from Appendix S, evidence should demonstrate analytical validity, clinical validity, or clinical utility, as appropriate to the choice of term in the code descriptor and consistent with CPT code criteria. That is more disciplined and more tactical. It appears designed to align Appendix S terminology directly with the evidentiary threshold implied by the descriptor claim. The likely purpose is to prevent overclaiming: an applicant should not be able to select a stronger Appendix S term than its evidence package can justify.

The treatment of assistive is also more carefully delimited in March. In February, assistive meant drawing attention to clinically relevant data without deriving a new parameter, making an interpretation, or providing conclusions, and required physician/QHP interpretation and report. It also said that an assistive device improves physician/QHP performance and that improvement should be substantiated by clinical evidence. March keeps the same basic idea, but it adds several guardrails. It now explicitly defines parameters as quantitative or categorical outputs such as an index, score, or classification. It states that assistive outputs are clinically supportive because they improve physician/QHP performance, and that the improvement should be a patient benefit substantiated by technical or analytical validation where the primary service output is unchanged. But it then adds a very important qualification: if the output uses language such as “likelihood of,” “suggestive of,” or “risk for,” those terms should be substantiated by clinical validation. This is one of the clearest March-over-February moves. February’s assistive language still left room for products to flirt with predictive or inferential terminology while claiming low-level assistive status. March closes that gap. The apparent intention is to stop the semantic creep by which quasi-interpretive outputs masquerade as simple detection aids.

In augmentative, the March draft becomes more exacting and more practical than February. February said augmentative output derives a quantitative or categorical parameter qualitatively different from the input, that it must be more than descriptive statistics, and that it does not include an interpretation or conclusion. It also allowed expression through common clinical scales or other metrics, or validated predictive/prognostic indices. That was already fairly strong. But March goes a step farther in several ways. First, it keeps the requirement that output be qualitatively different and more than mere summation, adding more emphatic language that it must provide something beyond adding, averaging, or otherwise reporting descriptive statistics. Second, March explicitly says that the designation of software output as augmentative is based on demonstration that it is clinically meaningful and distinct from the input function. Third, it refines the three-part test for clinical meaningfulness by tying the output directly to the CPT code characteristics, including typical patient, procedure description, and descriptor. This makes the test more CPT-specific and less abstract than February’s intended-use wording.

Another notable deletion is that the February draft included special language for Category III codes, stating that use of “augmentative” could be substantiated by the design of the product and the design of ongoing clinical trials intended to yield Category I-level validation. A parallel Category III accommodation also appeared in autonomous. That language is absent from the March 26 text. I think that is one of the most consequential edits in the whole comparison. It suggests that between February and March, the drafters pulled back from giving what might have looked like a special evidentiary lane for emerging technologies. The likely reason is concern that such language could be read as an invitation to claim a higher Appendix S category based on future evidence plans rather than current evidence. Removing it makes the document more conservative and more immediate: classification should reflect what the software output can substantiate now, not what trials may later show. That deletion is highly consistent with a general March pattern of tightening access to stronger labels.

March also reframes the question of physician work more pointedly. February said augmentative output may involve non-traditional physician/QHP interaction and that most augmentative services do not involve traditional interpretive physician work; the output may serve as a data element in E/M, a factor in surgical planning, or input to another interpretive service. March retains this logic but makes it even more explicit that physician work related to augmentative services may already be captured by existing codes. This is a subtle but important hardening. It sounds like the authors want to ensure that Appendix S does not become a Trojan horse for arguments that every sophisticated software output should carry newly recognized physician work or stand-alone reimbursement logic. The intention seems to be to separate the classification of the software service from the valuation of physician effort, keeping both questions analytically distinct.

The March version of autonomous is likewise a refinement rather than a full rewrite, but it is a meaningful refinement. February defined autonomous as automatic derivation of parameters and independent generation of interpretations or conclusions in accordance with clinical scales, practice guidelines, or direct demonstration of impact on patient care. It also required reporting of derived parameters for oversight, transparency, and explicability, and said recommendations for diagnoses or interventions should be based on parameters and reported within the context of epidemiologic data, practice guidelines, or evidence for clinical utility. March keeps much of that structure but makes some of the language more pointed. It says recommendations for definitive diagnostic conclusions, specific management, or interventions should be validated for clinical utility, and notes that these standards are especially important for Levels II and III. The drift is toward a more explicit gradient of evidentiary seriousness as the software’s practical authority increases. February hinted at that when it said higher autonomy levels require higher evidentiary standards due to increasing patient risk. March operationalizes the point more concretely within the main autonomous text.

The treatment of the three autonomous levels is also revealing. February described them in cleaner prose: Level I provides recommendations requiring physician/QHP judgment; Level II initiates management actions with compulsory alert and opportunity for override; Level III automatically initiates management unless or until physician/QHP action reverses it. March keeps the same staircase but rewrites the wording to emphasize outputs that include recommendations of definitive diagnosis and/or specific management, then outputs that include medical management actions, and finally outputs that automatically initiate management actions that continue unless the physician intervenes. This makes the levels feel slightly less like design categories and more like an escalating sequence of clinical control and workflow consequence. The likely intention is to help the Panel judge not just sophistication, but how much authority the software is exercising in the patient-care chain.

One of the most striking March additions is the new summary table distinguishing primary objective, required evidence of clinical meaningfulness, and whether physician/QHP interpretation and report is required. February’s summary table was more conventional: primary objective, independent diagnosis/management, analyzes data, requires interpretation, evidence of patient benefit required. March’s table is more doctrinal. It distinguishes assistive from augmentative/autonomous by stating that assistive does not require evidence of clinical meaningfulness in the same way, though it does require evidence of benefit to patient care, while augmentative and autonomous do require evidence of clinical meaningfulness. It also gives examples of language associated with each category, such as assistive outputs drawing attention to data and even including terms like “likelihood of” or “risk for,” while augmentative includes outputs “predictive of” or “prognostic of.” That table does a lot of work. It is almost a cheat-sheet for future coding debates. Compared with February, it shows the authors trying to convert Appendix S into something more adjudicative and scalable.

So what are the apparent intentions behind the March revisions relative to February?

First, to make Appendix S sound less like a statement about software products/devices and more like a statement about software services and outputs in CPT terms. February still carried some residual FDA/device flavor; March is more squarely CPT.

Second, to narrow the assistive lane and stop inferential or risk-bearing outputs from receiving low-level categorization without stronger validation. The “likelihood of / suggestive of / risk for” language is central here.

Third, to make augmentative a higher and more disciplined threshold, not merely “any score or classification,” but a clinically meaningful parameter linked tightly to the descriptor, typical patient, and procedure.

Fourth, to eliminate what may have been perceived in February as a too-generous Category III glide path. By deleting those passages, March seems to reject the idea that design and planned trials are enough to justify stronger Appendix S terminology.

Fifth, to align increasing autonomy with increasing demands for transparency, guideline anchoring, and clinical utility, especially as the software moves from informing decisions to initiating actions.

In terms of practical impact, I think March will make Appendix S more useful to CPT leadership and staff, but more demanding for applicants. February already set out a serious framework. March turns the screws. It will likely reduce the room for applicants to rely on broad AI rhetoric, product design, or future-study arguments. It will favor services with crisp output definitions, solid present-tense evidence, and a clear relationship to existing clinical practice. It will also likely make coding debates more explicit around the questions the Panel actually cares about: what is the output, how different is it from the input, what exactly is the clinical claim, what evidence supports that claim, and where does physician work sit, if anywhere?

My overall read is that the February version was a serious and already mature draft, but it still carried traces of an effort to accommodate innovation by describing device structure, codifiability, and Category III developmental pathways. The March version is noticeably more guarded, CPT-native, evidence-calibrated, and boundary-conscious. It is not trying to be friendlier to applicants. It is trying to give the Panel a sturdier vocabulary for saying, with more confidence and less ambiguity, this output is assistive, this one is augmentative, this one is autonomous, and this is the level of proof needed for each claim.

If you’d like, I can next turn this into a side-by-side change matrix with columns for 0204 text, 0326 text, likely reason for revision, and likely impact at the April 30 meeting.

Monday, March 30, 2026

BQ CRuSH COMMENT

 I have two comments relative to Genomics initiatives proposed within CRUSH.  The first involves nationalization of the MOLDX program.   The second involves stopping fraudulent genomics payments while not adding pointless burdens to the legitimate laboratory industry.  Before making those two comments, I will first briefly summarize Medicare genomics fraud.


1) MEDICARE GENOMICS FRAUD

In 2014, the AMA CPT issued a new coding system for genomic tests, creating hundreds of codes, some valued >$1000, many with no relevance to Medicare-page patients.  Because the codes often had no relevance to Medicare, MACs issued no LCDs on them.  A key difference between recent DME fraud and Genomics fraud is that DME fraud involves services (equipment) that was potentially valid - like electric wheelchairs.  The genomics fraud generally involved services that were intrinsically nonsensical in the Medicare population, making Medicare's vulnerability more outrageous.   In the first fraud phrase, illicit labs mostly in TX and FL billing >$6000 per patient on absurd genetic tests, usually code 81408 (x2 = $2000 x 2) and adjacent codes.  TX and FL apparently had no edits on this and autopaid > $1B in a few years. See my March 10 blog, https://www.discoveriesinhealthpolicy.com/2026/03/the-crush-initiative-and-medicares-bone.html


OIG and others forced a crackdow on codes in the 81408 series by CY2023.  However, it is possible to find labs that billed tens of millions on fraudulent codes in CY2022 and merely switched seamlessly to other, nearby nonsensical costly codes in 2023 and 2024 (see blog).  Uncontrolled codes included 81419, 81440, 81443. Again, many hundreds of millions of dollars flowed out uncontroled (see blog for bar charts to 2024).   


The level of nonsense includes Medicare's stopgap medically unlikelyl edits which allowed some of these nonsensical codes to be paid in multiples of 2 (from 2018 to 2026 ongoing). If someone had merely set the unikely edit to N=1, half a billion dollars would have been saved for 5 minutes work.  


At least in genomics fraud, these problems went on for years and were absurd.  AI and PhD computer science was not needed, a ten year old with Excel could have found this.


COMMENT ONE - MOLDX

Had the MolDx program been in place in TX and FL over a billion, likely nearly 2 billion, would have been saved 2018-2024.  MolDx is not strictly an antifraud program, but involves close review of the laboratory and its molecular validation and quality procedures, which most fraudulent labs could not pass.  In addition, MolDx had extensive dedicated staff to tracking coding and expenditures and policies, and nonsensical claims would have been stopped in weeks or a couple months, not 5-8 years.  Even implementing MolDx as-is would be highly effective and adding modest anti-fraud detection controls would close off remaining problems.   


COMMENT TWO - NORMAL GENOMICS LABS

Legitimate genomic  testing has risen greatly in importance, especially in cancer, the #2 cause of death in the Medicare population.  The legitimate lab industry is vastly different than the fraudulent genomics industry.  Fraud can be stopped - cut 99% - by measures that do not impact the legitimate lab industry.  The lab industry sometimes suffers under document collection - if asked to provide reams of distant hospital and clinician paperwork as part of audits and prior auth, requests that are nearly impossible to comply with may show a so-called error rate (sic) that is artifactual (relative to valid tests for real cancer patients).   


###

Thank you! Your comment has been submitted to Regulations.gov for review by the the Centers For Medicare & Medicaid Services.

Comment Tracking Number: mnd-h702-jjm1



Thursday, March 26, 2026

CMS Cuts Non Timed Services 2.5% - Why

 What's Up - PFS RULE Nov 2025

I just read that most services that are not 'timed services' (e.g. "30 minute office visit") had a 2.5% technology or practice expense cut in January 2026.  I'm attaching the November 2025 PFS final rule for Cy2027 which probably has all the answers.  ????