The standard RFP framework for evaluating AI medical coding vendors is inadequate. Auditable accuracy, architectural transparency, operational depth, and the line between rigorous claims and marketing theater require a different set of questions than traditional software RFPs surface.
The RFP framework that worked for evaluating traditional revenue cycle vendors does not work for evaluating AI medical coding vendors. The questions are different, and the answers require different kinds of evidence. Buyers using a traditional RFP miss the substantive distinctions.
What questions should an AI coding RFP include?
A rigorous RFP for autonomous medical coding has to surface five categories of evidence that standard software RFPs were never designed to surface. The first is auditable accuracy methodology. The right question is not "what is your accuracy" but "how is your accuracy measured, by whom, against what comparison standard, on what chart sample, and at what specialty stratification." A single accuracy number with no methodology behind it is not comparable to another vendor's single number with no methodology behind it. Both are marketing.
The second is specialty footprint versus production claims. The KLAS 2025 Autonomous Coding report documents that current production deployments are concentrated in radiology and emergency department coding, where documentation is structured and code distributions are narrow. A vendor that quotes a high automation rate without disclosing which specialties that rate covers, and in what proportion of the buyer's actual encounter mix, is presenting a number that does not transfer to the buyer's operation. Ask for accuracy and automation by specialty, in the proportions the buyer actually produces. Not the vendor's most-tested slice.
The third is architectural transparency. Does the system train on patient data, and from which patients? Models trained against historical human-coded examples inherit the structural ceiling of human inter-coder reliability, established in Peng et al., 2018 at 82.2% agreement between credentialed coders at four-digit ICD-10 specificity. A system trained that way cannot reliably exceed it. Other architectural questions: where does inference compute run; how are ICD-10-CM, CPT, and HCPCS specification updates handled; how are AHA Coding Clinic and NCCI updates incorporated; how is multi-tenant data isolation implemented; does customer data flow into model training, even indirectly.
The fourth is operational depth. Production medical coding is a regulatory and operational discipline, not a software problem. Who at the vendor holds CCS, CPC, CCS-P, or equivalent credentials, and what authority do they hold over coding decisions and audit findings? How is the post-go-live audit cadence structured, and who runs it? How does the vendor handle denial follow-up, payer-specific rules, and ICD/CPT specification updates between releases? Vendors that grew out of pure software backgrounds frequently lack the operational depth that determines whether automation works in production. The fifth, in the wake of the operational depth question, is compliance posture: BAA execution, U.S. data residency, audit log accessibility, SOC 2 Type II reporting, HIPAA compliance, and the documented controls a covered entity will be audited against itself.
What does "auditable accuracy" look like in practice?
Auditable accuracy is the property of an accuracy claim that an outside party can verify. The methodology that supports it is the credentialed-panel audit: a representative sample of charts is independently coded by two or more credentialed coders, the AI system codes the same charts in parallel, and disagreements are adjudicated against the source clinical documentation and current coding guidance. The published outputs are the initial agreement rate (the share of charts where the AI matched the human coder before adjudication) and the post-adjudication consensus rate (the share where the AI matched the panel-adjudicated ground truth after disagreements were investigated). Both numbers matter. The post-adjudication rate is the system's actual error rate once human coder errors are removed from the comparison.
A vendor-produced audit establishes baseline credibility. A buyer-conducted pilot audit, on a representative sample of the buyer's own charts, is a stronger evaluation signal. The two are complementary. The vendor audit demonstrates the system's performance against a population the vendor has worked with previously; the buyer pilot demonstrates the system's performance against the buyer's documentation patterns, payer mix, and specialty distribution. Buyers should expect to do both, and should not treat reference-customer accuracy claims as a substitute for either. Reference customers typically deploy a narrow slice of their encounter mix at first; ask for accuracy data on the full encounter mix the buyer would deploy, not the slice the reference customer ran the pilot on.
A defensible audit publishes its methodology in enough detail that an outside auditor could replicate it. That means sample size and how it was drawn; specialty stratification of the sample; credentialing level of the human reviewers; the adjudication process and who held final authority over disagreements; and the time horizon. A 500-chart single-snapshot audit is informative. An 18-month longitudinal audit across tens of thousands of encounters by two independent organizations, on the highest-volume and the highest-complexity codes in the field, is a different category of evidence. Both can be done, but they cannot be presented as equivalent.
The AccuCode audit posture, as a concrete example of what published methodology looks like, is independent audit by MedAxiom (the American College of Cardiology's cardiovascular collaborative) on cardiovascular surgery and evaluation-and-management codes, plus independent audit by Professional Consulting Services (PCS) on E/M and broader specialty coverage. Both audits were led personally by the most senior credentialed coders at each organization. Initial agreement of 95.6% with the credentialed coding panel on cardiovascular surgery (the floor of the specialty audit, intentionally chosen because it is among the most complex specialties in the codebook) rises to 99.1% after panel adjudication. On the higher-volume E/M codes, initial agreement is 98.4% and post-adjudication consensus is 99.6%. The full methodology is published on the medical coding product page; what makes those numbers auditable is the methodology, not the numbers themselves.
How do you distinguish rigorous claims from marketing theater?
The patterns that signal marketing theater are recurring and recognizable. The most common is the automation-rate equivocation. A vendor reports "90% automation" without specifying what the 90% is a percentage of. The honest reading frequently turns out to be: 90% of charts within certain workflows that the vendor is configured to attempt, where those workflows represent perhaps 20% of the buyer's actual encounter volume. The net automation rate of the buyer's total volume is then not 90% but 18%. The buyer who deploys at the 90% expectation and discovers the 18% reality has bought a different product than they were sold.
A "90% automation rate" scoped to narrow workflows can be 18% of the buyer's actual encounter mix. Always multiply the claimed rate by the scope.
Net automation of total encounter volume = claimed automation rate × scope of the workflows the claim covers. Both factors are required to interpret the claim.
A related pattern is accuracy without methodology. A single accuracy number with no audit methodology behind it is not comparable to any other number. The remedy is to insist on the methodology before accepting the number: who conducted the audit, how was the sample drawn, what was the credentialing level of the human reviewers, how were disagreements adjudicated, and what comparison standard was used. A vendor unwilling or unable to supply that methodology has not, in any operational sense, measured what they are reporting. A vendor whose accuracy is measured "against human coders" without panel adjudication is structurally capped at the human inter-coder reliability ceiling of roughly 82% (Peng et al., 2018), which means the published number is comparing the system to an 82%-reliable standard rather than to source documentation.
Two further patterns deserve scrutiny. The specialty-coverage gap: a vendor claims to handle "all specialties" but the audit data covers radiology, ED, and perhaps one or two other narrow specialties at production accuracy. Per the KLAS 2025 segment report, this is the actual current state of most production deployments. The remedy is to ask for audited accuracy by specialty in the proportions the buyer actually produces, and to treat anything else as untested. The vendor-funded audit: an audit conducted by the vendor, by a consulting firm the vendor has paid, or by a customer with whom the vendor has an active commercial relationship is not equivalent to an audit conducted by an independent third party with no financial relationship. Both can be informative. They are not the same evidence.
What rigor looks like, in contrast, is straightforward to describe. Independent audit by named credentialed reviewers with their full credentials disclosed. Published methodology that an outside auditor could replicate. Specialty stratification of accuracy results across the full encounter mix. Longitudinal accuracy reporting over a meaningful time horizon, not a single snapshot. Compliance documentation that a covered entity will be audited against itself. And operational leadership at the vendor with credentialed coding authority over the system's behavior in production. Buyers who insist on those properties end up with a smaller short list and a higher hit rate on deployments that survive go-live.
Quick answers to the questions buyers ask most often about this topic.
What's the most important question to ask an AI coding vendor?
How is your accuracy measured, by whom, against what comparison standard, and on what specialty sample? Any vendor whose accuracy figure was not produced by independent credentialed-panel audit against source documentation, on a representative sample of the buyer's actual encounter mix, is quoting a number that is not directly comparable to other vendors' numbers and not actionable in a buying decision. The methodology is the operative variable, not the headline percentage.
Should the audit be performed on the buyer's own data?
At the pilot stage, yes. Vendor-produced audits across all customers establish baseline credibility for the system's behavior on a population the vendor has worked with previously. A pilot audit on the buyer's own representative chart sample, in the buyer's specialty mix and against the buyer's payer rules, is the most reliable single evaluation signal. The two are complementary; neither is a substitute for the other.
What architectural questions matter most?
Whether the system is trained against historical human-coded examples (which structurally caps accuracy at the human inter-coder reliability ceiling of roughly 82%), whether the vendor trains on patient data and from which patients, where inference compute runs, how ICD-10-CM, CPT, and HCPCS specification updates are incorporated, how AHA Coding Clinic and NCCI changes propagate to production, and how multi-tenant data isolation is implemented. These are all questions with substantive answers that distinguish vendors.
Sources cited
- KLAS Research. Autonomous Coding 2025: A Promising Start for an Early Market. August 2025. First KLAS report formally recognizing autonomous coding as a distinct healthcare technology segment; documents current specialty concentration in radiology and emergency department settings. klasresearch.com/segment/autonomous-coding
- Peng, M., Eastwood, C., Boxill, A., Jolley, R.J., Rutherford, L., Carlson, K., Dean, S., Quan, H. (2018). Coding reliability and agreement of International Classification of Disease, 10th revision (ICD-10) codes in emergency department data. International Journal of Population Data Science, 3(1):445. DOI: 10.23889/ijpds.v3i1.445. Establishes the 82.2% inter-coder agreement ceiling (Cohen's kappa 0.82) for credentialed coders at four-digit ICD-10 specificity. ijpds.org/article/view/445
- American Health Information Management Association (AHIMA). Coding standards, audit methodology guidance, and CDI program references. ahima.org
- American Academy of Professional Coders (AAPC). CPC and CCS credentialing standards, audit methodology references, and code-set update guidance. aapc.com
- American Hospital Association. AHA Coding Clinic for ICD-10-CM/PCS. Quarterly official coding guidance that any production coding system, human or AI, must track. ahaonlinestore.com
- Centers for Medicare & Medicaid Services. National Correct Coding Initiative (NCCI) edits and Medicare claims coding guidance. cms.gov/medicare/coding-billing/ncci-edits
- MedAxiom (American College of Cardiology) and Professional Consulting Services (PCS). Independent audits of AccuCode AI coding accuracy on cardiovascular surgery, E/M, and broader specialty coverage, 2024-2025. Audit leadership: Nicole F. Knight (LPN, CPC, CCS-P, EVP Revenue Cycle Solutions, MedAxiom), Jammie Quimby (Director of Coding, MedAxiom), Scott Roper (MBA, CPC, COO AccuCode and President PCS), Tracye Enis (CPC, VP Corporate Compliance, PCS).
