US-Based Workforce · US Data Residency
Epic App Available Customer Login · Call (501) 830-1478
Product 02 · Clinical Quality & Compliance

Every measure,
abstracted.
Every submission,
defensible.

AccuCode abstracts quality measures from the complete medical record, calculates denominators and numerators per registry specification, and submits directly to CMS, The Joint Commission, and specialty registries. Your abstractors move from data entry to data governance.

Abstractor time 0 min · Fully Automated
Accuracy Exceeds Industry Standard
Submission Direct or Audited
The weight of evidence: Truth versus ConvenienceAn architectural line drawing of a balance scale tipped toward Truth, weighed down by a stack of evidence documents, while Convenience rises lightly on the other side.FIG. I · THE WEIGHT OF EVIDENCEConvenienceTRADITIONAL PROCESSTruthEVIDENCE BOUNDA C C U C O D ECLINICAL QUALITY ENGINE
Registry & measure coverage
CMS
eCQMs · Hospital IQR · VBP · HRRP · HACRP · MIPS · MVP
LIVE
TJC
ORYX measures · Perinatal · Stroke · VTE · Substance use
LIVE
NCDR
CathPCI · ICD Registry · AFib Ablation · Chest Pain-MI
LIVE
STS
Adult Cardiac · Congenital · General Thoracic
LIVE
AHA
Get With The Guidelines · Stroke · CAD · Heart Failure
LIVE
 
Additional registries added on customer request
-
The abstractor's workflow

Every element, cited and justified.

✓ CMS SEP-1 Sepsis Bundle · Single case abstraction · Live product shown

A typical CMS SEP-1 sepsis case generates the equivalent of 177 single-spaced pages of clinical documentation across notes, flowsheets, labs, medication records, and imaging reports. No human can read it cover to cover, hold the chronology in working memory, and reason across it accurately. AccuCode reads all of it, in seconds, and returns every field bound to the exact passage in the chart that supports it.

accucode.app / quality / sts / cases / 03-4419
EXPORT⋯
← QUEUE · 12
CMS SEP-1 · Case 03-4419
F · 68 yo · MRN ••••2847 · CABG x3 · DOS 03/28/2026
Fields abstracted
29 / 38
Demographics & history 7 / 7
Age at surgery
68 years
Sex
Female
BSA
1.74 m²
Prior MI within 21 days
No
Cardiac status 6 / 6
Ejection fraction (pre-op)
42 %
NYHA class
III
Angina status
Stable
Heart failure
Yes
Comorbidities 8 / 8
Diabetes
Yes, insulin
CKD (last eGFR)
48 mL/min
Dialysis dependent
No
Peripheral arterial disease
No
Operative details 8 / 11 · 3 pending review
Procedure type
CABG only
Number of distal anastomoses
3
IMA used
LIMA only
Post-op outcomes 0 / 6 · pending chart completion
Currently reviewing · Initial assessment
Initial lactate (at presentation)
CMS SEP-1 · Initial Lactate · Required · Numeric mmol/L
Abstracted value
4.2mmol/L
Confidence
94%
◆ Source evidence · 2 passages
ED Triage Note · 03/24/2026 · 0815
Pt arrives by EMS with SBP 88/52, HR 118, T 102.4, mental status altered. Two systemic inflammatory response criteria met. Suspected source: urinary. Lactate drawn.
Lab Results · 03/24/2026 · 0823
Patient with hypotension (SBP 88), tachycardia (HR 118), and altered mental status. Lactate 4.2 mmol/L at 0823. Sepsis bundle initiated.
⚡ SEP-1 spec guidance applied
Where multiple lactate measurements exist within the sepsis presentation window, use the earliest value that meets severe sepsis criteria. Subsequent measurements after intervention are recorded as re-assessment, not initial. Here: the 0823 measurement is the first to meet criteria after presentation.
The problem we solve

Quality data lives in
the unstructured text.

Not the fields.

A single
sepsis case
177
pages of clinical documentation

Hospital quality departments employ certified clinical data abstractors to read charts, apply registry specifications field by field, and key values into registry portals one entry at a time. The problem is that modern clinical documentation has outgrown what any human can comprehensively process.

How big a chart can be · click to expand

A typical CMS SEP-1 sepsis case generates roughly 177 single-spaced pages of structured and unstructured data spread across triage notes, hospitalist progress notes, nursing flowsheets, medication administration records, lab results, imaging reports, and discharge summaries. The information that actually matters for the abstraction, the moment sepsis criteria were first met, the time antibiotics were administered, the documentation of source identification, is rarely sitting in a structured field. It is buried in the text of clinical notes, in nursing observations, in pharmacy timestamps, and it requires reading and reasoning across all of them simultaneously to establish the chronology. Humans don't fail at this because they aren't trying. They fail because the task, performed comprehensively, exceeds what any individual can hold in working memory.

This is also why accuracy on registry abstraction has been stuck for two decades. The existing software category, rules-based extractors, looks at structured fields and dropdown entries. When the documentation matches the expected pattern, the extractor catches it. When it doesn't, which is most of the time, the extractor returns blank and asks the abstractor to find it manually. The result: software handles the easy data and humans handle the data that actually drives accuracy. The hard fields, the ones buried in clinical text, in nursing observations, in the unstructured narrative of how care actually unfolded, are exactly the fields the existing software category cannot read.

AccuCode reads the entire chart. Not the structured fields. The entire chart. Every triage note, every flowsheet entry, every progress note, every medication administration timestamp, every pharmacy record, every imaging report, every discharge summary, processed in parallel and reasoned across as a single unified clinical record. This is the architectural difference that makes accuracy go up dramatically, not by extracting better, but by being able to see the entire record in the first place. Most quality measure accuracy gaps are not abstractor errors. They are documentation that existed in the chart but was never read.

Most quality teams are perpetually behind. Abstractors burn out. The data that does get submitted is often incomplete because the timeline forces shortcuts. When registries reject submissions, which they do, frequently, for missing fields or specification violations, the rework cycle consumes additional weeks. This pattern repeats, quarter after quarter, at every US hospital that reports quality measures. The direct cost is enormous. The opportunity cost is worse.

AccuCode replaces the workflow entirely. Charts are abstracted automatically from ingestion through submission. The pipeline is fully automated. Your credentialed abstractors are not faster at the same work. They are doing different work: compliance sampling, exception handling, documentation improvement, and the governance role that a credentialed quality professional should hold. The work that remains is the work a credentialed abstractor is actually paid to do.

Why we're different

Extraction versus reasoning.

Most quality automation tools on the market today are rule-based extractors: pattern-match structured fields, fail silently when the pattern breaks, hand everything else back to the abstractor. AccuCode treats abstraction as the clinical reasoning task it actually is, the same task a certified abstractor performs at a different speed.

Most vendors

Extract what you can. Ask for the rest.

Rules-based tools look for specific structured fields in specific locations. When the documentation doesn't match the expected pattern (which is most of the time), the field is skipped and the abstractor is left to find it manually.

This is why "automation" saves abstractors maybe fifteen percent of their time. The work that's left is the work that was always hard.

Chart documentation: "Antibiotic order placed 0815. Medication
dispensed 0820 per pharmacy. Administered
0823 per nursing note."
→ EXTRACTION: Order time 0815. Wrong field.
vs.
Methodology · Iterative audit · Inter-rater reliability

How abstraction accuracy is
actually measured.

Most vendors compare their accuracy to humans and arrive at a number close to but slightly below human reliability. The question they don't ask: what is human reliability, really?

Clinical abstraction has no single AHIMA-equivalent benchmark for what credentialed human accuracy actually is. But three independent authoritative measurements all reach a similar floor. The federal CMS Hospital Inpatient Quality Reporting Validation Program, which re-abstracts hospital quality measure data to verify accuracy, sets its passing threshold at 75%.

What the rest of the literature shows · click to expand

Peer-reviewed inter-rater reliability studies of credentialed chart abstractors consistently produce Cohen's kappa values in the 0.51 to 0.84 range, classified by the methodology literature as "substantial" agreement, not "near perfect." AHRQ's own reliability testing of its evidence-grading instruments, conducted on credentialed raters reviewing the same material, found kappa values as low as 0.27. The implication is uncomfortable for the vendor category but unambiguous in the literature: credentialed human chart abstraction has a structural ceiling well below the 95% to 99% numbers vendors typically claim against it.

AccuCode's models were never trained on patient encounters and will not be trained on customer data. No proprietary content from any registry steward or measure developer was used in training. The architecture does not pattern-match against past human-abstracted examples. The system was built the way you would train a person to perform the same task: it reads clinical documentation and reasons from what is documented, then produces output that conforms to the current CMS Hospital IQR measure definitions, current Joint Commission core measure sets, and current AHA Get With The Guidelines specifications. The result is a system whose accuracy ceiling is not the noisy human ground truth of a training corpus, because there is no such corpus. The ceiling is the accuracy of the reasoning applied to the documentation and the measure logic in front of it.

The assumption that adding a human reviewer to an AI system improves accuracy is intuitive but empirically wrong in the published research. A 2024 randomized clinical trial published in JAMA Network Open found that LLM systems alone achieved 92% accuracy on diagnostic reasoning tasks, while physicians using the same LLM achieved only 76% accuracy, barely better than the 74% they achieved without it. The researchers concluded that adding physicians to AI did not significantly improve clinical reasoning. AccuCode's own data, described below, produces an even stronger version of the same finding for the narrower task of chart abstraction.

Revision 0 · lowest measure
95.6%
Initial system, post-adjudication
Baptist Health multi-stakeholder panel
Revision 3 · lowest measure
99.61%
Production system, after three audit cycles
CMS · TJC · GWTG, all measures
vs.
Published human inter-rater reliability for chart abstraction
Centers for Medicare & Medicaid Services
CMS Hospital IQR Validation Program Federal validation passing threshold for hospital-submitted quality measure data
75%
Agency for Healthcare Research and Quality
AHRQ EPC reliability testing Inter-rater kappa across credentialed raters reviewing the same evidence
κ 0.27
Peer-reviewed chart abstraction studies Inter-rater kappa among credentialed chart abstractors across measure types
κ 0.51–0.84

Over an audit period spanning more than eighteen months, AccuCode was tested against credentialed human abstractors at Baptist Health Arkansas across all measures in the three largest US quality reporting frameworks: CMS Hospital IQR, Joint Commission core measures, and the AHA Get With The Guidelines registry suite. Every data element on every chart, on every revision of the system, was adjudicated to consensus by a multi-stakeholder panel that included Baptist Health's senior clinical and quality leadership.

CMS · TJC · GWTG · Three revisions · Every data element adjudicated
Amanda Novack, MD
Infectious Disease Physician
Corporate Medical VP of Quality and Safety
Baptist Health Arkansas
Robert Furrey
Corporate Vice President
Clinical Performance Improvement
Baptist Health Arkansas
Caitlyn Wright, RN
Director of Clinical Informatics
AccuCode AI
Baptist Health Abstractor Team
Original abstractor team plus an independent secondary panel of Baptist Health credentialed abstractors

Two things to notice. First, even the worst-performing measure on the unrefined Revision 0 of the system, 95.6%, exceeds CMS's own federal validation threshold for hospital quality measure data by roughly twenty points. AccuCode's worst result, on its worst day, is far above what the federal regulator considers passing for credentialed human abstraction. Second, the production system achieves 99.61% on its worst measure, against a multi-stakeholder Baptist Health panel that a corporate medical vice presidents personally adjudicating every disputed data element. The remaining error rate on the hardest measure is less than four-tenths of one percent. This is why AccuCode operates as a fully automated abstraction pipeline. You have full audit access through the validation interface, and you should sample regularly for compliance and drift. But routing every chart through routine human review does not improve accuracy. The published literature, the federal validation program, and a 2024 JAMA randomized clinical trial all converge on the same finding: adding human review to a high-accuracy system reduces accuracy, not increases it.

Audit methodology · click to expand

The audit was structured as iterative multi-revision panel adjudication. The system produced abstractions for a sample of Baptist Health charts. Every data element was reviewed by the original Baptist Health abstractor, who flagged disagreements. Disagreements were then routed to Dr. Amanda Novack and to a panel that included the original Baptist abstraction team and AccuCode's internal clinical informatics team, lead Caitlyn Wright. Every flagged element was investigated against the source documentation and the relevant measure specification, with each adjudicated to consensus. Findings were used to refine the system. The refined version was then re-audited through the same process. Three full audit cycles were completed before any production-ready version of the system (Revision 3) is approved for general deployment.

Cohen's kappa and percent agreement were both calculated, in alignment with the methodology used in CMS's own Hospital IQR Validation Program and in the peer-reviewed inter-rater reliability literature. The "lowest measure" figures cited above (95.6% on Revision 0, 99.61% on Revision 3) refer to the worst-performing individual measure within the audit scope on each revision, post-adjudication. Performance on every other measure within the scope was higher.

After Revision 3 of each measure is approved for production, AccuCode performed an Inter-Rater Reliability (IRR) assessment for Baptist Health using the validated R3 output as the reference standard, in the same form that hospitals routinely use to assess their internal abstractor reliability. The observed range for credentialed human abstractor scores against R3 was consistent with the published literature on credentialed inter-rater reliability documented in the panel above, and with CMS's own 75% validation threshold.

Audit scope covered all measures in the CMS Hospital Inpatient Quality Reporting (IQR) program, all Joint Commission core measure sets, and the AHA Get With The Guidelines registry suite. The initial audit spanned more than eighteen months and tens of thousands of individually adjudicated data elements.

Sources cited

  1. Goh E, Gallo R, Hom J, et al. Large Language Model Influence on Diagnostic Reasoning: A Randomized Clinical Trial. JAMA Network Open. 2024;7(10):e2440969. jamanetwork.com/journals/jamanetworkopen/fullarticle/2825395
  2. Centers for Medicare & Medicaid Services. Hospital Inpatient Quality Reporting Program. cms.gov/medicare/quality/initiatives/hospital-quality-initiative/inpatient-reporting-program
  3. CMS FY 2027 Hospital IQR Program Guide (75% validation threshold detail). qualityreportingcenter.com (PDF)
  4. Berkman ND, Lohr KN, Morgan LC, et al. Reliability Testing of the AHRQ EPC Approach to Grading the Strength of Evidence in Comparative Effectiveness Reviews. Agency for Healthcare Research and Quality, 2012. ncbi.nlm.nih.gov/books/NBK98220
  5. Kennedy CC, Holroyd-Leduc J, Wong CL, et al. Examining intra-rater and inter-rater response agreement: A medical chart abstraction study of a community-based asthma care program. BMC Medical Research Methodology. 2008;8:29. bmcmedresmethodol.biomedcentral.com/articles/10.1186/1471-2288-8-29
Customer feature · 01
Regional integrated health system · 100+ access points across Arkansas

Validated by the team
responsible for the data.

Most quality vendors validate their abstraction accuracy internally and ask hospitals to trust the results. Baptist Health Arkansas took the opposite approach. Before readying AccuCode in production, Baptist's quality team conducted three full audit cycles of the system, with a corporate medical vice president personally adjudicating every disputed data element across CMS Hospital IQR, Joint Commission core measures, and the AHA Get With The Guidelines registry suite.

The adjudication panel was led by Amanda Novack, MD, Baptist's Corporate Medical VP of Quality and Safety, and Robert Furrey, Corporate VP of Clinical Performance Improvement. The team included Baptist's original abstractor staff, an independent secondary Baptist abstractor panel, and AccuCode's clinical informatics lead. Three revisions of the system were audited end-to-end until the production version reached 99.61% on its worst-performing measure.

The system generates abstractions directly from the chart; Baptist's quality team uses the validation interface for compliance sampling and the IRR reports AccuCode delivers as ongoing documentation that the production output continues to meet the consensus-validated standard.

"We didn't want a vendor that scored well on a controlled audit under ideal circumstances. We wanted to know how the system would perform when dealing with real-world medical documentation behavior against our own measurement standard, on our own data, over time. Three revisions and tens of thousands of adjudicated data elements later, we got our answer."
Robert Furrey Corporate Vice President, Clinical Performance Improvement
Baptist Health Arkansas
Iterative audit refinement: R0 baseline rising through three revisions to R3 production accuracyA line graphic showing accuracy improvement from R0 at 95.6 percent through R1 and R2 to R3 at 99.61 percent across three audit cycles at Baptist Health Arkansas.AUDIT REFINEMENT · CMS · TJC · GWTG95.6%99.61%R0R1R2R3BASELINEPRODUCTIONWorst-performing measure on each system revision · post-adjudication
Production accuracy
99.61%
Lowest measure on the production system after three full audit cycles. Every other measure in scope scored higher.
Audit scope
CMS · TJC · GWTG
All measures across CMS Hospital IQR, Joint Commission core measures, and the AHA Get With The Guidelines registry suite.
Adjudication panel
MD · VP · RN
Corporate Medical VP of Quality and Safety, Corporate VP of Clinical Performance Improvement, Director of Clinical Informatics, plus the Baptist abstractor team.
Audit volume
Tens of thousands
Individually adjudicated data elements across three revisions of the system, spanning 18+ months of audit work.
Full audit methodology, panel composition, and source citations are documented in the accuracy section above. IRR reports are delivered as part of the production service.
From quality directors and CMOs

The questions we hear most.

How do you handle registry specification updates?

Each registry is implemented from its current specification. Not from a generic template. When a registry updates its measure definitions, we update within the grace period and notify every affected customer. Specification versions are tracked, so submissions made under prior versions remain reproducible for audit purposes.

What happens when documentation is ambiguous or incomplete?

Fields with low confidence are flagged for abstractor review with a full explanation of why. When documentation is genuinely incomplete (a field that should exist per the specification but is not present in the chart), AccuCode marks the field as missing and provides the abstractor with the specification language and guidance on documentation improvement.

How is accuracy measured for abstraction specifically?

Field-level concordance with a blinded panel of senior certified abstractors coding the same charts independently. We report per-field, per-measure, and overall concordance. The relevant baseline is inter-abstractor agreement, which in most studies sits in the mid-80s range. We publish our methodology for every audit.

Do you submit directly to registries, or do we?

Either. We support direct submission via each registry's production API where available, or export to the registry's upload format for your team to submit manually. Most customers use direct submission for high-volume registries and manual upload for low-volume ones.

Can you handle both chart-abstracted and eCQM measures?

Yes. AccuCode supports both. For eCQMs, we generate QRDA Category I (patient-level) output. For chart-abstracted measures, we produce the registry-native format. Hybrid measure sets (which are increasingly common) are supported natively in one workflow.

What does a pilot look like?

Typically 50–100 historical charts for one registry, run in parallel with your abstractor team. We measure concordance, citation quality, and time-to-abstract on both sides. You see whether the accuracy and workflow meet your bar before committing to production deployment. Pilot runs on de-identified data under an executed BAA.

How do you handle PHI?

Every employee who can touch customer PHI is US-based. All infrastructure is US-region. Customer data is never used for model training under any circumstances. AccuCode does not train models on patient data of any kind, whether identified, de-identified, aggregated, or synthetic. There is no contractual path by which your PHI becomes training data, because the system does not learn from patient encounters at all. Full detail in our Security & Compliance documentation.

◆
Start here

Clear Your Backlog.
Keep it clear.

A 30-minute scoping call, a pilot on your active registry submissions, and a measured result before you commit. We'll show accuracy on your charts, not ours.

Request a pilot → Call (501) 830-1478
Business Hours: 8AM - 5PM PST