E&M Coding Audits: How to Know If Your Providers Are Billing Outliers (And Why It Matters)
Author Jacob Yates at Healthcare Compliance Pros
Every practice administrator has had the same quiet worry at
some point: is one of our providers coding too many high-level visits? For
decades that worry was hard to assess because there was no straightforward way
to see how a provider's billing compared with everyone else's. Now all that has
changed. The Centers for Medicare & Medicaid Services (CMS) and its network
of program integrity contractors now run data analysis across the Medicare
claims universe. Using that analysis, they can ascertain whose evaluation and
management (E&M) coding looks different from their peers.
No surprise here—commercial payers have also followed a
similar path. Understanding how that comparison works, and what it does
and does not prove, is the first step toward protecting a practice before a
letter ever arrives or an audit occurs.
What "Outlier" Actually Means in E&M Coding
The word "outlier" gets thrown around loosely in coding
conversations, but in the context of Medicare program integrity it has a
specific, statistical meaning. The CMS's Medicare Program Integrity Manual[1]
instructs Medicare Administrative Contractors (MACs) and Unified Program
Integrity Contractors (UPICs) to apply "well-established statistical
methods to claim information and other related data to identify potential
errors and potential fraud by claim characteristics,"[2]
rather than relying on a single national average as the benchmark for every
provider. In practice, this means a provider's E&M code distribution is
compared against a peer cohort defined by specialty, and often by geography and
practice type, rather than against the entire Medicare physician population as
one undifferentiated group.
This peer-cohort approach is visible in the CMS's
Comparative Billing Report (CBR) program. The CMS describes CBRs as a
reflection of a specific provider's billing and/or prescribing patterns, as
compared to their peers' patterns, within a service area commonly known to
result in improper payments. These reports are built around percentile-based
comparisons to state and national peers in the same specialty. In simpler
terms—an Internal Medicine physician will only be compared to another Internal
Medicine physician, and a Dermatologist will only be compared to other
Dermatologists.
CMS's regulations describe similar logic for MAC-issued
reports, which may target providers with the highest utilization for the
services they bill and show "how the provider varies from other providers
in the same specialty payment area or locality."[3]
These reports are commonly built around a threshold at or above the 90th
percentile for the metric being evaluated.
The fact of the matter is that a provider does not need a
single obvious red flag, such as an implausible diagnosis, to draw algorithmic
attention. If they are consistently in the upper range of a specialty's E&M
distribution, that can be enough on its own to place a provider into the pool
that is reviewed more closely. Remember: This is a pattern-based signal,
not proof of wrongdoing. The CMS has been very explicit that
receiving a comparative report is not an indication of wrongdoing and does
not require a response. That being said—the same data can still lead to a closer review as
part of routine program integrity work.
How Payers and CMS Detect Outliers Today
The Medicare audit detection program has grown to be more
centralized and analytically driven. UPICs are now at the center of Medicare
and Medicaid fraud, waste, and abuse detection. Their primary goal is to find
cases of suspected fraud, waste and abuse, review them thoroughly and in a
timely manner, and take immediate action to ensure that Medicare Trust Fund money
is not paid inappropriately. UPICs work across both the Medicare
fee-for-service and Medicaid data sets. In fact, the Department of Health and
Human Services (HHS) Office of Inspector General (OIG) reports they are
"the only program integrity contractors that safeguard both"[4]
programs. This gives them a much broader and more analytic view than a
single-program contractor could provide.
On top of the UPICs, the CMS's Fraud Prevention System is
also monitoring claims from all Part A, B, and DME and prioritizing individual
providers for review based on their risk scoring. Most frequently, alerts are focused
on individual providers and routed to UPIC analysts and MACs, giving
contractors a running risk score for every billing provider rather than an
occasional snapshot.
When we examine the CMS and OIG program integrity work
collectively there are six key data elements that recur when identifying
E&M outliers:
1.
The E&M level distribution itself - Meaning
how often a provider bills the highest-level codes relative to specialty peers.
2.
The Modifier 25 rate - Since that modifier lets
an E&M service bypass edits that would otherwise bundle it into a same-day procedure.
3.
The relationship between documented visit
duration and billed code level - Which can be cross referenced against EHR timestamps.
4.
A provider-level outlier score generated from
aggregated claims analysis.
5.
Documentation completeness is relative to the
level billed.
6.
The provider's denial and clean-claim rate
None of these six data elements function as a standalone
fraud determination, but together they can form a composite picture that both
government contractors and commercial payers can use to prioritize reviews. Although
technically different, commercial payer coding-edit programs describe similar
logic to that of the CMS. Most describe their coding-edit programs as a form of
evaluating the appropriateness of high-level E&M service levels and
identifying outlier providers who consistently over-code. Again, drawing on a
similar statistical-style approach as CMS has published for its own Medicare
contractors.
The Real Consequences of Being Flagged
Being identified as a statistical outlier does not
automatically produce a penalty. Typically, it starts with a chain of activities
that are worth understanding in advance and investigating more. The CMS's
Targeted Probe and Educate (TPE) program illustrates this process quite
clearly. The TPE is designed so MACs can use data analysis to identify
providers and suppliers with high claim error rates or unusual billing
practices. Once they identify these providers/suppliers, then they request a
sample of claims (typically 20 to 40 per "round") for review and education. Providers get up to
three "rounds," with at least 45 days between "rounds" to
improve. The CMS has made it abundantly clear that problems failing to improve
after three "rounds" are going to be referred to CMS for next steps. Those next
steps could consist of anything from a 100 percent prepay review,
extrapolation, referral to a Recovery Auditor, or other actions. Needless to
say, you do not want to end up at this point.
Now, let's break down a few of these potential mitigation
steps. First, extrapolation. This is the mechanism that turns a small sample
finding into a large financial demand (and it compounds quickly). The CMS'
Program Integrity Manual authorizes contractors to extrapolate whenever a
"sustained or high level of payment error" is established. Once an
overpayment estimate is calculated and reviewed, contractors are instructed
that "in most situations, the lower limit of a one-sided 90 percent
confidence interval should be used as the amount of overpayment to be demanded
for recovery"[5]
This figure is then predicted across the provider's entire
universe of claims for the audited service and period. So, a review of a dozen or
so claims can easily translate into a demand covering thousands of claims. Often,
these actions encompass claims well over a year after the services were actually
billed. The most severe escalations run through 100 percent prepayment review. This
means every single claim of that type must clear a contractor review before
payment. Think about the impact that could have on your organization. It
essentially halts the normal cash flow for that service line until sustained
compliance can be proven.
Common Causes of Outlier Status (That Aren't Actually Fraud)
As any statistician will tell you, statistical outliers can
show a pattern. However, that pattern by itself is not a verdict or
confirmation of inaccuracy or wrongdoing. Even the CMS frames it this way. Some
practices can push a provider's E&M distribution toward the upper end of
their specialty's curve. For example, a provider who concentrates or is
specially trained in complex chronic disease management or a subspecialty that
structurally requires more intensive medical decision making, will bill more
high-level codes than their average peer.
Now, a quick sidenote and history lesson. In 2021, changes
to E&M documentation requirements were made to reduce the reliance on rote
history and exam elements in favor of medical decision making or time as the
primary drivers of code selection.[6]
That flexibility is clinically valuable, however, if it is inconsistently applied
across a provider's notes, it can leave a real, medically necessary high-level
visit without the clear documentation needed to support the code on paper. The CMS's
2026 guidance notes that providers should not use the volume of documentation to
influence the specific level of service to be billed. Documentation should support
the level reported, not the other way around.
Another (and in my opinion, significantly increasing)
non-fraud driver is cloned or templated documentation. The use of generative AI
and other new "tools" in EHR systems continues to draw attention from federal
and private payers. Some of this problem exists because of how EHR systems are
built. To be perfectly clear—canned or templated EHR notes are helpful tools,
but canned or templated notes should never be used at the expense
of accurate and appropriately thorough documentation. Telling an auditor or
contractor that the "EHR told you to do it" will never be a valid argument, so use
these tools and resources with appropriate caution and understanding.
Now let's talk a little more specifically about "cloning"
and how federal regulators are interpreting this term. The OIG has described
"cloning," as documentation worded nearly identically to
previous entries and a fraud vulnerability tied to copy-paste functionality. The
CMS has warned EHR uses for over a decade that copy-and-paste documentation
"lacks the patient-specific information necessary to support services
rendered to each patient" and can lead to "coding from old or
outdated information that may lead to upcoding."[7]
Back in 2013, an OIG review discovered only about
one-quarter of surveyed hospitals had any policy governing copy-paste use,
despite their workforce having widespread reliance on the function.[8]
I suspect if this same review were conducting today this statistic would be
even higher. The use of AI—and worse, shadow AI use—is rampant in professional
organizations today, but it appears especially prevalent in healthcare. To be
clear, none of this means a provider using templates is committing fraud,
but it provides an explanation as to why legitimate high-level coding can look
indistinguishable from overdocumentation.
How to Self-Audit Before Payers Do
The CMS built tools specifically designed so practices could
run this type of benchmarking analysis internally before a payer
or contractor does it for them. The Comparative Billing Report program itself
is described by CMS as an educational tool to support providers' internal
compliance activities through peer comparisons at the state, specialty, and
national level.[9]
The CMS published a self-audit guidance framework for
internal monitoring and auditing as a part of the OIG's Seven Elements of an Effective
Compliance Program[10],
published in the early 2000's. It recommends practices regularly review and
identify their highest-risk billing areas and measure performance against those
risks, rather than reacting after an external trigger.[11]
A practical internal review process built on that framework generally covers
three primary checks.
1.
Pull the practice's own E&M code
distribution by provider and specialty and compare its shape to publicly
available Medicare specialty benchmarks, watching for a provider whose share of
high-level codes sits well above what CMS data would suggest is typical.
2.
Review modifier 25 use as a standalone metric. Since
federal audit activity on that modifier has been persistent for over two
decades, the OIG's original review found a 35 percent improper payment rate on
sampled modifier 25 claims. A subsequent OIG audit in 2025 found the
documentation for 22 of 24 sampled E&M services billed with modifier 25 (alongside
intravitreal injections) did not support the modifier's use.[12]
3.
Spot-check a sample of visit notes against
scheduling-system timestamps to confirm where time is used to select the
E&M level, that the documented time and actual visit duration are aligned.
What to Do If You're Flagged
Let's say tomorrow you receive a TPE notice, a prepayment
review notice, or a comparative billing report. Again (because the third time's
a charm) this is not in itself evidence of a compliance failure. CMS is
explicit that these communications require no immediate response beyond
internal awareness. The appropriate first step is to run the same
benchmarking analysis internally before responding, so any reply to the
contractor is grounded in the practice's own understanding of why its numbers
look the way they do. The TPE is structured around education rather than
immediate penalty in its first rounds, and CMS notes most providers who go
through the process actually do improve their accuracy.
Building a defensible documentation trail going forward
means closing the specific gaps identified during the probe, whether that involves
more consistent medical decision making or time documentation, tighter modifier
25 support, or even reduced reliance on cloned note language. Keeping a written
record of the corrective steps taken shows the CMS contractors and auditors,
you take this seriously. The OIG's long-standing compliance program guidance
for physician practices[13]
reiterates the importance of responding appropriately to detected offenses and
developing a corrective action plan, as
one of Seven Elements of an Effective Compliance Program.
If you or your practice is facing an extrapolated
overpayment demand, a referral to 100 percent prepayment review, or an active
UPIC investigation, you should bring in outside audit or legal support at that
stage. The financial exposure from extrapolation and the complexity of challenging
a sampling methodology on appeal both increase substantially once a case moves
past the TPE education phase.
How SENTRY Coding Intelligence Applies This Same Logic — For You
The detection logic described throughout this piece, such as
specialty-based peer benchmarking, modifier 25 rate monitoring, and
documentation vs. billing consistency checks, is publicly documented by CMS and
OIG. This means practices do not have to wait for a payer or contractor to
apply it first. In fact, it is better if you get ahead of any coding issues
first! Healthcare Compliance Pros proprietary SENTRY Coding Intelligence is
built to run this same analysis for healthcare organizations internally and on
a recurring basis, rather than as a one-time exercise. It visualizes
specialty-level benchmarks across E&M utilization, relative value units,
modifier density, time, and visit volume so a compliance officer, physician, or
billing specialist can see, provider by provider, where the practice's coding
profile sits relative to its peer cohort. It generates a score for each
provider so leadership can prioritize which patterns to review first, based on
both statistical deviation and financial exposure. Running the analysis
quarterly allows for the creation of an analytics, audit, and education cycle. And
a bonus! It mirrors the same probe-and-educate rhythm CMS itself uses, allowing
your organization to evaluate documentation and coding habits and make
necessary course corrections. Following this pattern allows for continual reinforcement
rather than only correcting issues after a payer letter arrives.
So, What Now?
The shift toward AI-assisted, peer-benchmarked audit
detection continues to be a frequent topic between CMS's own program integrity
manuals, its Fraud Prevention System, and a steady stream of OIG audit findings
and active work plan items. This includes the modifier 25 review OIG announced
in March 2026[14]. Practices
that understand how percentage-based peer benchmarking works, and incorporate
their own recurring internal audits, are better positioned to distinguish a
legitimate high-acuity coding profile from a documentation gap that needs
attention. Put forth the effort to get a baseline coding risk profile for your
practice and see where your providers stand before a payer or
Medicare contractor decides to audit you.
Frequently Asked Questions
What percentile of E&M coding is considered a payer
or CMS audit risk?
There is no single universal cutoff, however, CMS's own Comparative Billing
Report criteria have used the 90th percentile. Of course, this is relative to
state or national peers in the same specialty as a threshold for issuing an
educational report. Organizations with persistent billing in the 75th to 90th
percentile range (especially for high-level E&M codes) have also been
associated with increased scrutiny. Although, technically speaking, CMS treats
these thresholds as risk indicators rather than automatic triggers.
Does receiving a Comparative Billing Report or a TPE
notice mean my practice is being accused of fraud?
No. CMS has reiterated on multiple occasions and in publication that
receiving a comparative billing report is not an indication of wrongdoing. The
TPE is simply designed as an education-first program, with prepayment review or
extrapolation occurring when providers error rates do not improve after
multiple rounds of probe and education.
How far back can Medicare go when demanding an
extrapolated overpayment?
There is no single fixed lookback period stated for every case, but because
extrapolation projects a sample error rate across a provider's full claims
universe for the audited service and period, demands can cover a substantial
span of prior billing once a sustained or high level of payment error is
established.
Is modifier 25 use itself a red flag?
No, not on its own. The purpose of Modifier 25 is to allow payment for a
legitimate, significant, and separately identifiable E&M service performed
on the same day as a procedure. The OIG has continued to find high error rates
in modifier 25 documentation across specialties. However, it currently has an
active audit reviewing Medicare Part B claims for E&M services billed on
the same day as minor surgery without modifier 25. So, as of the
publishing of this blog post, both overuse and inconsistent use will draw
federal attention.
Can a practice legally run its own peer-benchmarking
analysis before a payer does?
Absolutely! The Medicare Part B utilization and specialty data, which underlays
the CMS' own benchmarking programs, is publicly available. The CMS and OIG
actively encourage internal monitoring and auditing as one of the core elements
of an effective compliance program.
[1] https://www.cms.gov/regulations-and-guidance/guidance/manuals/internet-only-manuals-ioms-items/cms019033
[2] Medicare
Program Integrity Manual; Chapter 2; Section 2(B)
[3]
Medicare Program Integrity Manual; Chapter 3; Section 2; Subsection 2(A)
[4] https://www.oversight.gov/sites/default/files/documents/reports/2022-11/OEI-03-20-00330.pdf
[5] https://www.cms.gov/files/document/r11797pi.pdf
[6] https://www.cms.gov/files/document/mln006764-evaluation-management-services.pdf
[7] https://www.cms.gov/files/document/ehrdecisiontable062816pdf
[8] https://oig.hhs.gov/oei/reports/oei-01-11-00570.pdf
[9] https://www.cms.gov/data-research/monitoring-programs/medicare-fee-service-compliance-programs
[10] https://oig.hhs.gov/documents/provider-compliance-training/945/Compliance101tips508.pdf
[11] https://www.cms.gov/files/document/selfauditfactfactsheet020116pdf
[12] https://oig.hhs.gov/reports/all/2025/medicare-payments-for-evaluation-and-management-services-provided-on-the-same-day-as-eye-injections-were-at-risk-for-noncompliance-with-medicare-requirements/
[13] https://oig.hhs.gov/compliance/physician-education/
[14] https://oig.hhs.gov/reports/work-plan/browse-work-plan-projects/evaluation-and-management-services-on-same-day-as-minor-surgery-with-no-modifier-25/