A coder at a six-physician internal medicine group kept a spreadsheet of every note she had to send back for clarification. In one quarter she had queried one physician 214 times, mostly for the same three things: a time statement that said "spent 40 minutes" with no indication of what counted, an assessment that listed "diabetes" with no type or complication, and a modifier 25 visit whose note did not separate the problem work from the procedure. She had emailed him about each one. He had answered each one. Nothing had changed, because 214 emails about individual notes do not add up to a conversation about a pattern.
Coding feedback for physicians fails in most practices for a simple reason: it arrives as individual corrections, at the wrong time, from someone the physician does not see as a peer, with no picture of the pattern and no agreement about what to change. The physician experiences it as nagging. The coder experiences it as being ignored. Both are right. This article is the monthly review we build with practices to replace that loop: how the sample is drawn, what the scorecard shows, how the conversation runs and why we insist on one change at a time.
A glossary line. "E/M" means evaluation and management, the office visit codes 99202 through 99215 that make up most of a primary care or medical specialty practice's revenue. "MDM" is medical decision making, the framework since 2021 that determines the level of an office visit from the number and complexity of problems, the data reviewed and the risk of management, unless the visit is coded by total time instead.
Key takeaways
- Feedback works when it describes a pattern across a sample, not an error in a single note.
- Ten charts per physician per month, chosen by a fixed rule, is enough to show the pattern and small enough to review carefully.
- The E/M distribution report, comparing each physician's mix of visit levels with the practice and the specialty, is the fastest way to start a useful conversation.
- The scorecard fits on one page, the conversation takes fifteen minutes, and it ends with one agreed change for the coming month.
- The coder's query log is a data source, not a delivery mechanism; use it to build the scorecard, not to send 214 emails.
The sample
Each month, for each physician, pull ten encounters from the previous month. Not the ten with denials; those tell you about payers. Not the ten highest levels; those tell you about outliers. A fixed rule such as every fifteenth encounter by date, or five random plus the five with the most coder queries, gives a mix that represents how the physician actually documents. Keep the rule the same every month so trends mean something.
The reviewer is a certified coder, internal or external, who compares the documentation with the codes billed and records for each chart: the E/M level documented versus billed, whether MDM or time was the basis and whether it was stated, the specificity of each diagnosis, whether any procedure or modifier was supported, and whether anything documented was not billed. That last item matters. Under-coding is a finding too, and physicians hear feedback about lost revenue more readily than feedback about compliance risk, even though the two are the same conversation.
The E/M distribution report
Alongside the sample, run each physician's E/M level distribution for the quarter: what percentage of established visits were 99212, 99213, 99214 and 99215, and the same for new visits. Put the practice average beside it, and the national specialty distribution that CMS publishes from Medicare claims data. A physician who bills 80 percent 99213 in a specialty where the norm is closer to half 99214 is either seeing simpler patients or under-documenting, and the ten-chart sample usually says which. A physician who bills 40 percent 99215 is either unusually complex or over-coding, and the sample says which.
The distribution report is not evidence of anything by itself. It is the opening question. "Your visits are mostly 99213 and the sample shows four that documented 99214-level decision making. Let's look at those four." That is a sentence a physician will engage with, because it is about their patients and their money, and it is not an accusation.
The one-page scorecard
| Section | What it shows | Example |
|---|---|---|
| Sample accuracy | Charts where documentation supported the billed level, over ten | 7 of 10 supported; 2 under-documented, 1 over-coded |
| Direction | Net effect of the three misses | Two 99214 visits documented as 99213 (under); one 99215 billed with 99214 documentation (over) |
| E/M distribution | Physician versus practice versus specialty, established visits | 99213: 74 percent (practice 58, specialty 45); 99214: 22 percent (practice 36, specialty 44) |
| Diagnosis specificity | Unspecified codes where a specific one was documented | E11.9 used in 6 of 10; 4 of those documented neuropathy or CKD (E11.40, E11.22) |
| Modifiers and procedures | Supported, unsupported, missed | Modifier 25 on 3 visits; 2 supported, 1 with no separate problem documentation |
| Time statements | Present and compliant when time was the basis | 3 visits coded by time; 1 stated activities and date, 2 said "40 minutes" only |
| Missed charges | Documented services not billed | Smoking cessation counseling documented twice, 99406 not billed |
| Last month's change | Did the agreed change happen? | Agreed: specify diabetes complications. Result: 4 of 6 diabetes visits now specific (was 1 of 7) |
| This month's change | One item, agreed in the meeting | Time statements: list the activities and confirm the date of service |
Everything on the page is a count or a short example. No paragraphs. No coding guideline citations; those go in an appendix the physician can read if interested. The page is designed to be read in three minutes before the conversation, and to be compared with last month's page.
The fifteen-minute conversation
The meeting is between the physician and one person, ideally the coder who reviewed the charts, or the practice's coding lead if the coder is uncomfortable in the role. It is scheduled, not grabbed in the hallway, and it is fifteen minutes. The order is fixed. First, what went well: last month's change and whether it happened. Second, the two or three charts that best illustrate the pattern this month, with the note on screen. Third, the direction of the misses and what they cost or risk. Fourth, agree on one change for the coming month. Fifth, write the change on the scorecard in the physician's words.
The tone matters and so does the framing. "The documentation supports a 99214 here, and it went out as 99213; that is $40 the practice did not collect for work you did" lands differently from "you under-coded this". "This note bills a 99215 and the MDM reads as moderate to me; help me see what I am missing" is a question, not a verdict, and physicians often have a clinical answer that changes the coder's read. When they do not, they usually agree without being told.
What kills the conversation is volume. Six findings, four handouts and a guideline quiz produce a physician who nods and changes nothing. One change a month, repeated until it sticks, produces a physician who changes twelve things a year and remembers them. The internal medicine coder's 214 queries were about three patterns. Three months of scorecards, one change each, would have addressed all of them.
The patterns we see most
Across specialties, the findings that appear on scorecards most often are a small set. Unspecified diagnoses where the note documents the specific condition, which affects risk adjustment and quality reporting as well as the claim. Time statements that give a number without the activities or the confirmation that the time was on the date of service. Assessments that list problems without the status (stable, worsening, at goal) that MDM depends on. Modifier 25 visits where the note does not separate the E/M work from the procedure. Data review that happened but was not documented: the outside records read, the independent historian, the test ordered and interpreted. And, increasingly, ambient scribe drafts signed without editing, which produce long notes with generic plans that support less than the physician actually did.
Each of those is a documentation habit, not a knowledge gap. Physicians know what a stable versus a worsening problem is. They do not always write it down, because nobody has shown them that the word "worsening" is the difference between a 99213 and a 99214. The scorecard shows them.
Training new physicians and residents
For a new physician joining the practice, the monthly review starts in month one with a larger sample, twenty charts, and a weekly ten-minute check-in for the first six weeks. For residents and new graduates who have never coded their own visits, we add practice cases in a training EHR, where they document a fictional visit, code it, and compare their code with the coder's, before the pattern sets in on real patients. Our coding courses use the same case method. The point is the same as the monthly review: a pattern seen across cases, discussed in a short conversation, changed one habit at a time.
Questions we hear
Should the coder or the practice manager deliver the feedback?
The person who read the charts, if they can hold the conversation. Second-hand feedback loses the specifics, and the specifics are what persuade. If the coder is junior or contract, pair them with the medical director for the first few months so the physician hears a peer endorsing the process.
What if a physician refuses to participate?
Start with the distribution report and the missed-charge line, which are about revenue, and let the medical director own the request. Most physicians engage when the first meeting shows them money they left on the table. A physician who refuses after that is a governance problem, not a coding one, and the owners have to decide how to handle it.
How long until the accuracy rate moves?
In our experience a physician who engages goes from six or seven supported charts in ten to nine within three to four months, and the E/M distribution drifts toward the specialty norm over two quarters. Physicians who were under-coding see the revenue effect first, which tends to keep them engaged for the compliance findings that follow.
What to do this week
- Write the sampling rule and pull ten charts per physician from last month.
- Run the E/M distribution report by physician for the last quarter and add the practice and specialty comparison.
- Build the one-page scorecard template from the table above and complete it for one physician.
- Schedule a fifteen-minute meeting with that physician, with the notes ready on screen, and agree on one change.
- Turn the coder's query log into a monthly count by pattern and stop sending per-note emails for patterns that are on a scorecard.
