Artificial intelligence is moving into medical documentation and billing faster than independent evidence can establish how accurately the tools perform, the Government Accountability Office warned in a new technology assessment.

The workload explains the appeal. GAO says U.S. clinicians average a 57-hour week, including seven hours of administrative work. Visit notes and billing review can continue after patients leave, adding to burnout pressure.

AI scribes listen during an appointment and produce a draft clinical summary for the clinician to review. Coding systems analyze records and suggest standardized diagnosis and service codes used in insurance claims; newer systems may assign codes more autonomously.

Graphic shows a 57-hour average clinician workweek including seven administrative hours.
GAO says U.S. clinicians average 57 hours of work each week, seven of them administrative.Boho News graphic from cited primary dataView source

Use is increasing. An American Medical Association survey cited by GAO found that the share of clinicians using AI for clinical documentation or medical coding rose from 21% in 2024 to 28% in 2026.

Efficiency findings are promising but incomplete. GAO cites one study in which clinicians reduced documentation time by 20%, or about two minutes per appointment. A developer reported more than 95% coding accuracy in one deployment, but vendor claims do not answer how systems perform across settings.

The evidence gap is central. GAO found few independent accuracy studies, and some products do not preserve recordings or transcripts that a health system could use to audit the generated note against the underlying conversation.

Graphic shows surveyed clinician use of AI for notes or coding rising from 21% to 28%.
The share of clinicians surveyed by the AMA using AI for documentation or coding rose seven percentage points from 2024 to 2026.Boho News graphic from cited primary dataView source

An inaccurate note can affect later care. An incorrect code can also produce overpayment or underpayment. Even a more complete record could raise spending if it consistently captures more billable diagnoses and services than existing workflows.

Privacy and consent add a separate risk. Ambient systems process sensitive conversations, retention practices vary and patients may not always understand when recording occurs or how the data will be used.

GAO does not call for banning the tools. It asks what performance information providers and policymakers need, and how agencies and insurers should oversee reimbursement. The immediate standard is human review backed by auditable evidence—not an assumption that a fluent draft is an accurate record.