Conformance checklist

Utilization-review AI evidence checklist

Eleven statutory duties for AI in utilization review, and the evidence each one requires you to be able to produce.

Why it exists

Maryland defines artificial intelligence, for utilization review, as a system that varies in its level of autonomy and produces outputs that can influence physical or virtual environments. That is a description of an agent — written into state insurance law in May 2025, months before the April 2026 revised US interagency model risk guidance placed generative and agentic AI expressly outside its scope. The controls for agentic decisioning in health insurance are being specified by state insurance regulators, not by federal model-risk supervision, and they bind the vendor conducting the review as directly as the payer. Each duty below is a yes or no a system can either evidence or cannot.

What it contains

01

Scope test — the three questions that determine whether the duties bind you, including the two that reach vendors

02

The eleven duties of § 15-10B-05.1(C), each with the evidence required and a test you can run

03

The ablation test for group-data reliance — the only version that produces evidence rather than assertion

04

The prohibition in (D), treated as an entitlement question rather than a process one

05

Reporting and production obligations, and the architectural precondition each one carries

06

Why the AI-involvement flag must be captured at decision time and cannot be reconstructed from logs

07

Closure rules — including why evidence produced under examination is not evidence

No email required, and nothing is recorded when you download. Use it, adapt it, argue with it — attribution is welcome, not a condition.

Conformance checklist · v1.0 · 4 August 2026

The eleven duties

Subsection (C) requires an entity in scope to ensure each of these. The test column is the part that matters: a duty you cannot test is a duty you are asserting.

ClauseWhat the statute requiresThe test
(C)(1)Determinations rest on the enrollee's own clinical history, circumstances or record.List every enrollee-specific input the tool saw for one determination. Access to the chart is not use of the chart.
(C)(2)Determinations are not based solely on a group dataset.Run the determination with and without the enrollee-specific inputs. If the output does not change, group data is deciding.
(C)(3)Criteria and guidelines comply with the requirements of the title.For a determination made six months ago, produce the criteria version then in force.
(C)(4)The tool does not replace the provider's role under § 15-10B-07.Report override rate and median reviewer dwell time. A role exercised in seconds at volume is a role in name.
(C)(5)Use does not result in unfair discrimination prohibited by law.Produce the latest disparity analysis and what changed because of it. Excluding protected attributes is not a finding.
(C)(6)The tool is fairly and equitably applied, including per HHS guidance.Is it applied to some populations and not others, and is that difference justified or just rollout sequencing?
(C)(7)The tool is open to inspection for audit or compliance review by the Commissioner.Request the inspection package with five days' notice. Did it exist, or was it assembled?
(C)(8)Written policies in the filed utilization plan state how the tool is used and what oversight is provided.Read the filed description to an engineer who built it. Divergence between filing and behaviour is the finding.
(C)(9)Performance, use and outcomes reviewed and revised at least quarterly, to maximise accuracy and reliability.Does the cadence survive a mid-quarter model version change made on the provider's schedule, not yours?
(C)(10)Patient data is not used beyond its intended and stated purpose, consistent with HIPAA.Query the retrieval layer under a purpose the data was not collected for. Does it refuse, or rely on restraint?
(C)(11)The tool does not directly or indirectly cause harm to an enrollee.Name the monitoring that would catch a systematic delay. Indirect harm does not show in determination accuracy.
(D)The tool may not deny, delay or modify health care services.Try to make it issue an adverse determination without a provider decision. If only process stops it, process is the control.

Reporting and production — and what each presumes about your architecture

ObligationSourceArchitectural precondition
Quarterly report stating, per adverse decision, whether an AI, algorithm or other software tool was used§ 15-10A-06, field added 2025The flag is per decision and must be captured at decision time. "The tool was in the pipeline" is a different claim from "the tool was used in making this decision."
Causal account where adverse decisions for a service type rise 10% in a year or 25% over three§ 15-10A-06(a)(2), added 2026Pre- and post-deployment comparability has to be designed in. It cannot be assembled once an examination is open.
Production of all documents related to an adverse decision, including those held by a private review agent acting for the carrier§ 15-10B-21, added 2026Production reaches through the vendor. Contractual audit rights are necessary and not sufficient — the records must exist in producible form.

How to use it

  1. Select one material historical adverse determination where a tool was involved — not a synthetic case, and not the best-documented one.
  2. Work the eleven duties in order, linking the record that proves each. Do not paste explanations into the evidence column.
  3. Mark anything missing, indirect or unreconstructible as a gap, with an owner and a date.
  4. Have the reconstruction performed by someone outside the team that built the system. The failure mode is familiarity, not competence.
  5. If a vendor conducts any part of the review, run the scope and production sections against them, not only against yourself.

Built from public standards and general practice. The instruments cited in this artifact are checked against their primary sources on an ongoing basis. Instruments move — several cited here changed inside the last year — so verify against the source before you rely on one. If you find something stale, tell me and I will correct it.

Related: the other artifacts · the diagnostic.

Next step

Want this applied to your estate?

These artifacts are general by design — and they are the method behind the Agent Estate Review: two to three weeks establishing what is actually running, what each thing is permitted to do, and where you could not evidence it if you were asked next week. You have just read how it works.