AI Governance 13 min read

A NIST AI RMF Review Process That Survives Legal Scrutiny

J

September 20, 2026

Ask ten companies whether they follow the NIST AI Risk Management Framework and nine will say yes. Ask for the review file, the risk register from six months ago, or the name of the person who signed off on residual risk for a given model, and the confidence tends to evaporate. That gap, between claiming a framework and being able to prove you ran it, is exactly where a plaintiff's attorney or a state regulator will go looking.

NIST AI RMF 1.0, published January 26, 2023, was never built to be audited against. There's no certification body behind it, no accredited assessors, and nothing resembling a pass or fail. That was a deliberate choice: NIST organized the framework around four functions, Govern, Map, Measure, and Manage, meant to be adapted rather than checked off.

But voluntary doesn't mean irrelevant once your AI system causes harm and someone asks what "reasonable" risk management looked like inside your organization. In my work helping companies build management systems that hold up under audit, I've watched the same pattern play out again and again: the framework was never the risk. The absence of evidence was.

Why a Voluntary Framework Still Ends Up in a Deposition

NIST AI RMF's voluntary status is precisely why it shows up in legal arguments instead of avoiding them. Courts and regulators evaluating whether a company acted reasonably don't need a mandatory rule; they need a recognized benchmark for what reasonable looks like, and NIST AI RMF has become the default answer. The FTC has used its Section 5 unfairness and deception authority against AI-related claims for years, including a September 2024 enforcement sweep it called Operation AI Comply, targeting companies that oversold what their AI products could actually do. None of those actions required a dedicated AI statute. They required a theory of what a reasonable company would have done, and reasonableness arguments pull in whatever recognized standard exists.

Colorado's original AI Act, SB 24-205, made the connection explicit rather than implicit. The 2024 law set a two-part affirmative defense for developers and deployers of high-risk AI systems. First, the company had to discover the violation through the internal review and feedback processes the statute required, and cure it. Second, the company had to otherwise be in compliance with a recognized risk management framework, with NIST AI RMF named as the clearest example of what qualified. Satisfy only the second prong and you don't have the defense; the statute required both. But SB 24-205 never took effect: its start date was pushed from February 2026 to June 2026, and Colorado then repealed and replaced it outright with SB 26-189, signed May 14, 2026 and effective January 1, 2027 — a law that does not carry forward the NIST-aligned affirmative defense. There is currently no codified Colorado safe harbor for NIST AI RMF alignment. What survives is the mechanism, not the statute: a legislature decided, if only for a while, that compliance with a named framework is what reasonable care looks like in court, and once that idea is in circulation, other states and regulators can pick it back up even after this particular bill is gone.

This is the part I want compliance and legal teams to sit with: a framework doesn't need enforcement teeth to create liability exposure. It just needs to become the reference point against which your conduct gets measured after something goes wrong.

Surviving legal scrutiny is not a synonym for avoiding a lawsuit. It means that when discovery happens, when a regulator issues a civil investigative demand, or when a plaintiff's expert reconstructs your AI governance timeline, what you produce holds together. Three things determine that, and none of them depend on how sophisticated your policy document sounds.

  1. Contemporaneous documentation. A risk assessment written the week after an incident, backdated in spirit if not in fact, is worse than no risk assessment at all, because it signals the process wasn't real.
  2. A named accountable owner for each function of the review — someone who can testify to what they did and when, not a shared inbox or a committee that never took minutes.
  3. Criteria applied consistently across systems, so opposing counsel can't show your risk tiers were built to justify whatever you'd already decided to ship.

I've sat across the table from auditors and FDA investigators, guiding companies through inspections and enforcement proceedings as their quality and compliance lead, not as their lawyer, and the documents that hold up share one trait: they were clearly produced to run the business, not to defend a future lawsuit. A slide deck built for a board meeting is not a review record. A vendor's marketing claim that their model was "tested for bias" is not your evidence; it's a claim you now have to verify and file yourself. If your internal review process only exists in someone's head or in scattered emails, you don't have a NIST AI RMF program. You have an intention.

The Four Functions, Built Into an Actual Review Cadence

NIST AI RMF's four functions are commonly treated as a reading list. They work better as the skeleton of a recurring internal review, each with its own owner, cadence, and output.

Govern

This is the charter layer: who owns AI risk decisions, what authority they have, and what happens when they're overruled. Write it down. Name the accountable executive, not a department, and define escalation triggers, for example, any AI system touching a regulated decision like credit, hiring, or a healthcare determination gets escalated to a named committee before deployment, not after a complaint arrives. This is the same discipline ISO/IEC 42001:2023 formalizes in clause 5.3, roles, responsibilities and authorities, and it's the first thing outside counsel asks for when trouble starts: who was supposed to catch this, and did they actually have the authority to stop it?

Map

Before you can manage AI risk you have to know what AI you have. Most companies I evaluate don't have a real inventory, they have the systems someone remembers. Build a living register: every model or AI-enabled feature in production, its intended use, the population it affects, and who owns it. Categorize by consequence, not by technical sophistication. A simple internal chatbot and a resume-screening model are not the same risk tier even if one happens to be more technically advanced.

Measure

This is where most internal reviews go soft, because measurement takes real work: performance testing, bias evaluation against defined criteria, and monitoring for drift after deployment. NIST's Generative AI Profile, released as NIST AI 600-1 in July 2024, extended the Measure function specifically for large language model risks like confabulation and content provenance, and it's worth building your testing criteria around it if any of your systems are generative. Whatever you measure, log the method and the result, not just the conclusion. "Passed bias testing" is not evidence. The test design, the data used, and the actual output are.

Manage

This is the decision layer: given the measured risk, what did you do about it, who approved the residual risk that remained, and what would trigger a re-review. Every AI system with a nontrivial risk tier needs a documented decision, not a default. Silence is not sign-off.

NIST AI RMF, ISO/IEC 42001, and the EU AI Act, Side by Side

Companies often ask me whether NIST AI RMF is enough on its own. The honest answer depends on what you're trying to survive: an internal audit, a regulator's inquiry, or a courtroom. Here's how the three most-referenced AI governance frameworks compare on the question that actually matters, whether they produce evidence a third party can independently verify.

NIST AI RMF ISO/IEC 42001:2023 EU AI Act
Nature Voluntary guidance Certifiable management system standard Binding regulation for in-scope activity
Third-party audit None built in Required for certification, by an accredited body Conformity assessment for high-risk systems
Core structure Govern, Map, Measure, Manage functions Harmonized-structure clauses plus Annex A controls Risk tiers: prohibited, high-risk, limited, minimal
Evidence produced Whatever you choose to document Auditable records under clause 9.2 (internal audit) and 9.3 (management review) Technical file and conformity documentation
Role in US litigation Cited as a reasonable-care benchmark where a legislature adopts one (Colorado's SB 24-205 was the model; it was repealed before taking effect and replaced by SB 26-189, which drops the affirmative defense) Increasingly required in vendor contracts and RFPs Not directly applicable, but shapes global vendor posture

The pattern in that table is the whole argument: NIST AI RMF tells you what good governance looks like, but nothing forces anyone else to check your work. ISO/IEC 42001 does, through clause 9.2's internal audit requirement and the accredited certification audit behind the certificate itself. That's exactly why more companies pair the two, NIST AI RMF for the substance of the risk management, ISO 42001 certification for the third-party proof that it actually happened. If your exposure runs through regulators or litigation rather than internal process improvement, the audit trail is the part that survives scrutiny, not the framework you cite.

Building the Review Process: A Practical Sequence

  1. Charter first. One page: accountable owner, review committee, escalation triggers, sign-off authority. Get it signed by someone with real authority, not drafted and left in a shared drive.
  2. Build the system inventory before you build anything else. You cannot review what you haven't listed.
  3. Assign risk tiers using criteria you'd be comfortable explaining to a regulator, consequence to individuals, reversibility of harm, degree of human oversight in the loop. For example: a consumer-facing chatbot with no decisional authority might land in Tier 1, reviewed annually, while an automated credit-decisioning model lands in Tier 3, reviewed quarterly with mandatory pre-launch sign-off.
  4. Set review cadence by tier: quarterly for high-consequence systems, at minimum annually for everything else, and immediately upon any material model change.
  5. Log every review as a decision record, not meeting minutes: what was assessed, what was found, what was decided, who decided it, and the date.
  6. Run an internal audit against your own program at least annually, modeled on the discipline of ISO 42001 clause 9.2, even if you never pursue certification. Someone who didn't build the process should be the one checking it.
  7. Hold a management review, matching clause 9.3's logic, where leadership actually looks at the audit findings and either commits resources to close gaps or documents why not.
  8. Set a records retention policy that outlasts the shortest statute of limitations you're exposed to, and put a litigation hold procedure in place before you need one, not during the discovery request.

That sequence takes most mid-sized companies two to three quarters to stand up properly. Rushing it produces the paper-trail-that-looks-fake problem I described earlier: reviews that appear all at once, dated suspiciously close together, covering systems that have been in production for years with no prior record.

Where I See Reviews Fall Apart

The failures I see most often aren't about ambition, they're about follow-through. A policy gets written and approved, and then nobody updates the risk register when a new model goes into production six months later. An accountable owner is named on paper but has no actual authority to delay a launch, so the review becomes a formality that happens after the decision is already made. Bias testing gets outsourced entirely to the vendor, and the vendor's report becomes the only evidence of measurement, with no independent verification of what was actually tested.

The failure I find most damaging is treating the review as a compliance artifact rather than an operating discipline. A review process that only produces documents nobody reads until a lawyer asks for them isn't a management system. It's a liability generator with extra steps, because it proves you knew what a proper process looked like and chose not to run it that way.

Frequently Asked Questions

Is compliance with the NIST AI RMF legally required?

No. NIST AI RMF 1.0 is voluntary guidance with no enforcement mechanism of its own. It becomes legally relevant when a regulator, court, or state statute treats it as the benchmark for reasonable AI risk management — which is exactly what Colorado's original AI Act, SB 24-205, did before it was repealed and replaced by SB 26-189 in 2026. The replacement law does not carry forward that affirmative defense, so no state currently codifies a NIST-alignment safe harbor, but the legislative template is still there for the next state that wants to use it.

Does NIST AI RMF require a formal internal audit?

No, the framework itself doesn't mandate an audit function. ISO/IEC 42001:2023 does, through clause 9.2 (internal audit) and clause 9.3 (management review). Many organizations borrow that audit discipline and apply it to their NIST AI RMF program even without pursuing ISO certification, because independent review is what makes the documentation credible later.

How often should an internal AI risk review happen?

Cadence should track consequence, not calendar convenience. High-risk systems, anything touching credit, employment, healthcare, or safety decisions, warrant quarterly review at minimum, with immediate re-review triggered by material model changes. Lower-risk systems can run on an annual cycle, but "annual" only counts if it's documented, not assumed.

What's the practical difference between NIST AI RMF and ISO/IEC 42001?

NIST AI RMF describes what good AI risk management looks like across four functions: Govern, Map, Measure, and Manage. ISO/IEC 42001:2023 is a certifiable management system standard, an accredited certification body audits your program against it and issues a certificate. NIST AI RMF gives you the substance; ISO 42001 gives you third-party proof that the substance is real.

What documentation actually matters if we're investigated or sued?

Contemporaneous decision records beat polished policy documents every time: who reviewed which system, what criteria they applied, what they found, and who signed off on the residual risk. A policy without dated, individually attributable review records reads, to a regulator or opposing counsel, as a policy nobody followed.

Where to Go From Here

None of this requires waiting for a federal AI law that may or may not arrive. NIST AI RMF, applied with real documentation discipline, already gives you a defensible position under general negligence and FTC unfairness theory, and it positions you ahead of whatever state framework legislation comes next; Colorado's now-repealed SB 24-205 shows exactly where those statutes tend to point when they do pass. Where I push clients further is toward ISO/IEC 42001 certification once the stakes of a given AI system justify third-party proof, because a certificate from an accredited body says something a self-authored policy never can. If you're building or auditing an AI governance program and want a second set of eyes on whether it would hold up under real scrutiny, an ISO 42001 gap assessment is the fastest way to find out where the paper trail breaks before someone else finds it for you.

Last updated: 2026-09-20

J

Jared Clark

Principal Consultant, Certify Consulting

Jared Clark is the founder of Certify Consulting, helping organizations achieve and maintain compliance with international standards and regulatory requirements.