Governance ¡ 5 minute read
AI Model Risk Management: Extending MRM to LLMs and Agents
AI model risk management extends the discipline of inventory, risk tiering, independent validation, ongoing monitoring, and governance to LLM applications and agents, with changes for their nature: the unit of validation is the configuration including prompts and retrieval, validation uses golden datasets and calibrated graders, provider updates are material changes, and monitoring covers quality, safety, and cost.
Model risk management gave regulated organizations a discipline for statistical models: know what you have, rate its risk, validate it independently, monitor it, and govern changes. LLM applications and agents fit that discipline in principle and break it in mechanics. Their behavior lives in prompts and retrieval as much as in the model; the model changes when the provider updates it; outputs are open-ended text and actions rather than scores. This guide covers how to extend each MRM component to AI systems, drawing on FISTA Solutions' AI enablement practice. The sector framework is in the AI controls for financial services whitepaper and the register that tracks residual risk in ai risk register. This article is general guidance, not legal or regulatory advice.
How do MRM components map to AI systems?
| Component | Classical MRM | AI systems |
|---|---|---|
| Inventory | Models with owners and uses | Configurations: model, prompts, retrieval, tools, gates, with owners and uses; includes bought AI |
| Tiering | Materiality and complexity | Consequence, autonomy level, data sensitivity, regulation, customer exposure |
| Development standards | Documented methodology and testing | Specification with acceptance criteria; golden datasets; evaluation gates; documentation |
| Independent validation | Conceptual soundness, outcomes analysis | Design review; evaluation by category; adversarial, safety, and fairness testing; challenge of thresholds |
| Change management | Re-validation on material change | Prompt, retrieval, tool, and provider model changes as material |
| Ongoing monitoring | Performance against benchmarks | Sampled quality, drift, safety, gate behavior, incidents, cost |
| Governance | Committee oversight; limits; documentation | Governance board; autonomy limits; model or system cards; audit trails |
What goes in the AI inventory?
Every AI system that informs or makes decisions, produces outputs used in processes, or interacts with customers, whether built or bought, with its owner, purpose, users, the configuration components, data sources, vendors, risk tier, validation status, and monitoring status. Embedded AI features in vendor software belong in the inventory. Inventory practice is in the ai governance checklist.
How should AI systems be tiered?
By consequence if wrong, autonomy level, reversibility of actions, data sensitivity, customer exposure, and regulatory coverage. A drafting assistant for internal memos is low tier; an agent that adjusts customer accounts is high tier; a model informing credit decisions is high tier with sector rules. Tier sets validation depth, monitoring intensity, and governance review. Autonomy mapping is in what is an autonomy level in ai.
How is independent validation performed?
By people independent of the builders, covering design and control review against the specification; evaluation on golden datasets by task category with calibrated graders, checking thresholds are appropriate to consequence; adversarial and safety testing including injection; fairness testing by group for systems affecting people; review of data lineage and permissions; and challenge of the specification and thresholds themselves. Validation reports become part of the system card. Evaluation design is in the ai evaluation checklist and fairness testing in the responsible AI implementation whitepaper.
What counts as a material change?
Changes to the base model or version including provider updates, system or task prompts, retrieval sources or settings, tools or permissions, guardrails, and autonomy levels. Each triggers re-validation proportionate to tier, with the change and its measured effect recorded. Pinning and contract notice terms make provider changes controllable. Change practice is in how to manage ai vendors and deployment gating in what is a canary deployment.
What does ongoing monitoring cover?
Sampled production outputs scored against thresholds by category; groundedness and safety signals; drift in input distributions and output patterns; gate approval and override behavior; incidents; cost per outcome; and user feedback, with thresholds that trigger review and restriction. Monitoring evidence feeds periodic re-validation. Monitoring design is in ai evaluation vs ai monitoring and the observability standard in the ai observability checklist.
How does governance apply?
A governance board approves high-tier systems on validation evidence, sets autonomy limits, reviews monitoring and incidents, and holds owners accountable; documentation follows a system card standard; audit trails allow reconstruction of any output or action; and internal audit provides assurance over the MRM program itself. Board structure is in ai governance board and documentation in what is a model card.
What must change from classical MRM?
Validation cannot rely on statistical benchmarks alone; it needs task-specific golden datasets and calibrated judgment. The model is not the unit; the configuration is. Changes are frequent and some are external. Outputs include actions, so gates and autonomy are part of the model. And cost is a risk dimension, because usage-based pricing can make a successful system unaffordable. MRM programs that recognize these adapt; those that force AI into classical templates produce validation that misses the real risks.
What does adapted MRM look like in practice?
A bank inventories its AI systems including vendor features, tiers a customer servicing assistant as high, and validates the full configuration independently: evaluation by intent category, injection testing, fairness testing on response quality across customer segments, and review of retrieval permissions. Provider updates are pinned and re-validated before adoption. Monitoring samples quality weekly and reports gate behavior. The governance board approves on evidence and reviews quarterly. Examiners receive the inventory, validation reports, and monitoring evidence. The sector detail is in the AI controls for financial services whitepaper.
How FISTA Solutions supports AI model risk management
FISTA Solutions delivers AI systems with the artifacts MRM requires: specifications, golden datasets, evaluation and adversarial results, fairness testing, system cards, change records, and monitoring, and helps clients adapt their MRM programs to configurations, provider changes, and agent actions. The AI enablement practice leads governance and evaluation, AI agents ship validation-ready, and forward deployed engineers embed with client model risk and validation teams. The record behind the approach is 150+ projects with 99.9% uptime.
To extend model risk management to the AI systems you actually run, message FISTA on WhatsApp, or read the AI controls for financial services whitepaper for the complete framework.
Share-ready article cover
Download the generated social format.
Clear answers
Questions raised by this field note.
Straightforward guidance for evaluating scope, fit, and the next step.
01Does model risk management apply to LLM applications?
Yes where they inform or make decisions, produce outputs used in business processes, or interact with customers. Regulators have indicated that existing model risk principles apply to AI, including generative AI, and organizations are expected to inventory, tier, validate, and monitor these systems.
02What is the unit of validation for an AI system?
The full configuration: base model and version, system and task prompts, retrieval sources and settings, tools and their permissions, guardrails, and autonomy levels. Validating the base model alone says little about how the configured system behaves on the organization's tasks.
03How is validation performed?
Through evaluation on golden datasets by task category with calibrated graders, adversarial and safety testing, fairness testing by group where decisions affect people, review of design and controls, and challenge of the specification and thresholds, performed by people independent of the builders.
04How are provider model updates handled?
As material changes: pin versions where possible, require change notice in contracts, re-run validation on every update before adoption, maintain a tested fallback, and record the change with its measured effect. Silent updates that alter behavior are a recognized risk.
05What does ongoing monitoring cover?
Sampled quality scored against thresholds, groundedness and safety signals, drift in inputs and outputs, gate approval behavior, incidents, cost per outcome, and user feedback, with thresholds that trigger review and, where necessary, restriction of the system.
Continue exploring
Related capabilities
Start with the hard problem
Need the outcome owned, not merely analyzed?
Tell us where delivery is constrained. Weâll map the fastest credible path from intent to verified production.