Agentic AI is moving into areas of banking where the consequences of an incorrect decision are much greater than they are in a typical enterprise workflow. Risk teams are beginning to explore AI agents that can investigate alerts, connect information across systems, identify relationships between entities, recommend next steps, and prepare regulatory documentation. For organizations that have spent years building rigorous controls around financial crime and risk operations, that creates an important question: how much autonomy can an institution give an AI system while maintaining appropriate oversight?
For Chief Risk Officers, Chief Compliance Officers, Model Risk Management leaders, BSA/AML Officers, and Internal Audit teams, the answer will depend less on whether an agent uses generative AI and more on how the agent is designed, governed, and deployed.
This distinction matters because an AI agent is different from a traditional model operating within a defined process. An agent can interact with multiple data sources, use different models or tools, interpret information, and determine what action to take next within the boundaries it has been given. Its behavior is therefore influenced by the broader workflow surrounding the model, not simply by the model itself.
That creates new governance considerations, but it does not mean financial institutions need to start from scratch. Existing model risk management, compliance, and audit principles provide a foundation. The challenge is extending those principles to account for the additional components that make an agentic system work.
For financial institutions considering agentic AI, the objective should be to increase the amount of work technology can handle while maintaining clear accountability for the decisions that matter.
Agentic AI Is Changing How Risk and Compliance Work Gets Done
Moving From Individual Model Outputs to End-to-End Workflows
Traditional AI deployments in financial crime often focus on a particular point in the workflow. A model may identify a suspicious transaction, assign an alert score, detect an unusual device, or identify a potentially suspicious relationship between accounts.
An AI agent can operate across several of these steps. It may retrieve customer information, examine transaction history, identify related entities through a Knowledge Graph, compare activity against known risk patterns, summarize relevant evidence, and recommend whether a case should be escalated.
That difference matters from a governance perspective because the risk associated with an agent comes from the interaction between several components. The underlying model is one part of the system. The data it accesses, the tools it can use, the instructions governing its behavior, the actions it is authorized to take, and the human controls surrounding those actions also influence the overall risk.
For banking risk leaders, this means governance needs to follow the complete workflow rather than stopping at the model boundary.
Where Agentic AI Fits Into Fraud and AML Operations
The most immediate applications are likely to appear in operational areas where analysts already spend significant time gathering information and performing repetitive investigative work.
An agent could assemble information from transaction systems, customer profiles, device intelligence, case management platforms, and network relationships before presenting an analyst with a consolidated view of the case. It could identify relevant signals, explain why those signals matter, and prepare a draft investigation summary.
In AML operations, similar capabilities could support alert triage, investigation preparation, entity research, and AI SAR generation. In fraud operations, agents could help investigate account takeover, payment fraud, application fraud, scams, and coordinated fraud networks.
The appropriate level of autonomy will depend on the use case. An agent that summarizes information for an analyst presents a different risk profile from one that recommends an account restriction or initiates an external filing.
Why Greater Autonomy Creates a Different Risk Profile
As an AI system takes on more responsibility, the institution needs greater visibility into what the system can access, what it can produce, and what actions it can take.
A governance framework should therefore distinguish between agents that provide information, agents that make recommendations, and agents that can execute actions. The controls surrounding each category should reflect the potential impact of an error.
This risk based approach is particularly important as regulatory expectations around AI continue to develop. Current federal model risk guidance has emphasized risk based model governance while also recognizing that generative and agentic AI are rapidly evolving areas requiring additional regulatory work. Banks should therefore avoid treating an agent as automatically covered or excluded based solely on its label and instead assess the specific risks created by how the technology is being used.
Model Risk Management Needs a Broader View of the AI System
Applying SR 11-7 Principles to Agentic AI Governance
SR 11-7 provides a foundation for model risk management through expectations around model development, implementation, use, validation, governance, policies, and controls. The framework is intended to be applied in a manner proportionate to the institution's size, complexity, and use of models.
For agentic AI, that means the governance process should consider how the agent operates within the broader system. Institutions should establish how the agent is tested, validated, monitored, documented, and approved for its intended use.
Calling an agent “SR 11-7 compliant” oversimplifies the issue. The more useful question is whether the institution has incorporated the risks associated with that agent into its existing model risk management and governance framework.
Looking Beyond the Underlying Model
An AI agent may rely on a foundation model or another machine learning model, but that model does not operate in isolation.
The agent's instructions, retrieval mechanisms, connected data sources, tools, permissions, decision thresholds, and escalation rules can all affect its behavior. A change to any of these components could materially alter the outcome of an investigation.
Model Risk Management therefore needs visibility into the complete architecture. Documentation should explain what the agent is designed to accomplish, which models and services it relies on, what information it can access, and which actions it is permitted to perform.
Defining What an AI Agent Is Allowed to Do
Clear boundaries are one of the most important controls for agentic systems.
Before deployment, institutions should define which tasks an agent can perform independently and which activities require human approval. An agent may be authorized to retrieve records and summarize evidence automatically while requiring analyst approval before closing an investigation or generating a final SAR.
These permissions should be explicit rather than assumed. They should also be reviewable as the agent evolves and as its responsibilities expand.
Classifying Agents Based on Their Risk and Level of Autonomy
Not every AI agent requires the same degree of governance.
An internal research assistant that retrieves policy information presents a different risk from an AML investigation agent that recommends case dispositions. An agent that can take action in a payment or account system introduces another level of operational risk.
Banks can establish an agent inventory that records each system's purpose, level of autonomy, data access, connected tools, business owner, model dependencies, validation status, and required human approvals. This creates a practical foundation for applying governance proportionately.
Human Oversight Must Be Built Into the Workflow
Deciding Where Human Judgment Is Required
Human oversight should be designed around the decisions that carry material regulatory, financial, or customer impact.
For example, an agent may be highly effective at collecting evidence and organizing an investigation but less appropriate as the sole decision maker for a complex case involving multiple jurisdictions or ambiguous customer activity.
The objective is to place human judgment where it adds the most value while allowing the agent to handle the work that can be performed consistently and efficiently.
Automating Investigation Work Without Automating Accountability
Automation can reduce the amount of manual work required to investigate an alert, but responsibility for the resulting decision remains with the institution.
An analyst should be able to see what the agent did, what evidence it considered, what recommendation it produced, and what information was missing or uncertain. This allows the analyst to challenge the output rather than simply accepting it.
For compliance leaders, this distinction becomes particularly important when AI is used to support regulatory processes. Human in the loop AML should mean more than placing an approval button at the end of an automated workflow. The human reviewer needs sufficient information and context to exercise meaningful judgment.
Escalating High Risk and Uncertain Cases
Agents should have defined escalation conditions for situations where confidence is low, evidence conflicts, or the potential impact of an incorrect decision is high.
A well designed workflow can allow an agent to recognize when a case falls outside its operating boundaries and route it to the appropriate analyst or specialist.
This provides a practical control against overreliance on automation. The agent does not need to resolve every case. It needs to recognize when human expertise is required.
Keeping Humans Accountable for AI Generated SARs
AI SAR generation illustrates the importance of maintaining human accountability.
An agent may be able to organize transaction information, identify relevant activity, and prepare a draft narrative. The final filing decision and submission should remain subject to the institution's established compliance controls and appropriate human review.
The analyst should also be able to verify the factual basis of the narrative against source records. This requires more than a well written draft. It requires evidence that allows the reviewer to determine whether the statements generated by the system are accurate and appropriately supported.
Explainability Should Show How the Agent Reached Its Recommendation
Why a Risk Score Alone Is Not Enough
Traditional model outputs often center on a score or classification. That can be useful for prioritization, but an agentic workflow requires a more complete explanation.
If an agent recommends escalating a case, the analyst needs to understand what information contributed to that recommendation and how the evidence was interpreted.
Explainable AI in banking therefore needs to extend beyond displaying a model score. The explanation should provide enough context for an experienced reviewer to understand and challenge the recommendation.
Identifying the Signals That Influenced the Agent
For a financial crime investigation, relevant signals could include unusual transaction velocity, device anomalies, shared IP infrastructure, unexpected changes to customer information, relationships between accounts, or activity associated with previously identified entities.
The agent should be able to identify which signals influenced its recommendation and distinguish between direct evidence and contextual information.
This gives analysts a clearer basis for determining whether the recommendation is supported by the facts of the case.
Giving Analysts a Clear View of the Agent’s Reasoning
A useful reasoning record should explain the sequence of investigative steps without forcing analysts to interpret technical model behavior.
The goal is not to expose proprietary model internals. The goal is to provide an operationally meaningful explanation of what the agent considered, what it found, and why that information led to its recommendation.
For risk and audit teams, this creates a much more useful governance artifact than a generic statement that an AI system “flagged” a case.
Connecting AI Recommendations to Supporting Evidence
Every material conclusion should be connected to the underlying evidence whenever possible.
If an agent identifies a relationship between accounts, the analyst should be able to trace that relationship to the relevant records. If the agent references a transaction pattern, the underlying transactions should be identifiable. If it prepares a SAR narrative, the claims within that narrative should be traceable to source information.
This connection between reasoning and evidence is fundamental to building confidence in AI assisted investigations.
Data Lineage Becomes More Important as Agents Access More Systems
Knowing What Information the Agent Retrieved
Agentic systems can potentially access a much broader range of information than a traditional model.
That makes data lineage increasingly important. Institutions should maintain visibility into which systems an agent accessed, what information it retrieved, and when that information was used.
Data access should also be governed through appropriate permissions. An agent should have access to the information necessary for its approved use case rather than unrestricted access to enterprise data.
Tracing AI Generated Statements Back to Source Records
When an agent produces a summary or recommendation, analysts and auditors should be able to identify the records that support material statements.
This is particularly important for AI generated compliance documentation. A reviewer should not have to reconstruct the agent's research manually to determine whether a statement in an investigation summary or SAR draft is supported by the underlying evidence.
Preserving Context Across the Investigation
Investigations often develop over time as new information becomes available.
An auditable agentic workflow should preserve the context that influenced the investigation at each stage. This includes relevant data, agent outputs, analyst decisions, escalations, and changes to the case.
Without that historical context, reconstructing why a decision was made months later becomes significantly more difficult.
Auditability Requires a Record of the Entire Agentic Workflow
What Should Be Captured in an Agentic Audit Trail
An audit trail should provide enough information to reconstruct the agent's activity.
Depending on the use case, this may include the inputs provided to the agent, models and versions used, prompts or instructions, data sources accessed, tools invoked, outputs generated, recommendations made, and human decisions that followed.
The appropriate level of logging should reflect the materiality and risk of the use case.
Recording Analyst Decisions and Overrides
Human decisions are an important part of the audit record.
If an analyst accepts an agent recommendation, that decision should be captured. If the analyst overrides it, the system should preserve the override and, where appropriate, the reason for the disagreement.
These records can become valuable governance data over time. Repeated analyst overrides may indicate that an agent's recommendations need to be recalibrated, its instructions need to be changed, or its use case needs to be reconsidered.
Making Historical Investigations Reconstructable
A strong audit trail should allow Risk, Compliance, and Internal Audit to reconstruct how an investigation progressed without relying on individual recollection.
This becomes especially important when an institution is asked to demonstrate how AI influenced a material decision. The organization should be able to show what the system knew at the time, what it produced, what the human reviewer saw, and what decision was ultimately made.
Using Audit Data to Identify Agent Performance Issues
Audit logs are also useful for ongoing governance.
Institutions can analyze patterns in agent behavior to identify recurring errors, unexpected outputs, frequent escalations, analyst overrides, or changes in performance.
This turns auditability into an operational control rather than a record that only becomes useful during an examination.
AI Governance Cannot Stop at Deployment
Monitoring Agent Performance in Production
Validation before deployment is only one part of effective governance.
Once an agent is operating in production, institutions should monitor whether it continues to perform as intended. Relevant measures may include recommendation accuracy, escalation rates, analyst override rates, processing times, data quality issues, and other metrics appropriate to the use case.
Monitoring should also account for changes in the environment in which the agent operates. Fraud patterns, customer behavior, transaction channels, and regulatory requirements can all change over time.
Detecting Changes in Agent Behavior
Agentic systems can be affected by changes to models, prompts, data sources, tools, and connected systems.
A governance framework should therefore establish mechanisms for identifying material changes in behavior. A system that begins producing materially different recommendations should trigger investigation even if no formal model version has changed.
Turning Analyst Feedback Into a Governance Signal
Analyst feedback provides an important source of information about real world agent performance.
When analysts consistently disagree with recommendations in a particular scenario, that pattern can reveal limitations that may not have appeared during initial validation.
Capturing and analyzing this feedback allows Model Risk Management and business owners to identify where the system needs improvement and whether changes require additional testing or approval.
Managing Changes to Models, Prompts, Data, and Tools
Change management becomes more complex when an agent depends on multiple components.
A governance process should identify which changes are material, who must approve them, what testing is required, and when the system needs to be revalidated.
This includes changes to underlying models as well as prompts, retrieval mechanisms, connected tools, data sources, permissions, and agent workflows.
A Consistent Governance Framework Makes Agentic AI Easier to Scale
Establishing Ownership Across Risk, Compliance, and MRM
Agentic AI governance cannot sit entirely within a technology organization.
Technology teams may own the implementation, while business teams own the use case and Model Risk Management evaluates model risk. Compliance and Legal may provide oversight for regulated processes, while Internal Audit independently evaluates the effectiveness of the control environment.
Clearly defined ownership prevents gaps from emerging between these groups.
Matching Controls to the Use Case
Governance should be proportional to the consequences of failure.
An agent supporting low risk research may require relatively limited controls. An agent supporting fraud decisions, AML investigations, or regulatory reporting requires more extensive oversight.
This risk based approach allows institutions to move forward with practical use cases without creating the same governance burden for every AI deployment.
Governing Third Party Models and AI Services
Many enterprise AI systems will rely on third party foundation models, AI services, or external data providers.
Vendor oversight should therefore extend to the components that influence the agent's behavior. Institutions should understand how those services are updated, what data they retain, how performance is evaluated, and what controls exist around security and availability.
Third party dependencies should be documented as part of the overall AI governance framework.
Giving Risk and Audit Teams Shared Visibility
A shared governance environment can make it easier for Risk, Compliance, MRM, and Internal Audit to work from the same information.
Rather than maintaining separate documentation across different teams, institutions can establish a common record of agent ownership, use cases, permissions, validation status, performance monitoring, changes, and audit evidence.
That shared visibility becomes increasingly valuable as the number of AI systems across the organization grows.
Questions Risk Leaders Should Answer Before Deploying an AI Agent
What Is the Agent Authorized to do?
Every agent should have a clearly documented scope of authority.
Risk leaders should understand which tasks the agent can perform independently, which activities require approval, and which actions are prohibited entirely.
What Decisions Must Remain With a Human?
The institution should identify decisions where human judgment is required and design the workflow so that those controls cannot be bypassed through automation.
What Data and Systems Can the Agent Access?
Access should be explicitly defined and appropriately controlled. Risk leaders should understand not only what information the agent needs but also what systems it can interact with and what actions it can perform within those systems.
Can Every Material Recommendation Be Traced to Evidence?
An analyst should be able to determine why an agent made a material recommendation and identify the evidence supporting that recommendation.
Can the Institution Reconstruct the Agent’s Actions?
The organization should be able to recreate the relevant history of an investigation, including the information available to the agent, its outputs, the actions it took, and the decisions made by human reviewers.
How Will the Agent Be Monitored and Revalidated?
Governance should define what performance indicators will be monitored, what constitutes a material change, and when additional validation or approval is required.
What Happens When an Analyst Disagrees With the Agent?
Disagreement should be treated as a normal part of a controlled workflow rather than as an exception to the process. Analysts need a clear mechanism for overriding recommendations, documenting their judgment, and escalating recurring issues.
Building a Responsible Path to Agentic AI Adoption
Start With Governed Use Cases
Banks do not need to begin with the most autonomous applications.
Investigation summarization, evidence gathering, case preparation, and analyst research can provide practical starting points while giving institutions an opportunity to establish governance processes around agentic systems.
As confidence and governance maturity develop, institutions can evaluate use cases involving greater levels of autonomy.
Establish Controls Before Increasing Autonomy
The controls surrounding an agent should mature alongside its capabilities.
Before expanding an agent's permissions, institutions should understand how it performs under normal conditions, how it behaves when information is incomplete or contradictory, and how effectively human reviewers can intervene.
This creates a measured path toward greater automation without allowing autonomy to outpace governance.
Treat Governance as an Operating Capability
Agentic AI governance should ultimately become part of how the institution manages technology and operational risk.
That means maintaining an inventory of agents, assigning ownership, documenting use cases, monitoring performance, managing changes, preserving audit trails, and periodically reassessing whether each agent remains appropriate for its intended purpose.
When these capabilities are built into the operating model, governance becomes easier to scale as AI adoption expands.
Ready to Explore Agentic AI for Risk and Compliance?
Agentic AI can give risk and compliance teams a way to handle growing investigation volumes while maintaining the oversight expected in highly regulated environments.
DataVisor Vera AI Agents combine AI driven investigation capabilities with Knowledge Graph intelligence, human in the loop controls, explainability, and auditability to support more efficient risk operations.
Download The CRO & CCO's Guide to Agentic AI to explore the governance considerations, operating model, and controls banks should evaluate before deploying AI agents.
Ready to see how agentic AI can fit into your risk and compliance operations? Request a demo of Vera AI Agents.
FAQ Section
What is agentic AI in banking?
Agentic AI refers to AI systems that can carry out multiple steps within a defined workflow rather than producing a single prediction or recommendation. In banking, an AI agent may gather information from multiple systems, analyze transaction and customer activity, identify relevant relationships, summarize evidence, recommend an action, and route the case for human review. The level of autonomy can vary significantly depending on the use case and the controls established by the institution.
How does SR 11-7 apply to agentic AI?
SR 11-7 provides a foundation for model risk management around the development, implementation, use, validation, governance, policies, and controls associated with models. For agentic AI, banks should consider how the underlying models and the broader agentic workflow fit within their existing model risk management framework. This includes evaluating the agent's intended use, level of autonomy, data access, outputs, decision impact, validation requirements, monitoring, and controls.
Are AI agents SR 11-7 compliant?
An AI agent should not be considered inherently “SR 11-7 compliant.” SR 11-7 is a model risk management framework that applies at the institution level, and the appropriate controls depend on how an AI system is developed, implemented, and used. Banks should assess the risks associated with each agent and determine how it should be governed within their existing model risk management program.
Does agentic AI need model validation?
Whether and how an agent requires model validation depends on its architecture, use case, materiality, and risk. A bank should evaluate the models and other components that materially influence the agent's behavior, including data, prompts, retrieval mechanisms, tools, and decision logic where appropriate. Validation should be proportionate to the potential impact of the system and should be supported by ongoing monitoring after deployment.
What is human in the loop for AML?
Human in the loop AML means that qualified personnel retain meaningful oversight over AI assisted compliance activities and decisions. An AI agent may gather evidence, prioritize alerts, summarize customer activity, or prepare an investigation or SAR draft, while a human remains responsible for reviewing material conclusions and making decisions that require professional judgment. Effective human oversight requires sufficient visibility into the evidence and reasoning behind an agent's recommendation.
Can AI agents generate SARs?
AI agents can assist with SAR generation by gathering relevant transaction information, identifying patterns, organizing supporting evidence, and preparing a draft narrative for analyst review. The institution should maintain appropriate human oversight over the final determination and filing process. AI generated SAR content should also be traceable to the underlying records so that compliance personnel can verify the accuracy of material statements before submission.
What does explainable AI mean for banking risk teams?
Explainable AI in banking means providing risk and compliance personnel with enough information to understand how an AI system reached a recommendation or conclusion. For an agentic workflow, this can include the signals considered, data sources accessed, investigative steps performed, evidence supporting the recommendation, and relevant uncertainty or limitations. A risk score by itself may not provide enough context for an analyst, auditor, or examiner to evaluate an AI assisted decision.
What should an AI agent audit trail contain?
An agentic AI audit trail should capture enough information to reconstruct how the system operated during a material workflow. Depending on the use case, this can include the data and sources accessed, models and versions used, prompts or instructions, tools invoked, outputs generated, recommendations made, analyst decisions, overrides, escalations, and relevant system changes. The level of detail should be proportionate to the risk and materiality of the agent's use.
How can banks make AI agents auditable?
Banks can make AI agents auditable by maintaining a record of the agent's inputs, actions, outputs, evidence, and human decisions throughout the workflow. Data lineage should connect material conclusions to their underlying source records, while versioning and change management should document changes to models, prompts, tools, and data sources. This allows Risk, Compliance, Model Risk Management, and Internal Audit to reconstruct an investigation or decision after the fact.
What controls should banks put around AI agents?
Controls should reflect the agent's intended use and level of autonomy. Common controls include defined permissions, restricted data access, human approval requirements, escalation rules, monitoring, validation, change management, audit logging, and clear ownership. Higher risk agents may require additional controls where they influence material financial, regulatory, customer, or compliance decisions.
How should banks classify AI agents by risk?
Banks can classify agents according to factors such as the materiality of their use case, level of autonomy, types of decisions they influence, data they can access, systems they can interact with, and potential impact of an incorrect output. An agent that summarizes information for an analyst generally presents a different risk profile from one that recommends account restrictions, influences transaction decisions, or supports regulatory reporting.
Who should own agentic AI governance in a bank?
Agentic AI governance should involve multiple functions rather than being owned exclusively by Technology. Business teams may own the use case, Model Risk Management may evaluate model risk, Compliance may oversee regulated applications, Technology may manage the underlying infrastructure, and Internal Audit may independently assess the control environment. Clear accountability across these functions is essential as agents become more deeply integrated into risk and compliance operations.
How should banks monitor AI agents after deployment?
Banks should monitor agents throughout their operational lifecycle rather than relying solely on predeployment testing. Monitoring can include recommendation accuracy, analyst override rates, escalation frequency, data quality, unexpected behavior, performance changes, and other measures appropriate to the use case. Banks should also establish thresholds that trigger investigation, remediation, additional validation, or governance review when an agent's behavior changes materially.
What happens when an analyst disagrees with an AI agent?
Analysts should have a clear and documented ability to override an agent's recommendation when they determine that it is incorrect or insufficiently supported. The disagreement and resulting decision can also provide valuable governance information. Patterns of repeated overrides may indicate that the agent requires additional testing, improved instructions, different data, or a narrower scope of use.
How does agentic AI change banking model governance?
Agentic AI expands the governance conversation beyond the underlying model. Banks need to consider the full system through which an agent retrieves information, interprets evidence, uses tools, generates recommendations, and interacts with human reviewers. Banking model governance therefore needs to account for the agent's architecture, autonomy, data access, connected services, change history, monitoring, and human controls alongside traditional model risk considerations.






