Summary
Agentic AI capabilities range from simple single-task prompts to open ecosystems where users build and deploy their own agents under governance guardrails. A four-level deployment framework helps firms evaluate where each workflow sits on that spectrum and calibrate the oversight, infrastructure, and rollback discipline that each tier demands.
Agentic AI capabilities aren’t always alike. Some only need a simple on-off deployment decision, while others take careful thought. The core factors are how much autonomy an agent has and what controls you have to build around it. They are a function of the agent’s intended outcome, the complexity of the tasks it completes, and the point in the loop where the human enters.
“53% of financial services executives reported their organizations are actively using AI agents in production, with 40% saying they have already launched more than ten. The complexity of these AI agents can span a broad spectrum. On one end are single-task agents, such as gen AI-powered digital assistants. On the other end are highly sophisticated, multi-agent systems that can work together to perform complex tasks, taking actions on behalf of users under their supervision.” — Googlei
To deliver outcomes while managing risk effectively, it’s helpful to think in terms of a four-level hierarchy, from simple to most advanced.
The simplest deployment is when a human asks the agent to do one thing, reads what comes back, and decides whether to use it. The interaction starts with a prompt and ends with an output without chaining. Humans are responsible for oversight throughout.
For example, an ops analyst receives rate reset notices from a loan administrator in PDF and asks an agent to parse them, extracting the key terms into a structured record for their loan system. The analyst reviews the extraction, confirms accuracy, and pushes it to their platform. The agent is fast and thorough. To control operational risk, a human analyst decides whether the output is right before recording it.
Almost every firm deploying AI has at least Level 1 capabilities, though Level 1 interactions aren’t necessarily agentic in the full sense. They can be as informal as having an enterprise account with a mainstream model provider. However, deployment considerations still play a part. Firms should aim to standardize and govern how these single-task interactions work. Otherwise, an undocumented prompt run ad hoc by whoever needs it is a dependency on individual judgment that is more challenging to audit or reproduce.
Doing this also creates a mindset of deployment discipline. If you can’t describe or manage how a single agent handles a single task, you can’t govern them at higher levels of complexity and risk.
Expanding the scope of an agent demands more deliberate deployment planning. As you advance, agents may execute a fixed sequence of steps, which humans review at defined checkpoints before the workflow advances. The process is controlled because the agent’s builder designs the sequence in advance, without giving the agent any latitude on whether or how to complete steps.
This level would apply to something like a cash break triage agent. Following a standard workflow, it moves through steps like pulling the break record, fetching the counterparty confirmation, and proposing a classification and probable cause. Then it delivers that output for a reconciliations analyst to approve to resolve the issue. The agent has no room for interpretation and can’t take action. It follows predetermined steps, and the analyst owns the transition from diagnosis to resolution.
This is where many firms sit today, and where many believe the technology should stay for now. Deciding where to place the human gates is a balance between too many checkpoints slowing the workflow below manual pace, and too few raising a higher risk of errors. It also means tracing and auditing every step (what the agent retrieved, what it proposed, who approved) because you need to know which step failed when outcomes are incorrect. Step-level observability provides the foundation for governing multiple agents passing work to one another, which is the next level of sophistication.
We’re seeing advanced adopters experimenting with deploying multiple agents, each using the relevant skills and tools to coordinate workflows between them. The outcomes are still controlled because the domain expert managing the agents determines the handoff sequence. The agents don’t negotiate with each other or change the sequence, and the human reviews at each handoff.
This level fits more complex scenarios, such as a voluntary corporate action flow. Key pieces of the workflow are appropriate for agent development. One might ingest notices and normalize the event details. Another that’s meant for determining entitlements would calculate positions and elections by fund. And a third agent could support communications by drafting client-facing election memos. Each has its own skills and tools, producing an output that the next one consumes. An ops reviewer is responsible for stepwise signoffs before the next agent picks up.
Such deployments take different shapes. Some firms add specialized agents around a Level 2 workflow that is now ready to expand beyond a single agent. Others use dedicated workflow engines that orchestrate purpose-built agents through a fixed, predetermined sequence. Either way, the architecture only holds if each agent’s reasoning and service calls are cleanly bounded, so you can upgrade or trace a failure to one agent without touching the others.
We’re seeing strategic interest and sporadic pilots at this level. As agent reliability increases, the human role changes. The idea would be a routing layer that decides which agents to invoke and in what order, where humans only step in if it breaches tolerance or confidence thresholds. The sequencing of decisions happens at runtime.
In theory, this level would allow agentic AI to take on entire workflows, such as post-trade exceptions. A router agent would receive breaks, failed settlements, margin-call mismatches, and issuer or admin notice discrepancies. After it evaluates each exception, it would invoke the appropriate specialist agents for reconciliations, collateral, pricing, tax lots, and more.
Agents coordinating at this level also need a shared protocol for passing work between them. Emerging standards like Model Context Protocol (MCP) and Agent-to-Agent (A2A) define how agents expose capabilities and exchange structured context, so a reconciliations agent can receive a break from the router with enough information to act without a human translating between them. Most of what this level requires is still ahead of what’s feasible today, which is why the deployments we see are still early-stage pilots. The router has to be built, versioned, monitored, and rolled back when it makes the wrong call — and that means solving for governance and observability, integration with systems of record, change management as specialist agents get upgraded, and graceful failover when the router takes a path no one expected.
Levels 1 through 4 don’t exhaust the question: the next frontier is an authorship story. As agent reliability increases and tooling matures, ops leads, fund controllers, and portfolio managers will build and deploy their own agents. The governance infrastructure it requires — identity controls, cost attribution, lifecycle management, policy boundaries — requires a thoughtful, coherent system. Firms that are serious about agentic AI should already be thinking about that layer before the authorship tools arrive and recreate exactly the bespoke Excel and key-person risk scenario they spent years moving away from.
Vera Shulgina
Vera is responsible for Arcesium's data strategy with a focus on driving value for clients through data solutions and data partner integrations.
Sources:
[i] Google, 2025. https://cloud.google.com/transform/new-research-shows-how-ai-agents-are-driving-value-for-financial-services
No spam. Just the latest releases and tips, interesting articles, and exclusive interviews in your inbox every week.