I'm not a technologist. I've spent almost forty years on the operations and sales side of financial services, which is about as regulated as industries get, and that's the lens I've watched every technology wave through. When I started, the fax machine was the most advanced thing in the office. Since then I've seen the internet arrive, the dot-com boom and bust, the move to SaaS and cloud, a blockchain cycle that mostly came and went, and now agentic AI. Every one of them went the same way. The technology showed up long before anyone had worked out how to run it safely.

What I hear from admins and architects these days isn't that the AI isn't good enough. It's closer to this: we can build the agent, but we can't tell you what it did, why it did it, or who signed off on it. That's a governance problem, not a technology problem. Our industry has solved governance problems before. We just haven't had to do it for something that makes its own decisions.

Stop treating agents like flows

When you bring on a new employee, you don't hand them a login and walk away. They get a role, limits on what they can see, and, in a regulated business (the only kind I've ever worked in), a record of what they touched and when. Automations never needed that kind of attention because they were predictable. Same trigger, same result, every time.

Agents don't work that way. Ask one the same question twice and you can get two different answers, both reasonable, depending on the context, the prompt or the model version. If you treat that like a Flow, where you build it, test it once and trust it forever, you're not being careful. You're hoping. Agents should be governed the way we govern people, not the way we govern automations.

I want to be blunt about one thing, because I see it mixed up all the time. This isn't about clean data. A firm can have a spotless data model and tightly scoped permissions and still get burned. The question was never whether the data was good. The question is whether you can prove what the agent did with it and who reviewed it. I've seen firms with excellent technology fail an audit for exactly that reason.

Prompt performance vs. agent performance

An advisor asking an AI to summarize a client's accounts is one thing. An agent that runs every morning, flags missing documents, drafts the client communication and routes it for approval is something else entirely. With the first, you're asking whether the answer was accurate. With the second, you're asking whether a worker did its job correctly, inside the role it was given, every day.

So prompt performance tells you if the AI gave a good answer, and agent performance tells you if the worker did its job over time. You measure that with things like completion rate, exceptions and how often people override it. Most firms govern the answers. Very few are governing the workers yet. If I could add one line to every firm's checklist, it would be this: scope it, log it, measure it, review it.

A scenario I'd bet is already playing out somewhere

Picture an agent drafting account update emails to clients, pulling from Salesforce and a linked document store. It does a good job and nobody complains. A few months in, compliance asks a simple question. Did that agent stay in scope, and did a person review its output before it went to the client? In most industries a bad answer to that is embarrassing. In ours it can be a reportable event, a fiduciary issue or a finding in your next exam. And if the answer is spread across API logs rather than sitting in one place tied to that agent, it's going to take a lot longer to find than you'd like. “We're not sure yet” is not something you want to say to a regulator.

What's Probably Already in Your Toolkit

I'm not the right person to walk you through configuring any of this. That's a conversation for your admin, not your CEO. But I do know which questions to ask:

  • Permission Sets and Permission Set Groups. Can you limit what the agent is allowed to touch the same way you would for a new hire?
  • Field Audit Trail. Does agent activity show up in the same detail as human activity, or is it buried under a generic “API” user?
  • Shield Event Monitoring. If you've licensed it, do you actually have somewhere to go to reconstruct what an agent did?

Ask your admin whether those three are in place. If the answer is “we'd have to check,” you have your answer. Just asking the question puts you ahead of most of the market, whatever platform you're on.

Where that stops being enough

For most day-to-day processes, that's enough. It stops being enough once agents work across several systems, say Salesforce plus a document repository, a compliance archive and an outside data feed. It also stops being enough when seeing what happened isn't good enough and you have to produce it, formatted, on demand. At RIAs and broker-dealers, that isn't a hypothetical. That's a normal Tuesday. Supervisors expect a human checkpoint there.

That's the gap we built Arcus Edge to close. It isn't a general-purpose agent platform. We built it only for SMB wealth management, RIA and broker-dealer firms on Salesforce, so permissions, logging and human review sit in one place instead of being pieced together after the fact.

My advice, for what it's worth

Don't start with the vendor question. Start by asking your own team whether this is already in place. I've watched technology get ahead of governance for forty years. In most industries that gap closes on its own eventually. In a regulated one it doesn't. It turns into something you have to answer for, in an exam, on someone else's schedule. When a deal stalls or a client walks, it's rarely because the AI got something wrong. It's because nobody could answer a simple question about what it did. Ask that question before an examiner does.

Get the Agent Governance Checklist

A practical, vendor-neutral checklist for governing AI agents in Salesforce — scope it, log it, measure it, review it.

Download the free checklist

About the author

Gerry Murphy is CEO and Co-Managing Partner at Arcus Partners. He has spent nearly four decades in financial services and FinTech, working where technology meets the capital markets, including senior leadership roles at SunGard (now FIS) and Fiserv, where he helped build the infrastructure modern financial services runs on. He writes about wealth management, data and artificial intelligence, and what they mean for the firms that have to put them to work.