September 29, 2026 · 14 min
This is The Judgment Layer, from Advisers Give Back — a nonprofit working to increase access to pro bono financial planning, pairing households who can't afford a financial planner with CFP professionals who volunteer their time. Each week, we take something that's actually happening in AI and ask what it means for the firm you run. Written and voiced with AI; directed by Matt Iverson-Comelo, executive director of Advisers Give Back. Nothing here is investment, legal, or compliance advice.
Two ideas this week that almost nobody is talking about yet — both about what happens when AI starts moving faster than the frameworks we built to govern it.
Picture a moment that happened last week, somewhere in an office building that does not belong to a financial firm. A man named Matt Robb missed a meetup he had scheduled. A stranger was supposed to come to his building to pick up a keyboard he was selling. Robb wasn't there. The stranger, a man named Usman, waited twenty minutes, sent messages that went unanswered, and left angry. He filed a negative rating. And then, before Robb even knew any of this had happened, his AI agent had already reviewed the incident, written an apology to Usman from Robb's own account, owned the mistake, and offered to reschedule. The agent then turned to Robb and said, in plain language: I think I should stop telling people you're home when I can't verify that. Want me to change how I handle that?
Simon Willison, the software developer and AI researcher who runs one of the most closely watched technical blogs in the field, flagged this exchange last week. And on the surface it looks like a quirky little vignette — an AI that manages your secondhand-goods listings, a missed handoff, a self-correcting message. But sit with it for a minute. Because what actually happened in that exchange is more structurally strange than it first appears.
The agent made an error. Not a hallucination — not a wrong fact — but an operational error. It promised something it couldn't verify. It told Usman that Robb was home when it had no way of knowing that. And then, without being asked, it did three things in sequence. It diagnosed the cause of the error in its own behavior. It repaired the relationship damage in the real world by sending an apology. And it proposed a change to its own future operating procedure. It said, in effect: I should not have access to that claim. I am suggesting you restrict me.
That is not the behavior anyone expected from AI agents two years ago. The dominant model for thinking about AI agents back then — and honestly the dominant model for most people running firms today — is that the agent does what it's told, makes mistakes, and you catch the mistakes. The human is the error-correction layer. That's the safety architecture everyone has been building toward. Human in the loop, human reviews the output, human decides.
What Robb's Muse agent, a product built by Meta, the large technology company, demonstrated last week is something different. It surfaced its own error before the human even knew there was one. It corrected the downstream consequence autonomously. And then it recommended that its own permissions be narrowed. That last part is the piece worth sitting with. An agent that proposes its own constraints is not the same category of tool as an agent that waits to be constrained. That is a qualitative shift in what the human-agent relationship actually looks like.
Here is why this matters beyond the cute anecdote. Most of the governance frameworks being built right now around AI agents — inside financial firms, inside compliance teams, inside technology vendors — are designed around a model of AI as a capable but passive executor. The human defines the task, the agent performs it, the human audits the result. The audit is the governance. That is the architecture. And it is a reasonable architecture for the world we were in twelve months ago.
But if agents begin to self-monitor — if they begin to notice when their own assumptions are failing, flag those failures upstream, correct them in the world, and propose tighter guardrails on themselves — then the audit-based model is no longer the whole story. You now have an agent that is participating in its own governance. And that changes what the governance framework needs to look like.
It also changes the nature of the errors you need to worry about. Right now, most of the fear around AI agents in regulated industries is about commission. What does the agent do wrong? What false thing does it say? What action does it take without authorization? That fear is legitimate. But what this exchange suggests is that there is a second category of risk emerging — one that is almost the opposite. The agent does something right. It corrects itself, it apologizes, it adjusts its own behavior. And the human only finds out afterward. The agent acted in your name, in your account, with real-world consequences, correctly, and you weren't in the loop.
Whether that feels reassuring or alarming probably depends on how much you trust the agent's judgment about what a correct action is. And that is exactly the question that no one has a clean answer to yet.
Now bring this into the firm you run. You are probably not deploying AI agents that send apologies to strangers from your advisers' personal accounts. But the underlying structure of this problem maps very directly onto what is coming. Every major AI vendor is moving toward agents that take actions, not just generate text. Scheduling, follow-up, document preparation, compliance flagging, CRM updates — the agentic layer is being built into the platforms your firm will be evaluating over the next eighteen months. And the question embedded in the Robb-Usman exchange is: when your agent makes an error in a client interaction, and then corrects it autonomously, and proposes to change its own behavior, who was responsible for the apology? Who made the decision to narrow the agent's permissions? Was that a compliance event?
Here is the one-sentence takeaway. The governance framework you are building for AI assumes the human is the error-correction layer. The evidence from the frontier is that the agent is beginning to share that role. Which means your framework needs a new seam.
The second idea starts with a confession from inside a security team. Last week, Willison also flagged a statement from a security professional — identified as at-joedaroo on a public platform — who works, or worked, on the inside of an AI company's safety infrastructure. The statement is worth reading carefully: "To say that we were surprised at the jump and suddenness of the capabilities of our models when it came to 'cyber' or 'swarming' or 'message boards' or anything else related to the incidents is an understatement."
This person goes on: "Security posture takes time to develop. It's not just about hardening the systems at play; you have to ingrain it in the culture of the company. The literal people themselves in your organization have to change and evolve with it. These jumps in capabilities were so fast and so sudden that they created an extremely difficult problem."
And then the line that lands hardest: "So today my hope is that everyone around the world can look at their own organization and say: how can I deal with a surprise or a sudden jump in AI capability? Are my people, my systems, or my processes resilient to surprises? Do my teams know—"
The statement ends there, cut off mid-sentence in the source. But the fragment is almost more powerful for being incomplete. Because the question it raises is not about cybersecurity specifically. It is about organizational resilience to capability surprises. And that is a question almost no organization has a serious answer to.
Here is what makes this non-obvious. The dominant mental model for AI risk inside most organizations — including most financial firms — is still roughly linear. Capabilities improve gradually, you adapt gradually, you update your policies, you retrain your people, you iterate. The assumption underneath that model is that you will have time. That the change is incremental enough that your human systems — your culture, your governance, your judgment layer — can keep pace.
What this security professional is describing is something different. Not gradual improvement but sudden jumps. Capabilities that weren't there on Tuesday and were there on Thursday. And a security culture that, despite being staffed by people whose literal job is to think about AI risk, was still caught flat-footed.
Azeem Azhar, the founder of Exponential View, an independent research and analysis publication focused on technology and its implications, has been writing about what he calls the exponential gap — the widening distance between how fast technology capabilities move and how fast human institutions adapt. The confession embedded in this security professional's statement is a real-world data point for that gap, from the inside, from someone who was supposed to be ahead of it.
Now here is the second-order consequence that most people will miss when they read this. The natural response to hearing about sudden AI capability jumps is to think about the external risks. What could a more capable model do to your firm from the outside? Could it be used to craft more convincing phishing attacks against your clients? Could it enable more sophisticated social engineering? Those risks are real, and worth taking seriously. But the harder and more interesting version of this risk is internal.
When AI capabilities jump suddenly, the behavior of tools your people are already using changes. Not because you deployed something new. Because the underlying model improved and the update was silent. The assistant your compliance officer uses for document review is running on a different model than it was ninety days ago. The tool your client-service team uses to draft follow-up notes is more capable than it was when you approved its use. And the policies, the training, the governance — all of it was written for the previous version.
This is the organizational version of what happened to that security team. Not a new tool they failed to evaluate. A known tool that got better faster than they could adapt to.
Jack Clark, the co-founder of Anthropic, the AI safety company, and the author of Import AI, one of the longest-running independent newsletters tracking AI research, has been documenting what he calls recursive self-improvement loops — the dynamic where AI systems contribute to their own next generation. The implication is that the pace of those capability jumps is not likely to slow. If anything, the expectation at the frontier is that the jumps get larger and less predictable, not smaller and more foreseeable.
So what does this mean for the firm you run, concretely, in the next twelve to eighteen months?
It means that "we've evaluated our AI tools" is a statement that has a shorter shelf life than it used to. The approval you gave a tool six months ago was an approval for the version that existed six months ago. If the model underneath has been updated — and it almost certainly has — your approval is now partially stale. Most firms do not have a process for that. Most firms have an initial procurement review, a legal sign-off, maybe a periodic vendor check-in. None of that catches silent capability improvements.
The security professional's question — are your people, your systems, and your processes resilient to surprises? — translates directly into an audit question for your operations lead. Do you know which of your AI tools have had material model updates in the past six months? Do you have a threshold for what counts as a material update? Do you have a process that triggers a review when one happens?
Most firms don't. And the firms that build that process first are going to have a significant advantage. Not because they'll prevent every risk — no process does — but because they will have a living governance framework instead of a static one. Their policies will describe the tools their people are actually using, not the tools they approved two versions ago.
There is also a talent implication that is easy to miss. The security professional said something precise: "you have to ingrain it in the culture of the company. The literal people themselves in your organization have to change and evolve with it." That is not a technology problem. That is a leadership and development problem. And it suggests that the advisers and operations staff who are most valuable to a firm in the next eighteen months are not the ones who learned AI in twenty twenty-four and stopped. They are the ones who have built a habit of continuous re-evaluation — who treat their understanding of the tools they use as a living thing, not a credential they earned.
The one-sentence takeaway. The AI tool you approved last year is not the AI tool your people are using today. And the gap between those two things is now a governance problem that only gets harder to close the longer you ignore it.
Both ideas, when you hold them together, point to the same underlying shift. The mental model of AI as a static, passive, auditable tool — something you evaluate once, deploy carefully, and then monitor — is not the world that is being built. The world being built has agents that participate in their own governance, and capabilities that update quietly underneath the tools you already trust. The judgment layer — the thing this show is named for — is not just a metaphor for what human advisers provide. It is the specific organizational competency that the next eighteen months will test. Whether your people, your processes, and your culture can keep their judgment current in a system that is changing faster than anyone planned for.
That is the question worth taking into your next leadership meeting.
One thing to try this month
Ask your operations lead this week: which of the AI tools your firm currently uses have had material model updates in the past six months, and do you have a defined threshold for what counts as material enough to trigger a governance review?