← All episodes

September 22, 2026 · 14 min

The Hollow Throughput

0:00 / 14:01

Transcript

This is The Judgment Layer, from Advisers Give Back — a nonprofit working to increase access to pro bono financial planning, pairing households who can't afford a financial planner with CFP professionals who volunteer their time. Every episode starts with a real development in AI and ends somewhere useful for the firm you run. Written and voiced with AI. Directed by Matt Iverson-Comelo, executive director of Advisers Give Back. Nothing here is investment, legal, or compliance advice.

Two ideas this week — both about what happens to human judgment when AI starts doing the work, and why the answer is not what most people think.

Picture a software engineer — call her a mid-level developer at a company she joined three weeks ago. She's sitting at her desk at eleven o'clock at night, not because she wants to be, but because the pressure to ship is relentless. And here is what her day actually looks like: she opens a ticket, she describes the problem to an AI coding assistant called Claude Code, she reads the output, she presses a button, she moves to the next ticket. She does not read the code. She is not sure her manager reads the code. She is not sure anyone does. She posted about this anonymously online last week, and what she said landed like a stone in still water. "Nobody knows anything here," she wrote. "Everyone, literally everyone, from an L1 to an L7 engineer here is doing the same thing. Talk to Claude."

Simon Willison, the developer and writer who runs one of the most careful technology observation blogs on the internet, flagged this post last week. He filed it under a tag he uses sparingly: "ai-misuse." That tag is doing a lot of work. Because here is what makes this story non-obvious: it is not a story about laziness. It is not even, at bottom, a story about AI replacing engineers. It is a story about what happens to an organization when the people inside it stop reading the work — and how fast that happens, and how invisible it is from the outside.

Think about what reading the work actually does in a knowledge organization. It is not just quality control. Reading is how you understand what you actually have. A specification, a plan, a piece of analysis — the act of a human moving through it carefully is also the act of that human forming a model of what the organization knows, what it has committed to, what the risks are. When you read a client plan, you are not just checking for errors. You are building the internal map that lets you give advice six months from now when circumstances change. The reading is the knowing.

What the engineer described is a firm that has outsourced the reading. The AI generates the work, and the humans move it forward without internalizing it. And from the outside — to management, to clients, to regulators — the outputs look normal. The code ships. The tests pass. The tickets close. The throughput numbers are excellent. What is missing is entirely internal, invisible, and accumulating.

Ethan Mollick, the Wharton professor who studies how people actually use AI and writes the newsletter One Useful Thing, published a piece this week about what he calls "the overhang" — the idea that most of us are sitting on enormous unrealized value from AI tools because we are not bringing our own deep knowledge, taste, and genuine judgment to the work. Mollick's frame is optimistic, about what we gain when we use these tools well. But turn his argument over and you see its shadow. If the payoff from AI comes from humans bringing their deep knowledge and taste and agency, then what happens when you systematically drain the organization of the conditions that produce deep knowledge? You don't get a faster version of the old firm. You get a firm that is generating outputs it does not understand.

Here is the second-order consequence that most people miss when they read that engineer's post: the danger is not bad code. It is undetectable confidence. The engineers are not confused or hesitant. They are moving fast. Management sees throughput and concludes there is no bottleneck. The gap between apparent competence and actual understanding is growing, silently, beneath a surface that looks completely normal. The organization doesn't know what it doesn't know. And there is no process left that would reveal this — because the reading, the mechanism of knowing, has been switched off.

Now turn that lens toward the firm you run.

Advisory firms are knowledge organizations in one of the most consequential senses imaginable. The thing you sell — genuinely — is judgment. Not information, not transactions, not even plans. Judgment. And judgment is built the same way it is built everywhere: through slow accumulation, through reading the work carefully, through the friction of actually understanding what you have committed to on a client's behalf.

Over the next twelve to eighteen months, AI tools inside your firm are going to get dramatically better at producing the outputs that look like judgment: the client summary, the planning scenario, the meeting prep document, the draft recommendation. The temptation — and the commercial pressure, from management watching throughput metrics — will be to treat those outputs as if they are judgment. To move faster. To do more reviews in less time. To close more tickets.

If that happens, what you get is not a more efficient version of your current firm. What you get is a firm where the advisers no longer hold the internal map. Where the plan exists in a document that was generated but never really inhabited by a human mind. Where a client asks a question six months after the plan was written and the adviser has to reach for the file rather than reaching for their understanding. The client will not notice this immediately. The compliance review will not catch it. The throughput numbers will look excellent. The erosion is interior, and it is cumulative.

The question for you, as a principal or a head of adviser development, is not whether to use these tools. You should. You will. The productivity gains are real. The question is whether you are designing your workflows so that your people are reading the work. Whether the AI output is the beginning of the thinking or the end of it. Whether your advisers are pressing enter — or whether they are genuinely engaging with what the machine produced and building on it from a place of deep knowledge.

Picture a head of adviser development at a firm with two hundred advisers. She has watched her team's output per person go up substantially over the past year. She is proud of this. But she runs a test — she asks three advisers to walk her through a client plan that their AI tools helped produce, without looking at the document. She wants to know what they actually hold in their heads. What she finds is going to tell her everything about whether her firm is getting faster or whether it is getting hollow.

One-sentence takeaway: the metric that matters is not how much your team produces. It is how much of what they produce they actually understand.

Now here is the thing about that engineer, pressing enter all night — she is operating in an environment where the outputs of AI are text, plans, code: things that look finished when they come out. And that is, until now, the only shape AI outputs have had. Text in, text out. The model speaks, and you read what it says.

Last week something structurally different appeared, and almost nobody outside technical circles noticed it or grasped what it changes. TypeSafe AI, a company building what it describes as machine-intelligence infrastructure, released a model called Jev. Simon Willison covered it carefully. Willison has been chronicling AI developments with unusual precision for years, and he flagged this one as a genuine category shift — not just a better model, but a different shape of model entirely.

Here is what Jev does, and why it matters. You give it unstructured information — the kind of messy, ambiguous text that describes a real situation. And instead of writing back to you in prose, it returns a structured set of decisions. Numbers. Confidence scores. Yes or no, with a probability attached. The category this situation belongs to, rated on a scale. Think of it as the difference between asking a colleague "what do you think about this client?" and getting a two-paragraph reflection — versus getting a structured evaluation form that says: complexity, high. Urgency, moderate. Risk flag, yes, confidence eighty-seven percent.

Maggie Appleton, a designer and technologist who thinks carefully about human-computer interaction, suggested the better name for this category is "decision models." And she is right that the name matters. Calling them System One models, as TypeSafe does, is a gesture toward cognitive science — fast, automatic processing — but "decision models" captures the functional thing more cleanly. These are not models that converse. They are models that decide. Or rather, that return the structured inputs a human or a system needs in order to decide.

There are two other features of Jev that Willison noted and that are easy to underweight. First: it is very fast. Not marginally faster — substantially faster than conversational models. Second: it is very cheap. TypeSafe charges only for the input, not the output, because the output is just numbers, and numbers are small. This matters because it changes the economics of where you can deploy intelligence. When a model is expensive and slow, you use it sparingly, for high-value tasks. When it is fast and nearly free to run, you can put it everywhere.

Now here is what most people miss when they read about Jev: this is not a better chatbot. It is something closer to a sensory layer. An advisory firm that deploys something like this is not adding a smarter assistant. It is adding the capacity to have structured intelligence running continuously across its operations — reading incoming information, flagging patterns, surfacing decisions — without a human having to ask a question.

Think about what that means for a firm in twelve to eighteen months, if this model category matures the way its first iteration suggests it might. You have a client base. Every day, information flows in: market movements, news events, account transactions, custodial data, emails, documents. Right now, that information mostly sits in systems, and a human has to know to look for it. A decision model running continuously over that information does not wait to be asked. It is, at every moment, answering questions no one thought to ask. Which client relationships have a pattern that looks like early disengagement? Which planning scenarios just became materially wrong because of something that happened in the market? Which incoming prospect inquiry contains signals that suggest a high-likelihood fit?

The outputs are not prose. They are structured. Confidence-weighted. Routed. And because they are structured, they are auditable in a way that prose AI output is not. You can look at why the model flagged something. You can tune what it flags. You can build accountability into the process in a way that is genuinely difficult when the output is a paragraph.

This is where the staffing implication becomes real. A decision model running at scale over client data is not replacing an adviser. It is replacing the version of an adviser that sits in front of a spreadsheet at eight in the morning, manually scanning for things that need attention. It is doing the triage so that the human can do the judgment. Which means the human needs to be extraordinarily good at judgment — and extraordinarily suited to the thing the model cannot do, which is bring depth to what surveillance-at-scale surfaces — for this arrangement to be stable and valuable.

What you are staffing for in eighteen months is not the same team you have today, even if the headcount is identical. You are staffing for people who can receive a structured flag — a confidence score, a decision prompt, a risk rating — and bring genuine depth to evaluating it. That is a specific kind of person. It is someone with deep knowledge of the client, of the planning landscape, of the emotional texture of a relationship. Someone who, to return to Mollick's frame, has the taste and the agency to use the machine's output and take it somewhere the machine cannot go.

The two ideas in this episode are connected by the same thread. AI is producing outputs — text, decisions, flags — and the question is always what the human brings to what comes next. The danger in the first story is humans who stop bringing anything, who press enter and move on. The opportunity in the second story is humans who bring everything — who take the structured decision the model returns and apply to it the thing no model has yet: the history of a relationship, the texture of a family's real situation, the judgment that is only available to someone who read the work, and kept reading, and never stopped.

The firms that understand this distinction — not just intellectually, but in how they hire, how they develop advisers, how they design workflows — are the ones that will look, in eighteen months, like they made a series of very smart choices. The firms that don't will look like they got faster. And fast and hollow are very easy to confuse. Until the moment when they aren't.

One thing to try this month

Ask three members of your team to walk you through a client plan their AI tools helped produce — without looking at the document. What they hold in their heads tells you whether your firm is getting faster or getting hollow.

Questions for your leadership team

  • In our current AI-assisted workflows, at what point does a human actually read and internalize the work — and if that point has been removed or compressed, what would we have to redesign to restore it?
  • If a decision model were running continuously over our client data and surfacing structured flags, what would we need to be true about our advisers' depth of knowledge for that system to make us better rather than just faster?
  • What metrics are we actually watching to evaluate AI's impact on our firm — and do any of them measure understanding, or do they all measure throughput?

Sources

  • Simon Willison, simonwillison.net — "Quoting voxium" (2026-09-21): https://simonwillison.net/2026/Sep/20/voxium/
  • Simon Willison, simonwillison.net — "Jev introduces a new shape of LLM - System One, aka Decision Models" (2026-09-21): https://simonwillison.net/2026/Sep/21/jev/
  • Ethan Mollick, One Useful Thing — "The Overhang" (2026-09-18): https://www.oneusefulthing.org/p/the-overhang
  • KEY INSIGHTS:
  • When AI generates the work and humans stop reading it carefully, organizations lose the interior map of what they know — outputs remain normal while understanding quietly collapses
  • The danger is not errors or bad outputs; it is undetectable confidence — teams moving fast on work they do not genuinely understand, with no remaining process that would reveal the gap
  • Mollick's "overhang" argument has a shadow: the value of AI is unlocked by human depth, taste, and agency — which means systematically removing those conditions destroys value even as throughput rises
  • Decision models like Jev represent a structural shift: not better chatbots but a sensory layer, returning confidence-weighted structured decisions continuously rather than conversational prose
  • Because decision models are fast and cheap, the economics allow them to run everywhere — transforming AI from a tool you consult to an ambient intelligence that surfaces what no one thought to ask
  • The staffing implication is not headcount reduction but capability reconfiguration: firms will need people who are extraordinarily good at receiving a structured flag and bringing genuine relational depth to evaluating it
  • The through-line: AI produces outputs; what matters is whether the human receiving those outputs brings genuine knowledge to what comes next — and whether the firm is designed to make that possible