Start with the number that should reframe every AI conversation in our industry.
According to MIT’s 2025 study, roughly 95 percent of enterprise generative AI pilots have delivered no measurable return. Tens of billions invested, and 19 in 20 projects produced nothing that showed up in the P&L.
It is tempting to read that as “AI is overhyped.” That is the wrong lesson.
The models are extraordinary. The failure is not intelligence; it is that a chatbot bolted onto a business does not do the business’s work.
MIT found the winning 5 percent shared a specific shape: they embedded deeply into real workflows, carried memory, and actually completed tasks rather than just answering questions.

The losers stayed stuck in slick demos that were brittle the moment they met a real process.
For insurance agencies, that gap is not abstract. It explains exactly why the last two years of “AI for insurance” have underwhelmed, and why that is about to change. There are four reasons AI has not, until now, actually run an agency.
The failure is not intelligence; it is that a chatbot bolted onto a business does not do the business’s work.
MIT found the winning 5 percent shared a specific shape: they embedded deeply into real workflows, carried memory, and actually completed tasks rather than just answering questions.
The losers stayed stuck in slick demos that were brittle the moment they met a real process.

For insurance agencies, that gap is not abstract. It explains exactly why the last two years of “AI for insurance” have underwhelmed, and why that is about to change.
There are three reasons AI has not, until now, actually run an agency.
Reason 1: It never really knew insurance, or your agency
A general model is a brilliant stranger.
It can define coinsurance and draft an email, but it does not know your book, your carriers, your renewal calendar, or the objections your producers hear every day.

Ask it “which of my clients renew in the next 90 days” and it has nothing, because it was never connected to the answer.
That is the difference between horizontal AI, trained broadly, and vertical AI, built and grounded in one industry and one agency’s data.
The 5 percent that worked in MIT’s study were not using smarter models than the 95 percent.
They were using models pointed at the right context. In insurance, context is everything: the same sentence means different things across personal lines and commercial, across carriers, across states.
Without that grounding, AI stays a clever assistant. With it, AI can be trusted with the actual work.
Agency data is genuinely hard, not just messy
It is worth being specific about why this grounding took so long, because “our data is messy” undersells the problem. Insurance data is structurally difficult in five ways.
-
Meaning shifts by line of business. A “renewal” in personal auto and a “renewal” in commercial property are different events on different clocks with different owners. A model that treats them as one concept gives confidently wrong answers.
-
Carriers do not agree with each other. Appetite, forms, naming, and submission requirements vary by carrier and change without notice. Any system that hard-codes one carrier’s vocabulary is wrong by the next quarter.
-
Rules are jurisdictional. Calling windows, disclosure requirements, and licensing all vary by state. A national assumption is a compliance problem waiting to happen.
-
The systems of record drift. Agencies customize their management system over years. Custom fields get repurposed, statuses accumulate, and the schema in the documentation stops matching the schema in production. An integration built against the documentation breaks against reality.
-
The richest data is unstructured. The truth about why deals are lost lives in call recordings, not in fields. Extracting it needs transcription plus domain judgment about what actually mattered in the conversation.
None of that is solved by a bigger model. It is solved by building for one industry and connecting to one agency, which is exactly what almost nobody did.
Reason 2: It could talk, but it could not act
The second barrier is deeper. Most AI in agencies has lived in a chat window.
You ask, it answers, and then you go do the work in your AMS, your dialer, and your carrier portals.

That is assistance, not automation. It still charges the tool tax, because a human still moves every output by hand.
Doing the work means acting across a fragmented stack: reading and writing the AMS, placing and logging calls, running a multi-touch follow-up, updating a record, and doing it reliably enough to trust.
This is the harder engineering problem, and it is precisely what separated MIT’s 5 percent.
The shift the whole market is now making has a name: from assist to act.
Gartner projects that 40 percent of enterprise applications will include task-specific AI agents by the end of 2026, up from less than 5 percent in 2025.
The category is crossing from software that answers to software that does.
Reason 3: It could not be trusted in a regulated business
The third barrier is the one insurance feels most sharply, and the one the rest of the market is learning the hard way.
Autonomy without control is a liability.
Gartner predicts that by 2027, 40 percent of enterprises will demote or decommission autonomous AI agents after governance gaps surface in production incidents
. In a licensed, regulated business, an AI that changes a customer record or sends a message on its own is not a feature; it is exposure.
So the answer for insurance was never full autonomy. It is an AI that does the work and asks permission.
Reads and analysis need no approval, because they change nothing.
Anything t
hat touches a customer, a policy, or a campaign waits for a human to say go. That approval-by-design model is what makes an AI worker safe to put inside an agency, and it is the piece most of the industry’s AI experiments skipped.
Reason 4: The pilot was never scoped to a number (for your agency)
The fourth reason gets almost no attention, and it may be the most common.
A large share of AI pilots could not have proven anything, because nobody decided in advance what they were trying to move.

Ask an owner how their AI pilot went and you often get a sentiment, not a measurement. People liked it. It saved time, probably. It was interesting. That is not a result, and it is indistinguishable from failure on a budget review.
Meanwhile the pilots that did land almost always had one number attached before switching anything on: minutes to first contact, percentage of calls answered, weeks to a new producer’s first close, retention on a specific segment.
This is the cheapest of the four barriers to fix, because it costs nothing. It is also the one entirely within your control, no matter which vendor you pick
Anatomy of a pilot that was always going to fail
Here is the pattern, and most owners will recognize at least part of it.
-
Month 1. Someone demos an impressive AI tool. It writes a good email in ten seconds. The agency buys three seats for the producers who seem most interested. No baseline is recorded, because the tool is obviously better than nothing.
-
Month 2. Two of the three producers use it for email drafting. Nobody uses it for the follow-up cadence, because the tool cannot see the lead list and cannot place the calls. Output still gets pasted into the management system by hand.
-
Month 3. The most enthusiastic producer leaves for another agency. Their seat sits unused. The tool’s context resets every session, so the remaining users have stopped explaining the agency and started accepting generic output.
-
Month 4. Renewal comes up. The owner asks what it did for the business. There is no baseline to compare against, no metric anyone agreed on, and no way to separate the tool’s effect from a good quarter. The honest answer is “we are not sure.”
-
Month 5. It gets cancelled, and the agency concludes that AI does not work for insurance.
Nothing in that story is a failure of the model. Four of the five months contain a structural failure that would have sunk any tool.
What just changed

For the first time, the three barriers are falling at once.
Frontier models got good enough to reason over messy agency data.
Vertical grounding lets them understand insurance and your specific book.
Agentic tool-use lets them act across the stack instead of chatting about it.
And approval-by-design gives owners control without giving up the automation.
None of these alone was enough. Together, they move AI from the 95 percent that assist to the 5 percent that actually do the work.
That is the real story behind the noise. Not another chatbot, but the arrival of an AI built for insurance that can be trusted with the work of the agency, with a human in command.
The winners in our industry will look like MIT’s 5 percent: deeply integrated, grounded in the domain, acting inside real workflows, and governed by approval rather than blind autonomy.
What this means for your agency
The practical takeaway is a buying filter.
Do not buy the 95 percent.
When you evaluate AI this year, insist on the pattern that actually works: it knows your agency without a tutorial, it acts inside your systems rather than drafting output you execute by hand, it remembers across conversations, and it asks your approval before anything reaches a customer.

That is the difference between an AI that demos well and an AI that shows up in your numbers.
The technology finally caught up to the promise.
The agencies that win the next few years will be the ones that recognize the difference, and hold their AI to the higher bar.
Editor’s note: This is the bar we built SUPERAGENT 3.0 to meet, and tomorrow, Tuesday, August 11, we unveil it. Our take on an AI Business Partner for insurance agencies, one that knows your agency, does the work across it, and asks before it changes anything, goes live in the launch keynote at 8:00 AM PST, 11:00 AM ET. It is free to watch: register at lp.getsuperagent.com/3-0-launch-sign-up.
Tags:
Insurtech News, Employee Retantion, Insurance, Insurance Agency, Blog, AI, SUPERAGENT, Outbound AI Agent, AI Agents
Aug 10, 2026, 2:42:07 PM
.png?width=2008&height=1432&name=Frame%201000007834%20(1).png)
.png?width=1512&height=1089&name=Frame%201000007849%20(1).png)
Comments