Why AI pilots fail before scaling: MIT's data
MIT NANDA's July 2025 study found just 5% of firms got task specific AI into production. Here is how to scope a first project that makes it.
If you run a Las Vegas dealership, law firm, restaurant group, or real estate office, here is the honest answer to why AI pilots fail before scaling: most first projects are scoped around a tool instead of one workflow, so they never change a number on the books. MIT NANDA's July 2025 report, The GenAI Divide, found 95% of organizations were getting zero measurable return from GenAI. Only 5% took task specific AI tools all the way to sustained productivity or profit and loss impact. Internal builds failed twice as often as externally partnered ones.
The report came out in July 2025, from research run January to June 2025. As of September 2026 it is still the largest sample AI pilot outcome study cited across vendor and analyst commentary, and we have not found a larger replacement. That makes it the best evidence an owner has for two decisions: whether to start at all, and how to scope a first project that reaches production instead of dying as a demo.
The 95% figure measures return, not pilots that crashed
The MIT NANDA team drew on three sources: a systematic review of over 300 publicly disclosed AI initiatives, structured interviews with representatives from 52 organizations, and survey responses from 153 senior leaders collected at four industry conferences between January and June 2025. Against $30 to $40 billion in enterprise GenAI investment, page 3 of MIT NANDA's State of AI in Business 2025 report states that 95% of organizations are getting zero return. Just 5% of integrated AI pilots are extracting millions of dollars in value.
That headline got repeated as "95% of AI projects fail," which is not quite what it says. Zero measurable return is a test of outcomes on the books. A pilot can run smoothly and get used every day and still land in the 95%, because nobody tied it to a number that moved.
That reading matters more to an owner than the scary version. The report rarely describes software that did not work. It describes projects that were never scoped to prove value, so when the pilot ended there was nothing to point at and no reason to keep paying.
Task specific tools lose three of four pilots; chat tools keep four of five
For a first project, the most useful page in the report is the funnel on page 6 of the MIT NANDA report PDF. It splits tools into two kinds: general purpose LLM tools such as ChatGPT and Copilot, and embedded or task specific enterprise tools built for one workflow. The table shows the share of organizations that reached each stage, with the timing and build findings from pages 7 to 8 underneath.
| Tool type | Investigated | Piloted | Successfully implemented, sustained productivity or P&L impact |
|---|---|---|---|
| General purpose LLM tools (ChatGPT, Copilot) | 80% | 50% | 40% |
| Embedded or task specific enterprise AI tools | 60% | 20% | 5% |
| Average time, pilot to full implementation | Mid market: 90 days | Enterprise (over $100M revenue): nine months or longer | Enterprise takes at least 3x as long |
Note: internal builds fail twice as often as externally partnered implementations.
Source: MIT NANDA, The GenAI Divide: State of AI in Business 2025, July 2025, pages 6 to 8. Percentages are shares of organizations.
The story is in the arithmetic between stages. For embedded tools, 20 of every 60 organizations that investigated reached a pilot, or one in three. Then 5 of every 20 piloting organizations reached successful implementation, or one in four. From first look to lasting impact, 5 out of 60 is about one in twelve.
General purpose tools run a very different funnel. Of 80 organizations that investigated, 50 piloted, and 40 of those 50 reached sustained impact. That is 80% of pilots and half of everyone who looked. So the embedded funnel collapses at the pilot stage: from pilot to production, embedded tools convert at 25% and general tools at 80%.
The plain reason is fit. A chat tool asks nothing of your systems, because a person just opens it and uses it. An embedded tool has to read your CRM, your case management system, or your reservation book, and it has to handle your messiest records correctly. DX's 2026 review of the GenAI Divide findings says internal builds frequently failed because of brittleness and poor workflow fit. Both are scoping failures, and you can see them before anyone writes code.
Mid market companies crossed in 90 days; enterprises took nine months
Pages 7 to 8 of the MIT NANDA report add a timing finding that should encourage smaller operators. Mid market companies that moved from pilot to full implementation averaged 90 days to do it. Enterprises, which the report defines as firms with over $100 million in annual revenue, took nine months or longer. That was true even though they ran more pilots and put more staff on AI.
Nine months is roughly 270 days, so the large firms took at least three times as long with more people. Extra approvals, systems, and stakeholders slow a project down more than extra headcount speeds it up. A Las Vegas business with one decision maker and one system of record looks more like the fast group.
Time is a real cost. Take our own example: a finished system worth $2,000 a month in saved labor or recovered revenue. On a 90 day path, month four is the first month of value. On a nine month path, it is month ten. The six months in between are $12,000 you never collect, before counting the staff hours spent in meetings about the pilot.
Partnered builds reached deployment about twice as often
The same pages of the report call it a myth that the best enterprises build their own tools. The report puts it bluntly: internal builds fail twice as often as externally partnered implementations. DX's 2026 review reads the data the same way, with externally sourced tools reaching deployment at roughly twice the rate of internal builds.
In a small business, an "internal build" rarely means an engineering team. Usually it is a sharp office manager connecting a chatbot to a web form through an automation platform, or a relative who codes. Those builds work in the demo. They break the first time a field name changes or a customer answers in a way nobody planned for. That is the brittleness DX describes.
We are an outside partner, so keep in mind that this section works in our favor. The data comes from MIT's sample, not from us, and one condition matters: the partner advantage holds only when the partner has shipped your workflow before. A vendor who learns your business while billing you is an internal build with extra steps. The real question is whether you need a custom AI assistant at all, or an existing product set up well.
Two first projects scoped to survive the pilot
The funnel points to a scoping standard. A first project should pass four tests before it starts:
- One workflow, one owner, one number. A named person owns it, and one metric with a measured baseline says whether it worked.
- It lives inside tools you already use. Staff should get the benefit without opening a new screen.
- The end date and kill criterion are written first. If the number has not moved by the date, the project stops.
- Whoever builds it has shipped this workflow before. Ask for their last one and talk to that customer.
Example one is a dealership service department. Every figure here is our assumption, not report data. Say the website sends 120 after hours service leads a month, and next morning callbacks book 30% of them. A project that books those leads overnight raises that to 50%. Twenty extra points on 120 leads is 24 more appointments. At an assumed $250 average repair order and 50% gross margin, that adds $6,000 in revenue and $3,000 in margin a month. With $12,000 to build and $500 a month to run, the net is $2,500 a month, which pays back the build 4.8 months after go live.
Now put that on the report's timelines. With 90 days to production, payback lands around month eight. With nine months, it lands around month fourteen. Then change the key assumption. If the booking rate rises only 5 points, that is 6 appointments, $750 in margin, and $250 net a month after running costs, which means a 48 month payback. The kill criterion exists for exactly this case, so write it as the booking rate you need by day 90.
Example two is a six person law firm, to show a smaller shop. Say the firm handles 60 consultations a month and a paralegal spends 20 minutes writing each intake summary into the case management system. If a tool cuts that to 5 minutes, saving 15 minutes on each of 60 consults adds up to 900 minutes, or 15 hours a month. At an assumed $35 an hour loaded cost, that is worth $525 a month.
At that size, the table argues against starting with an embedded build. The general purpose row, where 40 of 50 pilots reached sustained impact, is the lane a six person firm should try first: an existing chat tool, a written procedure, and a confidentiality review. With an assumed $60 a month in seats, the firm nets $465 a month. If the saved hours are still real at day 90, scope the embedded version with an AI strategy plan that starts from that measured baseline.
Where these numbers would not hold for you
The sample leans large. The report's 52 interviews and 153 surveyed leaders came from industry conferences, and it defines enterprise as over $100 million in revenue. It does not publish separate results for a 15 person business. Applying its conversion rates to your shop is an inference, and our worked examples are labeled as assumptions for that reason.
In the report, success means sustained productivity or profit and loss impact. A tool that makes your own week easier without moving a measured number counts as a failure. That is a strict bar. If all you want is a better drafting tool for yourself, the funnel is not your problem.
The research also dates from January to June 2025. Models and products have improved since then, and the 2026 review we cite backs up the findings without adding a new sample. If a larger, newer study appears, it should replace these numbers. And if you do not yet have a specific workflow in mind, none of this math applies. Find the workflow first, then the tool.
Questions owners ask
Do 95 percent of AI projects really fail?
Not exactly. MIT NANDA's July 2025 report found 95% of organizations were getting zero measurable return from GenAI. That is a statement about outcomes on the books, not about broken software. Many of those pilots worked technically but were never tied to a number. For task specific tools, 5% of organizations reached sustained impact, and most of the drop happened between pilot and production.
Is it better to build AI tools in house or hire an outside company?
The MIT NANDA data says internal builds fail twice as often as externally partnered ones. DX's 2026 review agrees and names brittleness and poor workflow fit as the reasons. The advantage holds only when the partner has shipped your exact workflow before. Before you sign anything, ask for their last deployment and a customer you can call.
How long should an AI pilot take before it goes live?
In the MIT NANDA report, mid market companies that succeeded went from pilot to full implementation in an average of 90 days. Enterprises took nine months or longer. A small business with one decision maker should aim for the shorter path, with a written deadline and a number that decides whether the project continues.
What to do this week
- Measure one baseline. Pick the workflow that costs you the most time or missed revenue and count it for a week: leads answered, intakes written, or bookings made. Without a baseline, you do not have a project.
- Run a manual version with a tool you already pay for. If a general purpose chat tool plus a written procedure can do a rough version, try that before buying anything embedded. That is the table's high survival row.
- Write the kill criterion before any vendor call. State the number, the date, and what happens if you miss it. Then ask every vendor, including us, when they last shipped this exact workflow.
If you want a second set of eyes on the scope, we start with a free 15-minute audit, which you can book from our AI consulting page.
Drafted with AI assistance, researched, edited, and fact-checked by Elias Musleh on September 28, 2026.
/ FREE 15-MIN AUTOMATION AUDIT
Want this handled for you?
We build the automation, run it, and hand you the numbers. Book a fifteen minute call and we will tell you straight whether it is worth doing for your business.
Book your free auditPrefer to skip the form? vendors@vegasbusinessai.com /702.773.8839