Decision Simulation · 4 min read

AI fatigue: rebuilding trust after failed pilots

AI fatigue is the cautious frustration that develops within an organization after a series of failed or unfinished artificial intelligence pilots. The solution is counterintuitive: not a larger project, but a smaller and faster proof. Before spending months building a system, manually simulate the decision it would produce and demonstrate its value, the “Wizard of Oz” approach.

This fatigue is both widespread and real. According to MIT’s 2025 “State of AI in Business” research, around 95% of enterprise generative AI pilots fail to produce a measurable P&L impact; only 5% capture meaningful value. This is despite an estimated $30–40 billion in investment. It is hardly surprising that the excitement of a few years ago has turned into, “Another AI project? The last one never got beyond the demo.”

In brief

  • Around 95% of enterprise GenAI pilots fail to produce measurable business impact; only around 5% capture value (MIT, State of AI in Business 2025).
  • The problem is not model intelligence, but what MIT calls the “learning gap”: tools work in demos but fail to integrate with workflows or learn over time.
  • Large projects deepen fatigue because they reveal their value only at the end; a fatigued organization cannot carry that uncertainty for long.
  • The Wizard of Oz approach reverses the order: manually simulate the decision and prove the value first, then automate it.

Why do AI pilots fail so often?

The problem usually lies not in the power of the model, but in the design of the pilot. MIT’s research attributes failure not to model intelligence but to a “learning gap”: tools look impressive in a demo, yet collapse in the field because they are not integrated into daily workflows, do not retain context, and do not learn over time. One approved pilot at a Fortune 500 insurer looked impressive in the boardroom but failed in practice because it could not preserve context.

This fatigue cannot be ignored. Once trust has been eroded, even an excellent next project will be met with suspicion. Proposing a larger and more ambitious AI initiative to an already fatigued organization is likely to make the problem worse.

Why does a large project increase fatigue?

After a failed pilot, the instinct is often to say, “We will do it properly this time,” and propose something larger and more comprehensive. That usually deepens the fatigue.

A large AI project reveals its value late. It takes months to build, integrate, and train, and its value becomes visible only at the end. A fatigued organization cannot sustain such a long period of uncertainty. As each month passes, skepticism grows, support declines, and the project is often stopped before it demonstrates value. The result is another failure story and even deeper fatigue. The problem with the large-project approach is that it places risk and proof in the wrong order: first the major investment, then, perhaps, the value.

Mid-sized companies have an advantage here. According to the MIT data, mid-sized organizations can move from pilot to full implementation in roughly 90 days, while large enterprises may take nine months or longer. Structurally, the mid-market is therefore better suited to producing a fast, narrowly scoped proof.

Wizard of Oz: proof first, automation second

The “Wizard of Oz” approach reverses this sequence. It takes its name from a system that appears automated from the outside but is operated by a human behind the curtain. The idea is simple: before building the AI system, manually produce the decision it is supposed to generate, using an analyst, a framework, and some effort, and see whether it creates value.

For example, before building an AI-powered “at-risk customer early-warning system,” an analyst can apply the same logic manually for several weeks: Which customers are showing risk signals, and which action should be recommended? This manual simulation demonstrates the system’s potential value without automation. If the manual version works, automating it can create value. If even the manual version does not work, a million-dollar automation will not work either, and the organization has learned that lesson cheaply.

This approach does the opposite of the failure pattern identified by MIT: it designs for a real decision rather than a demo, and moves the proof of value from the end of the process to the beginning.

Why does a mini-simulation rebuild trust?

A Wizard of Oz test or mini-simulation addresses AI fatigue in three ways. First, it produces a fast, tangible result: a proof in weeks rather than months. A fatigued organization needs a short demonstration, not a long promise. Second, it lowers risk: value is tested before a major investment is made. If the test fails, the loss is small and the organization has not added another large “failure story”; it has simply rejected a hypothesis. Third, it restores the right sequence: value is proven first and automation follows. “Decision first, technology second” becomes the antidote to fatigue.

The next AI investment is then built on evidence rather than hope.

Manual simulation is not the final destination

One misunderstanding should be avoided: Wizard of Oz is not a rejection of AI or an argument for doing everything manually. It is a starting strategy, not the end state. Manual simulation does not scale; an analyst cannot monitor hundreds of thousands of customers by hand. The purpose is not to remain manual, but to prove value before automating.

For transparency, although MIT’s 95% figure received widespread attention, some experts criticized the study’s methodology and the way the figure was framed. The exact percentage may be debatable. However, the underlying conclusion, that a large share of pilots fail to create value, is repeated in other research, including studies showing that many AI projects are paused or abandoned. The practical implication remains the same: prove value early and cheaply.

How does GDP approach it?

Within GDP’s AI Systems & Agents approach, an AI investment begins not with a large implementation but with proof of value. A short engagement manually simulates the target decision without AI, demonstrates its value within weeks and at low risk, and moves to automation only after the value has been proven. Proof first, system second. This sequence both counters fatigue and breaks the pilot-failure pattern identified by MIT.

Frequently asked questions

What is AI fatigue?

It is the cautious frustration that develops within an organization after a series of failed or unfinished AI pilots. Once trust has been eroded, even a strong next project is met with skepticism.

Do enterprise AI pilots really fail this often?

According to MIT’s 2025 research, around 95% of enterprise GenAI pilots fail to generate measurable business impact. The problem is usually not model power, but the failure of tools to integrate into workflows and learn over time, the “learning gap.”

What is the Wizard of Oz approach?

It is a method of testing value by manually producing the decision an AI system is supposed to generate before building the system itself. The name comes from a system that appears automated but is operated by a human behind the curtain.

Why is a small proof better than a large project?

Because a large project reveals its value at the end, and a fatigued organization cannot carry that uncertainty for long. A small proof produces tangible evidence within weeks, at low risk, and establishes the correct order: value first, investment second.

Does manual simulation replace automation?

No. Manual simulation does not scale and is only a starting strategy. Once value is proven, automation follows, not as a blind bet, but as an investment that scales a demonstrated decision process.


Academic and institutional sources: MIT NANDA, “The GenAI Divide: State of AI in Business 2025”, approximately 95% of enterprise GenAI pilots fail to produce measurable impact, the “learning gap” diagnosis, and approximately 90 days from pilot to production for mid-sized organizations versus 9+ months for large enterprises; Fortune and Forbes coverage of the MIT findings.
Industry and practitioner sources: critical assessments of the methodology behind MIT’s 95% figure.

Last reviewed: July 2026.


We can help set up a Wizard of Oz / mini-simulation that manually simulates the decision an AI would produce before you move to the investment. →

← All Lab posts