Last updated: July 26, 2026 · By Vishal Swami, Founder & Lead AI Reviewer, AISagely
AI mania is eviscerating global decision-making because organizations and individuals are outsourcing judgment to tools that were never built to carry it, then skipping the verification step that would catch the mistake. The result shows up in the failure data: nearly all generative AI pilots produce no measurable return, and the projects that do launch get canceled once the hype wears off.
Short answer: AI mania is eviscerating global decision-making by letting confident-sounding AI output replace the checks a real decision needs — a source, a second opinion, a small test before a big commitment. In my testing, the fix isn't avoiding AI tools; it's adding one deliberate verification step before any AI-assisted answer turns into an action, a purchase, or a policy.

I test AI tools for a living, and the pattern behind this headline isn't abstract — it's the same mistake I catch myself almost making every few weeks. Someone gets a fluent, well-formatted answer from ChatGPT, Claude, or Gemini, and the fluency itself reads as authority. The answer sounds finished, so it gets treated as finished, and the decision gets made on top of it without anyone checking whether the underlying facts were right. Multiply that by a few hundred thousand teams doing it inside real budgets, and you get the numbers below.
What's actually happening
Two data points from 2025 explain the headline better than any op-ed. MIT's NANDA initiative studied 300 enterprise generative AI deployments and found that 95% failed to produce a measurable financial return, despite an estimated $30–40 billion in enterprise spending, according to Forbes’ August 2025 coverage of the report. The report's own conclusion wasn't that the models are bad — it was that companies bought "generic tools, slick enough for demos, brittle in workflows," and skipped the governance and feedback loops that make a pilot durable.
Gartner reached a similar conclusion from a different angle. In a June 2025 prediction, the firm said more than 40% of agentic AI projects will be canceled by the end of 2027, and Gartner analyst Anushree Verma put the reason bluntly: most of these projects are "early-stage experiments or proof of concepts that are mostly driven by hype and are often misapplied," per MarTech’s reporting on the prediction. Gartner also estimates only about 130 of the thousands of vendors marketing "AI agents" have real agentic capability — the rest is what the firm calls agent washing.
Neither number is about the technology failing to work. It's about decisions getting made faster than they get checked — the exact mechanism that erodes good judgment at scale, whether it's a Fortune 500 rollout or a single person asking ChatGPT to plan a home renovation budget.
What you'll need
You don't need a new tool to fix this — you need a habit. Grab whatever AI assistant you already use (ChatGPT, Claude, Gemini, or a specialized tool), and pick one real decision you're currently leaning on AI for: a purchase, a hire, a strategy call, a health or financial choice. Have the original source material on hand if there is one — the report, the spreadsheet, the contract — since the whole point of the process below is comparing the AI's answer against something outside the AI. Fifteen minutes is enough for a personal decision; a team decision takes longer because more than one person needs to see the check before it goes live.
Step-by-step: how to avoid AI mania in your own decisions
1. Separate the answer from the confidence
Read the AI's response once for content, ignoring tone entirely. In my testing, a wrong answer and a right one from the same model read with identical confidence — same sentence structure, same certainty. If you can't tell the difference by tone, you can't use tone as your filter, so don't.
2. Ask the model to name its weakest claim
Prompt it directly: "Which part of that answer are you least certain about, and why?" This works better than asking "are you sure," because it forces the model to point at a specific claim instead of just reassuring you. It's not foolproof, but in my testing it surfaces at least one figure or assumption worth double-checking almost every time.
3. Check the one number that would change your decision
Every real decision has a number or fact that, if wrong, flips the outcome — a price, a deadline, a legal requirement, a statistic. Find that one thing and verify it against a primary source before moving on. Don't try to verify everything; that's how people give up on verification entirely.
4. Run it past a second, independent model
Paste the same question into a different AI tool — if you got the first answer from ChatGPT, check it against Claude or Gemini. Disagreement between the two doesn't tell you who's right, but it reliably tells you where to spend your remaining verification time. Agreement isn't proof either, since the same training gaps can produce the same wrong answer twice.
5. Pilot the decision before you scale it
This is the step MIT's report says most failed enterprise projects skipped. Before committing a full budget, team, or policy to an AI-assisted plan, run it at the smallest scale that still tells you something real — one department, one customer segment, one week. If the small version doesn't hold up, you've saved the cost of finding out at full scale.
6. Write down what you checked
A one-line note — "verified the Q3 figure against the vendor invoice; did not verify the market-size claim" — costs nothing and means the next person (including future you) knows what's already been checked and what hasn't. This is the single habit that separates the 5% of AI-assisted projects that hold up from the 95% that don't.
Example prompts you can copy
- Surface the weak point: "Before I act on this, tell me which claim in your answer you're least confident about and what would change if it's wrong."
- Force a source check: "Which of these numbers can you trace to a specific, named source, and which are you estimating? Be explicit about which is which."
- Stress-test a plan: "Give me the strongest argument against the recommendation you just made, as if you were advising the other side."
- Pilot-size a decision: "Scale this plan down to the smallest version that would still tell me whether it works, and tell me what to measure."
Common mistakes to avoid
The mistake I see most, and made myself early on, is treating a well-formatted answer as a verified one — a numbered list with a confident tone feels more trustworthy than a hedge-filled paragraph, even when the hedge-filled version is more accurate. Second is asking only one AI tool and stopping there; in my testing, running the same question through two models catches disagreements a single answer never flags. Third is skipping the small pilot and going straight to full-scale rollout, which is precisely the pattern behind the Gartner cancellation number above — teams commit before the workflow, data quality, and governance are actually in place. Fourth is verifying everything or nothing: trying to fact-check every sentence burns out fast, so people give up and check nothing instead, when picking the one decision-changing number is enough.
AI mania vs. a grounded AI workflow
| Behavior | AI mania (what fails) | Grounded workflow (what holds up) |
|---|---|---|
| Reading the answer | Confidence in tone = trust in content | Content is checked regardless of tone |
| Number of opinions | One model, one pass | Cross-checked against a second model or source |
| Scale of rollout | Full budget/team from the first pilot | Smallest version tested first, then scaled |
| Record-keeping | Nothing written down about what was checked | One-line note on what was and wasn't verified |
| Vendor claims | "AI agent" branding taken at face value | Capability checked against Gartner's ~130-of-thousands "real agent" estimate |
Tools that make this easier
The checklist above works with any assistant, but some tools make the verification step faster. If you're deciding which AI tool to trust with a task in the first place, my AI tool ratings page and my AI tool reviews hub cover where each one is strong and where it still guesses. For a small business weighing whether to run an AI pilot at all, my best AI tool for small business guide covers the lower-risk starting points instead of a full agentic rollout. If agentic AI specifically is what you're evaluating — the exact category Gartner flagged — my how to use ChatGPT agent mode guide walks through where it still needs a human checking each step. For the cross-checking step in the process above, my guides to using Claude and using Gemini cover getting a genuinely independent second opinion rather than just asking the same model twice. And if budget is the reason a team is tempted to skip the small pilot, my free AI tools roundup covers where you can test an idea at no cost before committing real spend.
My take
The 95% and 40% figures above aren't an argument against using AI — I use it every day and so does most of my audience. They're a measurement of what happens when speed replaces verification, at any scale from a single household budget to a multinational rollout. The fix costs almost nothing: one weak-point question, one second opinion, one small pilot before the big one, and one written line about what got checked. Skip all four and you're playing the same odds MIT and Gartner already measured. Do them and you're in the small group whose AI-assisted decisions actually hold up.
Frequently Asked Questions
Is AI mania actually measurable, or just a narrative?
It's measurable. MIT's NANDA report found 95% of enterprise generative AI pilots failed to produce a measurable financial return despite $30–40 billion in enterprise spending, and Gartner separately predicts more than 40% of agentic AI projects will be canceled by the end of 2027 — two independent studies pointing at the same underlying pattern.
Does this mean I should stop using AI tools for decisions?
No. Both studies point to skipped verification and rushed scaling as the cause, not the models themselves. Keep using AI for research and drafting; add the checks above — a weak-point question, a second model, a small pilot — before you act on the output.
How long does adding a verification step actually take?
For a personal decision, five to ten minutes: one follow-up prompt asking the model to name its weakest claim, and one check of the single number that would change your decision. Team decisions take longer only because more people need to see the check before it ships.
What's the easiest first step if I only do one thing?
Ask whichever AI tool you're using: "Which part of this answer are you least confident about?" It's a single follow-up prompt, and in my testing it surfaces a real gap often enough to be worth doing every time.
Are all "AI agent" products actually agentic?
No. Gartner estimates only around 130 of the thousands of vendors marketing AI agents have genuine agentic capability — the rest is rebranded existing software, what Gartner calls agent washing. Check for named, verifiable capabilities rather than the label.