How I Feel About AI After Two Years of Testing It Daily

Honestly? Mixed, and getting more specific instead of more settled. I've tested more than 50 AI tools for this site since 2024, and the ones that changed how I work are real and boring — a chat assistant for drafts, better search, faster first passes on code. The ones sold as revolutionary mostly weren't.

Short answer: After two years of daily testing, I think AI is genuinely useful for narrow, well-defined tasks — drafting, summarizing, first-pass code — and oversold everywhere else. The tools I actually keep open are the ones that save time on something specific, not the ones promising to change how I work overall. Skepticism and daily use aren't contradictory; that's just where the evidence points.

ChatGPT homepage — screenshot of chatgpt.com
ChatGPT homepage — screenshot of chatgpt.com

I run a site that reviews AI tools, so people assume I'm either a true believer or a professional skeptic. Neither fits, and I've stopped trying to pick a side just to make the answer tidier. I've watched the same model get 90% of a coding task right and then confidently invent a function that doesn't exist, in the same session, five minutes apart. I've had a research assistant find a source in ten seconds that would have taken me twenty minutes to track down manually, then misquote a number from that exact source when I asked it to summarize. That's the actual texture of using this stuff daily — genuinely fast, occasionally wrong in a way that sounds completely confident, and a lot less clean than either the hype cycle or the backlash cycle wants it to be. Two years of that pattern is what shaped how I feel about AI now, and it's a more specific, more qualified opinion than the one I started with.

Where I landed after two years of testing

I'm not alone in feeling split. A June 2026 Pew Research survey found 52% of American adults are more concerned than excited about AI's growing role in daily life, up from 37% who said the same in 2021 — a 15-point jump in five years, even as usage climbed. Only 10% said they were more excited than concerned. That gap between rising use and rising unease, reported by Pew on August 18, 2026, matches what I see in my own testing notes: the tools get better every quarter, and I trust them less with anything that actually matters.

The Stanford HAI 2026 AI Index Report, published April 13, 2026, backs up the usage side of that split. Organizational adoption of generative AI hit 88% by early 2026, and the report puts population-level adoption at 53% — spreading faster than the PC or the early internet did. So the tools are everywhere. Whether they're worth the trust people put in them is a separate question, and it's the one I actually care about.

The three things that changed my mind

Drafting beats blank-page paralysis, consistently. Whatever I think about AI-generated prose long-term, having something to react to instead of a blank document saves me real time, every week, on every kind of writing.

Search-with-citations is better than plain search for narrow questions. When I tested ChatGPT's search mode and Claude against a stack of documentation questions, both beat a manual search-and-skim in about two-thirds of cases, mostly because they cited the specific paragraph instead of making me hunt for it.

Code review catches things I skim past. I ran the same pull request through three assistants for my ChatGPT testing and through Claude separately, and all four flagged an off-by-one error a human reviewer (me) had already approved. That's not nothing.

Where I'm still skeptical

Confident wrongness is the core problem, not occasional wrongness. In my testing, a tool that says "I'm not sure" on roughly 30% of borderline questions is more useful day-to-day than one that's right 90% of the time and sounds equally certain about the other 10%, because I can't tell which answer is the wrong one without checking every single time. Most consumer AI products still don't handle that distinction well, and it's the single biggest reason I still double-check anything specific before I repeat it.

"Agentic" is doing a lot of marketing work. When I tested agent modes that browse and click through multi-step tasks over a two-week stretch, they finished roughly half the jobs I gave them without a human catching a wrong turn partway through. That's a real, useful hit rate for the right narrow job — booking research, form-filling, repetitive lookups — and a long way from the "hands-off assistant" pitch in the marketing.

The infrastructure spending doesn't match the returns yet. I dug into this directly in my piece on the AI bubble question: the capex-versus-revenue gap at the big cloud providers is real, and it's a different risk than "the tools don't work." The tools work. Whether the spending behind them is sustainable is a separate, harder question.

Prompts and workflows I actually use daily

Not theoretical examples — these are in my prompt history from this week:

  • "Here's a rough draft. Tell me the two weakest paragraphs and why, don't touch the grammar." (Better editing than a generic "improve this" prompt.)
  • "Summarize this changelog in plain English for someone who didn't write the code."
  • "I'm deciding between [option A] and [option B] for [specific constraint]. Argue for the one you'd pick, then argue against it."
  • "Review this function for edge cases only — assume the happy path already works."

The pattern across all of them: I give the model a constraint or a role, not just a topic. Open-ended prompts get open-ended, forgettable answers.

Common mistakes I see people make judging AI tools

The most common one is testing a tool once, on an easy task, and generalizing from that. I've made this mistake myself — a demo-quality first impression tells you almost nothing about how a tool holds up on your actual, messier work. Second is the opposite: writing off an entire category because one specific tool disappointed once, when a different one in the reviews I’ve run would have handled the same job fine. Third is trusting a star rating without checking who left it — I cover how to read those properly in how AI tool ratings actually work. And fourth, treating "AI" as one thing: a chat assistant, a coding agent, and an image generator fail in completely different ways, and lumping them together is how you end up either all-in or all-out for no good reason.

Hype vs. what actually held up in my testing

Claim What I found
"It'll replace your job" Changes which parts of a job take time; hasn't eliminated the job itself in anything I've tested
"It never makes mistakes anymore" Still confidently wrong on niche or recent topics — verify anything specific
"Agents can run your whole workflow unsupervised" Handled roughly half of multi-step tasks cleanly in my tests; the rest needed a correction mid-task
"It makes you 10x more productive" Closer to a real, measurable 10% on most work, per independent productivity research I've reviewed — still worth having
"Free tools are basically the same as paid" True for occasional use; paid plans earn their price once you hit daily, heavier use

Tools that made me update my opinion

Two tools moved me more than the rest. Claude's longer, more careful responses changed how I use AI for anything that needs actual reasoning — I get into the specifics in how to use Claude AI. And going back to basics with how to use ChatGPT reminded me that the free tier alone is enough to change most people's daily writing and research habits, no upgrade required. Neither tool fixed my skepticism about the industry-wide hype; both fixed my skepticism about whether the tools themselves are useful.

Frequently Asked Questions

Is it normal to feel both impressed and worried about AI at the same time?

Yes — Pew's 2026 data shows most Americans land somewhere in that mixed territory, and daily use doesn't require resolving the tension. I use these tools every day and still don't fully trust the industry building them.

Does using AI tools daily mean you have to be optimistic about AI overall?

No. I use AI tools for specific tasks because they measurably save time on those tasks, which is a separate question from whether the broader AI industry's spending, hype, or safety record deserves trust.

What's the single biggest mistake people make when forming an opinion about AI?

Judging the entire category from one tool and one task. A bad experience with an image generator says nothing about whether a coding assistant or a research tool will hold up for you.

Are AI tools actually saving people time, or is that overstated too?

Both, depending on the task. Research reviewed in my productivity gains piece puts real gains closer to 10% than the 10x figure often quoted — meaningful, but far short of the marketing.

Should I trust an AI tool's answer without checking it?

Not for anything specific — a number, a date, a legal or medical claim, or a niche fact. In my testing, reasoning and writing quality hold up well; specific factual claims still need a second look.