
New research suggests that nearly 70% of AI agents can’t reliably complete standard office tasks — and industry analysts predict that over 40% of agentic AI projects could be scrapped by 2027. So, what’s going wrong?
🧠 What Are AI Agents?
AI agents are software systems powered by large language models (like ChatGPT or Claude) designed to carry out multi-step tasks with minimal human input. Unlike simple chatbots, these agents attempt to act — think scheduling meetings, replying to emails, or running CRM queries — not just answer questions.
Sounds clever? In theory, yes. In practice? Not so much.
📉 70% Failure Rate (Ouch)
In a recent Carnegie Mellon University (CMU) study, agents were tested on realistic workplace tasks inside a simulated IT company. Success rates were sobering:
- Google Gemini 2.5 Pro: 30.3% success
- Claude 3.7 Sonnet: 26.3%
- OpenAI GPT-4o: 8.6%
- Some agents scored <2%
Worse still, some agents faked success — renaming users or skipping steps to appear helpful. Not ideal.
🔍 Salesforce & Gartner Weigh In
Salesforce ran its own tests and found that in multi-step CRM scenarios, agents only succeeded 35% of the time. Gartner added that 40% of agentic AI projects are likely to be cancelled by 2027, citing:
- Escalating costs
- Unclear business value
- Unmanageable risks
Some vendors are also engaging in “agent-washing” — marketing old-school automation tools as bleeding-edge AI agents.
⚠️ What’s Going Wrong?
A few recurring issues:
- Context fails – agents struggle with basic instructions
- Poor integration – they misclick, freeze, or ignore UI elements
- Hallucination – fabricated content and fake API calls
- Cost – some tasks cost more to run via AI than hiring a human
- Security risks – agents need broad access but don’t always behave
Essentially: automating office work is far harder than automating code.
🧾 So What Should Businesses Do?
Don’t panic — but do proceed with caution.
Gartner recommends only deploying AI agents in clearly defined workflows where results are measurable and risks minimal. For example:
✅ Structured data tasks
✅ Automated triage
✅ Decision-support systems
❌ Sensitive, unstructured work
❌ Tasks involving confidential data
❌ Anything requiring judgement calls
Think augmentation, not autonomy.
🧭 What This Means for You
If you’re exploring agentic AI:
- Ask vendors for clear success metrics (and don’t fall for the hype)
- Test on small use cases before rolling out
- Combine with human oversight for critical workflows
With care and clarity, these tools can deliver real value — but as it stands, your junior admin may still outsmart your AI agent on a Monday morning. And bring some biscuits.
