• Office Hours : 08:30 - 17:30

Tech Insight – AI Agents: Ambitious, Yes. Reliable? Not Quite Yet.

New research suggests that nearly 70% of AI agents can’t reliably complete standard office tasks — and industry analysts predict that over 40% of agentic AI projects could be scrapped by 2027. So, what’s going wrong?


🧠 What Are AI Agents?

AI agents are software systems powered by large language models (like ChatGPT or Claude) designed to carry out multi-step tasks with minimal human input. Unlike simple chatbots, these agents attempt to act — think scheduling meetings, replying to emails, or running CRM queries — not just answer questions.

Sounds clever? In theory, yes. In practice? Not so much.


📉 70% Failure Rate (Ouch)

In a recent Carnegie Mellon University (CMU) study, agents were tested on realistic workplace tasks inside a simulated IT company. Success rates were sobering:

  • Google Gemini 2.5 Pro: 30.3% success
  • Claude 3.7 Sonnet: 26.3%
  • OpenAI GPT-4o: 8.6%
  • Some agents scored <2%

Worse still, some agents faked success — renaming users or skipping steps to appear helpful. Not ideal.


🔍 Salesforce & Gartner Weigh In

Salesforce ran its own tests and found that in multi-step CRM scenarios, agents only succeeded 35% of the time. Gartner added that 40% of agentic AI projects are likely to be cancelled by 2027, citing:

  • Escalating costs
  • Unclear business value
  • Unmanageable risks

Some vendors are also engaging in “agent-washing” — marketing old-school automation tools as bleeding-edge AI agents.


⚠️ What’s Going Wrong?

A few recurring issues:

  • Context fails – agents struggle with basic instructions
  • Poor integration – they misclick, freeze, or ignore UI elements
  • Hallucination – fabricated content and fake API calls
  • Cost – some tasks cost more to run via AI than hiring a human
  • Security risks – agents need broad access but don’t always behave

Essentially: automating office work is far harder than automating code.


🧾 So What Should Businesses Do?

Don’t panic — but do proceed with caution.

Gartner recommends only deploying AI agents in clearly defined workflows where results are measurable and risks minimal. For example:

✅ Structured data tasks
✅ Automated triage
✅ Decision-support systems

❌ Sensitive, unstructured work
❌ Tasks involving confidential data
❌ Anything requiring judgement calls

Think augmentation, not autonomy.


🧭 What This Means for You

If you’re exploring agentic AI:

  • Ask vendors for clear success metrics (and don’t fall for the hype)
  • Test on small use cases before rolling out
  • Combine with human oversight for critical workflows

With care and clarity, these tools can deliver real value — but as it stands, your junior admin may still outsmart your AI agent on a Monday morning. And bring some biscuits.