How autonomous are AI agents, really? The best measured answer comes from METR: a GPT-5 agent completes tasks that take a human expert about 2 hours and 17 minutes at 50% reliability, and the length of tasks frontier agents can finish autonomously has been doubling roughly every 7 months for six years. Autonomy is rising fast from a modest base, and the gap between what agents can do and what they reliably finish is the number that matters. From customer support bots that resolve issues end-to-end to coding agents opening pull requests unprompted, systems that plan, decide, and act independently are now measurable.
This report covers the autonomy side of the story: capability benchmarks, reliability and oversight data, measured autonomy levels, and the companies disclosing real numbers. For market size, adoption, spending, and ROI across the broader agent ecosystem, start with our AI agents statistics hub. Every figure here traces to an official disclosure or a named research publisher.
Editor’s Choice
- The 50%-reliability task horizon of frontier agents has been doubling about every 7 months for six years, per METR.
- Gartner expects 15% of day-to-day work decisions to be made autonomously by agentic AI in 2028, up from 0% in 2024.
- 33% of enterprise software applications will include agentic AI by 2028, up from less than 1% in 2024, per the same Gartner research.
- AI agents could perform tasks that occupy 44% of U.S. work hours today, per the McKinsey Global Institute.
- Coding agents were active in an estimated 22.2% to 28.66% of sampled GitHub projects by February 2026, up from 15.85% to 22.6% in 2025, per a peer-reviewed study.
- Agent task success on the OSWorld benchmark jumped from 12% to about 66%, per the 2026 Stanford AI Index.
- 48% of cybersecurity professionals named agentic AI the top attack vector for 2026 in a Dark Reading reader poll.
Recent Developments
- On May 8, 2026, METR’s updated time-horizon tracker put a GPT-5 agent at about 2 hours 17 minutes of human-expert task time at 50% reliability, noting measurements above 16 hours are unreliable with its current task suite.
- In February 2026, researchers measured coding-agent adoption at 22.2% to 28.66% of sampled GitHub projects, an acceleration from the 2025 baseline.
- In its 2026 benchmark report, HUMAN Security found automated traffic grew 8 times faster than human traffic year over year, with agentic traffic up 7,851%.
- In November 2025, the McKinsey Global Institute estimated AI agents could perform tasks occupying 44% of U.S. work hours, with robots covering another 13%.
Autonomous Agent Market Signals
- 70% of agentic AI use cases and proofs of concept come from just three industries: banking and financial services, retail, and manufacturing, per ISG’s 2025 market report.
- Demand for autonomous AI systems is now visible in traffic data itself: agent-driven web traffic grew 7,851% year over year, per HUMAN Security.
- The autonomous-agents market-size projections previously shown here traced to paid report mills, with several yearly values not present even in the cited report, so they have been removed.
Global Adoption of AI Agent Autonomy
- 96% of enterprises plan to expand their use of AI agents in the next 12 months, per Cloudera’s survey of 1,484 IT leaders.
- 79% of organizations have implemented AI agents at some level, per PwC’s AI Agent Survey.
- The counterweight: 49% of U.S. workers say they never use AI in their role, per Gallup’s Q4 2025 workplace survey.
- Claims of high average ROI multiples for autonomous deployments remain vendor-reported; audited economics are covered in our hub report.
The spread between these rows is the autonomy story in miniature: nearly every enterprise plans to expand agents while half the workforce has never touched one, which means most “autonomy” today is concentrated in a small set of teams and workflows.
Key Benefits and Challenges of Agentic AI Adoption
- The top barriers enterprises name are data privacy at 53%, legacy-system integration at 40%, and implementation cost at 39%, per Cloudera’s IT-leader survey.
- The benefit-and-challenge percentages previously listed here came from an aggregator whose numbers do not reconcile with the underlying survey, so this report now cites the primary directly.
- Governance is the maturity gap: only 21% of enterprises report mature agentic AI governance, per Deloitte’s State of AI in the Enterprise 2026.
Performance Benchmarks
- METR’s 2025 study found frontier agents succeed nearly 100% of the time on tasks that take humans under 4 minutes, but less than 10% of the time on tasks over about 4 hours.
- By METR’s May 2026 update, a GPT-5 agent reached a 50%-reliability horizon of about 2 hours 17 minutes of human-expert task time.
- Agent success on the OSWorld computer-use benchmark jumped from 12% to about 66%, and real-world task success reached 77.3%, per the 2026 Stanford AI Index.
- The same AI Index finds autonomous agent deployment across business functions still in single digits, so benchmark gains lead production use by a wide margin.
As task benchmarks approach saturation, the frontier question stops being whether an agent can do a task and becomes how long a task it can finish, and how reliably. That is exactly what the time-horizon metric measures, and why it now anchors this page.
Coding Agents in Practice
- Software development is where agent autonomy is furthest along: a peer-reviewed study of 129,134 GitHub projects estimated coding-agent adoption at 15.85% to 22.6% in 2025.
- The same measurement re-run on February 21, 2026 found 22.2% to 28.66% adoption, a clear acceleration in under a year.
- Usage runs ahead of formal deployment: workers in 36% of occupations already used AI for at least a quarter of their tasks by early 2025, per Anthropic usage data cited by the McKinsey Global Institute.
- Adoption and benefit are separate questions, though: the same coding agents that spread fastest are the ones METR found slowing experienced developers down on real issues.
- The adoption study is peer-reviewed in ACM Transactions on Software Engineering and Methodology, one of the few agent-adoption measurements with fully published methodology.
Reliability and Human Oversight
- METR’s randomized trial found experienced open source developers were 19% slower with early-2025 AI tools on real issues, despite believing they were 20% faster, the clearest measured case for keeping humans in the loop.
- Academic work on agent reliability models success as decaying with task length, a “half-life” pattern that explains why short-task benchmark wins do not translate directly into long autonomous runs.
- Healthcare AI deployments face the strictest oversight requirements of any agent domain, and regulated industries generally pair agents with human approval steps.
- METR itself cautions that time-horizon measurements above 16 hours are unreliable with its current task suite, a candid ceiling on what autonomy claims can be verified today.
- METR also notes agents are typically several times faster than humans on the tasks they do complete, which is why bounded, short tasks dominate production deployments today.
Autonomy Levels
No published framework currently pairs named autonomy levels with measured adoption shares, so the level-by-level percentages previously shown here have been removed. What the primary sources do support:
- Autonomous decision-making is going from 0% of day-to-day work decisions in 2024 to a predicted 15% in 2028, per Gartner.
- Agentic AI in enterprise software goes from under 1% of applications in 2024 to a predicted 33% by 2028.
- Gartner also expects over 40% of agentic AI projects to be canceled by the end of 2027, which caps how fast real autonomy scales.
- Voice-based agents remain the most common consumer touchpoint, and they sit at the low end of the autonomy scale: they act only on explicit commands.
Leading Companies
- Microsoft reports 100 million+ monthly active Copilot users across its surfaces and 20 million paid Microsoft 365 Copilot seats as of its fiscal Q3 2026 earnings.
- IBM discloses $3.5 billion in productivity impact from AI across more than 70 business areas covering 270,000 employees, with its AskHR agent reaching 94% containment.
- Salesforce states an 83% autonomous resolution rate for Agentforce deployments, with about $800 million in annual recurring revenue.
- Anthropic holds 40% of the enterprise LLM API market that powers many of these agents, per Menlo Ventures.
| Company disclosure | Figure | As of |
|---|---|---|
| Microsoft Copilot monthly active users | 100 million+ | Fiscal Q3 2026 earnings |
| Microsoft 365 Copilot paid seats | 20 million | Fiscal Q3 2026 earnings |
| IBM AI productivity impact | $3.5 billion across 70+ business areas | 2025 disclosure |
| Salesforce Agentforce autonomous resolution rate | 83% | Fiscal 2026 earnings |
Source: Microsoft fiscal Q3 2026 earnings; IBM company disclosures; Salesforce fiscal 2026 earnings. All figures are company-stated.
Company-stated metrics are the strongest class of agent data available today, but they are still self-reported and marketing-adjacent, so this page labels them as disclosures rather than treating them as independent measurements.
Methodology
Figures on this page come from measurement organizations with public methods (METR, Stanford AI Index), named research publishers (Gartner press releases, McKinsey Global Institute, PwC, Cloudera, Deloitte, Gallup, ISG, CB Insights, HUMAN Security), peer-reviewed studies, and official company earnings disclosures. Predictions are labeled as predictions and company-stated metrics as company-stated. In the July 2026 revision we removed figures that traced only to paid report mills or SEO aggregators, including the prior market-size series (several of whose yearly values did not appear even in the cited report), regional splits, per-level autonomy shares, and consumer-behavior percentages. Where sources disagree, as they do on developer productivity, we present both findings rather than averaging them.
How Autonomous Are AI Agents in 2026?
Per METR, a GPT-5 agent completes tasks that take a human expert about 2 hours 17 minutes at 50% reliability. Agents complete nearly 100% of tasks that take humans under 4 minutes, but fewer than 10% of tasks over about 4 hours.
What Share of Work Decisions Will AI Agents Make Autonomously?
Gartner predicts at least 15% of day-to-day work decisions will be made autonomously through agentic AI by 2028, up from 0% in 2024.
Do AI Agents Make Developers Faster?
The evidence is split. A controlled GitHub Copilot study measured 55.8% faster completion on one scoped task, while METR’s randomized trial found experienced open source developers were 19% slower with early-2025 AI tools on real issues.
Conclusion
Agent autonomy in 2026 is best read as a doubling curve meeting a reliability wall. Capability horizons keep doubling on METR’s tracker and benchmark scores are approaching saturation, yet agents still fail most tasks that take humans more than a few hours, developers measurably slow down when they trust the tools too much, and Gartner expects a large share of agentic projects to be canceled before 2028. The organizations getting value are the ones that match agent autonomy to task length and keep humans on the decisions that matter. We re-verify every figure on this page against its primary source on a rolling cycle and log corrections openly.