← All posts
25,000 Agents. 11 Hours. 99.2% Success. The First AI Agent Endurance Benchmark.
June 9, 2026· 9 min read

25,000 Agents. 11 Hours. 99.2% Success. The First AI Agent Endurance Benchmark.

By Karsten Wade
*Nobody benchmarks agents for endurance. We did.* --- ## The Problem With Agent Benchmarks Every agent benchmark in the industry measures the same thing: burst capacity. How many agents can you fire at once? How fast can they complete? But production agent workloads don't look like that. Real swarms run for hours. Days. Weeks. The question isn't "can you handle 10,000 agents for 4 minutes?" — it's **"can you sustain thousands of agents around the clock without degradation?"** We ran 25,000 agents over 11.2 hours on Meta's Llama infrastructure to find out. --- ## The Results
Metric Result
Total Agents25,000
Success Rate99.2% (24,793/25,000)
Duration11.2 hours
Cycles5 x 5,000 agents
P50 Latency1,151 ms
P95 Latency2,078 ms
P99 Latency2,741 ms
99.2% success rate held **constant across all 5 cycles** — no degradation, no drift, no memory leaks, no connection pool exhaustion. The same reliability in hour 11 as hour 1. --- ## Why Endurance Matters Burst benchmarks tell you what a system can survive. Endurance benchmarks tell you what a system can **sustain**. Consider the difference:
Benchmark Type What It Proves What It Misses
Burst (minutes)Peak throughputMemory leaks, connection exhaustion, rate limit recovery
Endurance (hours)Sustained reliabilityNothing — this is what production looks like
Production agent swarms don't fire 10,000 requests and stop. They run monitoring loops, process queues, handle events, coordinate workflows — **continuously**. An endurance benchmark is the only honest measure of production readiness. --- ## Cycle-by-Cycle Breakdown Every cycle ran 5,000 agents across 4 Llama models with automatic wave pacing and retry handling.
Cycle Cumulative Agents Success Rate Elapsed
15,00099.2%2.2h
210,00099.1%4.5h
315,00099.2%6.7h
420,00099.2%9.0h
525,00099.2%11.2h
The success rate flatlined at 99.2% from cycle 1 through cycle 5. Zero degradation. That's the number you want to see. --- ## The Burst Tests (Same Day) Before the endurance run, we validated burst capacity at every tier:
Agents Success P50 Latency Duration Mode
50100%1,226 ms14sSimultaneous burst
500100%1,164 ms13 minWave (40/wave)
1,00099.7%1,204 ms26 minWave (40/wave)
5,00099.2%1,138 ms2h 15mWave (40/wave)
25,00099.2%1,151 ms11.2hEndurance (5 cycles)
P50 latency stayed between 1.1s and 1.2s at every scale — from 50 agents to 25,000. --- ## Four Models in Rotation Every agent was distributed evenly across four production Llama models:
Model Parameters Distribution
Llama 3.3 70B Instruct70B25%
Llama 3.3 8B Instruct8B25%
Llama 4 Maverick17B (128 experts)25%
Llama 4 Scout17B (16 experts)25%
Automatic load balancing across all four models with intelligent retry and rate limit management. No manual intervention during the entire 11-hour run. --- ## The Cost This benchmark processed an estimated 7.5 million tokens across 25,000 agents. Here's what that costs at market rates:
Platform 25K Benchmark Rate
AINative Agent Cloud$8.85$1.18/1M tokens
AWS Bedrock$16.88$2.25/1M tokens
Anthropic (Claude Sonnet)$67.50$9.00/1M tokens
OpenAI (GPT-4o)$93.75$12.50/1M tokens
**What does 24/7 sustained operation cost?**
Workload Agents/Day AINative Cost
This benchmark (11.2h)25,000$8.85
Full 24h operation~53,500$18.93
30 days continuous~1,600,000$568
Enterprise (30M agents/mo)30,000,000$10,620
Platform 30M Agents/Month vs AINative
AINative Agent Cloud$10,620Baseline
AWS Bedrock$20,2501.9x more
Anthropic (Claude Sonnet)$81,0007.6x more
OpenAI (GPT-4o)$112,50010.6x more
*Disclaimer: All pricing is based on published provider rates as of June 2026 and is subject to change. Competitor rates reflect Llama 70B class models where available. AINative rate is the blended average across all four benchmark models. See [ainative.studio/pricing](https://ainative.studio/pricing) for current rates.* --- ## The Competitive Landscape No other platform has published an endurance benchmark for AI agents. Here's what exists:
Platform Burst Benchmark Endurance Benchmark
CrewAI~50 agents, 56% successNone published
LangGraph2.70 RPSNone published
OpenAI Agents SDK"Not production-ready"None published
AWS Bedrock AgentCoreNo dataNone published
AINative Agent Cloud15,000 agents, 100%25,000 agents, 11.2h, 99.2%
--- ## What This Proves 1. **No degradation over time.** 99.2% success in hour 1 = 99.2% in hour 11. Connection pools, rate limit recovery, retry logic — all battle-tested. 2. **Consistent latency.** P50 stayed at ~1.15s across 25,000 agents. No latency creep, no queue buildup, no slow death. 3. **Multi-model resilience.** Four Llama models in rotation with automatic load balancing. If one model rate-limits, agents seamlessly retry on another. 4. **Production-grade infrastructure.** This isn't a demo. It's the same AINative Agent Cloud API available to every developer today. --- ## The Full Benchmark Series This endurance test is part of our ongoing agent scale benchmark series:
Benchmark Agents Result Link
Burst (DO)10,000100%, 4.4 minRead →
Burst (DO)15,000100%, 14,291 tool callsRead →
Endurance (Meta)25,00099.2%, 11.2 hoursYou're reading it
--- ## Try It Yourself AINative Agent Cloud is available today. - **Documentation**: [docs.ainative.studio](https://docs.ainative.studio) - **API**: [api.ainative.studio](https://api.ainative.studio) - **MCP Server**: `pip install zerodb-mcp` The benchmark script is open source. Verify our numbers yourself. --- *Published June 9, 2026 by the AINative Engineering Team*

Check your site's AX Score

Free scan, 6 categories, under 60 seconds. See how your site ranks on the agentic web.

Run a free audit →