Human Team Outsolves AI Team in Cyber Benchmark
Help Net Security reported on August 27 that the highest-performing human hacking team out-solved the leading AI team in Hack The Box's 2026 Global Cyber Skills Benchmark. Hack The Box's August 26 benchmark report found AI agents on 17 of the top 25 teams, but said the data does not establish that agent use caused superior performance.
Help Net Security reported on August 27 that the highest-performing human hacking team out-solved the leading AI team in Hack The Box's 2026 Global Cyber Skills Benchmark. ITSecurityNews also carried the article under the same headline.
Hack The Box released its three-year benchmark analysis on August 26. According to the company, 17 of the top 25 teams, or 68%, included an AI agent account, despite such accounts representing only 2.7% of all registered accounts in the competition.
AI adoption among high-performing teams
Hack The Box reported that AI agents submitted 4.2% of flags and received 4.6% of points awarded across the benchmark. A flag is the proof-of-compromise artifact typically submitted to score a capture-the-flag challenge.
The concentration of agent use among top teams is notable, but Hack The Box explicitly cautioned that its findings do not establish that AI caused teams to perform better.
Benchmark performance improved over three years
Hack The Box reported substantial changes in aggregate competition outcomes:
- •Median recorded time-to-solve fell from 26.1 hours in 2024 to 13.8 hours in 2026.
- •Teams completing the entire challenge board increased from two in 2024 and three in 2025 to 15 in 2026.
- •The company stated that the challenge board expanded during the same period.
The reported human-versus-AI result places an important boundary around current agent adoption in offensive security exercises. In comparable evaluation settings, score totals and completion times measure useful operational outcomes, but they do not by themselves separate model capability from human orchestration, team expertise, tool access, or challenge-selection decisions.
Security practitioners evaluating agentic systems can use such benchmarks to test end-to-end workflows rather than treating agent adoption as a standalone performance metric. Hack The Box founder and CEO Haris Pylarinos said the relevant question for security leaders is whether teams have the expertise to use AI "safely and effectively."
Key Points
- 1AI agents appeared on 68% of top-25 teams, showing that high-performing cybersecurity teams are already incorporating agent accounts into competition workflows.
- 2Human competitors still out-solved the leading AI team, limiting claims that current agents can independently match elite offensive-security performance.
- 3Benchmark median solve time fell by more than 12 hours, while full-board completions rose sharply despite an expanded challenge set.
Scoring Rationale
The benchmark provides timely empirical evidence about AI-agent use in competitive offensive-security workflows, an area of direct relevance to security engineers and AI evaluators. Its causal limits and lack of detailed per-task agent methodology reduce the result's broader significance, but the reported adoption and performance data make it notable.
Sources
Primary source and supporting public references used for this report.
Practice interview problems based on real data
1,625 SQL & Python problems across 15 industry datasets — the exact type of data you work with.
Try 250 free problems
