03Agentic AI Environments & Benchmarks
ENVIRONMENTS& BENCHMARKS
Where our agents are trained, stressed, and measured — realistic worlds with thousands of long-horizon tasks and transparent scoring.
Web automation
Opus WebArena
A sandboxed browser world with thousands of realistic multi-step tasks across commerce, CMS, and dashboards.
4,820 tasksSoftware engineering
Terminal World
Containerized repos with failing tests; agents must read, patch, and verify across long horizons.
2,140 reposDevOps & SRE
OpsSim
Live incident simulations with logs, metrics, and traces — agents triage and remediate under SLOs.
960 incidentsMulti-agent
Negotiation Grid
Cooperative and adversarial games measuring strategy, theory-of-mind, and tool use.
1,500 scenarios// BENCHMARK RESULTS
ModelWebArenaTerminalOpsSimOverall
Vega Reasoner
Atlas + Vega
GPT-class baseline
Open baseline