Arena started as a research project out of UC Berkeley's Sky Computing Lab in 2024, letting users compare model responses and choose the better answer. At the time, models had more limited capabilities than they do today, and evaluations focused on the quality of their answers. Arena used E2B to build WebDev Arena, where users could compare two model-generated websites side by side. As models became more powerful, the partnership grew with E2B: from web development evaluations to Code Arena and, later, Agent Arena's evaluations of agents doing everyday professional work.
Today, Arena has tens of millions of monthly visitors and runs up to 600,000 E2B sandboxes each day (up from 300,000-400,000 at the time of interview). Its users build websites, conduct research, and generate reports and presentations, turning everyday work into a side-by-side test of how models actually perform. Supporting those workflows means provisioning isolated computers on demand, keeping them running through hours of work, and preserving the workspace when a user returns to continue.
“We have a massive platform of tens of millions of monthly visitors, which makes us one of the largest AI consumer apps in the world. And E2B provides the actual sandbox infrastructure to make that all possible.”
The challenge
Every evaluation needs its own environment
Arena has multiple types of evaluations on their platform: Agent, Chat, Code, Image, and Video. With Agent and Code, agents need to write code to create apps, install dependencies, start processes, and build a project that users could interact with.
Users can submit arbitrary prompts and upload files, and agents can execute commands that fail or leave processes running. Having a security boundary around each evaluation is essential so that one session's code and workspace could not interfere with another.
That isolation also makes the results trustworthy. Users are judging what an agent can build and accomplish, and its output should reflect its own performance. If not secured, files or processes left behind by another session could change the outcome, undermining the comparison.
A model release can change traffic overnight
Arena's traffic follows the frontier. When a new model appears on the platform, users arrive to discover what it can do, sometimes before the model is publicly available elsewhere. That makes demand difficult to forecast: the team has to prepare for launches whose timing and popularity it cannot always predict.
Startup speed mattered even outside those surges. A fast, responsive user experience is critical to keeping users engaged on the platform. The team wanted the lowest possible latency between a user submitting a prompt and the agent getting to work. Time spent provisioning a machine or installing basic dependencies delays that experience before the agent has accomplished anything. Arena needed environments that were ready quickly, both during normal traffic and when a new model drew a sudden crowd.
At the same time, the work inside those environments was getting longer. Some generations finished in minutes; others involved hours of research, tool use, code execution, and artifact creation. Even then, the user might not be done. They could submit another prompt right away or return days later to keep iterating. Arena needed sandboxes that preserved files and runtime state throughout a generation and between sessions, and that could be restored quickly without the cost of leaving inactive machines running.
Building a runtime would pull engineers away from evaluations
Arena considered building its own infrastructure with Docker and Kubernetes. That would have made the team responsible for provisioning, isolation, scaling, lifecycle management, and ongoing security maintenance. On top of that, each new product capability would potentially mean more infrastructure work before it could ship.
Arena's core work was building evaluations: observing how models and agents perform, collecting feedback from real users, and turning those signals into useful comparisons. Building and maintaining a cloud runtime would require a substantial engineering commitment alongside that work.
The solution
A full computer behind each agent session
When a user submits a task that requires execution, Arena provisions an E2B sandbox: a full Linux computer inside a Firecracker microVM, with its own kernel and an isolated workspace. The agent can execute code, install dependencies, manipulate files, and run processes without exposing another evaluation's environment to that work.
That freedom is essential to testing what agents can actually accomplish. In Code Arena, models build and run applications that users interact with and compare. In Agent Arena, users delegate research, coding, and other knowledge work that can require hours of execution. Arena returns the finished work to the user and evaluates both the result and the steps taken to produce it.
For Agent Arena, that means combining user feedback with execution signals: whether the agent completes the task, recovers from Bash errors, or attempts to call tools that don't exist. These signals are only useful if Arena can distinguish an agent's mistakes from problems in its execution environment. A Bash command that fails because the agent wrote it incorrectly tells Arena something about the agent's ability. A command that fails because another session modified a file or left a conflicting process running would distort that assessment. E2B's microVM isolation keeps each evaluation's files and processes separate, helping Arena accurately attribute the observed behavior to the agent being tested.
Ready for the first prompt and every follow-up
Arena uses E2B Templates to give agents a prepared workspace from the start. Code Arena launches with Next.js, Node.js, npm, Tailwind, and shadcn already installed, while Agent Mode uses a custom template with dependencies for knowledge work and coding tasks. Agents can begin working on the user's request without repeatedly installing foundational packages.
Once work begins, E2B keeps that environment available as the task evolves. The sandbox runs throughout long generations, and snapshots and pause/resume preserve files, processes, and runtime state between active periods. Whether a user sends a follow-up immediately or returns days later, the agent can pick up where it left off without rebuilding its workspace or leaving a machine running while idle.
Thousands of additional sandboxes before launch day
E2B gave Arena the capacity to meet sudden spikes in demand. Ahead of GPT-5's public release, users flooded Code Arena to try the model and compare the websites and applications they generated. E2B powered the sandboxes behind those comparisons, letting Arena handle the surge without the worry of being able to accommodate their users.
“I remember when GPT-5 was about to come out, we were testing a ton of different checkpoints. A lot of people were coming to Code Arena to experience the model firsthand. They were creating a ton of different comparisons of websites and different apps, and E2B was the backbone behind all of that. So when we had this massive surge, we didn't have to worry about, 'Hey, can we actually handle this?'”
That ability to scale mattered again ahead of Arena's June 2026 Agent Mode launch. After months of product development, the Arena team realized they needed capacity for thousands of additional concurrent sandboxes to handle the expected traffic the next day. Arena contacted E2B, and within hours, that additional capacity was ready to support the launch and the traffic that followed.
“It was mission critical that we had enough sandbox compute for the day of launch and the expected traffic surge we were going to get.”
An execution platform without a dedicated infrastructure team
E2B took on the sandbox infrastructure that Arena would otherwise have needed a dedicated team to build and operate. The initial integration was running in under two hours, giving Arena a working foundation for code execution without months of infrastructure development.
“What we do best is creating evaluations. The underlying infrastructure would have taken us many, many months and many, many senior engineers to actually build out. E2B was basically that team from day one.”
The results
With E2B, Arena has expanded from comparing model-generated websites to evaluating agents across coding and knowledge work, while keeping its engineering effort focused on the evaluation platform.
Under two hours to the initial implementation. Arena began prototyping with E2B without first building its own secure execution infrastructure.
Approximately 600,000 sandboxes per day. Arena reports tens of millions of E2B sandboxes launched over the lifetime of the platform.
Thousands of additional concurrent sandboxes within 24 hours. Within just a few hours, E2B provided the compute needed to launch Agent Mode to a platform with tens of millions of monthly visitors and support the traffic that followed.
Looking ahead
Arena expects agents to take on increasingly long tasks. Today's multi-hour workflows could extend to days or potentially weeks as users delegate more substantial projects. Evaluating those agents will require environments that can support the work for as long as it takes.
E2B has supported Arena from its first web development evaluations in a research lab to hundreds of thousands of sandboxes a day. The next generation of agents will need new ways to prove themselves, and E2B will continue to power the environments where Arena makes that possible.