Skip to content
GitHub

Product

Benchmarks Partners

Resources

Docs Blog

Sandbox Benchmarks - Provider Leaderboard

Google Cloud Run logo Agents, sandboxes, & vibe-coded apps are powered by Cloud Run.

Sandbox Benchmarks

A leaderboard of common benchmarks for each of our sandbox providers.

Powered by Namespace · 4 vCPU / 16GB RAM · Northern Virginia, US
Last run: August 14, 2026

Performance Over Time

Composite Score

Detailed Metrics

Provider
Score
Median
P95
P99
Success
Daytona34.90.22s1.05s1.07s37%
CreateOS96.50.34s0.38s0.38s100%
Northflank96.00.38s0.42s0.42s100%
Arker95.70.42s0.44s0.49s100%
Blaxel94.20.49s0.72s0.74s100%
Isorun95.00.49s0.51s0.51s100%
Declaw94.20.54s0.64s0.64s100%
Archil93.70.59s0.68s0.69s100%
Lightning AI90.40.72s1.29s1.35s100%
Vercel55.50.74s26.00s31.00s100%
Mosaic92.10.76s0.84s0.85s100%
Cloud Run90.10.79s1.23s1.35s100%
Modal91.50.81s0.89s0.95s100%
E2B89.70.83s1.10s1.14s99%
Beam87.01.18s1.45s1.51s100%
Runloop81.31.21s2.86s2.88s100%
Superserve85.51.22s1.78s1.80s100%
Tenki82.51.59s1.98s2.00s100%
Tensorlake78.72.08s2.20s2.21s100%
OpenComputer48.02.12s7.12s18.45s88%
Sail68.52.98s3.42s3.42s100%
Cloudflare46.84.84s6.00s6.11s100%
Upstash29.46.02s8.63s8.63s100%
Sandbox00.014.83s20.66s21.59s100%
CodeSandbox0.035.12s53.46s53.78s88%
Run Cloud0.044.96s118.01s118.04s100%

Want to see a provider added?

Let us know on X

Methodology

What We Measure

Every benchmark measures Time to Interactive (TTI) — the elapsed time from calling compute.sandbox.create() to the first successful runCommand() inside the sandbox.

Each provider is tested with 100 iterations per run. Benchmarks run automatically via GitHub Actions on a recurring schedule. All results are committed to the public benchmarks repo.

Burst Test: All sandboxes are launched concurrently in a single burst.

How We Score

The Composite Score is a weighted blend of timing metrics multiplied by the success rate. Each metric is scored against a fixed 10-second ceiling: 100 × (1 − value / 10,000ms), so a 200ms median scores 98 and anything ≥10s scores 0.

The weighted timing score is then multiplied by the success rate (0–1), so providers that fail frequently are penalized proportionally.

  • Median: 60% — primary signal for typical experience
  • P95: 25% — tail latency / consistency
  • P99: 15% — extreme tail latency

Sandbox Benchmarks FAQs

Have another question? Email us.

A sandbox is anywhere you can run code in isolation. It could be a VM, bare metal, a container, anywhere with compute resources.
PartnersLatitude