Sandbox Benchmarks - Provider Leaderboard
Sandbox Benchmarks
A leaderboard of common benchmarks for each of our sandbox providers.
Powered byPerformance Over Time
Composite Score
Detailed Metrics
Provider | Score | Median | P95 | P99 | Success |
|---|---|---|---|---|---|
| Daytona | 34.9 | 0.22s | 1.05s | 1.07s | 37% |
| CreateOS | 96.5 | 0.34s | 0.38s | 0.38s | 100% |
| Northflank | 96.0 | 0.38s | 0.42s | 0.42s | 100% |
| Arker | 95.7 | 0.42s | 0.44s | 0.49s | 100% |
| Blaxel | 94.2 | 0.49s | 0.72s | 0.74s | 100% |
| Isorun | 95.0 | 0.49s | 0.51s | 0.51s | 100% |
| Declaw | 94.2 | 0.54s | 0.64s | 0.64s | 100% |
| Archil | 93.7 | 0.59s | 0.68s | 0.69s | 100% |
| Lightning AI | 90.4 | 0.72s | 1.29s | 1.35s | 100% |
| Vercel | 55.5 | 0.74s | 26.00s | 31.00s | 100% |
| Mosaic | 92.1 | 0.76s | 0.84s | 0.85s | 100% |
| Cloud Run | 90.1 | 0.79s | 1.23s | 1.35s | 100% |
| Modal | 91.5 | 0.81s | 0.89s | 0.95s | 100% |
| E2B | 89.7 | 0.83s | 1.10s | 1.14s | 99% |
| Beam | 87.0 | 1.18s | 1.45s | 1.51s | 100% |
| Runloop | 81.3 | 1.21s | 2.86s | 2.88s | 100% |
| Superserve | 85.5 | 1.22s | 1.78s | 1.80s | 100% |
| Tenki | 82.5 | 1.59s | 1.98s | 2.00s | 100% |
| Tensorlake | 78.7 | 2.08s | 2.20s | 2.21s | 100% |
| OpenComputer | 48.0 | 2.12s | 7.12s | 18.45s | 88% |
| Sail | 68.5 | 2.98s | 3.42s | 3.42s | 100% |
| Cloudflare | 46.8 | 4.84s | 6.00s | 6.11s | 100% |
| Upstash | 29.4 | 6.02s | 8.63s | 8.63s | 100% |
| Sandbox0 | 0.0 | 14.83s | 20.66s | 21.59s | 100% |
| CodeSandbox | 0.0 | 35.12s | 53.46s | 53.78s | 88% |
| Run Cloud | 0.0 | 44.96s | 118.01s | 118.04s | 100% |
Want to see a provider added?
Methodology
What We Measure
Every benchmark measures Time to Interactive (TTI) — the elapsed time from calling compute.sandbox.create() to the first successful runCommand() inside the sandbox.
Each provider is tested with 100 iterations per run. Benchmarks run automatically via GitHub Actions on a recurring schedule. All results are committed to the public benchmarks repo.
Burst Test: All sandboxes are launched concurrently in a single burst.
How We Score
The Composite Score is a weighted blend of timing metrics multiplied by the success rate. Each metric is scored against a fixed 10-second ceiling: 100 × (1 − value / 10,000ms), so a 200ms median scores 98 and anything ≥10s scores 0.
The weighted timing score is then multiplied by the success rate (0–1), so providers that fail frequently are penalized proportionally.
- • Median: 60% — primary signal for typical experience
- • P95: 25% — tail latency / consistency
- • P99: 15% — extreme tail latency
Sandbox Benchmarks FAQs
Have another question? Email us.