Dax Benchmark

A real-world workload benchmark inspired by this tweet from dax to test potential sandbox providers for OpenCode.

Per-Phase Duration Breakdown

Median duration per phase of the clone + install + typecheck cycle.

Detailed Metrics

Provider
Phases
Success
Total (med)
Prepare
Bun DL
Bun Unpack
Clone
Install
Typecheck
Miosa7/7100%39.32s1.37s0.63s0.44s2.41s11.22s21.37s
Isorun7/7100%39.54s1.15s0.28s0.43s1.21s8.82s22.81s
Namespace7/7100%40.22s4.01s0.29s0.43s1.22s8.11s20.40s
Blaxel7/7100%45.60s1.21s0.43s0.48s1.78s10.80s23.98s
CreateOS7/7100%49.17s2.73s0.45s0.48s2.54s11.95s26.69s
Upstash7/7100%50.44s3.86s0.26s0.62s1.68s9.49s25.01s
Mosaic7/7100%59.22s1.21s0.38s0.62s2.44s14.69s28.70s
Arker7/7100%62.55s12.04s0.34s0.62s2.18s11.18s32.18s
Sail7/7100%69.39s4.32s0.84s0.90s3.56s16.32s34.43s
Vercel7/7100%69.68s12.96s0.28s0.68s1.88s11.51s36.17s
E2B7/7100%71.97s7.72s0.66s0.72s2.89s14.65s34.38s
Tenki7/7100%75.34s6.07s0.67s0.63s4.14s19.02s40.13s
CodeSandbox7/7100%81.23s0.47s0.60s3.25s6.55s57.70s
Tensorlake7/7100%92.64s19.27s0.30s1.10s3.03s17.68s44.90s
Modal7/7100%97.71s8.54s0.30s0.70s2.37s19.78s58.97s
Runloop7/7100%114.37s16.19s0.35s0.77s10.31s28.34s52.96s
Superserve7/7100%119.82s11.23s0.43s1.12s5.67s19.24s78.88s
Daytona7/7100%141.26s8.31s0.34s0.68s4.11s37.19s34.93s
Declaw7/7100%172.63s7.77s0.74s0.72s4.15s25.08s122.56s
OpenComputer3/70%Failed6.25s
Cloud Run2/70%Failed
Sandbox02/70%Failed
Archil0/70%Failed
Beam0/70%Failed
Cloudflare0/70%Failed
Hopx0/70%Failed
Lightning AI0/70%Failed
Microsandbox0/70%Failed
Northflank0/70%Failed
Run Cloud0/70%Failed

Performance Over Time

Want to see a provider added?

Let us know on X

Methodology

What We Measure

Each run executes a cold clone + install + typecheck cycle of the opencode repository inside a fresh sandbox: install system packages and Node.js, download and unpack Bun, shallow-clone opencode at a pinned commit, run bun install, then bun typecheck.

Each provider runs 1 iteration per scheduled benchmark run. Benchmarks run via GitHub Actions and results are committed to the public benchmarks repo.

How We Score

The Composite Score is the median number of the 7 benchmark phases (prepare, cache clear, bun download, bun unpack, clone, install, typecheck) a provider completed before failing, if any — shown as a fraction like 7/7.

The other metrics (Total, Prepare, Clone, Install, Typecheck) are real-world phase durations in seconds — providers are ranked by median duration (lower is better) when one of those is selected. The benchmark requires curl, root/sudo, and a package manager (apt, dnf, or apk) inside the sandbox — providers without those capabilities fail early, which shows up as a low phase count.

PartnersLatitude