Methodology

How the Numbers Are Made

Every benchmark is reproducible: open code, fixed workloads, results published in full

1

One workload, every provider

The same script, payload, and model hit every provider, addressed the way its own API expects. Nobody gets a tuned path.

2

Repeated on a schedule

Every suite runs many iterations per provider, automated on a recurring schedule, so results reflect steady state rather than one lucky request.

3

Scored on outcomes, not one lucky run

Every family has its own scoring function, but none reward a single fast run. Latency is blended from median and tail percentiles, completion is counted in phases, and coverage or quality is measured where those matter. All scored families multiply by success rate so providers that fail often lose real points.

4

Raw results published

Every run is committed as JSON to the public benchmarks repo, next to the code that produced it. Nothing is cherry-picked or held back.

Trusted by the best

Independent benchmarks of every major sandbox provider, sponsored by industry leaders.

LatitudeGoogle Cloud RunBrowserbaseTigrisNeonGitbookNamespace

Start Benchmarking

We provide independent benchmarking for all compute providers.