Modal rebuilt its core platform for sandboxes, isolated environments for running code, from the ground up.

The new system is designed for millions of concurrent sandboxes and tens of thousands of launches per second. In a demonstration, the company launched one million sandboxes concurrently in under a minute.

Several scheduling servers process creation requests in parallel. Each selects a worker using in-memory cached data and asks it to launch the sandbox directly through RPC, a remote procedure call. Workers publish their state periodically, and the scheduling servers consume it asynchronously, so data stores are not used when a sandbox is created.

In the test, the median time until a sandbox was ready was below half a second, and Modal attributes long delays to kernel and network contention when many sandboxes start at once and plans to reduce them. The new system is already available in beta and should soon handle all of Modal’s sandbox scheduling.

Claim check:

  • Over the past few months, Modal rebuilt its core sandbox platform from the ground up; the new system is designed for millions of concurrent sandboxes and tens of thousands of creations per second. (confirmed by the publication itself: evidence; «Over the last few months, we’ve rebuilt our core sandbox platform from the ground up for both scale and reliability. On our new system, users can run millions of sandboxes concurrently and create tens of thousands of sandboxes per second.»)
  • In a demonstration, Modal ran one million sandboxes concurrently and created them all in under a minute. (confirmed by the publication itself: evidence; «As a demonstration of what our platform is capable of, we’ve run a million sandboxes concurrently, creating all 1 million in under a minute.»)
  • Several scheduling servers process creation requests concurrently, select a worker using in-memory cached data, and directly request a launch from the worker through RPC. (confirmed by the publication itself: evidence; «Rather than a single, serialized scheduler, we run a fleet of scheduling servers which handle sandbox creation requests concurrently. To handle a creation request, a scheduling server runs a fast scheduling algorithm against in-memory cached data. Once a scheduling server decides which worker to create a sandbox on, it contacts the worker directly via RPC to request that a sandbox is created.»)
  • Workers publish their state periodically, scheduling servers consume it asynchronously, and data stores are not used when a sandbox is created. (confirmed by the publication itself: evidence; «Workers publish their state periodically into a Redis stream. The scheduling servers consume this state asynchronously and use it to make scheduling decisions. We have no data stores in the critical path of sandbox creation at all.»)
  • In the test, the median time until a sandbox was ready was below half a second; Modal attributes long delays to kernel and network contention and plans to reduce them. (confirmed by the publication itself: evidence; «Sandbox start times on our new system (the latency from when the client first tries to create a sandbox, to when the sandbox can run user code) are less than half a second at the median, and remain solid at scale. We attribute much of this tail to kernel and network contention (including the rtnl lock contention mentioned previously) when many sandboxes start simultaneously on the same worker, and we’re working to reduce it.»)
  • The new system is already available in beta and should soon handle all of Modal’s sandbox scheduling. (confirmed by the publication itself: evidence; «Soon this new system will back all sandbox scheduling at Modal, but it’s already available in Beta.»)

Publications:

Primary sources:

score 70.5 out of 100 · kind: guide