DeepSeek's DSec Sandbox Platform Handles 3 Million Agentic Training Runs a Day

DeepSeek Elastic Compute (DSec)

DeepSeek Elastic Compute (DSec) is a production sandbox platform that unifies FnCall, container, microVM, and full-VM backends under one SDK, coordinating placement and lifecycle across clusters. Co-designed with the reinforcement learning framework, it decouples stateful rollout execution from preemptible GPU training and mitigates reward hacking. A single unit spans about 160 nodes, serving roughly 3 million sandboxes daily, supporting over 380,000 concurrent sandboxes and more than 5,000 creations per second.

A single production-scale unit of DSec spans around 160 nodes, serving about 3 million sandboxes per day; in production, it supports over 380,000 concurrent sandboxes and sustains over 5,000 sandbox creations per second.
  1. flowerlad

    It seems every DeepSeek paper/patent has a huge number of authors, and this one is no exception. They couldn't even fit everyone on the page, there are 31 others not shown. This could be an asset protection strategy (i.e., human assets). Imagine if there were only 3 authors. Those authors may get hired away by competitors. If you list every employee on every paper then competitors don't know who to lure away.

  2. samayashar

    As models get better, safe-and-secure sandboxes/environments are going to be the way forward. With the recent rise in cases where models can somehow gain access to the internet and blast past the sandbox, it's very important to have all the resources contained within the sandbox with no access to the outside world.

    DSec is a good step in this direction!

  3. vblanco

    380.000 concurrent sandboxes on 160 Epyc based server nodes. Crazy stuff

  4. piterrro

    12 sandboxes per code is insane, I wonder how many of these sandboxes are idle at a time. Depending on the tasks assigned the resource requirements are different. Compare an agent doing pdf conversion and one responding to a simple question. One is cpu bound the other is mostly network wait.

    This is an interesting problem from infra perspective since you cannot predict the workload. On a bigger scale you may get away with forecasts.

    Im waiting for tech that elastically allocates cpu/mem without restarting a container.

  5. erulabs

    Appears to be similar to what Google is building with ax https://github.com/google/ax

  6. cloudengineer94

    Very similar to Google Ax, looks like it's Deepseek answer to Google

  7. throwaway7783

    Is this like agent substrate?

  8. jerrygenser

    I wonder if they are signalling that if they can do this for training, then they can create an style agent swarm to hack anyone with 380k concurrent agents.

More from this day

2026-09-26