DeepSeek's DSec: 380K concurrent agentic training sandboxes on 160 EPYC CPU nodes
- DeepSeek released DSec (DeepSeek Elastic Compute), a sandbox infrastructure built for agentic training at scale, as arXiv paper 2609.22978 with a 131 person author list.
- One DSec scale unit spans nearly 160 CPU nodes with 30K cores and about 250 TB of DRAM, and it manages petabytes of container layers and images.
- A single scale unit serves about 3M sandbox instances on a typical day, with peak concurrency near 380K and a creation rate above 5,000 instances per second.
Hacker News opinions
Is there any lab more innovative than DeepSeek right now? Imagine what they would ship with the compute Anthropic and OpenAI sit on.
I would argue the constraints are what produce that creativity. Give them unlimited compute and they might not be the same lab.
Short term maybe less compute, long term almost surely more. No energy crunch, nobody blocks their data centers, and the only real gap is a domestic chip on Nvidia's level, which I would bet gets closed in a year or so.
Other companies are plenty innovative, they just do not write papers about what they do. Publication volume is not the same as innovation.
380,000 concurrent sandboxes running on 160 EPYC server nodes. That is an absurd amount of parallel agentic environment.
Divide it out and that is only 12 sandboxes per core, which is the number that actually surprised me.
Full disclosure, I have not read the paper yet, just scrolled past the author list. 131 authors on one systems paper might be a record.
Hard to beat the ATLAS Higgs paper though. That one ran pages 25 to 32 with around 5,150 authors plus five pages of affiliations.
To me the interesting part is not the topic, it is how 131 people coordinated to produce one paper. Maybe it is lab level socialism where everyone gets equal credit regardless of input on this specific work.
I genuinely do not get why half this thread is about the author list. Large scale physics and biology experiments always publish like this, and every GPT release from OpenAI had papers with long author lists.
My read is that the long list is asset protection. With 3 authors, rivals know exactly who to poach. With 131, good luck figuring out who actually built it.
That theory does not survive clicking the link. arXiv truncates the abstract page, and the PDF lists all 131 authors plus the 31 not shown.
Does not matter either way, competitors will just approach everyone on the list anyway.
Could also roll into an whole team aqui hire, the way Nvidia picked up Groq people. Buying 100 engineers at once is borderline impossible for a normal company.
This looks close in spirit to what Google is building with ax (github.com/google/ax), though DSec is aimed at agentic training environments rather than general agent tooling.
Stand back and digest the unit numbers: 30K cores, about 250 TB DRAM, petabytes of layers and images, 3M sandboxes a day, 5,000+ creations per second. AI agents, not people, are turning into the biggest consumers of cloud compute.