Command Palette
Search for a command to run...

Nvidia Launches Dynamo Snapshot, Cutting Kubernetes Inference Startup Time to Under 5 Seconds

aiai-infrastructureai-inference-platforms 3 posts · 2 accounts

Nvidia introduced Dynamo Snapshot, a Kubernetes tool designed to speed startup for inference workloads by reducing startup time from minutes to under five seconds. The company said the software addresses cold starts in production inference deployments and, in one cited workload, cut gpt-oss-120b restore times to less than five seconds.

The launch targets a common problem in scaling LLM inference replicas elastically on Kubernetes, where demand can shift quickly and cold-starting workloads leaves GPUs idle. Nvidia said Dynamo Snapshot uses concurrent weight restoration over a high-speed interconnect, along with Linux native AIO and parallel memfd restoration, to accelerate CRIU restore performance.

From the sources (3 posts)

@nvidiaai

Introducing Dynamo Snapshot, our approach for fast startup for inference workloads on Kubernetes, which reduces startup time from minutes to under 5 seconds. In production inference deployments demand fluctuates over time. Cold-starting in

@nvidiaai

You can read the full deep dive here:

@hut_

Scaling LLM inference replicas elastically on Kubernetes usually means brutal cold starts. My team and a large group of engineers across NVIDIA just launched NVIDIA Dynamo Snapshot to solve this—slashing gpt-oss-120

Preview built on a synthetic news corpus (16 weeks, Apr–Jul 2026). Impact calls are model reads, not price data.

About Archive