The First Token Is the Hardest
Cold starts, runtime stability, and the physics of serving large models
We spent twenty years making web servers boring. Stateless, small, fast to boot, and cheap to leave idle. Large models break every one of those assumptions at once, and most of the instability people blame on "AI" is really physics we haven't budgeted for.
Here's a scene most platform teams have lived through. An endpoint has been quiet all night. At 8:02 a.m. someone opens the product, types a question, and watches a spinner for three minutes. By 8:05 the dashboard is fine, the on-call engineer sees nothing wrong, and the user has already left. Nothing crashed. The system did exactly what we built it to do. It just took three minutes to become a system.
A container running a web API is typically tens of megabytes and boots in well under a second. A container serving a 70-billion-parameter model has to pull an image that is often over ten gigabytes, initialize eight GPUs and the links between them, move about 140 GB of weights into high-bandwidth memory, compile kernels, capture execution graphs, and warm up before it can emit a single token. That is a gap of four to five orders of magnitude in state that has to exist before the first byte of response. Every pattern we inherited from stateless compute, like scale-to-zero, aggressive bin packing, and "just add a replica," was tuned for the small end of that gap.
This piece is my attempt to lay the whole problem out flat. I don't think "cold starts are slow" is one problem. I count at least eight dimensions, each with its own unit of measure, and they compound. Fixing one while ignoring the others is how teams end up with a fast boot and a system that still falls over at 9 a.m.
| Dimension | The question | Unit |
|---|---|---|
| Time | How long until a new replica can answer? | seconds per cold start |
| Bytes | How much has to move, and over which link? | GB ÷ GB/s |
| Traffic | How often does a request find nobody home? | P(cold) |
| Load | What happens to latency as warm replicas fill up? | 1 / (1 − ρ) |
| Memory | How many conversations fit at once? | KV bytes per token |
| Neighbors | How much does everyone else slow me down? | tokens/s per stream |
| Numerics | Do I get the same answer twice? | reduction order |
| Fleet | How often does a user feel the worst case? | 1 − pⁿ |
1. Anatomy of a cold start
A "cold start" is really seven separate waits stacked end to end. Some are about getting a machine, some are about moving bytes, and some are about rebuilding runtime state that was identical the last time this model booted:
Figure 1 models that sum for four increasingly prepared deployments. The worst case, a brand-new GPU node pulling everything from object storage, lands around seven minutes. The best case, with weights already pinned in host memory and compiled graphs restored from cache, is about twenty seconds. That is a 19× difference on identical GPUs running an identical model.
Two things stand out to me. First, the GPU is idle for almost the entire timeline. Every bar except the last one is disk, network, CPU, or a scheduler. You're renting the most expensive silicon in the building so it can wait on I/O.
Second, the dark "compile + CUDA graphs" segment is recomputation. Serving engines capture execution graphs for a set of batch sizes so each decode step skips per-kernel launch overhead, and many stacks also run a compiler pass. That work depends only on the model, the GPU type, and the parallelism layout, and none of those changed since the last boot. When a stage produces the same output every time, it belongs in a cache, not on the critical path.
2. Bytes over bandwidth
Strip away orchestration and the fetch and load stages reduce to one line of physics:
Here is bytes per parameter (2 for bf16, 1 for fp8) and is the slowest link on the path. For our reference model, GB. The table of possible links spans more than two orders of magnitude:
Over a single object-store stream the model takes about 140 seconds to arrive. Over local NVMe it takes about 20. From pinned host memory each GPU only needs its own 1/8 shard, about 17.5 GB, over its own PCIe link, which is well under a second in theory. In practice, deserialization and allocation overhead push that into single-digit seconds. That is also why a zero-copy, memory-mappable format like safetensors [13] matters more than it looks.
The uncomfortable implication is that model size grows cold start linearly unless you change tiers. A team that moves from an 8B to a 405B model on the same infrastructure has quietly multiplied its transfer floor by 25. ServerlessLLM [3]attacks exactly this. It treats the GPU server's local memory and storage hierarchy as a checkpoint cache, uses a loading-optimized format, and schedules requests toward the servers that already hold the bytes. The general lesson is locality-aware scheduling: route the request to where the weights already are, rather than moving weights to where the request landed.
3. How often you pay: the traffic dimension
A cold start only hurts when a request actually waits on one. Suppose a replica scales to zero after an idle timeout and requests arrive as a Poisson process at rate . The gap between arrivals is exponentially distributed, so the chance a request arrives after the replica has gone cold is
An endpoint that sees one request every five minutes, with a ten-minute timeout, sends of its users into a cold start. One request every twenty minutes means . That long tail of quiet endpoints is not hypothetical. In Microsoft's characterization of production serverless traffic, most applications were invoked far less than once per minute on average, and invocation rates varied by many orders of magnitude between applications [4]. Model endpoints are typically lumpier still: internal tools, per-customer fine-tunes, evaluation jobs, and features that only wake up during business hours.
There's a subtler result hiding in that formula that I think most autoscaling policies miss. The rate of cold starts, not the probability, is what costs you GPU time and pages the on-call:
The endpoints that generate the most cold starts are neither your busiest nor your quietest. They are the ones whose traffic rate is about one request per idle timeout. With a 10-minute timeout, an endpoint at roughly six requests an hour can churn through more than two cold starts every hour, forever. The busy endpoint never goes cold, and the dead one rarely wakes. The middle is where your fleet flaps.
4. Spikes: where cold start becomes a queue
Scale-from-zero is the visible case. The more damaging case is scale-from-one during a traffic spike, because every second of cold start becomes queued requests. Let traffic jump to rate against one warm replica with service rate . The autoscaler reacts immediately and requests a second replica, which becomes useful after . The backlog is a triangle:
With and requests per second, a 30-second cold start peaks at 60 queued requests and is fully recovered a minute after the spike. A 180-second cold start peaks at 360 queued requests and takes six minutes to recover. By Little's law [5], the average wait grows in proportion to the queue.
The linear model is optimistic, because real clients don't wait politely. If client timeouts are shorter than the cold start, users and SDKs retry. Each retry is a new arrival, so the effective rises exactly when capacity is lowest, and the triangle turns into a ramp. This is the mechanism behind most "the model provider was down" incidents I've seen: a slow boot, short timeouts, and retry logic combine into a retry storm. The fix is less about faster GPUs and more about making cold start shorter than your client timeout and giving clients a clear signal to back off.
5. Warm isn't the same as stable
Say you solve cold starts completely and every replica is always warm. The next instability is the one finance asks for. GPUs are expensive, so the pressure is to run them hot. Queueing theory has been unambiguous about the cost of that since the 1970s [6]. For the simplest model, a single server with random arrivals and service times (M/M/1), the mean time in the system is
where is the service time and is utilization. At 50% utilization, requests spend twice their service time in the system. At 80% it's 5×, at 90% it's 10×, and at 95% it's 20×.
LLM serving is kinder than M/M/1 in one way and crueler in another. Continuous batching lets one replica serve many requests at once, which raises effective . But service time varies enormously, since one request asks for 20 tokens and the next asks for 4,000. Variance is what feeds a queue. In the Pollaczek–Khinchine form for general service times, mean waiting time scales with , where is the coefficient of variation of service time:
Output lengths with a coefficient of variation of 2 make queueing delay 2.5× worse than the same average load with fixed-length work. So the right utilization target for a model fleet is lower than it is for a web fleet, not higher. If your capacity plan says "run GPUs at 90% to justify the spend," your latency plan has already been decided for you.
6. Memory: the KV cache is the real capacity
GPU memory holds two things: weights, which are fixed, and the key/value cache, which grows with every token of every active conversation. The per-token cost comes straight from the architecture:
The leading 2 is for keys and values. Grouped-query attention, with 8 KV heads instead of 64, already makes this 8× smaller than it would otherwise be [8]. Our node has GB of HBM. Reserve about 10% for headroom, subtract 140 GB of weights and some activation and graph workspace, and roughly 400 GB is left for KV. That buys about 1.2 million tokens of live context, shared by everyone on the node.
That's why context length is a capacity decision, not a product checkbox. The same hardware serves 596 concurrent 2k-token chats or 9 concurrent 128k-token document sessions. When demand exceeds what fits, the engine has to preempt a sequence, either swapping its cache out or throwing it away and recomputing it later. The user sees a stall with no error. PagedAttention [2]was a major step because it removed the fragmentation that used to waste much of this memory, so more of the budget holds real tokens. It can't change the arithmetic of the budget itself.
7. Your latency depends on your neighbors
Web requests on the same server are mostly independent. LLM requests in the same batch are not. Decoding one token for every sequence in the batch requires reading all the weights from HBM once, plus each sequence's entire KV cache. At small and medium batch sizes that step is limited by memory bandwidth, not compute [9], so the time per step and each stream's speed are roughly
Here is weight bytes, is the context length of sequence , and is KV bytes per token. Alone on the node with 4k of context, a stream tops out near tokens per second. Put 64 such streams in the batch and each drops to about 119 tokens per second while aggregate throughput rises to about 7,600. At 256 streams, each gets about 55. A quick compute check confirms this is still bandwidth-bound: 256 tokens × ~140 GFLOP ≈ 36 TFLOP per step, about 4.5 ms of the node's dense bf16 peak, versus about 18 ms spent reading memory.
This is the property I find most underappreciated. A user's tokens-per-second depends on how many strangers share their batch and how long their conversations are.The same prompt at the same hour can stream at 190 or 55 tokens per second, and nothing is broken either way. Operators trade per-user speed for aggregate throughput every time they raise the batch limit. That's a legitimate trade, but it should be a stated SLO, not an accident.
Prefill makes it worse. Processing a new 32k-token prompt is a large, compute-bound burst. If the scheduler puts it into the same iteration as a hundred decoding streams, every one of those streams stalls for that iteration, and users watching output see the text freeze mid-sentence. Sarathi-Serve [10] splits prefills into chunks that fit alongside decodes. DistServe [11]moves prefill and decode onto separate GPU pools and optimizes goodput, meaning requests that meet their latency target, rather than raw throughput. Both start from the same observation: the two phases have opposite hardware profiles, and mixing them naively turns one user's long prompt into everyone else's latency spike.
8. Numerics: same prompt, different answer
Stability isn't only about time. Teams are routinely surprised that the same prompt at temperature zero can produce different outputs on different calls. The root cause is that floating-point addition isn't associative:
GPU kernels sum long vectors in parallel, and the order of those sums can depend on how the work is split. The common assumption is that concurrency randomness in atomics is to blame. He and colleagues at Thinking Machines argue that the dominant cause in typical inference engines is different: kernels that aren't batch-invariant, where the reduction strategy for your request changes with batch size [12]. Batch size changes with server load, which you don't control. So your output depends, in the last few bits, on who else is using the service. Greedy decoding amplifies last-bit differences: once two logits swap order, the sequences diverge and never reconverge.
Across a model upgrade, or a change in tensor-parallel layout, or a new GPU type, the differences get larger. That's why evaluation results don't transfer cleanly across deployments that should be "the same model." Batch-invariant kernels exist and cost some throughput. I'd treat them like strong consistency in a database: something you enable on the paths that need reproducibility (evals, audits, caching, regression tests), not a global default you either always pay for or never get.
9. The fleet: tail at scale, now with agents
Dean and Barroso's observation from 2013 is more relevant to AI systems than to the search backends it was written about [1]. If each call is slow with probability , a request that makes calls hits at least one slow call with probability
In 2013, was the fan-out of a search query across index shards. In 2026, it's the number of model calls behind one user action: a retrieval step, a reranker, a guardrail, a planner, and a dozen tool-calling turns in an agent loop. A 20-step agent whose every call meets a p99 target still gives 18% of users a p99 experience somewhere in the run. At 100 calls, 63% do. In a sequential agent, the cost is worse than in a parallel fan-out because slow steps add up instead of overlapping. Every dimension above, from cold starts and queueing to KV preemption and noisy neighbors, feeds that , and agents raise it to a power.
What to actually measure
Most model dashboards I review show GPU utilization, average latency, and error rate. By the argument above, all three can look healthy while users suffer. These are the numbers I'd put on the first screen:
| Metric | Definition | Why it matters |
|---|---|---|
| Cold start, by stage | Wall time for provision, pull, init, fetch, load, compile, and warmup, recorded separately | A single total hides which tier to fix |
| Cold-hit rate | Share of requests that waited on a replica starting up | The user-facing cost of your scale-to-zero policy |
| TTFT p50 / p99 / p99.9 | Time to first token, including queue time | What users experience as "is it working?" |
| Inter-token latency p99 | Gap between consecutive streamed tokens | Catches prefill interference and batch pressure |
| Preemptions / min | Sequences evicted or recomputed because KV memory ran out | Silent latency and wasted compute |
| Goodput | Requests per second that met their latency target | Throughput that doesn't count failures as work |
| Output drift | Rate at which identical greedy requests produce different tokens | Stability of the answer, not just the server |
A playbook, one dimension at a time
Tier the bytes
Keep hot checkpoints in host RAM or local NVMe on GPU nodes, and use parallel ranged reads for everything else. Memory-map safetensors instead of unpickling. The goal is to make weight fetch a copy between local tiers, never a download from a region.
Snapshot the runtime, not only the weights
Persist compile caches and captured CUDA graphs alongside the checkpoint, keyed by model, GPU type, and parallelism layout. The dark bars in Fig. 1 are pure recomputation of something that hasn't changed since the last boot.
Warm pools sized by math
Endpoints with λ ≈ 1/T produce the most cold starts per hour. Keep a minimum replica there, let very quiet endpoints stay cold, and pre-warm on predictable signals like time of day or an upstream deploy.
Admit, don't just queue
Past about ρ = 0.8, an extra queued request mostly adds latency for everyone behind it. Shed or redirect load based on token budgets and deadlines, and give clients honest backpressure instead of silent timeouts that turn into retry storms.
Separate prefill from decode
Chunk long prefills so they can't stall every decoding stream in the batch, or split them onto their own pool. Measure inter-token latency separately from time to first token.
Budget KV cache like money
Price context length explicitly. Cap it per tier, share prefixes across requests, and alert on preemptions. At 128k context, a 70B node that fits hundreds of 2k chats fits about nine sequences.
Pin numerics where answers matter
Use batch-invariant kernels and a fixed parallelism layout for evaluations, audits, and anything cached or compared. Pay for determinism only on the paths that need it.
Hedge the tail
For fan-out and agent loops, send a backup request after the p95 delay and cancel the loser. A small increase in load buys a large cut in tail latency.
Where this leaves us
The web taught us to treat compute as stateless and disposable, and it worked because the state was small. A model server is the opposite: a few hundred gigabytes of state that takes minutes to assemble, shared by strangers whose presence changes your latency and even your answer. We keep applying stateless patterns to it and calling the results "AI being flaky."
My view is that runtime stability for models is a state-placement problem. It's about where the weights live, where the compiled runtime lives, and where each conversation's KV cache lives, and about routing work to the state instead of rebuilding the state for the work. The teams that get this right will ship systems that feel instant at 8:02 a.m. The others will keep buying faster GPUs and wondering why the spinner is still there.
I'm still working through how this changes once models are routinely swapped per request, which is the same statefulness question I keep coming back to in my notes on sidecar context architectures. If you run inference at scale and your numbers disagree with my model, I want to see them. That's the most useful email I can get.
References
- [1]Dean, J., & Barroso, L. A. (2013). The tail at scale. Communications of the ACM, 56(2), 74–80. link
- [2]Kwon, W., et al. (2023). Efficient memory management for large language model serving with PagedAttention. Proceedings of the 29th ACM Symposium on Operating Systems Principles (SOSP '23). link
- [3]Fu, Y., et al. (2024). ServerlessLLM: Low-latency serverless inference for large language models. 18th USENIX Symposium on Operating Systems Design and Implementation (OSDI '24). link
- [4]Shahrad, M., et al. (2020). Serverless in the wild: Characterizing and optimizing the serverless workload at a large cloud provider. USENIX Annual Technical Conference (ATC '20). link
- [5]Little, J. D. C. (1961). A proof for the queuing formula L = λW. Operations Research, 9(3), 383–387.
- [6]Kleinrock, L. (1975). Queueing Systems, Volume 1: Theory. Wiley.
- [7]NVIDIA. H100 Tensor Core GPU datasheet (SXM5: 80 GB HBM3, 3.35 TB/s). link
- [8]Grattafiori, A., et al. (2024). The Llama 3 herd of models. arXiv:2407.21783. link
- [9]Pope, R., et al. (2022). Efficiently scaling transformer inference. arXiv:2211.05102. link
- [10]Agrawal, A., et al. (2024). Taming throughput-latency tradeoff in LLM inference with Sarathi-Serve. OSDI '24. link
- [11]Zhong, Y., et al. (2024). DistServe: Disaggregating prefill and decoding for goodput-optimized large language model serving. OSDI '24. link
- [12]He, H., & Thinking Machines Lab (2025). Defeating nondeterminism in LLM inference. Connectionism. link
- [13]Hugging Face. Safetensors documentation. link