Measuring Cold Starts on a 27B Hugging Face Endpoint

October 1, 2026
8 min read
By Rahat Kabir

Contents

Why I did this

I had $10 of Hugging Face compute credits and a question I couldn’t find a straight answer to: everyone recommends “scale to zero to save money,” but almost nobody measures what that actually costs when the traffic comes back.

So I turned it into a weekend experiment. I deployed a dedicated Inference Endpoint, scaled it to zero three times, timed every recovery, and paid attention to everything the hourly bill doesn’t show. This post is the full write-up — method, raw numbers, and the cost math, including one claim from my first draft that turned out to be wrong.

The short version

  • I deployed Qwen/Qwen3.8-27B (revision 1d4bf0f2) on 1x Nvidia A100 (80 GB) via Hugging Face’s vLLM catalog recipe ($2.50/hr, AWS us-east-1), scaled it to zero three times, and timed recovery.
  • In my tests, waking this deployment took ~6.6–10 minutes before a successful answer completed (median 6.6 min). Warm answers completed in about 2 seconds (TTFT median 1.49 s).
  • Every wake-up began with 36–50 consecutive HTTP 503s — no request queue, hard failures until the replica was healthy. (The docs describe 502s; this deployment served 503s.)
  • Boot time is billed. Estimated startup compute cost was ~0.28–0.42 per wake. This recipe’s automatic 15-minute idle timeout adds a tail of billed idle time to every isolated request (worked example below).
  • Total experiment cost: ~$1.60 (my estimate from billed minutes; the dashboard Usage panel is the authoritative number).

How Hugging Face Inference Endpoints actually work

An endpoint is a managed container, not a magic API. HF provisions a VM at your chosen vendor/region, pulls a serving image (here vllm/vllm-openai:v0.27.1), downloads the weights from the Hub, health-checks it, and exposes a stable URL. You never touch Kubernetes, CUDA, or weight storage.

The lifecycle is pending → initializing → running, plus two “off” states: scaledToZero and paused.

Billing applies per replica-minute while initializing or running. The two off states cost $0 — but they are not the same:

  • scaledToZero auto-wakes on the next request, with a cold start.
  • paused stays down until you explicitly resume it.

Scale-to-zero triggers on idleness, not utilization: this recipe fires after 15 minutes with no requests (configurable per deployment). And critically: there is no queueing during initialization — requests arrive and fail immediately. Client-side retry is your job.

The setup

ItemValue
ModelQwen/Qwen3.8-27B, revision 1d4bf0f2 (Hub catalog recipe)
EnginevLLM 0.27.1 (managed container, port 8000, OpenAI-compatible /v1)
HardwareAWS us-east-1, 1x Nvidia A100, 80 GB VRAM, 11 vCPUs, 145 GB RAM
Engine flags (endpoint config via hf endpoints describe)--gpu-memory-utilization 0.95, --max-model-len 262144, MTP speculative decoding (num_speculative_tokens: 3), tool parser qwen3_coder
Rate$2.50/hr per replica, billed by the minute
AuthPrivate (personal HF token, Authorization: Bearer)
ClientA small Python script — streaming chat completions against /v1/chat/completions

Procedure: deploy → 10 warm streaming requests → hf endpoints scale-to-zero (verifying the scaledToZero state before each run) → poll until the first successful response, counting every 503 → repeat ×3 → pause.

How “boot time” is defined here — and its limits

  • Boot time = from the first request sent after scale-to-zero until the first successful response fully completed (streaming, short prompt, 64 max tokens). This includes polling delay (10 s interval) and the final response, so it slightly overstates the time when the endpoint became ready. It is not “time to first HTTP 200.”
  • The script stores its own run timestamps in local time; run 2’s boot time was reconstructed from separate UTC shell timestamps (first 503 → first 200) after my polling tool hit a 9-minute timeout while the endpoint was still booting. Runs 1 and 3 were measured by the script directly. The two methods agree within ~1 s where they overlap (run 1: script 398.4 s vs timestamps ~399 s).
  • The script only speaks OpenAI-compatible streaming chat — it will not measure other endpoint types.

The results

Warm steady state (n=10)

metricvalue
TTFT median1.49 s (range 1.37–2.35 s)

Cold boots (n=3)

runboot timefailed requestsTTFT after bootboot cost
1398 s36 × 5032.25 s$0.28
2601 s *~50 × 5032.29 s$0.42
3399 s36 × 5032.37 s$0.28

* reconstructed from UTC timestamps (see above). All runs: identical config, same endpoint, same hour.

What n=3 can and can’t tell you

  1. Startup time varied by 50%. Two boots came in at ~399 s, one at 601 s. With three runs I cannot establish whether later boots benefit from caching, nor explain the outlier (plausible suspects: VM allocation, image pull, weight-download contention — the logs don’t say). The practical takeaway is defensive: don’t assume reboots get faster; size your retries for ~10 minutes.
  2. The failures are hard, not slow. Every request during initialization failed immediately with 503 (the docs describe 502 — on this deployment, this day, it was 503). There is no accept-and-queue behavior; availability is entirely the client’s problem.
  3. First request after boot sat near the upper end of the warm range. TTFT 2.25–2.37 s vs the 1.49 s warm median (warm range 1.37–2.35 s) — roughly 0.8 s above median. Whatever the residual warm-up is, it is marginal next to the 6.6–10 minute wall. I did not isolate its cause.
  4. Multimodal status is unclear. The catalog lists image-text-to-text, and one startup logged “no registered multimodal processor; running in text-only mode.” I tested only text, so I can’t say whether images work — that warning alone doesn’t establish that they never do.
  5. Engine choice decided whether the model ran at all. HF’s default engine (transformers toolkit) crashed on boot — its pinned transformers version didn’t recognize the new qwen3_5 architecture (ValueError: ... model type 'qwen3_5' but Transformers does not recognize this architecture, from the failed deployment’s boot log). The vLLM catalog recipe ran it. On bleeding-edge models the “verified recipe” isn’t a convenience; it’s the difference between running and not running.

The cost math

Billing applies while initializing or running, by the minute. There are two distinct regimes, and they have different math.

Manual scale-to-zero (you trigger it right after use, no idle tail):

cost per wake  = boot_time × rate          # 399–601 s × $2.50/hr ≈ $0.28–$0.42
break-even gap = boot_time                 # ~6.6–10 min — below this, staying warm is cheaper AND faster

Automatic scale-to-zero with a 15-minute idle timeout (this recipe’s default): after every request the endpoint stays warm for 15 more billed minutes ($0.625) before shutting down. An isolated request then costs roughly:

boot ($0.28–$0.42) + idle tail ($0.625) ≈ $0.90–$1.05   vs   $2.50/hr for staying warm

As a simplified estimate, savings vs staying warm only materialize when traffic gaps exceed roughly 22–25 minutes — the 15-minute idle tail plus the 6.6–10-minute boot — where a “gap” runs from the previous completed response to the next request, the same clock the idle timer uses. Below that, the idle tail makes auto-scale-to-zero equal to or more expensive than never scaling down. The naive “break-even = boot time” claim — which my first draft of this post got wrong — only holds for the manual regime.

And the bill never shows the other cost: every wake is 6.6–10 minutes during which the endpoint simply does not exist for your users.

The engineering takeaway: the 503 problem

Since the server queues nothing during initialization, the client owns availability:

for attempt in range(max_retries):
    r = requests.post(f"{url}/v1/chat/completions", headers=auth, json=payload, timeout=420)
    if r.status_code == 200:
        return r
    time.sleep(10)   # 503 while a new replica initializes (docs say 502 — retry both)

When I’d use each option

traffic patternbest option
steady trafficmin replicas ≥ 1, never scale to zero
bursty, gaps well over the ~22–25 min thresholdautomatic scale-to-zero + client retry with generous timeouts
dev/demo box you forget aboutpause (scaledToZero still wakes — and bills — on traffic)
occasional single calls, no infraserverless Inference Providers (pay per token) instead

Caveats

n=3 boots, one model, one region (AWS us-east-1), one vendor, one day, one endpoint. Boot times depend on weight size, image state, VM allocation, and region; run 2 proves the variance is real for an identical config. The 503-vs-502 observation is scoped to this deployment. The total cost (~$1.60) is my reconstruction from billed minutes — the dashboard Usage panel has the exact figure.