# Running models locally

_Topics at every level · sourced_

Running open-weight models on your own hardware: what fits, quantization, and what you give up.


## Guided worked example · Engineering & technical work

Fictional scripted fixture, not a measured run. No model or external actions execute.

**Overview:** Follow a proposed offline task through resource, quality, and privacy tradeoffs. Inspect whether a candidate local setup fits the actual workload instead of equating local execution with suitability.

**Assumptions:** Hardware capacity, model files, context length, and supporting software all matter. Local inference does not automatically mean every part of an application stays offline.

**Design choices:** Test representative inputs on the intended device and inspect network behavior where offline operation matters. Compare usability and maintenance effort as well as response quality.

**Request:** Estimate whether a local summarizer fits an offline laptop.

**Starting evidence:** Illustrative memory: 16 GB available; weights 8 GB, runtime 3 GB, context/cache 6 GB.

**Action and control:** Add estimated components: 17 GB. Weights alone are not the workload budget.

**Stage records (authored, not executed):**

### Input record

Illustrative memory: 16 GB available; weights 8 GB, runtime 3 GB, context/cache 6 GB.

What changed: Establish the facts supplied for this version of the task.

### Design note

Test representative inputs on the intended device and inspect network behavior where offline operation matters. Compare usability and maintenance effort as well as response quality.

What changed: Choose an approach before treating a proposed result as accepted.

### Proposed work

Add estimated components: 17 GB. Weights alone are not the workload budget.

What changed: Turn the request and evidence into the next action or transformation.

### Result record · illustrative

Estimate exceeds 16 GB. Reduce model/context and benchmark actual hardware. No throughput is claimed.

What changed: Inspect the result of the authored example; this is not an executed model run.

### Verification plan

Clearly estimated memory budget, quantization comparison, offline data-flow review, and a benchmark plan without fabricated throughput.

If the result falls short:
If memory or performance is inadequate, reduce workload, choose a suitable model, or revise the deployment plan. Do not present an estimate as a measured device result.

What changed: Separate what needs checking from what the illustration establishes.

### Adaptation handoff

Use this for private drafting, field work, or offline assistance. Your device and workload determine the tradeoff; the example's resource assumptions are not universal specifications.

What changed: Decide which assumptions, tools, and controls should change for your own task.

**Sample result:** Estimate exceeds 16 GB. Reduce model/context and benchmark actual hardware. No throughput is claimed.

**Change something — Consider only the 8 GB weight file:** Apparent fit ignores 9 GB of runtime/cache estimates.

**Decision:** Does a fitting weight file prove the workload fits?

**Answer:** No; include runtime and context overhead.

**Why:** Model weights, context, and runtime overhead all consume memory; local execution alone does not guarantee private handling.

**Review criteria:** Clearly estimated memory budget, quantization comparison, offline data-flow review, and a benchmark plan without fabricated throughput.

**Recovery:** If memory or performance is inadequate, reduce workload, choose a suitable model, or revise the deployment plan. Do not present an estimate as a measured device result.

**Adapt it:** Use this for private drafting, field work, or offline assistance. Your device and workload determine the tradeoff; the example's resource assumptions are not universal specifications.

Running a model locally means downloading an open-weight model's parameters and running them on
your own hardware instead of calling a hosted API. llama.cpp, one runtime built for this,
describes itself in four words ("LLM inference in C/C++") and states its goal as inference
"with minimal setup and state-of-the-art performance on a wide range of hardware - locally and in
the cloud"[2]. Why: privacy (nothing leaves the machine), cost at real volume (no
per-token bill), working offline, and control over exactly which model version runs. What it
costs instead: hardware you buy and keep running, quality capped by what fits on that hardware,
and the upkeep of a system component you now run yourself rather than a vendor.

This topic is not a level on the ladder; it applies at every level. What changes by level is how
much context and how many concurrent requests the hardware has to hold, which is this page's own
"At each level" section below.

This page is sourced, not measured: the sizes and requirements below come from the model makers'
and runners' own pages, and nothing here has been timed on hardware and scored.

## Practical guidance

If privacy, not price, is why you want this: a desktop app that runs a model on your own computer
needs no account and sends nothing anywhere else. Look for a downloadable application rather than
a website. Ollama, one such tool, states plainly that "Local models are
always free"[1]; llama.cpp and LM Studio are two more. Install one, let it suggest a
model sized for your computer, and run the actual questions you would otherwise send to a chat
app before trusting it with anything real.

What to expect first: most models you can run locally are compressed to fit consumer hardware, a
step called quantization. llama.cpp's own documentation puts it plainly: quantization "reduces the
precision of model weights (e.g., from 32-bit floats to 4-bit integers), which shrinks the model's
size and can speed up inference," at a cost the same page is careful about: it "may introduce some
accuracy loss which is usually measured in Perplexity (ppl) and/or Kullback–Leibler Divergence
(kld)," minimized "by using a suitable imatrix file"[3]. What that means for you: the
model you install is smaller and less capable than the full-size version you may have read about,
and how much was cut varies by download. Run the same handful of questions you already know good
answers to through it and through a hosted chat app, side by side; if the local one is noticeably
worse, that is the real trade for keeping the data on your machine, not a setup mistake.

The setting most worth checking is whatever the app calls "context size" or "memory": a smaller
number there means the model forgets more of what you told it earlier in the conversation, in
exchange for needing less of your computer's memory to run at all. If it keeps losing track of
something you said two messages back, that setting is usually why, and raising it is the fix if
your machine has room.

A hosted model is the honest answer once none of this is actually about privacy: a per-message
price you likely will not notice, nothing to install or keep updated, and no ceiling set by what
your own machine can run. Local inference earns its place when something specific truly cannot
leave the machine, not as a default upgrade from a chat app that already works fine.

## Implementation details

The example estimates how much memory a model needs from four things: how many parameters it
has, how many bits each weight is stored at, how long the context window is, and how many
requests are being served at once. Every number it prints is an estimate, and the demo says so
above the table: it counts only weights and a key/value cache, the two costs every runtime
accounts for in some form, and leaves out activation memory and a framework's own overhead, both
of which only add to the real figure. Use it to size hardware with room to spare, not to predict
what a runtime will report.

`examples/local_inference/run.py` (lines 46-52)

```python
def weights_bytes(params: int, bits_per_weight: float) -> int:
    """Parameter count times bits per weight, converted to bytes. `bits_per_weight` is the
    caller's own average figure: 16 for fp16, roughly 4 for a 4-bit quantization such as Q4_K_M.
    A real quantized file usually averages a little above the bit width in its name, because not
    every tensor is quantized the same way (llama.cpp's quantize tool can leave the output tensor
    unquantized, for instance). This function models none of that; it takes the average given."""
    return round(params * bits_per_weight / 8)
```

The weights are the simple half: parameter count times bits per weight, in bytes. A named
quantization is not exactly its nominal bit width in practice, because not every tensor is
quantized the same way: llama.cpp's own quantize documentation offers `--leave-output-tensor` to
"leave output.weight un(re)quantized," and says a multimodal projector is usually kept in a
high-quality format instead[3]. So `bits_per_weight` is an average the caller supplies,
not a number this function looks up, and the real file usually averages a little above the bit
width in its name.

The key/value cache is the half that depends on how the model is used, not just what it is. Every
token the model has already seen leaves behind one key and one value vector in every layer, and
they stay there for the rest of the sequence, which is why context length costs memory before a
single token of it is used:

`examples/local_inference/run.py` (lines 55-63)

```python
def kv_cache_bytes(
    *, context_length: int, num_layers: int, num_kv_heads: int, head_dim: int, bytes_per_value: int = 2, num_sequences: int = 1
) -> int:
    """The key/value cache: two tensors (key and value) per layer, each sized
    context_length x num_kv_heads x head_dim at bytes_per_value bytes, times how many sequences
    are served at once. `num_kv_heads` is the key/value head count from the model's own config,
    which grouped-query attention makes smaller than the attention head count."""
    per_sequence = 2 * num_layers * context_length * num_kv_heads * head_dim * bytes_per_value
    return per_sequence * num_sequences
```

Two per layer for the key and the value; `num_kv_heads` times `head_dim` for how wide each one
is; `context_length` for how many of them accumulate; `bytes_per_value` for the precision they
are held at. The one term worth checking twice is `num_kv_heads`. Grouped-query attention lets
several attention heads share one key/value head, so a model's key/value head count is often a
fraction of its attention head count, and passing the larger number quietly inflates the whole
estimate. It comes from the model's own config, not from this function.

The cache then multiplies by how many sequences (concurrent requests) are held at once: the cost
of serving several users from one running model rather than one at a time. That multiplication is
the naive ceiling, and servers are built to avoid paying it in full: vLLM documents "Efficient
management of attention key and value memory with PagedAttention" and "Continuous batching of
incoming requests" among its own features[4]. This example models neither, so a real
server with either should need less than the number printed here for the same concurrency, while
still needing more for everything else.

`python -m examples.local_inference --demo` prints the same illustrative 8B-parameter shape four
ways (full precision against roughly 4-bit, a short context against a long one, one user against
eight) so each variable's effect is visible on its own:

`examples/local_inference/README.md` (lines 11-11)

```text
python -m examples.local_inference --demo
```

`tests/test_example_local_inference.py` checks the arithmetic against numbers computed by hand,
never against a real model: 1,000,000 parameters at 16 bits is exactly 2,000,000 bytes, and a
small cache shape works out to exactly 640 bytes. It also pins the two proportionalities the
sizing questions on this page turn on (doubling the context doubles the cache, eight sequences
cost eight times one) and that a smaller `bits_per_weight` only ever shrinks the weights term,
leaving the cache alone, because quantizing weights does nothing about what a conversation is.

## When you do not need this

Skip running a model locally when a hosted API's per-call price is not the bottleneck, when the
task needs a model larger than anything the available hardware can run well, or when there is no
one available to keep a local server patched and running. A hosted API is a service someone else
operates; a local one is a service you now operate.

Move to local inference once volume makes a per-call bill add up faster than hardware would cost
over the same period, once data cannot leave the machine at all, once the task works offline, or
once a model small enough to run locally already passes the eval that matters for the task.

## Speculative decoding

Quantization makes a model fit. Speculative decoding is the other lever a local server has, and it
changes what the same model costs per token rather than what it weighs. A small draft model
proposes the next few tokens, the large one checks them in a single pass, and a sampler keeps the
ones the large model would have produced itself and throws the rest away. vLLM's own guide states
the regime it is for: to "reduce inter-token latency under medium-to-low QPS (queries per second),
memory-bound workloads"[5]. Both runtimes this page cites document it. llama.cpp's
server lists "Speculative decoding" among its features, with `--spec-draft-model` documented as the
"draft model for speculative decoding (default: unused)" and `--spec-draft-n-max` as the "number
of tokens to draft for speculative decoding (default: 3)"[6].

No decision moves, at any level, which is why this is a topic and never a rung. The output is
meant to be the one the large model would have given on its own. What it costs is memory, in the
terms this page's own estimator already uses: a second set of weights and a second key/value
cache, which llama.cpp exposes as their own settings for the draft model's cache type and its
layers in VRAM[6]. Size for both, or the drafting that was supposed to buy speed takes
the memory the context window needed.

You do not need it if your bottleneck is fitting the model at all, or if the hardware is already
saturated by concurrent requests rather than waiting between tokens.

The failure mode to know is that "identical output" is a guarantee with edges. vLLM states that
"Speculative decoding sampling is theoretically lossless up to the precision limits of hardware
numerics", and, in the same section, that "variations in generated outputs with and without
speculative decoding can occur due to following factors", naming floating-point precision and
batch size[5]. So a golden set that passed before it was turned on is worth re-running
after, rather than assumed.

## Failure modes

### Hardware sized for the wrong model

- **How to notice it:** A self-hosted setup that ran a smaller model comfortably starts missing its latency target, or stops fitting in memory at all, once the task needs a larger local model to pass the same eval.
- **How to test for it:** Run the actual eval the product needs to pass on the smallest local model that could plausibly work before committing to hardware, the same test ops names for this failure, not just on whichever model happened to be handy.

### Context length quietly exceeds what was sized for

- **How to notice it:** A local server sized for a given context window starts failing or truncating once a real conversation or a retrieved document set grows past it, because the KV cache for a longer context is larger than what was planned for.
- **How to test for it:** Compute the KV cache size at the longest context the product is actually expected to reach, not just a typical one, using the same arithmetic this page's own estimator does.

### A quantization is chosen for size without checking what it costs in quality

- **How to notice it:** A smaller quantization is picked because it fits the available hardware, without measuring whether it still passes the task it needs to pass.
- **How to test for it:** Run the same eval against the full-precision and the quantized version of a model and compare the scores directly, rather than assuming a smaller file is "close enough."

### One more concurrent user than the hardware was sized for

- **How to notice it:** A local deployment handles the first several concurrent requests fine and then fails or slows sharply on the next one, because the KV cache for each additional sequence was not budgeted for.
- **How to test for it:** Compute the memory needed at the maximum number of concurrent requests the deployment is expected to serve, the way this page's own estimator's num_sequences parameter does, not just at one.

### A runtime's own overhead is left out of a hardware estimate

- **How to notice it:** A memory estimate covering only weights and a KV cache undershoots what a real runtime actually needs, because activation memory and the runtime's own overhead are real costs this kind of estimate does not include.
- **How to test for it:** Compare this page's own estimator's number against a real runtime's reported memory usage on the same model and context length, and treat the gap as a floor to add, not a rounding error to ignore.

## At each level

- [Conventional software](/gradient_ascent/levels/0/): nothing to run locally: this topic's own concerns
  start at level 1.
- [Direct prompting](/gradient_ascent/levels/1/): the smallest case (one short prompt, one short
  answer), the least memory this page's own estimator will ever report for a given model.
- [Added context](/gradient_ascent/levels/2/): a retrieved passage set or a large pasted document
  fills the context window before the model answers at all, so context length, not just model
  size, starts to drive the memory number. This is the same window
  [RAG](/gradient_ascent/techniques/rag/)'s own prompt assembly fills.
- [Tool use](/gradient_ascent/levels/4/): tool definitions and their results are extra tokens in
  the same context window, so a tool-heavy [function
  calling](/gradient_ascent/techniques/function-calling/) run costs more context, and so more KV cache, than the same question answered
  with no tools at all.
- [Agent loops](/gradient_ascent/levels/5/): a long-running loop's transcript keeps growing across
  turns, so the KV cache a [single agent](/gradient_ascent/techniques/single-agent/) run needs
  grows with it: a context budget sized for one exchange is not sized for a whole run.
- [Teams of Agents](/gradient_ascent/levels/6/): several agents worked at once on one machine
  are exactly the "several sequences" case this page's own `num_sequences` parameter counts:
  [a lead and its workers](/gradient_ascent/techniques/orchestrator-workers/) running locally
  multiplies the KV cache cost by however many are active together, not just by one.
- [Always-on agents](/gradient_ascent/levels/7/): load is sustained rather than one burst per
  question, so the number to size for is how many sequences are live at the same time on an
  ordinary day, which is a different question from how large the largest single request is.

## Practices

- Size from the model's own config, not from its name: layer count, key/value head count and
  head dimension are what the cache formula needs, and grouped-query attention makes the third
  of those smaller than the head count people usually quote.
- Estimate at the longest context and the highest concurrency the system is meant to reach, not
  at a typical one. Both terms are linear, so the worst case is easy to compute and easy to skip.
- Treat any estimate, this one included, as a floor. Measure what the runtime actually reports
  and keep the gap; that gap is the activation memory and overhead nobody's formula counts.
- Decide quantization with an [eval](/gradient_ascent/techniques/evals/) on your own task, and
  record the exact quantization you ran, not "4-bit".
- Compare against what the hosted alternative costs over the same period before buying anything:
  that arithmetic belongs to [cost optimization](/gradient_ascent/techniques/cost-optimization/),
  and hardware is a fixed cost that does not care whether you use it.

## Run it

**What to monitor.** Memory actually in use against what was sized for, and how many concurrent requests
  are being served against how many the hardware was budgeted for: both from this page's own
  estimator, checked against the real runtime's own reported usage.

**Cost at volume.** Local inference trades a per-call bill for hardware you buy and keep running;
  whether that trade pays depends on real, sustained volume, not on how the hardware looks on
  paper before anything is deployed.

**How it fails in production.** Context grows past what the KV cache was sized for, or one more concurrent
  request arrives than the hardware was budgeted to hold, and the failure looks like the model
  being slow or crashing rather than a sizing problem.

**What to log.** Context length actually used per request, concurrent request count over time, and
  memory actually consumed, so a sizing failure traces back to which of the two terms (weights
  or KV cache) was undersized, rather than reading as a generic slowdown.

## Try it

1. **Use it.** If you have a local model tool installed (or are willing to install one), check what context size or "memory" setting it exposes, and what it says that costs in hardware. If it doesn't say, that silence is itself worth noting.
2. **Build it.** Run python -m examples.local_inference --demo from the repo root and compare the KV cache column across the four rows. Then add a fifth shape of your own with a 128K context and one user. Going from 32K to 128K is four times the context: predict the KV cache figure before you run it, then check whether the printed number is four times the 32K row.
3. **Either lane.** Pick a model you know the approximate parameter count of. Using this page's estimator (or your own arithmetic from weights_bytes' formula), compute its memory footprint at full precision and at roughly 4-bit quantization, and compare the difference against what hardware you actually have.


## Sources

1. [Ollama](https://ollama.com/) — Ollama (accessed 2026-09-19)
2. [llama.cpp](https://github.com/ggml-org/llama.cpp) — ggml-org (accessed 2026-09-19)
3. [quantize (tools/quantize/README.md)](https://github.com/ggml-org/llama.cpp/blob/master/tools/quantize/README.md) — ggml-org (accessed 2026-09-19)
4. [vLLM documentation](https://docs.vllm.ai/en/latest/) — vLLM (accessed 2026-09-19)
5. [Speculative Decoding](https://docs.vllm.ai/en/latest/features/speculative_decoding/) — vLLM (accessed 2026-09-19)
6. [llama.cpp server (tools/server/README.md)](https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md) — ggml-org (accessed 2026-09-19)


Last reviewed 2026-09-19.
