Topics at every level

Running models locally

Running open-weight models on your own hardware: what fits, quantization, and what you give up.

Sourced

Concept at a glance

Run the model on hardware you control.

SequenceConceptual illustration
Run the model on hardware you control.Model weights leads to Your hardware. Your hardware leads to Local response. Model size, precision, and context compete for memory on the same machine.Model weightsChoose size and precisionYour hardwareRuntime, memory, computeLocal responseMeasure speed and qualityRun the model on hardware you control.Model weights leads to Your hardware. Your hardware leads to Local response. Model size, precision, and context compete for memory on the same machine.Model weightsChoose size and precisionYour hardwareRuntime, memory, computeLocal responseMeasure speed and quality
Read the connections in words
  • Model weights → Your hardware: Runtime, memory, compute.
  • Your hardware → Local response: Measure speed and quality.
Key idea

Model size, precision, and context compete for memory on the same machine.

A focused engineering & technical work example. Additional perspectives appear where they provide a useful contrast.

GUIDED WORKED EXAMPLE Fictional fixtures · scripted outputs · no live model or external actions

Running models locally: see it in practice.

Executing model inference on locally controlled hardware rather than a hosted model endpoint.

What you’ll walk through

Follow a proposed offline task through resource, quality, and privacy tradeoffs. Inspect whether a candidate local setup fits the actual workload instead of equating local execution with suitability.

The task in this version

Estimate whether a local summarizer fits an offline laptop.

What you’ll learn to check

Clearly estimated memory budget, quantization comparison, offline data-flow review, and a benchmark plan without fabricated throughput.

The setting makes the example concrete. Carry the underlying pattern into your own work; adapt the sources, tools, and level of oversight to your task.

Engineering & technical workAn authored case with its own evidence, changed condition, and decision.
The task in this example

Estimate whether a local summarizer fits an offline laptop.

Authored case. Select any record below; nothing is sent to a model.
FOLLOW THE EXAMPLE1 / 6
Interpret this honestlySample evidence, not your actual data.No real messages, tools, training, or hardware operations run.The sequence illustrates the concept; it is not a recorded agent trace.
THE VISIBLE WORKStarting evidence
Input record
AUTHORED TEACHING RECORD · NOT A LIVE RUN
Illustrative memory: 16 GB available; weights 8 GB, runtime 3 GB, context/cache 6 GB.

What changed: Establish the facts supplied for this version of the task.

WHY THIS MATTERS

What this case assumes

Hardware capacity, model files, context length, and supporting software all matter. Local inference does not automatically mean every part of an application stays offline.

1 / 6

Apply this to your project

Describe your task to your own model and use Running models locally as a reference. Ask whether it fits, which alternatives meet the same automation needs, and how you would implement and check the result.

Go deeper: practical guidance, failure modes, and implementation

Running a model locally means downloading an open-weight model’s parameters and running them on your own hardware instead of calling a hosted API. llama.cpp, one runtime built for this, describes itself in four words (“LLM inference in C/C++”) and states its goal as inference “with minimal setup and state-of-the-art performance on a wide range of hardware - locally and in the cloud”[2]. Why: privacy (nothing leaves the machine), cost at real volume (no per-token bill), working offline, and control over exactly which model version runs. What it costs instead: hardware you buy and keep running, quality capped by what fits on that hardware, and the upkeep of a system component you now run yourself rather than a vendor.

This topic is not a level on the ladder; it applies at every level. What changes by level is how much context and how many concurrent requests the hardware has to hold, which is this page’s own “At each level” section below.

This page is sourced, not measured: the sizes and requirements below come from the model makers’ and runners’ own pages, and nothing here has been timed on hardware and scored.

Practical guidance

If privacy, not price, is why you want this: a desktop app that runs a model on your own computer needs no account and sends nothing anywhere else. Look for a downloadable application rather than a website. Ollama, one such tool, states plainly that “Local models are always free”[1]; llama.cpp and LM Studio are two more. Install one, let it suggest a model sized for your computer, and run the actual questions you would otherwise send to a chat app before trusting it with anything real.

What to expect first: most models you can run locally are compressed to fit consumer hardware, a step called quantization. llama.cpp’s own documentation puts it plainly: quantization “reduces the precision of model weights (e.g., from 32-bit floats to 4-bit integers), which shrinks the model’s size and can speed up inference,” at a cost the same page is careful about: it “may introduce some accuracy loss which is usually measured in Perplexity (ppl) and/or Kullback–Leibler Divergence (kld),” minimized “by using a suitable imatrix file”[3]. What that means for you: the model you install is smaller and less capable than the full-size version you may have read about, and how much was cut varies by download. Run the same handful of questions you already know good answers to through it and through a hosted chat app, side by side; if the local one is noticeably worse, that is the real trade for keeping the data on your machine, not a setup mistake.

The setting most worth checking is whatever the app calls “context size” or “memory”: a smaller number there means the model forgets more of what you told it earlier in the conversation, in exchange for needing less of your computer’s memory to run at all. If it keeps losing track of something you said two messages back, that setting is usually why, and raising it is the fix if your machine has room.

A hosted model is the honest answer once none of this is actually about privacy: a per-message price you likely will not notice, nothing to install or keep updated, and no ceiling set by what your own machine can run. Local inference earns its place when something specific truly cannot leave the machine, not as a default upgrade from a chat app that already works fine.

Implementation details

The example estimates how much memory a model needs from four things: how many parameters it has, how many bits each weight is stored at, how long the context window is, and how many requests are being served at once. Every number it prints is an estimate, and the demo says so above the table: it counts only weights and a key/value cache, the two costs every runtime accounts for in some form, and leaves out activation memory and a framework’s own overhead, both of which only add to the real figure. Use it to size hardware with room to spare, not to predict what a runtime will report.

examples/local_inference/run.py · lines 46–52
def weights_bytes(params: int, bits_per_weight: float) -> int:
    """Parameter count times bits per weight, converted to bytes. `bits_per_weight` is the
    caller's own average figure: 16 for fp16, roughly 4 for a 4-bit quantization such as Q4_K_M.
    A real quantized file usually averages a little above the bit width in its name, because not
    every tensor is quantized the same way (llama.cpp's quantize tool can leave the output tensor
    unquantized, for instance). This function models none of that; it takes the average given."""
    return round(params * bits_per_weight / 8)

The weights are the simple half: parameter count times bits per weight, in bytes. A named quantization is not exactly its nominal bit width in practice, because not every tensor is quantized the same way: llama.cpp’s own quantize documentation offers --leave-output-tensor to “leave output.weight un(re)quantized,” and says a multimodal projector is usually kept in a high-quality format instead[3]. So bits_per_weight is an average the caller supplies, not a number this function looks up, and the real file usually averages a little above the bit width in its name.

The key/value cache is the half that depends on how the model is used, not just what it is. Every token the model has already seen leaves behind one key and one value vector in every layer, and they stay there for the rest of the sequence, which is why context length costs memory before a single token of it is used:

examples/local_inference/run.py · lines 55–63
def kv_cache_bytes(
    *, context_length: int, num_layers: int, num_kv_heads: int, head_dim: int, bytes_per_value: int = 2, num_sequences: int = 1
) -> int:
    """The key/value cache: two tensors (key and value) per layer, each sized
    context_length x num_kv_heads x head_dim at bytes_per_value bytes, times how many sequences
    are served at once. `num_kv_heads` is the key/value head count from the model's own config,
    which grouped-query attention makes smaller than the attention head count."""
    per_sequence = 2 * num_layers * context_length * num_kv_heads * head_dim * bytes_per_value
    return per_sequence * num_sequences

Two per layer for the key and the value; num_kv_heads times head_dim for how wide each one is; context_length for how many of them accumulate; bytes_per_value for the precision they are held at. The one term worth checking twice is num_kv_heads. Grouped-query attention lets several attention heads share one key/value head, so a model’s key/value head count is often a fraction of its attention head count, and passing the larger number quietly inflates the whole estimate. It comes from the model’s own config, not from this function.

The cache then multiplies by how many sequences (concurrent requests) are held at once: the cost of serving several users from one running model rather than one at a time. That multiplication is the naive ceiling, and servers are built to avoid paying it in full: vLLM documents “Efficient management of attention key and value memory with PagedAttention” and “Continuous batching of incoming requests” among its own features[4]. This example models neither, so a real server with either should need less than the number printed here for the same concurrency, while still needing more for everything else.

python -m examples.local_inference --demo prints the same illustrative 8B-parameter shape four ways (full precision against roughly 4-bit, a short context against a long one, one user against eight) so each variable’s effect is visible on its own:

examples/local_inference/README.md · lines 11–11
python -m examples.local_inference --demo

tests/test_example_local_inference.py checks the arithmetic against numbers computed by hand, never against a real model: 1,000,000 parameters at 16 bits is exactly 2,000,000 bytes, and a small cache shape works out to exactly 640 bytes. It also pins the two proportionalities the sizing questions on this page turn on (doubling the context doubles the cache, eight sequences cost eight times one) and that a smaller bits_per_weight only ever shrinks the weights term, leaving the cache alone, because quantizing weights does nothing about what a conversation is.

When you do not need this

Skip running a model locally when a hosted API’s per-call price is not the bottleneck, when the task needs a model larger than anything the available hardware can run well, or when there is no one available to keep a local server patched and running. A hosted API is a service someone else operates; a local one is a service you now operate.

Move to local inference once volume makes a per-call bill add up faster than hardware would cost over the same period, once data cannot leave the machine at all, once the task works offline, or once a model small enough to run locally already passes the eval that matters for the task.

Speculative decoding

Quantization makes a model fit. Speculative decoding is the other lever a local server has, and it changes what the same model costs per token rather than what it weighs. A small draft model proposes the next few tokens, the large one checks them in a single pass, and a sampler keeps the ones the large model would have produced itself and throws the rest away. vLLM’s own guide states the regime it is for: to “reduce inter-token latency under medium-to-low QPS (queries per second), memory-bound workloads”[5]. Both runtimes this page cites document it. llama.cpp’s server lists “Speculative decoding” among its features, with --spec-draft-model documented as the “draft model for speculative decoding (default: unused)” and --spec-draft-n-max as the “number of tokens to draft for speculative decoding (default: 3)”[6].

No decision moves, at any level, which is why this is a topic and never a rung. The output is meant to be the one the large model would have given on its own. What it costs is memory, in the terms this page’s own estimator already uses: a second set of weights and a second key/value cache, which llama.cpp exposes as their own settings for the draft model’s cache type and its layers in VRAM[6]. Size for both, or the drafting that was supposed to buy speed takes the memory the context window needed.

You do not need it if your bottleneck is fitting the model at all, or if the hardware is already saturated by concurrent requests rather than waiting between tokens.

The failure mode to know is that “identical output” is a guarantee with edges. vLLM states that “Speculative decoding sampling is theoretically lossless up to the precision limits of hardware numerics”, and, in the same section, that “variations in generated outputs with and without speculative decoding can occur due to following factors”, naming floating-point precision and batch size[5]. So a golden set that passed before it was turned on is worth re-running after, rather than assumed.

Failure modes

Hardware sized for the wrong model

How to notice it
A self-hosted setup that ran a smaller model comfortably starts missing its latency target, or stops fitting in memory at all, once the task needs a larger local model to pass the same eval.
How to test for it
Run the actual eval the product needs to pass on the smallest local model that could plausibly work before committing to hardware, the same test ops names for this failure, not just on whichever model happened to be handy.

Context length quietly exceeds what was sized for

How to notice it
A local server sized for a given context window starts failing or truncating once a real conversation or a retrieved document set grows past it, because the KV cache for a longer context is larger than what was planned for.
How to test for it
Compute the KV cache size at the longest context the product is actually expected to reach, not just a typical one, using the same arithmetic this page's own estimator does.

A quantization is chosen for size without checking what it costs in quality

How to notice it
A smaller quantization is picked because it fits the available hardware, without measuring whether it still passes the task it needs to pass.
How to test for it
Run the same eval against the full-precision and the quantized version of a model and compare the scores directly, rather than assuming a smaller file is "close enough."

One more concurrent user than the hardware was sized for

How to notice it
A local deployment handles the first several concurrent requests fine and then fails or slows sharply on the next one, because the KV cache for each additional sequence was not budgeted for.
How to test for it
Compute the memory needed at the maximum number of concurrent requests the deployment is expected to serve, the way this page's own estimator's num_sequences parameter does, not just at one.

A runtime's own overhead is left out of a hardware estimate

How to notice it
A memory estimate covering only weights and a KV cache undershoots what a real runtime actually needs, because activation memory and the runtime's own overhead are real costs this kind of estimate does not include.
How to test for it
Compare this page's own estimator's number against a real runtime's reported memory usage on the same model and context length, and treat the gap as a floor to add, not a rounding error to ignore.

At each level

  • Conventional software: nothing to run locally: this topic’s own concerns start at level 1.
  • Direct prompting: the smallest case (one short prompt, one short answer), the least memory this page’s own estimator will ever report for a given model.
  • Added context: a retrieved passage set or a large pasted document fills the context window before the model answers at all, so context length, not just model size, starts to drive the memory number. This is the same window RAG’s own prompt assembly fills.
  • Tool use: tool definitions and their results are extra tokens in the same context window, so a tool-heavy function calling run costs more context, and so more KV cache, than the same question answered with no tools at all.
  • Agent loops: a long-running loop’s transcript keeps growing across turns, so the KV cache a single agent run needs grows with it: a context budget sized for one exchange is not sized for a whole run.
  • Teams of Agents: several agents worked at once on one machine are exactly the “several sequences” case this page’s own num_sequences parameter counts: a lead and its workers running locally multiplies the KV cache cost by however many are active together, not just by one.
  • Always-on agents: load is sustained rather than one burst per question, so the number to size for is how many sequences are live at the same time on an ordinary day, which is a different question from how large the largest single request is.

Practices

  • Size from the model’s own config, not from its name: layer count, key/value head count and head dimension are what the cache formula needs, and grouped-query attention makes the third of those smaller than the head count people usually quote.
  • Estimate at the longest context and the highest concurrency the system is meant to reach, not at a typical one. Both terms are linear, so the worst case is easy to compute and easy to skip.
  • Treat any estimate, this one included, as a floor. Measure what the runtime actually reports and keep the gap; that gap is the activation memory and overhead nobody’s formula counts.
  • Decide quantization with an eval on your own task, and record the exact quantization you ran, not “4-bit”.
  • Compare against what the hosted alternative costs over the same period before buying anything: that arithmetic belongs to cost optimization, and hardware is a fixed cost that does not care whether you use it.

Run it

What to monitor

Memory actually in use against what was sized for, and how many concurrent requests are being served against how many the hardware was budgeted for: both from this page's own estimator, checked against the real runtime's own reported usage.

Cost at volume

Local inference trades a per-call bill for hardware you buy and keep running; whether that trade pays depends on real, sustained volume, not on how the hardware looks on paper before anything is deployed.

How it fails in production

Context grows past what the KV cache was sized for, or one more concurrent request arrives than the hardware was budgeted to hold, and the failure looks like the model being slow or crashing rather than a sizing problem.

What to log

Context length actually used per request, concurrent request count over time, and memory actually consumed, so a sizing failure traces back to which of the two terms (weights or KV cache) was undersized, rather than reading as a generic slowdown.

Try it

  1. Use it

    If you have a local model tool installed (or are willing to install one), check what context size or "memory" setting it exposes, and what it says that costs in hardware. If it doesn't say, that silence is itself worth noting.

  2. Build it

    Run python -m examples.local_inference --demo from the repo root and compare the KV cache column across the four rows. Then add a fifth shape of your own with a 128K context and one user. Going from 32K to 128K is four times the context: predict the KV cache figure before you run it, then check whether the printed number is four times the 32K row.

  3. Either lane

    Pick a model you know the approximate parameter count of. Using this page's estimator (or your own arithmetic from weights_bytes' formula), compute its memory footprint at full precision and at roughly 4-bit quantization, and compare the difference against what hardware you actually have.

How it connects

Before, after and instead of this

Read first

Optional: products, tools, and models

6 current examples · Products, tools, and models that demonstrate this concept. A selection, not a ranking.

In practice

Run a model on a workstation

Choose weights and precision that fit the available memory, then measure performance on your own workload.

Out there

Named products, tools and models

Tools6
  • llama.cppopen source · runs models locally
  • LM StudioElement Labs · runs models locally
  • MLXApple · training and inference on Apple hardware
  • OllamaOllama · runs models locally
  • SGLangSGLang community · model serving runtime
  • vLLMopen source · model server

Names listed 09/19/2026. 223 of 223 registry entries have been checked against the maker's own page; the registry marks the rest as unchecked.

Where this comes from

Primary sources

  1. Ollama · Ollama (accessed 09/19/2026)
  2. llama.cpp · ggml-org (accessed 09/19/2026)
  3. quantize (tools/quantize/README.md) · ggml-org (accessed 09/19/2026)
  4. vLLM documentation · vLLM (accessed 09/19/2026)
  5. Speculative Decoding · vLLM (accessed 09/19/2026)
  6. llama.cpp server (tools/server/README.md) · ggml-org (accessed 09/19/2026)

Last reviewed 09/19/2026. Pages unreviewed for 90 days are flagged for another pass. Markdown version of this page