LLM Local Inference Sheet · technical notes

Technical Kitchen

Apple M3 Pro, 18 GB · llama.cpp b11246 · MLX-LM 0.31.3 · Ollama 0.9.0 · data as of 2026-10-01

Results

Qwen3-14B fits on an 18 GB M3 Pro even at 16k context: context grows the KV cache, not the weights. The harness's own timer agrees with llama.cpp's and Ollama's internal timers to within 0.9% on long prompts, so client-side numbers can stand in for server ones.

All 63 measured runs carried memory pressure: they started with 3.5–11.2 GiB available (median 5.9 GiB), and 47 peaked at pressure level 2 or higher. Absolute speeds are therefore pessimistic; use them to compare runs, not as absolute numbers.

Where GPU memory goes

Context grows the KV cache, not the weights: Qwen3-14B at 16k uses 11.1 of 13.3 GiBllama.cpp b11246 GPU memory at exit, one request. A q8_0 KV cache nearly halves the cache.WeightsKV cacheCompute buffers0 GiB4 GiB8 GiB12 GiBQwen3-1.7B, Q8_01k contextQwen3-1.7B, Q8_0, 1k context: weights 1,743 MiBQwen3-1.7B, Q8_0, 1k context: KV cache 140 MiBQwen3-1.7B, Q8_0, 1k context: compute buffers 45 MiB1.9 GiB4k contextQwen3-1.7B, Q8_0, 4k context: weights 1,743 MiBQwen3-1.7B, Q8_0, 4k context: KV cache 476 MiBQwen3-1.7B, Q8_0, 4k context: compute buffers 48 MiB2.2 GiB16k contextQwen3-1.7B, Q8_0, 16k context: weights 1,743 MiBQwen3-1.7B, Q8_0, 16k context: KV cache 1,820 MiBQwen3-1.7B, Q8_0, 16k context: compute buffers 103 MiB3.6 GiBQwen3-8B, Q4_K_M1k contextQwen3-8B, Q4_K_M, 1k context: weights 4,789 MiBQwen3-8B, Q4_K_M, 1k context: KV cache 180 MiBQwen3-8B, Q4_K_M, 1k context: compute buffers 89 MiB4.9 GiB4k contextQwen3-8B, Q4_K_M, 4k context: weights 4,789 MiBQwen3-8B, Q4_K_M, 4k context: KV cache 612 MiBQwen3-8B, Q4_K_M, 4k context: compute buffers 92 MiB5.4 GiB16k contextQwen3-8B, Q4_K_M, 16k context: weights 4,789 MiBQwen3-8B, Q4_K_M, 16k context: KV cache 2,340 MiBQwen3-8B, Q4_K_M, 16k context: compute buffers 123 MiB7.1 GiB16k, q8_0 KVQwen3-8B, Q4_K_M, 16k, q8_0 KV: weights 4,789 MiBQwen3-8B, Q4_K_M, 16k, q8_0 KV: KV cache 1,243 MiBQwen3-8B, Q4_K_M, 16k, q8_0 KV: compute buffers 133 MiB6.0 GiBQwen3-14B, Q4_K_M1k contextQwen3-14B, Q4_K_M, 1k context: weights 8,579 MiBQwen3-14B, Q4_K_M, 1k context: KV cache 200 MiBQwen3-14B, Q4_K_M, 1k context: compute buffers 123 MiB8.7 GiB4k contextQwen3-14B, Q4_K_M, 4k context: weights 8,579 MiBQwen3-14B, Q4_K_M, 4k context: KV cache 680 MiBQwen3-14B, Q4_K_M, 4k context: compute buffers 126 MiB9.2 GiB16k contextQwen3-14B, Q4_K_M, 16k context: weights 8,579 MiBQwen3-14B, Q4_K_M, 16k context: KV cache 2,600 MiBQwen3-14B, Q4_K_M, 16k context: compute buffers 138 MiB11.1 GiBMetal GPU limit: 13.3 GiB
results-public/mac-capacity, mac-kvquant-q8 · llama.cpp memory breakdown printed at exit, one run per cell

Weights are fixed per model, so the context length decides whether a model fits. For Qwen3-8B the measured KV cache matched the estimate from layer geometry exactly, so the fit of an untested context can be computed before running it.

How far the client timer can be trusted

On long prompts, client and server prefill timing agree within 0.9%Gap between the prefill rate derived from client time to first token and the server timer, one request per cell0%2%4%6%8%1% gapPerformance runs, prompts of 1k–16k tokensQwen3-1.7B · 1k in / 128 out · llama.cppQwen3-1.7B · 1k in / 128 out · llama.cpp: 0.88% gap0.88%Qwen3-4B · 1k in / 256 out · OllamaQwen3-4B · 1k in / 256 out · Ollama: 0.52% gapQwen3-4B · 1k in / 256 out · llama.cppQwen3-4B · 1k in / 256 out · llama.cpp: 0.49% gapQwen3-8B · 1k in / 256 out · OllamaQwen3-8B · 1k in / 256 out · Ollama: 0.49% gapSmolLM3-3B · 1k in / 256 out · llama.cppSmolLM3-3B · 1k in / 256 out · llama.cpp: 0.47% gapQwen3-8B · 1k in / 256 out, rerun · llama.cppQwen3-8B · 1k in / 256 out, rerun · llama.cpp: 0.39% gapQwen3-1.7B · 4k in / 128 out · llama.cppQwen3-1.7B · 4k in / 128 out · llama.cpp: 0.38% gapQwen3-8B · 1k in / 128 out · llama.cppQwen3-8B · 1k in / 128 out · llama.cpp: 0.37% gapQwen3-8B · 1k in / 256 out · llama.cppQwen3-8B · 1k in / 256 out · llama.cpp: 0.33% gapQwen3-14B · 16k in / 128 out · llama.cppQwen3-14B · 16k in / 128 out · llama.cpp: 0.28% gapQwen3-14B · 1k in / 128 out · llama.cppQwen3-14B · 1k in / 128 out · llama.cpp: 0.20% gapQwen3-1.7B · 16k in / 128 out · llama.cppQwen3-1.7B · 16k in / 128 out · llama.cpp: 0.12% gapQwen3-14B · 4k in / 128 out · llama.cppQwen3-14B · 4k in / 128 out · llama.cpp: 0.08% gapQwen3-8B · 4k in / 128 out · llama.cppQwen3-8B · 4k in / 128 out · llama.cpp: 0.08% gapQwen3-8B · 4k in / 128 out, q8_0 KV · llama.cppQwen3-8B · 4k in / 128 out, q8_0 KV · llama.cpp: 0.06% gapQwen3-8B · 16k in / 128 out · llama.cppQwen3-8B · 16k in / 128 out · llama.cpp: 0.04% gapQwen3-8B · 16k in / 128 out, q8_0 KV · llama.cppQwen3-8B · 16k in / 128 out, q8_0 KV · llama.cpp: 0.03% gapShort prompts and task suitesQwen3-4B · task suite · OllamaQwen3-4B · task suite · Ollama: 7.53% gap7.53%Qwen3-4B · 128 in / 256 out · llama.cppQwen3-4B · 128 in / 256 out · llama.cpp: 4.78% gapQwen3-8B · task suite · OllamaQwen3-8B · task suite · Ollama: 4.38% gapQwen3-4B · proofreading · llama.cppQwen3-4B · proofreading · llama.cpp: 3.08% gapQwen3-8B · proofreading · llama.cppQwen3-8B · proofreading · llama.cpp: 2.04% gapSmolLM3-3B · task suite · llama.cppSmolLM3-3B · task suite · llama.cpp: 1.86% gapQwen3-4B · task suite · llama.cppQwen3-4B · task suite · llama.cpp: 1.82% gapQwen3-8B · task suite · llama.cppQwen3-8B · task suite · llama.cpp: 1.24% gap
results-public/report/results.csv · all measured llama.cpp and Ollama cells at one request (25); prefill proxy = input tokens / client TTFT

On a long prompt, time to first token is almost pure prefill, so the client figure can stand in for the server timer, which MLX-LM does not expose. On short prompts and the task suites, HTTP and scheduling overhead is a larger share of a small number. With several requests in flight, the client figure also includes queueing, so it is not compared there by design.

Environment gotchas

The laptop itself was the biggest source of measurement risk.

GotchaWhat we sawHow the harness handles it
Rosetta shellThe agent shell reported x86_64 on an M3 Pro (sysctl.proc_translated = 1); universal binaries started from it would run translatedEvery backend starts with arch -arm64; doctor warns; manifests record binary architecture
Memory pressure~4 GB free and 11.6 GB swap before any model loaded; runs swapped in 8–18 GiBEach run records pressure level and swap deltas; runs above "normal" or with > 256 MiB swapped are flagged
Metal working setMetal lets the GPU use 13,639 MiB of 18 GBLogged from llama.cpp and MLX; capacity cells measured against it
Disk36 GB free at the start; swap growth consumed moreA cell refuses to start below 5 GiB free; Ollama's duplicate blob copies were removed after use
Heat~85 °C under hours of full GPU loadpmset -g therm showed no throttling; thermal state is not yet logged per run

Runtime gotchas

Each runtime has defaults that would silently change a benchmark; all of these are set explicitly and checked in the logs.

RuntimeDefault or behaviourWhy it mattersWhat we do
llama.cpp b11246--fit on adjusts unset arguments to fit memoryContext or offload can change without notice--fit off plus explicit context and layers
llama.cpp b11246--cache-ram 8192: an 8 GiB host prompt cacheReuses prefixes and eats RAM on an 18 GB machine--cache-ram 0 and cache_prompt:false per request
llama.cpp b11246-np auto turns on a unified KV cachePer-slot context becomes unclearExplicit -np = concurrency, --no-kv-unified, total context = slots × per-slot
llama.cpp b11246Model-loader facts appear only at -lv 4No offload or buffer facts in default logs-lv 4; GPU memory read from the exit breakdown
MLX-LM 0.31.3A request with a seed is not batchableConcurrency silently serializesThe adapter never sends a seed
MLX-LM 0.31.3/health answers before the model loads"Ready" would be wrongReadiness = a completed 1-token generation
Ollama 0.9.0Template detection fails on the official Qwen3 GGUFPrompt differs from the intended oneRaw mode with prompts rendered from the pinned template
Ollama 0.9.0Reuses shared prompt prefixes; no switch to disableTTFT looks better than it is when prompts share a system promptReuse noted; quality-run latency marked not like-for-like
Ollama 0.9.0Repetition guard aborts some greedy outputs without done:trueTruncated stream could pass as successCounted as errors (2/30 and 6/30 at 4 users)
Ollama 0.9.0Does not know the SmolLM3 architectureModel never loadsDetected from the log and recorded as unsupported in seconds

Measurement details

Metric definitions are strict, and every label says how a number was obtained.

Design principles

The harness separates process control from measurement and keeps raw data, so every published number can be recomputed without rerunning a model.