Apple M3 Pro, 18 GB · llama.cpp b11246 · MLX-LM 0.31.3 · Ollama 0.9.0 · data as of 2026-10-01
Results
Qwen3-14B fits on an 18 GB M3 Pro even at 16k context: context grows the KV cache, not the weights. The harness's own timer agrees with llama.cpp's and Ollama's internal timers to within 0.9% on long prompts, so client-side numbers can stand in for server ones.
All 63 measured runs carried memory pressure: they started with 3.5–11.2 GiB available (median 5.9 GiB), and 47 peaked at pressure level 2 or higher. Absolute speeds are therefore pessimistic; use them to compare runs, not as absolute numbers.
Where GPU memory goes
results-public/mac-capacity, mac-kvquant-q8 · llama.cpp memory breakdown printed at exit, one run per cell
Weights are fixed per model, so the context length decides whether a model fits. For Qwen3-8B the measured KV cache matched the estimate from layer geometry exactly, so the fit of an untested context can be computed before running it.
How far the client timer can be trusted
results-public/report/results.csv · all measured llama.cpp and Ollama cells at one request (25); prefill proxy = input tokens / client TTFT
On a long prompt, time to first token is almost pure prefill, so the client figure can stand in for the server timer, which MLX-LM does not expose. On short prompts and the task suites, HTTP and scheduling overhead is a larger share of a small number. With several requests in flight, the client figure also includes queueing, so it is not compared there by design.
Environment gotchas
The laptop itself was the biggest source of measurement risk.
Gotcha
What we saw
How the harness handles it
Rosetta shell
The agent shell reported x86_64 on an M3 Pro (sysctl.proc_translated = 1); universal binaries started from it would run translated
Every backend starts with arch -arm64; doctor warns; manifests record binary architecture
Memory pressure
~4 GB free and 11.6 GB swap before any model loaded; runs swapped in 8–18 GiB
Each run records pressure level and swap deltas; runs above "normal" or with > 256 MiB swapped are flagged
Metal working set
Metal lets the GPU use 13,639 MiB of 18 GB
Logged from llama.cpp and MLX; capacity cells measured against it
Disk
36 GB free at the start; swap growth consumed more
A cell refuses to start below 5 GiB free; Ollama's duplicate blob copies were removed after use
Heat
~85 °C under hours of full GPU load
pmset -g therm showed no throttling; thermal state is not yet logged per run
Runtime gotchas
Each runtime has defaults that would silently change a benchmark; all of these are set explicitly and checked in the logs.
Template detection fails on the official Qwen3 GGUF
Prompt differs from the intended one
Raw mode with prompts rendered from the pinned template
Ollama 0.9.0
Reuses shared prompt prefixes; no switch to disable
TTFT looks better than it is when prompts share a system prompt
Reuse noted; quality-run latency marked not like-for-like
Ollama 0.9.0
Repetition guard aborts some greedy outputs without done:true
Truncated stream could pass as success
Counted as errors (2/30 and 6/30 at 4 users)
Ollama 0.9.0
Does not know the SmolLM3 architecture
Model never loads
Detected from the log and recorded as unsupported in seconds
Measurement details
Metric definitions are strict, and every label says how a number was obtained.
Time to first token: request dispatch to the first chunk that carries generated text. Role-only, metadata and usage chunks are skipped. It includes HTTP, queueing and prefill.
Time per output token: (last text chunk − first text chunk) / (output tokens − 1). Labelled token_level only when the number of text chunks equals the server's token count; otherwise chunk_derived. Null below 2 tokens.
Throughput: output tokens of successful measured requests divided by the window from first dispatch to last stream end. mean_inflight shows how full the window was.
Prompt verification: prompts hit exact token counts after chat templating with each model's pinned tokenizer. llama.cpp's /apply-template rendered byte-identical prompts, and server prompt-token counts matched on all three runtimes.
Output length: performance runs send ignore_eos where supported; elsewhere early stops are counted and a full-length subset is reported. Nothing is padded.
Cache control: baseline runs disable prefix reuse (llama.cpp cache_prompt:false + --cache-ram 0, MLX cache size 0) and verify it through cached_tokens or llama.cpp's cache_n.
Memory on unified memory: reported side by side, never summed. llama.cpp's exit breakdown (e.g. 5,094 MiB = 4,789 weights + 216 KV + 89 compute), the MLX allocator peak, the system GPU counter delta, and RSS. phys_footprint is wrong for llama.cpp because mmap'd weights are excluded (0.4 GiB for a 4.7 GiB model).
Cross-checks: client TTFT matched llama.cpp's native prefill time within about 1%, and the KV buffer matched the geometric estimate (144 KiB per token for Qwen3-8B) exactly.
Statistics: medians with bootstrap CIs; p95 marked exploratory below 200 samples; quality differences from paired bootstraps over the same items.
Design principles
The harness separates process control from measurement and keeps raw data, so every published number can be recomputed without rerunning a model.
Adapters launch, the client measures. A backend adapter only starts a server, waits until it is ready, collects metadata and stops it. One async client measures every runtime, so a metric means the same thing everywhere.
Raw events are the source of truth. Every streamed chunk is stored with a monotonic perf_counter_ns timestamp in events.jsonl.gz. Reports recompute metrics from these events and never trust cached summaries.
One server per cell. A cell is runtime × artifact × workload × concurrency. Each gets a fresh server, so memory is released between cells and load time is observed every time. Cells run strictly one at a time.
States are results.unsupported, gated, untested, out_of_memory, timeout and failed are recorded outcomes, not missing rows. Mock-server runs are labelled synthetic and never reach a public report.
Pin everything. Model revisions and file sha256, chat-template hashes and runtime builds (llama.cpp b11246, mlx-lm 0.31.3) are fixed. Runs resume by configuration identity and never overwrite earlier ones.