LLM Local Inference Sheet · Apple M3 Pro · 18 GB unified memory · macOS 26.6.2

Which local LLM setup works on this Mac?

Measured serving performance, task quality and proofreading for Qwen3 and SmolLM3 on llama.cpp, MLX-LM and Ollama. Every figure is recomputed from saved raw stream events.

All Mac runs measured under memory pressure: absolute numbers are pessimistic – cells measured Llama-3.1-8B gated
Recommended
Qwen3-4B ~4-bit
llama.cpp for one user, MLX-LM for several
Best quality score
–
–
Qwen3-8B at 4 users
–
total output tok/s, MLX-LM vs llama.cpp
Largest model that fits
–
–
Profiling investigation

llama.cpp stops scaling between 4 and 8 concurrent requests

Qwen3-4B, 128 tokens in / 256 out. Metal switches matmul kernels by batch size, and K-quant weights land on a slower path for batches of 4–8. Above 8 a matrix kernel takes over.

llama-server (end to end)MLX-LM serverllama-batched-bench (no HTTP)
Shaded band: batches routed to mul_mv_ext for K-quants. Source: ggml-metal-ops.cpp at build b11246.
Runtime comparison

At 4 users, MLX-LM nearly doubles throughput

Qwen3-8B, 1,024 tokens in / 256 out. With one user the three runtimes are within about 10%. llama.cpp and Ollama run the identical GGUF file.

llama.cppMLX-LMOllamalighter bar = 1 user, solid = 4 users
Model selection

Qwen3-4B matches 8B on quality at half the latency

60 held-out tasks: JSON extraction, request routing and CV incident reports. Budget: score ≥ 0.80 and ≤ 3 s per task.

llama.cppMLX-LMOllama
Ollama's latency here benefits from prefix-cache reuse it cannot disable. SmolLM3 does not load on Ollama 0.9.0.
Capacity

14B fits at 16k context, but waits minutes for the first token

Time to first token as the prompt grows, one user, 5 measured requests per point.

Qwen3-1.7BQwen3-8BQwen3-14Bllama.cppMLX-LM
–
KV-cache quantization

8-bit KV cache halves its memory

Qwen3-8B Q4_K_M on llama.cpp, KV sizes from llama.cpp's own buffer report.

f16 KVq8_0 KV
Quality impact of q8_0 KV was not measured yet.
Key findings

What the data supports

Concurrency 4–8 buys nothing on llama.cpp

Throughput fell from 3 to 4 users and jumped 63% at 9, matching the kernel thresholds in the source.

8B adds no measurable quality

Paired bootstrap over 60 items: no significant difference between 4B and 8B, at twice the cost for 8B.

GGUF and MLX 4-bit score the same

No significant quality difference on either model, despite different quantization schemes.

For proofreading, 8B is worth it

On JFLEG grammar correction, 8B beats 4B by 0.043 F0.5 (significant), unlike on the JSON tasks.

Same file, same answers

llama.cpp and Ollama on the byte-identical GGUF scored within 0.004.

Proofreading

Qwen3-8B corrects grammar better, at 1.7× the time

748 learner sentences from JFLEG, each with 4 human corrections, scored with ERRANT F0.5 (the standard grammar-correction metric).

llama.cppMLX-LM
Paired bootstrap, 8B − 4B: +0.043 F0.5 on llama.cpp (95% CI +0.025 to +0.061), +0.018 on MLX.
Proofreading detail

Both sizes still edit a third of correct sentences

Keep rate: share of 150 already-correct sentences returned unchanged. Style: 40 formal/casual rewrites that keep every fact.

For a Grammarly-style tool, over-editing is the main weakness to fix with a minimal-edit prompt.
All performance cells

Qwen3 and SmolLM3, 1,024 tokens in / 256 out

Quality detail

Scores by task

Compatibility

Model × runtime status