LLM Local Inference Sheet · Apple M3 Pro · 18 GB unified memory · macOS 26.6.2
Which local LLM setup works on this Mac?
Measured serving performance, task quality and proofreading for Qwen3 and SmolLM3 on llama.cpp, MLX-LM and Ollama. Every figure is recomputed from saved raw stream events.
All Mac runs measured under memory pressure: absolute numbers are pessimistic– cells measuredLlama-3.1-8B gated
Recommended
Qwen3-4B ~4-bit
llama.cpp for one user, MLX-LM for several
Best quality score
–
–
Qwen3-8B at 4 users
–
total output tok/s, MLX-LM vs llama.cpp
Largest model that fits
–
–
Profiling investigation
llama.cpp stops scaling between 4 and 8 concurrent requests
Qwen3-4B, 128 tokens in / 256 out. Metal switches matmul kernels by batch size, and K-quant weights land on a slower path for batches of 4–8. Above 8 a matrix kernel takes over.
llama-server (end to end)MLX-LM serverllama-batched-bench (no HTTP)
Shaded band: batches routed to mul_mv_ext for K-quants. Source: ggml-metal-ops.cpp at build b11246.
Runtime comparison
At 4 users, MLX-LM nearly doubles throughput
Qwen3-8B, 1,024 tokens in / 256 out. With one user the three runtimes are within about 10%. llama.cpp and Ollama run the identical GGUF file.
llama.cppMLX-LMOllamalighter bar = 1 user, solid = 4 users
Model selection
Qwen3-4B matches 8B on quality at half the latency
60 held-out tasks: JSON extraction, request routing and CV incident reports. Budget: score ≥ 0.80 and ≤ 3 s per task.
llama.cppMLX-LMOllama
Ollama's latency here benefits from prefix-cache reuse it cannot disable. SmolLM3 does not load on Ollama 0.9.0.
Capacity
14B fits at 16k context, but waits minutes for the first token
Time to first token as the prompt grows, one user, 5 measured requests per point.
Qwen3-1.7BQwen3-8BQwen3-14Bllama.cppMLX-LM
–
KV-cache quantization
8-bit KV cache halves its memory
Qwen3-8B Q4_K_M on llama.cpp, KV sizes from llama.cpp's own buffer report.
f16 KVq8_0 KV
Quality impact of q8_0 KV was not measured yet.
Key findings
What the data supports
Concurrency 4–8 buys nothing on llama.cpp
Throughput fell from 3 to 4 users and jumped 63% at 9, matching the kernel thresholds in the source.
8B adds no measurable quality
Paired bootstrap over 60 items: no significant difference between 4B and 8B, at twice the cost for 8B.
GGUF and MLX 4-bit score the same
No significant quality difference on either model, despite different quantization schemes.
For proofreading, 8B is worth it
On JFLEG grammar correction, 8B beats 4B by 0.043 F0.5 (significant), unlike on the JSON tasks.
Same file, same answers
llama.cpp and Ollama on the byte-identical GGUF scored within 0.004.
Proofreading
Qwen3-8B corrects grammar better, at 1.7× the time
748 learner sentences from JFLEG, each with 4 human corrections, scored with ERRANT F0.5 (the standard grammar-correction metric).
llama.cppMLX-LM
Paired bootstrap, 8B − 4B: +0.043 F0.5 on llama.cpp (95% CI +0.025 to +0.061), +0.018 on MLX.
Proofreading detail
Both sizes still edit a third of correct sentences
Keep rate: share of 150 already-correct sentences returned unchanged. Style: 40 formal/casual rewrites that keep every fact.
For a Grammarly-style tool, over-editing is the main weakness to fix with a minimal-edit prompt.