LLM Local Inference Sheet · proofreading

Proofreading on a Mac: Qwen3-8B vs Qwen3-4B

Apple M3 Pro, 18 GB · llama.cpp b11246 · MLX-LM 0.31.3 · greedy decoding, reasoning off · data as of 2026-10-01

Can a model running on a laptop proofread English well enough that you no longer need to paste your text into ChatGPT or another cloud service? We measured Qwen3-8B and Qwen3-4B on a MacBook Pro against human corrections: fixing grammar, leaving correct text alone, and rewriting tone.

Results

Qwen3-8B is the better proofreader: 0.718 grammar-correction F0.5 against 0.675 for Qwen3-4B on llama.cpp (+0.043, 95% interval +0.023 to +0.062), at 1.17 s vs 0.70 s per sentence. Both are usable as a suggestion tool, neither as blind auto-correct: both change about a third of sentences that were already correct.

Use Qwen3-8B when correction quality matters. Use Qwen3-4B for faster interactive suggestions that a person accepts or rejects one by one.

Model and runtimeGrammar F0.5PrecisionRecallCorrect text untouchedTone rewrites passSeconds per sentence, p50 / p95
Qwen3-8B Q4_K_M, llama.cpp0.7180.7520.60963%100%1.17 / 2.24
Qwen3-8B MLX 4-bit, MLX-LM0.7070.7440.59265%88%1.28 / 2.35
Qwen3-4B MLX 4-bit, MLX-LM0.6890.7280.56860%98%0.79 / 1.36
Qwen3-4B Q4_K_M, llama.cpp0.6750.7190.54265%92%0.70 / 1.32

Each configuration saw the same 748 learner sentences, 150 already-correct sentences and 40 tone rewrites.

Examples

Both models on llama.cpp, same prompt, greedy decoding. Highlighted: words the model added or changed; struck through: words it removed.

OriginalQwen3-4BQwen3-8B
Both fixed it
for example, in math club, teachers usually teach us how to solve a math problem using vary ways.Human: For example, in math club, teachers usually teach us how to solve a math problem using various ways.for For example, in math club, teachers usually teach us how to solve a math problem using vary various ways.✓ matches a human correctionfor For example, in math club, teachers usually teach us how to solve a math problem using vary various ways.✓ matches a human correction
As they are involved with study, they donot get enough time to look after their community.Human: As they are involved with study, they do not get enough time to look after their community.As they are involved with study, they donot do not get enough time to look after their community.✓ matches a human correctionAs they are involved with study, they donot do not get enough time to look after their community.✓ matches a human correction
That's why he is a legend in these days and people repect him.Human: That's why he is a legend in these days and people respect him.That's why he is a legend in these days and people repect respect him.✓ matches a human correctionThat's why he is a legend in these days and people repect respect him.✓ matches a human correction
Only Qwen3-8B fixed it
One of this important element is internet.Human: One of these important elements is the internet.One of this important element is internet.✗ error left in placeOne of this these important element elements is the internet.✓ matches a human correction
They were drove settlements in other domain.Human: They were driven to settlements in other domains.They were drove drive settlements in other domain.✗ only wrong editsThey were drove driven to settlements in other domain domains.✓ matches a human correction
In the world oil price very high right now.Human: In the world, oil prices are very high right now.In the world, the oil price is very high right now.✗ only wrong editsIn the world, oil price prices are very high right now.✓ matches a human correction
If they really interested in these courses, it's all right.Human: If they are really interested in these courses, it's all right.If they really interested in these courses, it's all right.✗ error left in placeIf they are really interested in these courses, it's all right.✓ matches a human correction
Only Qwen3-4B fixed it
It is necessary a breakthrough event or product to succeed.Human: It is necessary for a breakthrough event or product to succeed.8B's fix reads fine too, but no human wrote it, so the score counts it as wrong.It is necessary for a breakthrough event or product to succeed.✓ matches a human correctionIt is necessary to have a breakthrough event or product to succeed.✗ only wrong edits
I thought it is kind of ridicurous, but we have to do as tour guide said.Human: I thought it was kind of ridiculous, but we have to do as the tour guide said.8B also changed "have" to "had", which fits "thought", but no human made that change.I thought it is was kind of ridicurous ridiculous, but we have to do as the tour guide said.✓ matches a human correctionI thought it is was kind of ridicurous ridiculous, but we have had to do as the tour guide said.~ partly fixed
Neither fixed it
Confidence comes naturally since perfection is its along.Human: Confidence comes naturally since perfection is alongside it.Confidence comes naturally since perfection is its along own.✗ only wrong editsConfidence comes naturally since perfection is its along alone.✗ only wrong edits
Nevertheless, this city is having a big change.Human: Nevertheless, this city is experiencing a big change.Nevertheless, this city is having a big change.✗ error left in placeNevertheless, this city is having a big change.✗ error left in place
They have to contact variety fields to find out what their interests are.Human: They have to contact a variety of fields to find out what their interests are.Both wrote "various fields", which reads fine; the humans wrote "a variety of fields".They have to contact variety various fields to find out what their interests are.✗ only wrong editsThey have to contact variety various fields to find out what their interests are.✗ only wrong edits
Already-correct text
These are enjoyable tasks that make your mindset look forward to the various experiences.These are enjoyable tasks that make your mindset look forward to the various experiences.✗ changed correct textThese are enjoyable tasks that make your mindset look forward to the various experiences.✓ left as it was
Ignorance of striped bass will abbreviate our earnings, also it has some astonishing effects on our health.8B fixed a comma splice that the human correction left in; the score counts it as over-correction.Ignorance of striped bass will abbreviate our earnings, also it has some astonishing effects on our health.✓ left as it wasIgnorance of striped bass will abbreviate our earnings, ; it also it has some astonishing effects on our health.✗ changed correct text
People tend to choose other medias, and that is why literature is in danger.Counted as over-correction, but "media" is a real fix: the human correction kept the error.People tend to choose other medias media, and that is why literature is in danger.✗ changed correct textPeople tend to choose other medias media, and that is why literature is in danger.✗ changed correct text
Tone rewrite, formal to casual
Dear Aisha, thank you for your report. Could we please discuss the figure of 2500 units at 16:00 on Monday? The meeting will take place in the Berlin office.Hey Aisha, thanks for the report! Could we chat about the 2500 units figure at 16:00 on Monday? The meeting is in Berlin.✗ lost a factHey Aisha, thanks for the report! Could we chat about the 2500 units figure at 4 PM on Monday? The meeting is in the Berlin office.✓ facts kept, tone changed
Tone rewrite, casual to formal
hey Priya, thx for the deck!! looks awesome. tiny thing: slide 7 says 2500 users but i think it's way more. chat at 15:30?Hello Priya, thank you for the deck! It looks excellent. A minor point: slide 7 states 2500 users, but I believe the number is significantly higher. Would you like to chat at 15:30?✓ facts kept, tone changedDear Priya,
Thank you for the deck; it looks excellent. I noticed a minor detail: slide 7 states 2500 users, but I believe the actual number is significantly higher. Could we schedule a chat at 15:30?✓ facts kept, tone changed

Grammar prompt: You are a careful proofreader. Correct grammar, spelling and punctuation errors in the user's text and make it read fluently. Keep the meaning and the author's wording; change as little as necessary. If the text is already correct, return it unchanged. Reply with only the corrected text, without quotes or comments.

Tone prompts: Rewrite the user's message in a relaxed, friendly, casual tone, as a short chat message to a colleague. Keep every fact exactly: names, dates, times, numbers, amounts and places. Reply with only the rewritten message. and Rewrite the user's message in a formal, polite business tone. Keep every fact exactly: names, dates, times, numbers, amounts and places. Do not use contractions, slang or emoji. Reply with only the rewritten message.

Learner sentences and human corrections are quoted from JFLEG (Napoles et al., 2017), CC BY-NC-SA 4.0. The tone messages are synthetic.

Applied observations

What happens to a learner sentence

Qwen3-8B matches a human correction on 44% of sentences, Qwen3-4B on 40% (llama.cpp)Outcome per learner sentence, compared at edit level with the closest of its 4 human correctionsMatches a human correctionPartly fixedErrors left untouchedOnly wrong editsQwen3-8B Q4_K_M, llama.cppQwen3-8B Q4_K_M, llama.cpp: matches a human correction 43.9% (328 of 748)44%Qwen3-8B Q4_K_M, llama.cpp: partly fixed 47.3% (354 of 748)47%Qwen3-8B Q4_K_M, llama.cpp: errors left untouched 2.8% (21 of 748)Qwen3-8B Q4_K_M, llama.cpp: only wrong edits 6.0% (45 of 748)Qwen3-8B MLX 4-bit, MLX-LMQwen3-8B MLX 4-bit, MLX-LM: matches a human correction 44.0% (329 of 748)44%Qwen3-8B MLX 4-bit, MLX-LM: partly fixed 46.9% (351 of 748)47%Qwen3-8B MLX 4-bit, MLX-LM: errors left untouched 3.9% (29 of 748)Qwen3-8B MLX 4-bit, MLX-LM: only wrong edits 5.2% (39 of 748)Qwen3-4B MLX 4-bit, MLX-LMQwen3-4B MLX 4-bit, MLX-LM: matches a human correction 40.2% (301 of 748)40%Qwen3-4B MLX 4-bit, MLX-LM: partly fixed 48.9% (366 of 748)49%Qwen3-4B MLX 4-bit, MLX-LM: errors left untouched 4.9% (37 of 748)Qwen3-4B MLX 4-bit, MLX-LM: only wrong edits 5.9% (44 of 748)Qwen3-4B Q4_K_M, llama.cppQwen3-4B Q4_K_M, llama.cpp: matches a human correction 39.7% (297 of 748)40%Qwen3-4B Q4_K_M, llama.cpp: partly fixed 46.9% (351 of 748)47%Qwen3-4B Q4_K_M, llama.cpp: errors left untouched 6.6% (49 of 748)7%Qwen3-4B Q4_K_M, llama.cpp: only wrong edits 6.8% (51 of 748)7%
results-public/proofread · 748 JFLEG sentences per configuration, each compared with the closest of its 4 human corrections

Correct text gets edited

Of 150 sentences that needed no changes, every configuration changed 52–60. The edits are small: 63% of those changed sentences carry a single edit. Some are real fixes, because a few human corrections still contain errors. In a tool that shows suggestions, the rest is noise a person dismisses quickly; in an auto-correct that applies them silently, it rewrites a third of what you wrote.

Tone rewrites keep facts, with different failure habits

All configurations pass 88–100% of the 40 rewrites.

Which differences are real

8B corrects grammar better on both runtimes; the other gaps are small or noiseDifference in each score with its 95% interval; blue = interval excludes zeroGrammar F0.508B − 4B, llama.cpp, Grammar F0.5: +0.043 (+0.023 to +0.062)+0.0438B − 4B, MLX-LM, Grammar F0.5: +0.018 (+0.001 to +0.036)+0.0184B: llama.cpp − MLX-LM, Grammar F0.5: -0.014 (-0.027 to -0.001)-0.014Correct text untouched08B − 4B, llama.cpp, Correct text untouched: -0.020 (-0.080 to +0.033)-0.0208B − 4B, MLX-LM, Correct text untouched: +0.053 (-0.007 to +0.113)+0.0534B: llama.cpp − MLX-LM, Correct text untouched: +0.053 (+0.013 to +0.100)+0.053Tone rewrites pass08B − 4B, llama.cpp, Tone rewrites pass: +0.075 (+0.000 to +0.175)+0.0758B − 4B, MLX-LM, Tone rewrites pass: -0.100 (-0.225 to +0.000)-0.1004B: llama.cpp − MLX-LM, Tone rewrites pass: -0.050 (-0.150 to +0.050)-0.0508B − 4B, llama.cpp8B − 4B, MLX-LM4B: llama.cpp − MLX-LM
results-public/proofread · paired bootstrap, 3,000 resamples of the same items for both sides; 95% intervals

An interval that excludes zero means the difference is unlikely to be chance. The 8B grammar gain holds on both runtimes, but on MLX-LM it only just clears zero. The two 4B builds use different 4-bit schemes, and this is the only place the weight format showed a measurable effect.

How it is scored

A model's edits are compared with the edits human correctors made, using ERRANT, the standard scorer for grammatical error correction.

  1. Data: the JFLEG test set (CC BY-NC-SA 4.0), 748 sentences written by English learners, each corrected independently by 4 people. Pinned revision 8b9fb0e6b3; the dataset is downloaded, not redistributed.
  2. Prompt: the model gets the sentence with an instruction to fix grammar, spelling and punctuation, change as little as needed, and return only the corrected text.
  3. Edits: ERRANT aligns the original with the model's output and lists the edits (words replaced, inserted or deleted). It does the same for each human correction. All texts are tokenized the same way, so spacing around punctuation never counts as an edit.
  4. Matching: a model edit counts as correct (TP) when a human made the same edit at the same position. Extra edits are false positives (FP); human edits the model missed are false negatives (FN). Each sentence is scored against whichever of its 4 human corrections fits the model best.
  5. Totals: TP, FP and FN are summed over all 748 sentences, then turned into one score:
P = TP / (TP + FP)    R = TP / (TP + FN)    F0.5 = 1.25 · P · R / (0.25 · P + R)

F0.5 weights precision twice as much as recall: for a proofreader, a wrong change is worse than a missed one.

Correct text untouched is the share of the 150 already-correct sentences (JFLEG human corrections fed back as input) that come back with zero ERRANT edits. Tone rewrites are 40 synthetic messages rewritten casual → formal or formal → casual; one passes only if every fact is kept (names, times, numbers, places; equivalent forms such as 14:45 = 2:45 PM count), the text actually changed, and a formal rewrite has no slang, emoji or contractions.

Intervals come from a paired bootstrap: draw the items at random with replacement, recompute both sides' score on that same sample, record the difference, and repeat 3,000 times. The middle 95% of those differences is the interval. Pairing removes the noise of which items happen to be easy.

Caveats

Reproduce: python scripts/prepare_proofread.py, then llm-sheet evaluate --config configs/experiments/proofread.yaml --download. Scores recompute from the saved model outputs with llm-sheet evaluate --runs results/proofread.