Can a model running on a laptop proofread English well enough that you no longer need to paste your text into ChatGPT or another cloud service? We measured Qwen3-8B and Qwen3-4B on a MacBook Pro against human corrections: fixing grammar, leaving correct text alone, and rewriting tone.
Qwen3-8B is the better proofreader: 0.718 grammar-correction F0.5 against 0.675 for Qwen3-4B on llama.cpp (+0.043, 95% interval +0.023 to +0.062), at 1.17 s vs 0.70 s per sentence. Both are usable as a suggestion tool, neither as blind auto-correct: both change about a third of sentences that were already correct.
Use Qwen3-8B when correction quality matters. Use Qwen3-4B for faster interactive suggestions that a person accepts or rejects one by one.
| Model and runtime | Grammar F0.5 | Precision | Recall | Correct text untouched | Tone rewrites pass | Seconds per sentence, p50 / p95 |
|---|---|---|---|---|---|---|
| Qwen3-8B Q4_K_M, llama.cpp | 0.718 | 0.752 | 0.609 | 63% | 100% | 1.17 / 2.24 |
| Qwen3-8B MLX 4-bit, MLX-LM | 0.707 | 0.744 | 0.592 | 65% | 88% | 1.28 / 2.35 |
| Qwen3-4B MLX 4-bit, MLX-LM | 0.689 | 0.728 | 0.568 | 60% | 98% | 0.79 / 1.36 |
| Qwen3-4B Q4_K_M, llama.cpp | 0.675 | 0.719 | 0.542 | 65% | 92% | 0.70 / 1.32 |
Each configuration saw the same 748 learner sentences, 150 already-correct sentences and 40 tone rewrites.
Both models on llama.cpp, same prompt, greedy decoding. Highlighted: words the model added or changed; struck through: words it removed.
| Original | Qwen3-4B | Qwen3-8B |
|---|---|---|
| Both fixed it | ||
| for example, in math club, teachers usually teach us how to solve a math problem using vary ways.Human: For example, in math club, teachers usually teach us how to solve a math problem using various ways. | ||
| As they are involved with study, they donot get enough time to look after their community.Human: As they are involved with study, they do not get enough time to look after their community. | As they are involved with study, they | As they are involved with study, they |
| That's why he is a legend in these days and people repect him.Human: That's why he is a legend in these days and people respect him. | That's why he is a legend in these days and people | That's why he is a legend in these days and people |
| Only Qwen3-8B fixed it | ||
| One of this important element is internet.Human: One of these important elements is the internet. | One of this important element is internet.✗ error left in place | One of |
| They were drove settlements in other domain.Human: They were driven to settlements in other domains. | They were | They were |
| In the world oil price very high right now.Human: In the world, oil prices are very high right now. | In the world, the oil price is very high right now.✗ only wrong edits | In the world, oil |
| If they really interested in these courses, it's all right.Human: If they are really interested in these courses, it's all right. | If they really interested in these courses, it's all right.✗ error left in place | If they are really interested in these courses, it's all right.✓ matches a human correction |
| Only Qwen3-4B fixed it | ||
| It is necessary a breakthrough event or product to succeed.Human: It is necessary for a breakthrough event or product to succeed.8B's fix reads fine too, but no human wrote it, so the score counts it as wrong. | It is necessary for a breakthrough event or product to succeed.✓ matches a human correction | It is necessary to have a breakthrough event or product to succeed.✗ only wrong edits |
| I thought it is kind of ridicurous, but we have to do as tour guide said.Human: I thought it was kind of ridiculous, but we have to do as the tour guide said.8B also changed "have" to "had", which fits "thought", but no human made that change. | I thought it | I thought it |
| Neither fixed it | ||
| Confidence comes naturally since perfection is its along.Human: Confidence comes naturally since perfection is alongside it. | Confidence comes naturally since perfection is its | Confidence comes naturally since perfection is its |
| Nevertheless, this city is having a big change.Human: Nevertheless, this city is experiencing a big change. | Nevertheless, this city is having a big change.✗ error left in place | Nevertheless, this city is having a big change.✗ error left in place |
| They have to contact variety fields to find out what their interests are.Human: They have to contact a variety of fields to find out what their interests are.Both wrote "various fields", which reads fine; the humans wrote "a variety of fields". | They have to contact | They have to contact |
| Already-correct text | ||
| These are enjoyable tasks that make your mindset look forward to the various experiences. | These are enjoyable tasks that make your mindset look forward to | These are enjoyable tasks that make your mindset look forward to the various experiences.✓ left as it was |
| Ignorance of striped bass will abbreviate our earnings, also it has some astonishing effects on our health.8B fixed a comma splice that the human correction left in; the score counts it as over-correction. | Ignorance of striped bass will abbreviate our earnings, also it has some astonishing effects on our health.✓ left as it was | Ignorance of striped bass will abbreviate our earnings |
| People tend to choose other medias, and that is why literature is in danger.Counted as over-correction, but "media" is a real fix: the human correction kept the error. | People tend to choose other | People tend to choose other |
| Tone rewrite, formal to casual | ||
| Dear Aisha, thank you for your report. Could we please discuss the figure of 2500 units at 16:00 on Monday? The meeting will take place in the Berlin office. | Hey Aisha, thanks for the report! Could we chat about the 2500 units figure at 16:00 on Monday? The meeting is in Berlin.✗ lost a fact | Hey Aisha, thanks for the report! Could we chat about the 2500 units figure at 4 PM on Monday? The meeting is in the Berlin office.✓ facts kept, tone changed |
| Tone rewrite, casual to formal | ||
| hey Priya, thx for the deck!! looks awesome. tiny thing: slide 7 says 2500 users but i think it's way more. chat at 15:30? | Hello Priya, thank you for the deck! It looks excellent. A minor point: slide 7 states 2500 users, but I believe the number is significantly higher. Would you like to chat at 15:30?✓ facts kept, tone changed | Dear Priya, Thank you for the deck; it looks excellent. I noticed a minor detail: slide 7 states 2500 users, but I believe the actual number is significantly higher. Could we schedule a chat at 15:30?✓ facts kept, tone changed |
Grammar prompt: You are a careful proofreader. Correct grammar, spelling and punctuation errors in the user's text and make it read fluently. Keep the meaning and the author's wording; change as little as necessary. If the text is already correct, return it unchanged. Reply with only the corrected text, without quotes or comments.
Tone prompts: Rewrite the user's message in a relaxed, friendly, casual tone, as a short chat message to a colleague. Keep every fact exactly: names, dates, times, numbers, amounts and places. Reply with only the rewritten message.
and Rewrite the user's message in a formal, polite business tone. Keep every fact exactly: names, dates, times, numbers, amounts and places. Do not use contractions, slang or emoji. Reply with only the rewritten message.
Learner sentences and human corrections are quoted from JFLEG (Napoles et al., 2017), CC BY-NC-SA 4.0. The tone messages are synthetic.
Of 150 sentences that needed no changes, every configuration changed 52–60. The edits are small: 63% of those changed sentences carry a single edit. Some are real fixes, because a few human corrections still contain errors. In a tool that shows suggestions, the rest is noise a person dismisses quickly; in an auto-correct that applies them silently, it rewrites a third of what you wrote.
All configurations pass 88–100% of the 40 rewrites.
An interval that excludes zero means the difference is unlikely to be chance. The 8B grammar gain holds on both runtimes, but on MLX-LM it only just clears zero. The two 4B builds use different 4-bit schemes, and this is the only place the weight format showed a measurable effect.
A model's edits are compared with the edits human correctors made, using ERRANT, the standard scorer for grammatical error correction.
8b9fb0e6b3; the dataset is downloaded, not redistributed.F0.5 weights precision twice as much as recall: for a proofreader, a wrong change is worse than a missed one.
Correct text untouched is the share of the 150 already-correct sentences (JFLEG human corrections fed back as input) that come back with zero ERRANT edits. Tone rewrites are 40 synthetic messages rewritten casual → formal or formal → casual; one passes only if every fact is kept (names, times, numbers, places; equivalent forms such as 14:45 = 2:45 PM count), the text actually changed, and a formal rewrite has no slang, emoji or contractions.
Intervals come from a paired bootstrap: draw the items at random with replacement, recompute both sides' score on that same sample, record the difference, and repeat 3,000 times. The middle 95% of those differences is the interval. Pairing removes the noise of which items happen to be easy.
Reproduce: python scripts/prepare_proofread.py, then llm-sheet evaluate --config configs/experiments/proofread.yaml --download. Scores recompute from the saved model outputs with llm-sheet evaluate --runs results/proofread.