Is your cold email actually good?
Paste it below. We score it against the playbook that books the meetings: the subject tests, the 80-word cap, AI-tell vocabulary, offer structure, risk reversal and the ask. Then apply the playbook's mechanical fixes in one click. Runs entirely in your browser, nothing is uploaded.
Runs entirely in your browser. Nothing is sent or stored.
Playbook fixes
What this cold email grader does
This grader scores a cold email out of 100 against eleven mechanical checks and returns a letter grade with the specific items to fix. It reads copy only, it runs on the same standard applied to live client campaigns, and it never sends or stores your draft.
What does the grader check?
Eleven things: subject line, an 80-word cap, follow-up shape, em dashes, fluff and template artifacts, AI-tell vocabulary, specificity, the mechanism, the offer and its risk reversal, the ask, and reader focus against spam tone. Each is a rule the tool actually applies rather than a general language-model impression of your writing.
What does your score mean?
A score of 85 or above means the draft no longer fails the things that reliably suppress replies. It is a floor rather than a ceiling: it says nothing about whether the address is real, whether the domain is authenticated, or whether you targeted the right person.
What should you do with a low score?
Fix the flagged items in order of weight, starting with the offer, which carries the single largest contribution. Then re-grade. If a draft grades well and replies stay flat, the problem has moved to the list or the sending setup, and no copy change will reach it.
Why these checks, and not a generic AI score
Most cold-email graders score your draft against a general language model, and several are trained on a public 2000s-era corporate email corpus that has nothing to do with cold outreach. A grammar-and-sentiment score built on internal company mail cannot tell you whether a stranger will reply.
This grader's checks come from a different place: the same standard we apply to our own client campaigns before they go live. Across the 429,763 cold emails in our published campaign book, the drafts that cleared these checks - one concrete mechanism, a single light ask, short, plain text, no spam vocabulary, merge fields that actually resolve - are the ones that produced a 2.12% median reply rate at a 1.32% bounce rate. The number you get here is the number we would put on a draft before spending a real sending domain on it.
That is the difference between a grammar checker and a grader built on what measurably gets replies - and it is why the eleven checks below are weighted the way they are.
What the cold email grader checks
Every email you paste is run through
11 checks and scored out of
100 with a letter grade. These are the rules the tool actually applies, not a summary
of them. It runs entirely in your browser: nothing is sent anywhere and nothing is stored,
and merge fields such as {{first_name}} are masked before scoring so
capitalisation checks cannot trip on them.
The 11 checks
- Subject line. Must pass 2 of 3: what is this, who is it for, what does it promise.
- Length. A hard cap of 80 words. Over it the check fails outright, and the fix is to cut anything that is not identification, mechanism or the ask.
- Follow-up shape. In follow-up mode the playbook shape is one sentence, 20 to 28 words, with no greeting and no signoff.
- Em dashes. A hard rule of zero, including the
--form. Use a colon, period, comma or parentheses instead. - Fluff and artifacts. 480 phrases that delay the point or read as template or chatbot output, split into 247 hard flags and 233 softer ones.
- AI-tell vocabulary. 580 flagged terms across three severity tiers of 197, 298 and 85 entries, plus 76 assistant-voice phrases such as "great question" or "you're absolutely right". A clean email takes zero penalty.
- Specificity. Flags vague or salesy claims and looks for a definite, near-money statement carrying a number. 23 weak outcome phrases are flagged against 16 strong ones, alongside 12 superlatives and 6 jargon phrases.
- Mechanism. Name one mechanism in 3 to 7 words. A mechanism that reads like feature soup is downgraded rather than passed.
- The offer and risk reversal. The offer has to be concrete. A follow-up that reduces itself to "is now a good time?" fails.
- The ask. Ends best on a light value-check question with an easy out. Asking a stranger straight for a meeting is a warning, not a pass: match the gradient of reply first, then resource, then time.
- Reader focus and spam tone. Counts "you" against "I" and "we", and screens 1,130 weighted trigger phrases across five categories: money, overpromise, shady, unnatural and urgency. Risk-reversal clauses are exempt, because "you only pay when deals close" is the offer, not a trigger.
Eleven checks run on any one email. Follow-up mode swaps in the follow-up versions of the shape and mechanism checks in place of the first-touch ones.
How the score is graded
- A, 85 and above. "Sharp. This one should pull replies."
- B, 70 to 84. "Solid. Tighten the flagged items and it flies."
- C, 55 to 69. "Workable, but it is leaving replies on the table."
- D, below 55. "This will struggle. Run the copy doctor."
The score is clamped to the 0 to 100 range. Weighting is deliberately uneven: naming a concrete offer carries the single largest contribution, and length, subject line, specificity, the ask and reader focus each carry their own. A high score is not a promise of replies. It means the email no longer fails the things that reliably suppress them, which is a floor, not a ceiling.
What the grader flags most often
Across the drafts people paste in, the same handful of failures repeat. They are listed here in roughly the order they cost you the most, because the weighting is deliberately uneven.
- No concrete offer. This carries the single largest contribution to the score. A draft that describes a service but never states what the reader gets, and on what terms, cannot score well no matter how clean the prose is.
- Over the 80-word cap. A hard fail rather than a deduction. Anything that is not identification, mechanism or the ask is what to cut first.
- A subject line that passes only one of three tests. It has to answer at least two of: what is this, who is it for, what does it promise.
- Asking a stranger straight for a meeting. That is a warning, not a pass. The reply gradient runs reply, then resource, then time - asking for time first skips two steps.
- Assistant-voice and AI-tell vocabulary. 580 flagged terms across three severity tiers, plus 76 assistant phrases such as "great question". A clean email takes no penalty at all here, so this only bites drafts written or polished by a model.
- Em dashes. A hard rule of zero, including the double-hyphen form. A colon, period, comma or parentheses does the same work without the tell.
- Vague outcome claims. 23 weak outcome phrases are flagged against 16 strong ones. A definite statement carrying a number beats an adjective every time.
Follow-up drafts fail differently. In follow-up mode the shape check expects one sentence of 20 to 28 words with no greeting and no signoff, and a follow-up that reduces itself to "is now a good time?" fails the offer check outright.
What to do with your score
The number is a floor, not a ceiling. It tells you the email no longer fails the things that reliably suppress replies; it does not promise replies. What to do next depends on the band you land in.
- A, 85 and above. Send it. Spend your remaining effort on the list and the sending infrastructure, which is where the next marginal reply comes from.
- B, 70 to 84. Fix the flagged items individually rather than rewriting. At this band the draft is usually one deduction away from an A.
- C, 55 to 69. Something structural is missing, almost always the offer or the ask. Apply the playbook fixes, then re-grade before editing by hand.
- D, below 55. Start again from the mechanism: name what you do in three to seven words, then build the email outward from that one sentence.
A score cannot see the two things that decide most outcomes: whether the address is real and whether the sending domain is authenticated and warmed. If replies stay flat on a draft that grades well, the problem has moved somewhere the copy cannot reach.
Why cold emails get ignored
Most cold emails are ignored because they ask a stranger for time before giving them a reason to care, not because of grammar or tone. The measurable pattern is that drafts naming one concrete mechanism and making one light ask outperform longer, vaguer ones.
For scale: across 429,763 emails sent over 115 campaigns for five client programmes between April and August 2026, the median campaign drew a reply from 2.12% of contacted leads and the pooled figure was 2.58%. The middle half of campaigns that contacted 500 or more leads sat between 1.38% and 2.97%, with a top decile of 3.79%. A good email is not one that gets a 20% reply rate; it is one that reaches the upper part of a band that narrow.
That spread is mostly list and market rather than prose, which is the honest framing: copy decides whether the right person replies, not whether an indifferent one does. What copy reliably costs you is the downside - a draft that reads as automated, asks for a meeting cold, or buries the mechanism in feature language gives an interested reader a reason to stop. See what a realistic reply rate looks like before judging a draft by its response.
Length, CTA and personalisation
How long should a cold email be?
Under 80 words for a first touch. The grader treats that as a hard cap rather than a preference: past it the check fails outright, because the extra words are almost always context the reader has not yet agreed to care about. Cut anything that is not identification, mechanism or the ask. A follow-up is shorter still - one sentence of 20 to 28 words, no greeting, no signoff.
What makes a good cold email CTA?
The smallest ask that still moves things forward. Asking a stranger straight for a meeting is the largest request you can make in a first email, so the grader treats it as a warning rather than a pass. The gradient that converts is a reply first, then a resource, then time - each step cheaper for the recipient than the one after it. A light value-check question with an easy out beats a calendar link.
Does personalisation beat generic copy?
Specificity beats personalisation. A merge field that inserts a company name proves nothing; a sentence that names a mechanism in three to seven words and attaches a number does. The grader flags 23 weak outcome phrases against 16 strong ones for exactly this reason. It also masks merge fields before scoring, so a broken tag never inflates a capitalisation check - though a tag that fails to resolve in the real send is its own problem, covered in merge fields and their failure modes.
The full standard these checks come from is written up in the cold email copywriting rules.
The rest of the cold email toolkit
Grading the copy is one step. These free tools cover the other things that decide whether a cold email lands and gets a reply - the list, the infrastructure, the recipient's environment. All run in your browser; none needs a sign-up.
Rather we just write the ones that get replies?
We build and run the whole campaign, copy, infrastructure and booking, and get paid on the revenue we help you close.
Apply to work with usOther complimentary tools: Spam Checker · Deliverability Checker · ROI Calculator
Questions about grading cold email
Is my email uploaded or stored anywhere?
No. The grader runs entirely in your browser. Nothing is sent to a server and nothing is stored, which is also why it works with the tab offline once the page has loaded.
What score should I be aiming for?
85 or above. Below 70 there is normally a structural problem rather than a wording one, and editing sentences will not move it.
Does a high score mean I will get replies?
No, and it is worth being plain about that. The score means the draft no longer fails the checks that reliably suppress replies. Deliverability, list quality and targeting all sit outside what any copy grader can see.
Why does it flag em dashes?
Because they have become one of the most reliable tells that text came out of a language model, and a reader who suspects that stops reading. A colon, period, comma or parentheses carries the same meaning without the signal.
Why is asking for a meeting penalised?
It is a warning rather than a fail. Asking a stranger for time is the largest request you can make in a first email. Matching the gradient - a reply first, then a resource, then time - converts better because each step is cheaper for the recipient than the one after it.
Does it work for follow-ups as well as first touches?
Yes. Switching to follow-up mode swaps in the follow-up versions of the shape and mechanism checks, because a good follow-up is a different object from a good first email: one sentence, no greeting, no signoff.
Where does the 80-word cap come from?
From the campaign data behind the published benchmark set rather than from a style preference. Past that length the reply rate falls faster than the extra information helps, so the check treats it as a hard cap rather than a nudge.
Why does my score change when I switch to follow-up mode?
Because a good follow-up is a different object from a good first email, and it is graded as one. Follow-up mode swaps in the follow-up shape check, which expects a single sentence of 20 to 28 words with no greeting and no signoff, and the follow-up mechanism check in place of the first-touch versions. A draft that scores well as an opener can therefore score badly as a follow-up, and that is the tool working rather than a bug.
Can I grade an email somebody else wrote?
Yes, and it is one of the more useful ways to use it. The checks are mechanical rather than stylistic, so they apply equally to an agency draft, a template you inherited, or copy a model produced. The AI-tell and assistant-voice checks are particularly worth running over anything that came out of a language model, because those are the tells a recipient notices first.
What do the playbook fixes actually change?
They apply the mechanical corrections only: removing em dashes, cutting flagged fluff and AI-tell vocabulary, and trimming toward the word cap. They will not invent an offer or a mechanism for you, because those are claims about your business that a tool has no business fabricating. If the offer check is what failed, that part is yours to write.
Does the grader look at deliverability at all?
No. It reads copy and nothing else. Authentication, domain age, warmup and list quality decide whether the email arrives, and they are checked by the deliverability and spam tools rather than here. A draft can score an A and still never reach an inbox if the sending setup is wrong.