Why your word count disagrees with theirs
Two honest counters can differ by a few percent on the same text, and the cause is a hyphen, a dash or an apostrophe rather than a mistake in either tool.
Your word processor says 1,482. The submission form says 1,531. The limit is 1,500, and both numbers are honest. Neither tool is broken and neither is lying to you. They are counting different things and giving the result the same name. I have watched writers spend an afternoon trimming a piece against a number that was never going to agree with the other number, and the fix is not better editing. It is knowing which rule each tool follows, so that you can predict the gap instead of arguing about it.
There is no single correct number
There is no standards body that defines what a word is. Nothing specifies whether a hyphen joins two words or separates them, whether a number counts, or whether a symbol does. Every counter is a short program that somebody wrote, and that program encodes a rule: where to cut the text, what to ignore, what to keep. When your number and theirs come apart, you are looking at a seam where their rules part company. Neither rule is an error.
The simplest rule is the whitespace split: break the text at every space, tab and line break, and count what falls out. That is roughly what a word processor does, and it is why a hyphenated compound comes back as a single word. The other common rule is the boundary counter, which also breaks at hyphens, dashes, slashes and punctuation, then discards the fragments that are only symbols. Most counters sit somewhere between those positions, which is why the difference is usually a couple of percent, and occasionally a great deal more. The counter on this site is the first kind, and it says so on the page.
The gap is predictable once you know the rules
Almost every small disagreement traces back to a handful of characters. Learn how your own tool treats each of them and you can work out the difference in your head before you start cutting sentences.
A hyphenated compound is the classic case. up-to-date is one token to a whitespace
splitter and three to a counter that treats the hyphen as a boundary. Most tools report one, so if you
are over a limit, the strict counter is the one that will decide whether you are over it.
An em dash set closed, as in word—word, is one token to a whitespace splitter and
two to a counter that breaks at the dash. Open it up to word — word and the
disagreement reverses rather than disappearing. The whitespace rule now sees three tokens, because a
dash sitting between spaces is not itself whitespace, while the boundary counter still sees the same
two words and throws the dash away. The counter on this site splits on runs of whitespace and counts a
lone dash as a word, and its page tells you that. So adding a space to a dash can push one total up
while leaving another untouched, which is the kind of thing that makes writers decide a tool is
broken. The dashes in this article are set closed, and they will still be counted by someone.
The slash behaves the same way with more possible outcomes. and/or is one token to a
whitespace splitter, two to a tool that breaks at the slash and reports and and
or separately, and three to a tool that also counts the slash itself as a word.
Apostrophes are the one worth checking on your own tool before you trust a total. don’t,
it’s and y’all are one word to any sensible rule, but a counter
that treats all punctuation as a boundary reports two. The mistake runs in one direction: it inflates
your count, and an inflated count is what pushes writers into deleting good sentences for no reason at
all.
I packed a sentence on purpose to show the size of the effect. Here it is:
Our style guide—the up-to-date one—says don’t use and/or. It has 2026 rules and a £4.50 price tag or 50% off.
Count it by whitespace and you get 20 words. Count it with hyphens, dashes and slashes acting as boundaries, keeping the apostrophe inside the word, and you get 25, because five boundaries that the whitespace rule left alone have turned into cuts. That is a quarter of the sentence, in a sentence I wrote deliberately to make the point. Ordinary prose is nowhere near that dense, which is why the gap you are chasing is a few percent rather than a quarter. The mechanism is identical. Only the density of triggers changes.
| What you typed | Counted as | Why the tools differ |
|---|---|---|
| up-to-date | one, or three | A whitespace splitter keeps the hyphens inside the token. A boundary counter returns up, to and date as separate words. |
| word—word | one, two or three | A closed em dash stays inside the token until a counter treats the dash as a boundary. Spaced out, a whitespace splitter counts the dash itself as a word while a boundary counter discards it. |
| don’t | one, or two | The apostrophe is part of the word to nearly every tool. A counter that cuts at all punctuation reports two, which pushes the total up. |
| 2026 | one | Numerals are words to almost every counter. A year is one token with or without other punctuation attached to it. |
| and/or | one, two or three | One for a whitespace splitter, two for a counter that breaks at the slash, three for a counter that also treats the slash as a word. |
| % | zero | A symbol with no letters or digits attached is ignored by every counter I have used. Written closed against a number, as in 50%, it becomes part of one word. |
| Characters, 1,500 words | 7,500 without spaces, about 9,000 with | The difference is the spaces themselves. A form that asks for characters and a tool that reports characters with no spaces are roughly a sixth apart. |
| 今天天气很好 | one | Chinese and Japanese do not put spaces between words, so a whitespace splitter reads a whole paragraph as a single token. Tools that support the script use a character or segmentation rule instead, and the totals are not comparable with the ones on the English side of the page. |
The left column is what you typed and the right column is the rule that changes the answer. Every row is a case where both counters can be right about the same string, which is why the argument about whose number is correct has no winner.
Numerals and standalone symbols
A year is a word to almost every counter, because digits are tokens in the same way letters are.
£4.50 is less settled: some tools count it as one word, some cut at the decimal point and
report two, and the currency symbol is hanging off the front either way. A % or a
$ or a # standing on its own is not a word to anyone, because there is
nothing for it to attach to.
The practical consequence is that a document full of figures will show a wider disagreement between tools than an essay will. Figures are exactly where the rules diverge, so a report with a heavy table, a page of statistics or an invoice-like layout is the wrong place to expect counters to meet.
The text that is not in the body
When a gap is a few percent, the cause is punctuation. When the gap is large, hundreds of words or a tenth of the document, stop looking at punctuation and go looking for text that is on the page but outside the body.
Headings, table cells, captions, footnotes, endnotes, text boxes, comments and tracked changes are each included by some tools and left out by others. Word processors tend to be the least consistent of the group here: the same file can show you a text box or a margin note on screen and then hand you a count that leaves it out, or folds it in, depending on the version and the settings in front of you. A missing text box is not a five-word argument. It is several hundred words, and no amount of comma-watching will find it. This is the first thing to check whenever the totals are far apart rather than close: ask what the document contains that is not a paragraph in the main flow, then count those pieces by hand.
The document furniture counts as well
A title page, a contents list, a bibliography, an appendix. Some tools count everything the file holds. Some let you exclude front matter or reference lists, and then quietly forget which mode they are in. On a tight limit this is often the whole margin: a bibliography of forty entries is several hundred words, and whether it lands inside or outside the total is the difference between a submission that fits and one that does not. If nobody can tell you, put the section in a separate file and count both halves.
Character counts come with and without spaces
Character counts look unambiguous and are not, because a space is a character. English prose averages a little over five characters per word once you include the punctuation stuck to words, and adding the space between them brings it to roughly six. So a 1,500-word piece is around 7,500 characters without spaces and about 9,000 with them. The difference is the number of spaces, which works out at about a sixth of the with-spaces total.
If a form asks for characters and says nothing more, assume it means characters with spaces. If the tool in front of you reports characters without spaces, expect its number to be roughly seventeen percent lower than the one the form will measure. One sixth is a lot of text to remove for a definition, so establish which figure you are being shown before you cut anything.
Scripts with no spaces at all
Chinese and Japanese do not put spaces between words. A whitespace splitter therefore reads an entire
paragraph as one token: 今天天气很好 is one word to that rule, and a page written the same
way is also one word. Tools that handle these scripts use a different rule entirely, either counting
characters or running a segmentation step that guesses where word boundaries fall, and the numbers
they produce have almost nothing in common with the numbers on the English side of a bilingual
document.
For these languages, "word count" means something different, and the honest position is that the number belongs to the tool that produced it. Comparing a Chinese count with an English count is comparing units that merely share a name. I would not use a word count as a payment measure or a limit across scripts at all, and if a client insists on one, ask them for the tool so that both sides are counting the same thing.
Decide which number matters before you write
Most limits arrive with a tool attached: the journal's submission system, the client's content management system, the exam board's checker. That tool is the definition that counts, and the fastest way to end the discussion is to ask which one will be used. If the limit came from a client, ask the client. It is one sentence in an email and it removes the problem rather than solving it.
When nobody can tell you, aim under the limit instead of at it, and use the highest of the counts you can produce. If the strictest counter you have says 1,450 for a 1,500-word limit, you are inside it under any rule anyone is likely to apply. Fifty words of margin is a short paragraph, and it costs far less than a rejected submission.
Rates deserve the same care and rarely get it. Words per minute, price per word, reading time: each is a count divided into something, so a count you cannot reproduce is a rate nobody can check. The reading time on this site's counter is stated at a fixed 225 words per minute, which is a rule you can see and argue with. A bare "about six minutes" is not.
Where the tools fit. The Word Counter reports the figures separately, words on a whitespace split, characters with spaces, characters without, sentences and paragraphs, which is what you need in order to audit a disagreement rather than guess at it. Paste the text in and write all of the numbers down before you change a word, so that when the total moves you can see which measure carried the movement.
The Case Converter handles the step that comes before a fair comparison. When a paste of the same text has to be weighed against the original, normalise the capitalisation first: a heading pasted in title case and the same heading in sentence case hold the same words, and putting both into one case makes the counts comparable rather than merely close.
A count is a gauge, not a verdict
A count measures length, and length is real. A 3,000-word piece is a different promise from a six-hundred-word one, and a limit is usually shorthand for "do not send me an essay". That is what the number is good for: steering. It is not a measure of quality, and every counter described here is blind to whether a sentence works. A piece that reaches 1,500 exactly by padding is worse than one that lands at 1,470 with something to say, and readers can tell which one they are holding.
So: find out which counter will do the checking, learn how it treats hyphens, dashes and apostrophes, look outside the body when the gap is large, and leave yourself a paragraph of margin. Then put the number away and read the piece aloud, which remains the best length check anyone has come up with.