
Harriet Cole
My whole subject is that a word count is not one number, and the definition used changes it by a large enough margin to matter when a limit applies. Splitting on whitespace counts space-separated tokens, which treats a hyphenated word as one and a number with a decimal point as one. Splitting on a word-boundary regular expression splits both, and splits contractions in ways nobody expects. Counting characters is simpler and more ambiguous still, since whether spaces count changes the total by a few percent and whether a character means a code point or a byte depends on the encoding. Sentence counting is the least reliable of the set. Splitting on a period followed by a capital breaks on abbreviations and on decimal numbers, so 3.14 and e.g. each produce a sentence boundary that does not exist. Any figure produced this way carries an error rate worth knowing about. Paragraph counting needs a rule about blank lines, because some formats separate paragraphs with a blank line and others do not. Reading time is the most quoted and least examined. It comes from a words-per-minute figure, that figure is a population average, and it assumes continuous focused reading. Technical text read carefully is slower, and skimming is faster. I state the rate on every page rather than presenting the minutes as a measurement. I also cover what the count is used for, because a limit is enforced somewhere and knowing whether the limit counts markup changes the answer substantially. Where a limit applies, I would rather the tool counted with the same definition the rule was written against than quietly applying its own.
About ToolSura
ToolSura offers 80+ free, privacy-first online tools that run 100% in your browser — no uploads, no logins. Learn more about our mission →