Word & Character Counter
Count whitespace tokens, Unicode code points, sentences and paragraphs
Paste text and inspect word, character, and paragraph counts first; scenario comparisons are available in Advanced mode.
Text stays in this page; no drafts are saved. Whitespace tokens are not Chinese word segmentation, and code points are not visual characters. Reading time uses Latin-letter words and CJK code points only.
The full guide also includes pitfalls, worked examples, snippets, FAQs, and related tools for checking results or troubleshooting.
About this tool
Paste a draft to see whitespace-separated tokens, Unicode code points with and without whitespace, CJK code points, sentence estimates and blank-line paragraphs. These are explicit counting conventions: a Chinese sentence without spaces is one token, while its Han characters are counted separately; a combined emoji may contain several code points. Frequent words come from Latin-letter sequences with common English stop words removed. Reading and speaking time use Latin-letter word counts and CJK code points at fixed rates, rather than predicting a particular reader. The English Flesch estimate appears only for ASCII text with more than ten Latin-letter words; it uses a simple syllable heuristic. Text remains on this page and is not saved as a draft.
Scenario Recipes
Compare a mixed-script count
Goal: Understand the counter before applying a length limit.
- Paste Hi π δΈ.
- Compare whitespace tokens, Unicode code points and CJK code points.
Result: 3 tokens, 6 code points including spaces, 4 without whitespace and 1 CJK code point.
Suggested Workflow
Production Snippets
Worked example
text
Hi π δΈ β tokens 3; code points 6; CJK 1
π¨βπ©βπ§βπ¦ β code points 7Failure Clinic (Common Pitfalls)
A family emoji occupies several code points
Cause: The counter does not segment grapheme clusters.
Fix: Use the receiving editorβs rule when a limit refers to visual characters.
Frequently Asked Questions
What does the main word count mean?
It counts nonempty chunks separated by whitespace. For example, βHi π δΈβ has three tokens, six Unicode code points and one CJK code point. Punctuation-only chunks also count as tokens.
Are characters the same as visible symbols?
No. Characters are counted as Unicode code points, not UTF-16 units, bytes or grapheme clusters. π is one code point; the family emoji π¨βπ©βπ§βπ¦ contains seven.
How are sentences and paragraphs counted?
The sentence estimate splits on . ! ? and their Chinese equivalents. Paragraphs are separated by blank lines; CRLF and CR are normalized for these boundaries. Abbreviations and punctuation can affect the sentence estimate.
How are reading time and readability estimated?
Reading uses 200 Latin-letter words/minute and 400 CJK code points/minute; speaking uses 130 and 250. Other scripts are not modeled. Flesch uses an English syllable heuristic and is hidden for non-ASCII or short text.
Why can another editor show a different count?
Editors may use different token, punctuation and grapheme rules. For a submission limit, follow the receiving platformβs rules. This tool does not judge writing quality or SEO value from length.
Keep browsing