Text Counter
Count characters, words, bytes, and more in real time.
π How to Use
Type or paste the text you want to analyze.
All counts update in real time as you type: characters, words, lines, bytes, and sentences.
Click π Copy Text to copy the input, or ποΈ Clear to reset.
About the Text Counter
The Text Counter measures a string across four axes — characters, words, lines, and bytes — updating in real time as you type. Each axis answers a different question, which matters whenever limits are involved: a tweet is bounded by character count, an essay by word count, a source file by line count, and a network payload by byte count.
How it works
The character count uses the JavaScript String.prototype.length property, which returns the number of UTF-16 code units — not the number of visible symbols. Most text fits in a single code unit (one per ASCII letter, one per common CJK ideograph), but any code point above U+FFFF is encoded as a surrogate pair of two code units. That is why an emoji such as π reports .length === 2 even though a reader perceives one glyph. To count what humans actually see — grapheme clusters, which also cover Hangul jamo combinations and emoji joined with a Zero-Width Joiner (ZWJ) — you need Intl.Segmenter with the grapheme granularity.
Words are split on the regular expression /\s+/ and the non-empty segments are counted. Lines are split on \n. Bytes come from TextEncoder, which encodes the string as UTF-8: ASCII takes 1 byte, Latin-1 letters take 2, Hangul syllables take 3, and most emoji take 4. Everything runs locally in the browser.
Common use cases
- Checking length against a tweet limit of 280 characters or a meta description of ~155
- Counting words for essays, abstracts, and cover letters with a fixed quota
- Measuring the UTF-8 byte size of a payload before sending it over a network
- Estimating reading time at roughly 200–250 words per minute
- Verifying line counts when a style guide caps file length
Worked example
Take the string Hi μλ
π (a space, two ASCII letters, a space, two Hangul syllables, a space, and one emoji). The counts come out as follows:
Characters (.length): 8 // H,i,space,μ,λ
,space,surrogate-pair
Bytes (UTF-8): 13 // 2 + 1 + 6 + 1 + 4
Words (\s+ split): 3 // "Hi", "μλ
", "π"
Lines (\n split): 1
Graphemes (Segmenter): 6 // what the user sees
Notice the gap between .length (8) and graphemes (6): the surrogate pair collapses into one perceived character, so a strict .length check misreports emoji text.
Frequently asked questions
Why does an emoji count as two characters?
JavaScript strings are stored as UTF-16, and the .length property counts code units, not visible symbols. Characters outside the Basic Multilingual Plane (such as emoji) are encoded as a surrogate pair of two code units, so .length reports 2. A grapheme counter using Intl.Segmenter reports 1 because it counts what a user perceives as a single character.
How are words counted?
Words are counted by splitting the text on one or more whitespace characters (the regular expression /\s+/), then counting the non-empty segments. Punctuation attached to a word is not stripped, so hello, counts as one word.
What is the difference between characters and bytes?
Characters counts UTF-16 code units via string .length. Bytes counts the size of the UTF-8 encoding using TextEncoder. An ASCII letter is 1 byte, a Korean Hangul syllable is typically 3 bytes, and an emoji is usually 4 bytes.
How are lines counted?
Lines are counted by splitting the input on the newline character (\n). An empty input reports 0 lines, while a single line with no trailing newline reports 1.
Is my text uploaded to a server?
No. Every count is computed locally in your browser as you type. The text never leaves your device, so the tool is safe for drafts, messages, and sensitive content.
ν μ€νΈ μΉ΄μ΄ν°λ?
ν μ€νΈ μΉ΄μ΄ν°λ λ¬Έμμ΄μ ν¬κΈ°λ₯Ό λ¬Έμ, λ¨μ΄, μ€, λ°μ΄νΈλΌλ λ€ κ°μ§ κΈ°μ€μΌλ‘ μΈ‘μ νλ©°, μ λ ₯ν λλ§λ€ μ€μκ°μΌλ‘ κ°±μ ν©λλ€. κ° κΈ°μ€μ μλ‘ λ€λ₯Έ μ§λ¬Έμ λ΅νλ―λ‘ κΈμ μ μ νμ΄ κ±Έλ¦° μν©μμ μλ―Έκ° μμ΅λλ€. νΈμμ λ¬Έμ μλ‘, μμΈμ΄λ λ¨μ΄ μλ‘, μμ€ νμΌμ μ€ μλ‘, λ€νΈμν¬ νμ΄λ‘λλ λ°μ΄νΈ μλ‘ μ νλ©λλ€. λ΄ μν©μ μ΄λ€ κΈ°μ€μ΄ ν΄λΉνλμ§ μλ κ²μ΄ μ΄ λꡬ νμ©μ μ λ°μ λλ€.
μλ λ°©μ
λ¬Έμ μλ JavaScriptμ String.prototype.length μμ±μ μ¬μ©νλ©°, μ΄κ²μ 보μ΄λ κΈ°νΈμ κ°μκ° μλλΌ UTF-16 μ½λ λ¨μ(code unit)μ κ°μλ₯Ό λ°νν©λλ€. λλΆλΆμ ν
μ€νΈλ μ½λ λ¨μ νλλ‘ ννλ©λλ€(ASCII κΈμ νλλΉ 1κ°, μΌλ°μ μΈ νμ/νκΈ μμ νλλΉ 1κ°). νμ§λ§ U+FFFFλ₯Ό λλ μ½λ ν¬μΈνΈλ μλ‘κ²μ΄νΈ μ(surrogate pair)μ΄λΌλ λ κ°μ μ½λ λ¨μλ‘ μΈμ½λ©λ©λλ€. κ·Έλμ π κ°μ μ΄λͺ¨μ§λ μ½λ μ¬λμ΄ ν κΈμλ‘ λ³΄λλΌλ .length === 2λ‘ λμ΅λλ€. μ¬λμ΄ μ€μ λ‘ μΈμνλ λ¨μμΈ κ·Έλν ν΄λ¬μ€ν°(grapheme cluster) — νκΈ μλͺ¨ μ‘°ν©μ΄λ ZWJ(ν μλ κ²°ν©μ)λ‘ μ΄μ΄μ§ μ΄λͺ¨μ§κΉμ§ ν¬ν¨ — λ₯Ό μΈλ €λ©΄ grapheme λ¨μμ Intl.Segmenterλ₯Ό μ¬μ©ν΄μΌ ν©λλ€.
λ¨μ΄λ μ κ·μ /\s+/(νλ μ΄μμ 곡백 λ¬Έμ)λ‘ λΆν ν λ€ λΉ μΈκ·Έλ¨ΌνΈλ₯Ό μ μΈνκ³ μ
λλ€. μ€μ \nμΌλ‘ λΆν ν©λλ€. λ°μ΄νΈλ λ¬Έμμ΄μ UTF-8λ‘ μΈμ½λ©νλ TextEncoderμμ κ°μ Έμ΅λλ€. ASCIIλ 1λ°μ΄νΈ, λΌν΄-1 κΈμλ 2λ°μ΄νΈ, νκΈ μμ μ 3λ°μ΄νΈ, λλΆλΆμ μ΄λͺ¨μ§λ 4λ°μ΄νΈμ
λλ€. λ¬Έμ₯μ μ’
κ²° ꡬλμ μΌλ‘ λΆν ν©λλ€. λͺ¨λ κ³μ°μ λΈλΌμ°μ μμμ λ‘μ»¬λ‘ μ΄λ£¨μ΄μ§λλ€.
μμ£Ό μ°λ κ²½μ°
- νΈμ 280μ μ νμ΄λ λ©ν μ€λͺ μ½ 155μ κΈ°μ€μ λ§λμ§ νμΈνκΈ°
- μ ν΄μ§ λΆλμ μμΈμ΄, μ΄λ‘, μκΈ°μκ°μμ λ¨μ΄ μ μΈκΈ°
- λ€νΈμν¬λ‘ μ μ‘νκΈ° μ μ νμ΄λ‘λμ UTF-8 λ°μ΄νΈ ν¬κΈ° μ¬κΈ°
- λΆλΉ μ½ 200–250λ¨μ΄ κΈ°μ€μΌλ‘ μμ μ½κΈ° μκ° κ³μ°νκΈ°
- μ€νμΌ κ°μ΄λκ° νμΌ κΈΈμ΄λ₯Ό μ νν λ μ€ μ νμΈνκΈ°
μ¬μ© μ
Hi μλ
πλΌλ λ¬Έμμ΄(곡백, ASCII λ κΈμ, 곡백, νκΈ μμ λ κ°, 곡백, μ΄λͺ¨μ§ ν κ°)μ μκ°ν΄ λ΄
μλ€. κ²°κ³Όλ λ€μκ³Ό κ°μ΅λλ€.
λ¬Έμ μ (.length): 8 // H,i,곡백,μ,λ
,곡백,μλ‘κ²μ΄νΈ-μ
λ°μ΄νΈ (UTF-8): 13 // 2 + 1 + 6 + 1 + 4
λ¨μ΄ (\s+ λΆν ): 3 // "Hi", "μλ
", "π"
μ€ (\n λΆν ): 1
κ·Έλν (Segmenter): 6 // μ¬μ©μκ° λ³΄λ κΈμ μ
.length(8)μ κ·Έλν(6) μ¬μ΄μ μ°¨μ΄λ₯Ό μ£Όλͺ©νμΈμ. μλ‘κ²μ΄νΈ μμ λ μ½λ λ¨μκ° μΈμ§μ ν κΈμλ‘ ν©μ³μ§κΈ° λλ¬Έμ, μ격ν .length κ²μ¬λ μ΄λͺ¨μ§κ° ν¬ν¨λ ν
μ€νΈλ₯Ό μλͺ» μ
μ μμ΅λλ€.
μμ£Ό 묻λ μ§λ¬Έ
μ΄λͺ¨μ§κ° μ λ κΈμλ‘ μΉ΄μ΄νΈλλμ?
JavaScript λ¬Έμμ΄μ UTF-16μΌλ‘ μ μ₯λλ©° .length μμ±μ 보μ΄λ κΈ°νΈκ° μλλΌ μ½λ λ¨μ μλ₯Ό μ
λλ€. κΈ°λ³Έ λ€κ΅μ΄ νλ©΄(BMP) λ°μ λ¬Έμ(μ΄λͺ¨μ§ λ±)λ λ κ°μ μ½λ λ¨μλ‘ μ΄λ£¨μ΄μ§ μλ‘κ²μ΄νΈ μμΌλ‘ μΈμ½λ©λλ―λ‘ .lengthκ° 2λ₯Ό λ°νν©λλ€. Intl.Segmenter κΈ°λ°μ κ·Έλν μΉ΄μ΄ν°λ μ¬μ©μκ° ν κΈμλ‘ μΈμνλ λ¨μλ₯Ό μΈλ―λ‘ 1μ λ°νν©λλ€.
λ¨μ΄λ μ΄λ»κ² μΈλμ?
ν
μ€νΈλ₯Ό νλ μ΄μμ 곡백 λ¬Έμ(μ κ·μ /\s+/)λ‘ λΆν ν λ€ λΉ μΈκ·Έλ¨ΌνΈλ₯Ό μ μΈνκ³ μ
λλ€. λ¨μ΄μ λΆμ ꡬλμ μ μ κ±°νμ§ μμΌλ―λ‘ hello,λ ν λ¨μ΄λ‘ μΉ΄μ΄νΈλ©λλ€.
λ¬Έμ μμ λ°μ΄νΈ μμ μ°¨μ΄λ?
λ¬Έμ μλ λ¬Έμμ΄μ .lengthλ‘ UTF-16 μ½λ λ¨μλ₯Ό μ
λλ€. λ°μ΄νΈ μλ TextEncoderλ‘ UTF-8 μΈμ½λ©νμ λμ ν¬κΈ°λ₯Ό μΈ‘μ ν©λλ€. ASCII κΈμλ 1λ°μ΄νΈ, νκΈ μμ μ λ³΄ν΅ 3λ°μ΄νΈ, μ΄λͺ¨μ§λ λκ° 4λ°μ΄νΈμ
λλ€.
μ€μ μ΄λ»κ² μΈλμ?
μ
λ ₯μ μ€λ°κΏ λ¬Έμ(\n)λ‘ λΆν νμ¬ μ
λλ€. λΉ μ
λ ₯μ 0μ€, νν μ€λ°κΏμ΄ μλ ν μ€ μ
λ ₯μ 1μ€λ‘ λμ΅λλ€.
μ ν μ€νΈκ° μλ²λ‘ μ μ‘λλμ?
μλλλ€. λͺ¨λ μΉ΄μ΄νΈλ μ λ ₯νλ μ¦μ λΈλΌμ°μ μμ λ‘μ»¬λ‘ κ³μ°λ©λλ€. ν μ€νΈλ κΈ°κΈ°λ₯Ό λ λμ§ μμΌλ―λ‘ μ΄μ, λ©μμ§, λ―Όκ°ν λ΄μ©μλ μμ νκ² μ¬μ©ν μ μμ΅λλ€.