SMS Character Counter
The real SMS math, live: encoding detection (GSM-7 vs Unicode), the correct 160/153 or 70/67 budget, and the segment count that determines what a message actually costs to send.
Type a message to see its SMS segments.
Read the guide: SMS segments and GSM-7, explained — the full GSM 03.38 arithmetic walked through message by message, from the 160/161 boundary to the one-emoji flip. More in all guides.
What this checker counts
The SMS counter does not count characters against a cap. It implements the GSM 03.38 arithmetic: the engine walks your message character by character against the GSM-7 alphabet table. Basic-set characters cost one septet. The extended set — € [ ] { } ~ ^ \ | — costs two. The first character outside both sets switches the entire message to UCS-2, where length is counted in UTF-16 code units instead, and most emoji cost two. The budget then follows from the encoding: GSM-7 fits 160 septets in a single message, or 153 per segment once concatenation headers are needed; UCS-2 fits 70, or 67 concatenated.
The result panel reports three things, in this order: the segment count with the detected encoding in brackets, the units used, and how many units are left before the next segment opens. The Characters card beneath it is a different number on purpose. It counts Unicode code points, the way the other checkers on this site do, so a message containing extended characters or emoji shows more units than characters. The units figure is the one the segment arithmetic runs on; the Characters card is there so you can see the gap between the two.
Where the budgets come from
The 160/153 and 70/67 budgets, the alphabet table and the two-septet extended set all come from GSM 03.38 — a stable standard rather than a product decision, which is why this page states its arithmetic without hedging. The figures here are current as of mid-2026; should the standard or the engine's table ever change, the page copy and the checker are updated together. What the standard does not fix is what happens after the message leaves your gateway, which the described-not-modeled section below covers. Until you have verified a consequential send against your provider, treat any third-party counter, including this one, as a drafting aid rather than a guarantee.
Worked examples, computed at build time
Each example below is an original draft run through the same engine the live checker uses; no figure in this section is typed in by hand. Paste any of them into the box above and the panel will match.
One plain segment. Your table for two is confirmed for Thursday at 7pm. Reply C to cancel or call us to change the time.
Every character is in the basic set, so septets equal characters: 101 units, GSM-7, 1 segment, 59 left before the next segment.
The extended set, priced. Weekend sale: 20% off everything over €40 with code {SAVE} ~ ends Sunday
The Characters card reads 72, but €, the two braces and the tilde each cost two septets, so the panel reads 76 units — 4 extended characters, 4 extra septets. Still 1 segment, with 84 left.
Exactly full, then one more. Reminder: your dental appointment is tomorrow at 10:30am. Please arrive 10 minutes early and bring your insurance card. Reply YES to confirm or NO to change it.
160 septets, GSM-7: 1 segment with 0 left — it fills the single segment to the last septet. The same text with one trailing space after the final full stop is 161 septets and 2 segments. Because concatenation headers now take room, each segment holds 153, so the panel shows 145 left before a third segment opens rather than one over.
One emoji flips the encoding. Your parcel is out for delivery today between 2pm and 6pm. No signature needed.
79 septets, GSM-7, 1 segment, 81 left. Append a space and a 🎉 and the emoji is outside GSM-7, so the whole message is recounted in UTF-16 units — 82, of which the emoji is two — and the budget drops to 67 per segment: 2 segments. The Characters card reads 81, because the emoji is one code point; the two figures differ by exactly the emoji's second UTF-16 unit.
The curly-apostrophe flip. You're booked for 9am tomorrow. Reply STOP to opt out. versus
You’re booked for 9am tomorrow. Reply STOP to opt out.
Identical length — 54 and 54 units — and only the apostrophe differs. The first is GSM-7 with 106 left of 160; the second is UCS-2 with 16 left of 70. Same message, different budget, visible only through the encoding label.
Accents are a table lookup. Café confirmé à Zürich für Señor Peña
é, à, ü and ñ are all basic-set, so this is GSM-7 at 37
septets. The list is exact rather than general: à is in it and á is not, so
Olá is UCS-2 at 3 units, and the em dash in
Sale ends Sunday — last chance does the same (UCS-2). Whether a character fits GSM-7 is
membership in a table, not a judgement about how it looks.
A two-segment message, plus one emoji. Hi Sam, your order #4821 has shipped and should arrive within three working days. Track it from the link in your confirmation email. If anything looks wrong with the delivery address, reply to this message before 5pm today and we will update it before dispatch.
261 septets: 2 segments at 153 each, 45 left. Append a 🎉 and it is 264 UTF-16 units at 67 per segment — 4 segments. One emoji, 2 extra segments on every copy sent.
What is described, not modeled
The methodology page draws a line between rules that run in the engine and rules that are only stated. For SMS, the modeled part is the full GSM-7/UCS-2 segmentation shown above. Described, not modeled: national-language shift tables and carrier-side transcoding exist at the margins and are not modeled; carriers can and do vary. In practice that means the count here is the standard's count for your exact characters, and a gateway between you and the handset may encode the message differently from the table on this page. This checker does not try to predict that — simulating it would require guessing at behavior we cannot verify from the outside — so it reports the standard, shows the encoding it detected, and leaves the final word on a large send to your SMS provider's own segment count.
Reading a borderline draft
Check the encoding label first. If it says UCS-2 and you expected GSM-7, one character flipped it: a curly apostrophe or quote, an em dash, an emoji, or an accented letter outside the basic set. Replacing that single character restores the 160 budget, and the examples above show the size of the swing — the apostrophe pair goes from 16 left to 106 left with no other change.
Then read the units against the per-segment figure, not against 160. Once a message is over the single-segment ceiling, every segment holds 153 (or 67 in UCS-2), so "left before the next segment" is measured against the concatenated budget. That is why the full-segment example reads 145 left at 161 units rather than one over: the second segment arrived with its own room. Trailing whitespace counts too — a space or line break after the final character is a septet like any other, which is how a message that reads as exactly full in one place is two segments in another.
Finally, verify at the provider, not the handset. You cannot verify segmentation by texting yourself, because phones reassemble concatenated messages invisibly — the split shows up only on the bill. The arithmetic on this page is the standard's; the figure you are billed is your provider's, and for a large send the two should be compared before it goes out.
Common misreadings
- Reading the Characters card as the budget. It counts code points. The budget is spent in septets or UTF-16 units, and the two diverge the moment an extended character or an emoji appears — 72 characters against 76 units in the sale example.
- Expecting accents to flip the encoding, or expecting emoji not to. The basic set carries a specific list of accented letters, so a draft like the Zürich example stays GSM-7; every emoji is outside the set, so a single one always flips the message.
- Treating 160 as fixed after concatenation. The second segment does not add 160 on top of the first. Both drop to 153, which is why 161 septets leaves 145 spare rather than one over.
- The curly-apostrophe flip. The house example: a curly apostrophe pasted from a word processor is the classic cause of a silent switch to the 70 budget. The checker shows the detected encoding precisely so this is never silent — compare the apostrophe pair above.
- Assuming one extra character costs one more segment at the old rate. It costs a segment at the concatenated rate, and the parcel example shows the sharper version: one emoji moved a one-segment message to 2 by changing the budget it was measured against, not just by adding two units.
Frequently asked questions
Why is SMS not just "160 characters"?
Because 160 assumes the compact GSM-7 alphabet. A single character outside it — most emoji, curly quotes, many accented letters — switches the whole message to Unicode (UCS-2), where a segment holds only 70. And once a message needs multiple segments, each segment shrinks (153 or 67) to make room for stitching headers.
What are segments, and why do they matter?
Long messages are split into segments that carriers transmit separately and phones reassemble. Senders are billed per segment — the 261-character order update in the worked examples above is 2 segments as plain GSM-7, and adding one emoji makes it 4 (264 Unicode units ÷ 67). For bulk SMS, segment math is money.
Which characters secretly cost two?
Within GSM-7, the extended set — € [ ] { } ~ ^ \ | — costs two "septets" each. And watch pasted text from word processors: curly apostrophes (’) are NOT in GSM-7 and silently flip the entire message to the 70-character Unicode budget. This counter shows the detected encoding so the flip is visible.
Do links and phone numbers count normally?
Yes — SMS has no shorteners or special counting; every character is just a character. That is why serious SMS senders use short domains and strip words: at 160-or-70, each character is budget.
Why does the Characters card disagree with the units figure?
The Characters card counts Unicode code points, like every other checker on this site. The units figure is what the segment arithmetic runs on: septets in GSM-7, where extended characters cost two, or UTF-16 code units in UCS-2, where most emoji cost two. The sale example above is 72 characters but 76 septets; the parcel notice with one emoji is 81 characters but 82 units.
Do accented letters force the Unicode budget?
Not all of them. The GSM-7 basic set includes a specific list of accented letters, so "Café confirmé à Zürich für Señor Peña" stays GSM-7 at 37 septets. The list is exact rather than general: à is in it and á is not, so "Olá" is UCS-2. Whether a character fits is a table lookup, and the detected encoding tells you the outcome.
What does "left before the next segment" mean?
It is the room remaining in the budget you are currently in, not distance from 160. The full-segment example above sits at exactly 160 septets with 0 left; add one trailing space and it becomes 2 segments of 153, so the panel reads 145 left before a third segment opens, not one over.
Can I check the split by texting myself?
No. Phones reassemble concatenated messages invisibly, so the split shows up only on the bill. The arithmetic here follows the GSM 03.38 standard; national-language shift tables and carrier-side transcoding exist at the margins and are not modeled, and carriers can and do vary, so confirm billing behavior with your SMS provider before a large send.
Segmentation follows the GSM 03.38 standard (160/153 septets for GSM-7 with extended characters costing two; 70/67 UTF-16 units for Unicode). Carrier behavior can vary at the margins. Nothing you type is transmitted or stored. See the methodology page.