SMS segments and GSM-7, explained

An SMS is not 160 characters long. It is 160 septets long while every character comes from one small alphabet, 70 UTF-16 units long the moment any character does not, and shorter per part once it has to be split. This guide walks through that arithmetic — the GSM 03.38 rules the SMS character counter implements — on worked messages whose figures were computed at build time by the engine that runs the counter, and ends with what the counter deliberately does not model.

What a segment is

A single SMS carries a fixed amount of data. A message that fits is one segment; one that does not is split into several, each carrying a concatenation header so the parts can be reassembled in order. Those headers take room out of the same fixed envelope, which is why the allowance per part drops as soon as a message needs more than one: from 160 septets to 153 per segment in GSM-7, and from 70 UTF-16 units to 67 in UCS-2. Segments are what is actually sent, so for a large send the question is never "how many characters" but "how many segments" — and the second is not a simple function of the first.

Two encodings, two kinds of unit

GSM 03.38 defines the GSM-7 alphabet, and a message stays in it only while every character is in the table. The basic set costs one septet per character: unaccented Latin letters, digits, space, line break, common punctuation, and a specific selection of accented letters (é, ü, ñ, à and ö among them). The extended set — € [ ] { } ~ ^ \ | — costs two septets per character. The first character in neither set changes the encoding of the entire message to UCS-2, where length is counted in UTF-16 code units instead. Most emoji are one code point but two UTF-16 units, so in UCS-2 they cost two; the code-points guide explains why those two counts differ.

The flip does not penalize the offending character alone: it re-prices every character in the message, plain letters included, against the 70/67 budget. That is why a single emoji can turn a one-segment message into 3 — the last sentence of Example 4 does exactly that to the 160-septet message of Example 3 — and why the counter shows the detected encoding first.

The arithmetic, step by step

The engine runs this on every keystroke, in order:

  1. Walk the message one character at a time: a basic-set character adds one septet, an extended-set character two.
  2. The first character in neither set stops the walk. The message is UCS-2, and its length is its count of UTF-16 code units — the whole message, not just the part after the flip.
  3. Pick the budget from the encoding: 160 single or 153 per concatenated segment for GSM-7; 70 or 67 for UCS-2.
  4. Count segments: zero for an empty message, one within the single-segment allowance, otherwise the unit count divided by the concatenated allowance, rounded up.
  5. Report the room left: the capacity of the segments in use (the single allowance for one segment, otherwise segments times the concatenated allowance) minus the units used — how many more units fit before another segment is needed.

Each example below shows those five outputs in that order.

Worked examples

Example 1 — plain text. Your parcel is out for delivery today between 9am and 1pm. Reply R to reschedule.

GSM-7 · 81 septets · 1 segment of 160 · 79 to spare

Every character is in the basic set, so septets equal characters: 81 characters, 81 septets, 79 to spare in a single segment. In GSM-7 this is the only case in which "characters" and "SMS length" are the same number; a UCS-2 message matches too only when nothing in it takes two UTF-16 units (see Example 6).

Example 2 — extended characters cost two. Spring sale: save €15 on orders over €60 [code SPRING] until Sunday.

GSM-7 · 72 septets · 1 segment of 160 · 88 to spare

68 characters but 72 septets: the difference of 4 is one extra septet for each of the 4 extended characters (2 euro signs, 2 square brackets). On its own, a single € reports 2 septets. The encoding is still GSM-7 — extended characters are expensive, but they do not flip the message.

Example 3 — the 160/161 boundary. Reminder: your dental appointment is on Tuesday at 10:30am at Northgate Dental Care, 22 Mill Lane. Please arrive ten minutes before your time. Reply C to cancel

GSM-7 · 160 septets · 1 segment of 160 · 0 to spare

Add one closing full stop and the engine reports:

GSM-7 · 161 septets · 2 segments of 153 · 145 to spare

The full stop did not merely add one septet of cost; it added a second segment. Every part of a concatenated message now carries a header, so the allowance per part fell from 160 to 153 while total capacity rose to 306 — which is why a 161-septet message shows 145 septets to spare: it is already paying for two segments.

Example 4 — one emoji flips the encoding. Your parcel is out for delivery today between 9am and 1pm. Reply R to reschedule. 🎉

UCS-2 · 84 UTF-16 units · 2 segments of 67 · 50 to spare

Example 1 plus a space and a 🎉. In characters it grew by 2, to 83 code points; in units by 3, because the emoji is two UTF-16 units. The larger change is the flip: the 🎉 is outside both GSM-7 sets, so the whole message became UCS-2, the single-segment allowance dropped from 160 to 70, and 84 units is over it — 2 segments of 67 where there was 1. The same emoji appended directly to Example 3's 160-septet message — 1 character added, no space — gives 162 units and 3 segments.

Example 5 — the 70/71 boundary. Happy birthday from all of us at Northgate Dental! Have a great day 🎂

UCS-2 · 70 UTF-16 units · 1 segment of 70 · 0 to spare

The 🎂 makes this UCS-2, so the budget is 70: 69 code points but 70 UTF-16 units, the emoji counting two. It fits exactly, with 0 to spare. One more exclamation mark gives:

UCS-2 · 71 UTF-16 units · 2 segments of 67 · 63 to spare

Example 3's mechanism at the other budget: 2 segments of 67, capacity 134, 63 units to spare.

Example 6 — the curly apostrophe. Thanks for your order! It's on its way and should arrive by Thursday. We'll text you again when it's out for delivery.

GSM-7 · 118 septets · 1 segment of 160 · 42 to spare

Paste the same message from a word processor that has curled the apostrophes — Thanks for your order! It’s on its way and should arrive by Thursday. We’ll text you again when it’s out for delivery. — and the engine reports:

UCS-2 · 118 UTF-16 units · 2 segments of 67 · 16 to spare

The unit count is identical — 118 characters, 118 units either way, since a curly apostrophe is a single UTF-16 unit — but the budget is not. The straight apostrophe is basic-set; the curly one is in neither set, so the message is UCS-2 and 118 units needs 2 segments of 67. Nothing is visibly different, which is why this is the classic silent flip and why the counter prints the detected encoding rather than leaving you to infer it.

Example 7 — accents in and out of the table. Bestätigt: Menü für zwei um 20 Uhr. Viele Grüße und à demain!

GSM-7 · 61 septets · 1 segment of 160 · 99 to spare

ä, ü, ß and à are all basic-set, so this is GSM-7 at one septet each. Change the sign-off from "à demain" to "à bientôt" and the ô, which is in neither set, flips it:

UCS-2 · 62 UTF-16 units · 1 segment of 70 · 8 to spare

Accents are not the trigger; membership in the table is, and the way to know is to let the engine walk the message and read the encoding it reports.

Reading a borderline draft

Read the encoding first, the segment count second, the room remaining third. GSM-7 with units equal to characters means every character is basic-set; units above characters means the difference is the number of extended characters. UCS-2 means something is outside the table, and units exceed characters by one for every emoji that takes two UTF-16 units. The room remaining is the margin: zero means the message exactly fills its segments, and the next unit opens another.

The common misreadings all use the wrong unit. Treating 160 as a character limit misses the two-septet extended characters, so 68 characters can be 72 septets. Assuming any accent flips the encoding leads to needless rewriting, since é, ü, ñ and à are basic-set; assuming none does misses ô and its kind. Forgetting that a line break is a character undercounts: See you at 7. Bring the signed forms. is 37 characters and 37 septets. The quietest is the curly apostrophe of Example 6, which changes nothing you can see and everything the budget is measured against. When a flipped message surprises you, find the character responsible and replace it; the encoding readout tells you when you have succeeded.

What the counter does not model

The arithmetic above is the GSM 03.38 standard — a stable specification, which is why this site states it without hedging — and the counter stops where that arithmetic stops. National-language shift tables and carrier-side transcoding exist at the margins and are not modeled, and carriers can and do vary. The counter says nothing about delivery, display on the recipient's device, or what a particular provider will do with it: it reports the segment count the standard produces for the text you typed, and the encoding behind it. Until you have verified a consequential send against the platform itself, treat any third-party counter — including this one — as a drafting aid rather than a guarantee. The methodology page lists what is modeled and what is described across every checker.

Frequently asked questions

Why is my message two segments when it is under 160 characters?

Because 160 is a budget of septets, not characters, and it applies only while the message stays in GSM-7. Extended characters (€ [ ] { } ~ ^ \ |) cost two septets each, so the sale message ("Spring sale: save €15 …") is 68 characters but 72 septets. And one character outside the table flips the whole message to UCS-2, where a segment holds 70 (67 concatenated): the order confirmation ("Thanks for your order! …") is 118 septets and 1 segment with straight apostrophes, 118 units and 2 segments with curly ones.

Does an emoji really cost two?

In UCS-2, yes: length is counted in UTF-16 code units and most emoji occupy two. The parcel-delivery message ("Your parcel is out …") is 81 septets in GSM-7; add a space and a 🎉 and the engine reports 84 UTF-16 units — 3 more, one for the space and two for the emoji — in 2 segments of 67. The larger cost is the flip: every other character is now priced against 70/67 instead of 160/153.

Do accented letters push a message to UCS-2?

Only the ones outside the GSM-7 table. The booking confirmation ("Bestätigt: Menü für zwei …") contains ä, ü, ß and à and reports GSM-7: 61 septets, 1 segment. Change its sign-off to "à bientôt" and the ô, which is in neither set, flips it to UCS-2: 62 UTF-16 units with 8 remaining in a 70-unit segment.

Why does a segment hold 153 instead of 160 once a message is split?

Each part of a concatenated message carries a concatenation header, and the header takes room out of the same fixed-size envelope. In GSM-7 the allowance drops by 7 septets, from 160 to 153; in UCS-2 by 3 units, from 70 to 67. That is why the dental reminder with its closing full stop ("Reminder: your dental appointment …"), at 161 septets, is 2 segments with 145 septets to spare, not one segment plus one character.

Can the counter tell me what my carrier will do with the message?

No. It implements the GSM 03.38 arithmetic — the alphabet walk, the two-septet extended characters, the UCS-2 flip and the 160/153 and 70/67 budgets — and reports the segment count that produces. National-language shift tables and carrier-side transcoding are not modeled, and carriers can and do vary. Treat the segment count as a drafting figure and verify a consequential send against the platform itself.

How do I get a flipped message back to GSM-7?

Find the character that caused the flip and replace it with a basic-set equivalent. If the readout says UCS-2 and you expected plain text, the usual cause is a curly apostrophe or quotation mark pasted from a word processor. Straightening the 3 apostrophes in the order confirmation ("Thanks for your order! …") takes it from UCS-2 and 2 segments to GSM-7 and 1 segment at the same 118 units.

Every figure on this page was computed at build time by the engine behind the SMS character counter, following the GSM 03.38 standard (160/153 septets for GSM-7 with extended characters costing two; 70/67 UTF-16 units for Unicode). Carrier behavior can vary at the margins. Nothing you type is transmitted or stored. See the methodology page.