Methodology & Accuracy
Every checker on CharLimit.net runs on one small, tested engine. This page states exactly what that engine computes, which platform behaviors it models versus merely describes, and where the honest uncertainty lives — because a limit checker that hides its own margins is not much of a checker.
This page is the reference; the guides take each part of it further. Counting characters in code points covers what a character is here and why platforms disagree on joined emoji. SMS segments and GSM-7 works the segment arithmetic across more messages than this page has room for. X's weighted counting, explained and Truncation and folds cover the quirks that are described rather than modeled, and which direction each one skews the count you see. Platform limits as of mid-2026 collects the caps in one table with the update promise attached, and Modeled versus described sets out the doctrine behind that distinction. The guides index lists them all.
The unit of counting: Unicode code points
Limit checks count Unicode code points — the convention most platforms document for their caps. Concretely: a precomposed é is 1 code point; written as e plus a combining accent it is 2, and the limit check counts the 2 that are present (the letter breakdown normalizes first and reports 1 letter either way); a CJK character is one; and an emoji like 🎉 is one code point even though it occupies two UTF-16 units in JavaScript's internal representation. Multi-part emoji are the known wrinkle: a family emoji such as 👨👩👧 is a single visible glyph built from several code points joined by invisible joiners, and this counter reports the code points. Platforms themselves disagree on such sequences, which is one reason the final authority on a borderline draft is always the platform's own composer.
SMS: a genuinely different algorithm
The SMS counter does not count code points against a cap — it implements the GSM 03.38 arithmetic. The engine walks your message character by character against the GSM-7 alphabet table: basic-set characters cost one septet, the extended set (€ [ ] { } ~ ^ \ |) costs two, and the first character outside both sets switches the entire message to UCS-2, where length is counted in UTF-16 code units instead (most emoji cost two). Segment math follows from the encoding: GSM-7 fits 160 septets in a single message or 153 per segment once concatenation headers are needed; UCS-2 fits 70, or 67 concatenated. The counter shows the detected encoding explicitly so an invisible flip — a curly apostrophe pasted from a word processor is the classic cause — is never silent. National-language shift tables and carrier-side transcoding exist at the margins and are not modeled; carriers can and do vary.
Modeled versus described
Some platform rules run in the engine; others are stated on the page but deliberately not simulated. The distinction is drawn where simulation would require guessing at server-side behavior we cannot verify from the outside:
- Modeled: code-point counts against each platform's published cap, and the full GSM-7/UCS-2 segmentation above.
- Described, not modeled: X's weighted counting (every URL fixed at 23, emoji and CJK weighted 2 — per X's developer documentation as of mid-2026), Instagram's ~125-character feed fold (as of mid-2026), YouTube's device-dependent title truncation, and Google's pixel-width snippet cut. Each page explains the quirk and tells you which direction it skews the count you see here.
The secondary word counts shown on the stat cards use Unicode-aware word detection: runs of letters and digits, with internal hyphens and apostrophes keeping a word whole.
Worked figures
None of the figures in this section is typed into the page. Each one is computed when the site is built, by calling the same engine the checkers run on the sample text shown beside it. A change to the engine therefore changes these numbers on the next build, and the two boundary samples are pinned: if either ever drifts off the budget it is meant to sit on, the build fails rather than publishing a figure that no longer illustrates what the sentence around it says. Where a cap is named it is the licensed figure as of mid-2026; where a count, a remaining figure, a segment count or a unit total appears, it is the engine's output.
Limit checks at each cap
A limit check is a subtraction. The engine counts the code points in the text, subtracts them from the cap, and reports the count, the cap, the remaining figure and an over flag. It inspects nothing else: every code point, including spaces and each line-break character, counts one. A bare line feed is 1 code point, and because the text box delivers every line break as a single line feed, a pasted line break counts the same.
- 280 (X).
Shipping the new release-notes page today: fewer screenshots, more changelog, and a search box that actually finds things.is 122 code points, leaving 158 of the 280. No URL and no emoji are present, so the described weighting has nothing to act on. - 2,200 (Instagram).
Morning light on the harbour wall, two hours before the market opened. We walked the whole length of the pier with coffee going cold and nobody else around. Back next week with the rest of the roll.is 198 code points, leaving 2,002 of the 2,200. The cap is a long way off for a caption of this length; the ~125-character feed fold is the figure that governs what is seen first, and it is described, not modeled. - 100 (YouTube).
How I organise a small workshop with no spare wall spaceis 56 code points, leaving 44. A longer title shows the other side of the flag:Rebuilding a vintage bicycle from the frame up: stripping paint, replacing bearings, and the first ride down the laneis 117 code points, over the 100 by 17, so the engine reports a remaining figure of -17 and sets the over flag. YouTube's device-dependent title truncation is a separate matter from the cap and is not modeled. - ~160 (meta description).
Plain-language notes on sharpening kitchen knives at home: which stones to start with, how to hold an angle, and how often a blade really needs it.is 147 code points, leaving 13 of the 160 preset. The tilde is deliberate: Google's pixel-width snippet cut is described, not modeled, and the preset is a code-point budget rather than a measurement of width. - Emoji. 🎉 on its own is 1 code point in a limit check, even though it occupies 2 UTF-16 units. The joined family emoji 👨👩👧 is 5 code points and 8 UTF-16 units, and the checker reports the 5 — this is the sequence on which platforms themselves disagree.
SMS segments at the boundaries
The SMS figures come from the alphabet walk described above, so the unit changes with the encoding: septets under GSM-7, UTF-16 code units under UCS-2. The two single-segment budgets, 160 and 70, are the boundaries the pinned samples sit on exactly.
- A plain message.
Your table for two is booked for Friday at 7pm. Reply C to cancel.is GSM-7: 66 septets, 1 segment, 94 septets left in the 160. - The extended-set surcharge.
Use code {SAVE10} at checkout ~ valid until Sundayis 50 code points but 53 septets: the braces and the tilde are extended-set characters at two septets each, so the message spends 3 septets more than its character count suggests. The encoding stays GSM-7 because every character is in one of the two GSM-7 sets. - 160 and one more.
Reminder: your appointment is tomorrow at 10:30. Reply YES to confirm or call the clinic to rebook. Please arrive 10 minutes early and bring your insurance cardis 160 septets exactly: 1 segment, 0 remaining. Add a closing full stop and it is 161 septets, which is 2 segments of 153 each. The per-segment capacity dropped as well as the segment count rising, which is why the remaining figure is now 145: 2 segments of 153 hold 306, less the 161 used. - 70 and one more.
We’re running late. The van should reach you at around half past sevencarries a curly apostrophe, so the whole message is counted in UCS-2 from its first character onward: 70 units exactly, 1 segment, 0 remaining. The same closing full stop makes it 71 units: 2 segments of 67, with 63 units free across the pair (2 × 67 = 134).
The encoding flip
- One apostrophe.
Thanks for your order! It's on its way and should be with you by Thursday.andThanks for your order! It’s on its way and should be with you by Thursday.differ in one character and nothing else. The first, with a straight apostrophe, is GSM-7: 74 septets, 1 segment, 86 septets spare. The second, with the curly apostrophe a word processor substitutes, is UCS-2: 74 UTF-16 units, 2 segments of 67. Nothing about the message got longer; the single-segment budget shrank from 160 to 70, which this 74-unit message no longer fits, so it is sent as 2 segments of 67. - One emoji. Append 🎉 to the plain message above and it becomes UCS-2: 68 code points but 69 UTF-16 units, because the emoji alone costs 2. It is still 1 segment — the message is short enough that the flip did not add one — but the headroom fell from 94 septets to 1 unit. The flip is not always a jump in segments; it is always a change of budget, and the segment count follows from the length.
Where the numbers come from
SMS behavior follows the GSM 03.38 standard — a stable specification, which is why this site states its arithmetic without hedging. Platform caps, by contrast, are product decisions: the 280, 2,200, 100, and ~160 figures reflect each platform's published rules as of mid-2026 and are stated with their history where it matters (X's limit was 140 until 2017; paid tiers now post longer). When a platform changes a limit, the page copy and preset are updated — and until you have verified a consequential send against the platform itself, treat any third-party counter, including this one, as a drafting aid rather than a guarantee.
Where the checker stops
The checker stops at the count. It does not shorten links, weight emoji or CJK characters, measure the rendered width of a snippet, predict where a platform will fold or truncate a caption or a title, or model carrier-side transcoding of a message. Each of those would require either guessing at behavior that happens on a server we cannot observe, or asserting figures that the platform can change without notice. The site states those behaviors instead — as of a date, in plain words, with the direction of skew — and updates the statement when the platform changes. That is a narrower promise than simulating the platform, and a promise the site can actually keep.
How to read a borderline draft
A draft is borderline when the remaining figure is small relative to the quirks the platform applies on its own side. The useful question is which direction each described quirk pushes. On X, a pasted URL is counted here at face value but fixed at 23 by X, so the figure here reads high for a link longer than 23 code points and low for a shorter one. The saving, or the cost, per link is the difference between its length and 23, not the whole link: a draft that reads over by less than the combined saving on its longer links may still fit, and a draft that reads under with a short link in it has less room than it appears to. Emoji and CJK characters are weighted 2 by X and counted 1 here, so a draft that reads under with many of either may not fit. On Instagram the 2,200 is a cap this checker counts against directly, while the ~125-character feed fold is what a reader sees first — a caption can be well inside the cap and still have its point below the fold. On YouTube the 100 is the field's cap; where a title is truncated depends on the device and is not modeled. For a meta description the ~160 preset is a code-point budget, and Google's snippet cut is by pixel width, which this checker does not measure.
SMS is the exception, because there is no platform-side quirk to describe: the arithmetic is the standard's. A borderline SMS draft is read by its encoding line first and its segment count second. If the encoding reads UCS-2 and you expected GSM-7, look for the character that caused it — a curly apostrophe or quotation mark, an em dash, an emoji — before shortening anything, because replacing that one character restores the 160 budget, whereas trimming words inside the 70 budget may not change the segment count at all. If the encoding is what you expected, the remaining figure is exact to the septet or unit, as the boundary samples above show. The reservation that remains is the one stated earlier: national-language shift tables and carrier-side transcoding are not modeled.
Common misreadings
- Reading the character count as the SMS cost. The stat card counts code points; the segment panel counts septets or UTF-16 units. For a plain message they agree. For a message with extended-set characters they do not, and the gap is the surcharge the extended example above makes visible. The segment panel is the one that reflects what is sent.
- Assuming 160 after a flip. Once a single character outside the GSM-7 sets is present, the single-segment budget is 70, not 160, for the whole message — not just for the character that caused it. The encoding line is shown precisely so that this cannot pass unnoticed.
- Carrying the SMS unit into a limit check. Most emoji — the astral ones, such as 🎉 — cost two in a UCS-2 message and one in a code-point limit check: 🎉 itself is 2 UTF-16 units on the segment panel and 1 code point in a limit check, and both are correct, because the two tools count in different units. The unit is stated on each page; the number only makes sense alongside it.
- Treating a pasted link as a hard over on X. This checker reads a URL at face value and X fixes every URL at 23, so the figure here reads high for a link longer than 23 code points and low for a shorter one; the saving, or the cost, per link is the difference between its length and 23, not the whole link. The direction of the skew is known once the link's length is; the exact figure X will reach is not modeled, and the composer is the final authority.
- Reading a tilde figure as a cap. The two tildes on this site mean different things. The ~125 is a description of where Instagram folds a caption in the feed; the engine does not enforce it. The ~160 is a preset the meta-description checker does count against, and its tilde marks that Google's pixel-width snippet cut is described, not modeled: a description that fits the preset can still be shortened in display, and this page says so rather than simulating it.
- Expecting a joined emoji to count as one. A family emoji is one visible glyph and several code points, and the checker reports the code points — 5 for the example above. Platforms disagree on such sequences, so a draft that leans on joined emoji close to a cap is one to confirm in the composer rather than here.
Tested against pinned cases
The engine is a set of pure, typed functions with an automated test suite that pins the cases that have historically broken counters of this kind: GSM-7 alphabet membership (including the characters most often gotten wrong — Ü and § are basic-set, curly quotes are not, form feed is extended), the one-segment/two-segment boundary at exactly 160 and 161 septets, extended characters priced at two septets, the encoding flip to UCS-2 on a single emoji with UTF-16-unit counting, astral emoji as one code point in limit checks, and Unicode word detection across scripts. A build cannot ship with a failing case. If you believe a count is wrong, the contact page explains exactly what to include; confirmed issues are fixed in the engine and locked in with a new test.
Privacy as a design constraint
Everything you type or paste is processed locally in your browser. Nothing is transmitted, logged, or stored — there is no server-side counting endpoint at all, which you can confirm in your browser's network tab while typing.
Platform caps are stated as of mid-2026 and are updated when a platform changes a limit; the worked figures above are recomputed by the engine on every build. Nothing you type is transmitted or stored.