Modeled versus described

Every number on a CharLimit.net checker is one of two kinds. Either the engine computed it from the text you typed, or the page is telling you about a rule it did not run. The methodology page draws that line in two short lists under the heading "Modeled versus described"; this guide explains why the line sits where it does, how to use the distinction when a draft is within a few characters of a cap, and what the site does to keep its own claims honest. It is part of the guides series, and every worked figure below was computed at build time by the engine the checkers run.

Two kinds of statement on a checker page

A modeled rule runs in the engine. When you type into the X/Twitter counter, the engine counts Unicode code points and subtracts them from the cap; when you type into the SMS counter, it walks your message through the GSM 03.38 arithmetic. The result changes with every keystroke and is checked against pinned test cases. A described rule is a sentence on the page: it tells you that the platform counts something differently from the meter, and which direction the difference pushes the figure you see, but the meter itself does not move. A modeled number is one the engine stands behind; a described rule is a documented quirk the engine deliberately does not pretend to reproduce.

What the engine models

Two things, exactly. The first is the code-point count against each platform's published cap: as of mid-2026, 280 for a standard X post, 2,200 for an Instagram caption, 100 for a YouTube title, and the ~160-character drafting budget for a meta description. Those figures are product decisions, stated as of a date and updated in page copy and preset when a platform changes one. The second is the full GSM-7/UCS-2 segmentation for SMS, which follows a stable standard. Here is the first kind at work on a short draft against the X cap:

Doors open at nine. Bring the printed ticket. — 45 code points, 235 remaining of 280, inside the cap.

Nothing about that result depends on X: the engine spreads the string into code points, counts them, and subtracts the total from the cap, which is why one algorithm serves every preset on the landing checker and only the cap changes.

The SMS counter is the other modeled algorithm, and it is genuinely different. Take a draft with one emoji in it: Launch day 🎉. Against a code-point cap it is 12 code points, because 🎉 is one code point even though JavaScript stores it as two UTF-16 units. Run through the SMS algorithm, the same string is UCS-2: 13 UTF-16 units, 1 segment, with 57 units of room left in a 70-unit budget. The first character outside the GSM-7 alphabet switched the whole message to UCS-2, where length is counted in UTF-16 units and most emoji cost two. Two modeled algorithms, one string, two unit counts, each correct for its question. The SMS segments guide works through the full arithmetic.

What is only described

Four rules, as the methodology page lists them, are stated on the relevant pages but not simulated: X's weighted counting (every URL fixed at 23, emoji and CJK weighted 2, per X's developer documentation as of mid-2026), Instagram's ~125-character feed fold, YouTube's device-dependent title truncation, and Google's pixel-width snippet cut. Each page explains the quirk and the direction it skews the count you see here, but the meter counts code points regardless. The truncation and folds guide covers the three display quirks; the weighting rule gets its own treatment below.

The link rule is a clear contrast between the two categories. Here is a draft that ends in a pasted URL: Full notes: https://example.com/notes/launch-day-recap. The engine reports 54 code points, 226 remaining of 280. The URL on its own is 42 characters as written, and the engine charged every one of them. X, as of mid-2026, would charge it 23 whatever its length. That rule is described on the X page and in the weighted-counting guide, not modeled, so the honest statement is about direction: on a link-heavy draft the figure here is higher than X's, and the margin you see is smaller than the margin you have. Emoji and CJK characters push the other way, since each weighs 2 under X's rule and 1 in a code-point count, and the page says so.

Why the line is drawn where it is

The rule is short to state: a behavior is modeled when it can be verified from outside the platform, and described when simulating it would require guessing at server-side behavior that cannot be verified. GSM 03.38 is a stable, published specification; the 160/153 and 70/67 budgets and the two-septet extended set are not opinions, and the engine implements them. A published cap can likewise be checked against the platform's documentation and its composer. The four described quirks fail that test in different ways. Device-dependent truncation and a pixel-width cut have no single value to compute. The Instagram fold is stated as an approximation. X's weighting is documented, but applying it means deciding, for every sequence in a draft, what X's servers will count, and that is exactly where platforms disagree. A joined family emoji is a single visible glyph built from several code points; this counter reports the code points, and platforms themselves differ on such sequences. Modeling every platform's handling of every sequence would mean guessing where a guess cannot be checked, so the page states the rule and its direction and leaves the final word to the composer.

The engine also stops where its own guarantees stop, and says so. It refuses text longer than 2,000,000 characters, and a limit check accepts caps from 1 to 100,000. The code-points guide explains what a character is on this site before any cap is applied.

How to read a borderline count

A borderline draft is one whose remaining figure is within a handful of characters of zero. The modeled count says exactly where you stand in code points; two title drafts against the YouTube cap show the two states.

Repotting a Root-Bound Fiddle-Leaf Fig Without Shocking It: Soil Mix, Pot Size, and Aftercare — 93 code points, 7 remaining of 100, inside the cap.

Now extend the same title by a trailing phrase:

Repotting a Root-Bound Fiddle-Leaf Fig Without Shocking It: Soil Mix, Pot Size, and Aftercare and Watering — 106 code points, -6 remaining of 100, over the cap. A negative remaining figure is the number of characters to cut.

Then ask whether any described rule applies to this draft. For a title, the cap is modeled and the truncation is described: where a results page cuts the title depends on the device, so the YouTube title counter tracks the hard 100 and leaves the fold judgment to you. For an X post, the question is whether the draft contains URLs, emoji, or CJK text; if it does, the skew runs in a known direction and the page says which. For an Instagram caption, the ~125-character fold is a display behavior rather than a cap, so a caption well inside 2,200 can still fold early. For a meta description, ~160 is a drafting budget standing in for a pixel-width cut. When the margin is comfortable and no described rule applies, the count is the answer. When it is a few characters and a described rule applies, the count is a drafting aid and the platform's own composer is the final authority: paste the draft in and read its meter before a consequential send.

The common misreadings

Mistaken readings come from treating a described rule as if it were modeled, or a modeled number as if it covered a rule it does not. The house example is the SMS encoding flip. Two versions of the same message differ by one character. We'll confirm by text., with a straight apostrophe, is GSM-7: 22 septets, 1 segment, 138 septets of room in a 160-septet budget. We’ll confirm by text., with the curly apostrophe a word processor substitutes, is UCS-2: 22 UTF-16 units, still 1 segment, but with 48 units of room in a 70-unit budget. The character count did not change; the budget did. This flip is modeled, and the counter prints the detected encoding so it is never silent. The misreading is to watch only the character stat and miss the encoding label beside it.

The other misreadings are the mirror image. Reading the X counter's remaining figure as X's own count on a post full of links ignores a described rule and overstates how tight you are. Reading a plain count of code points as X's count on an emoji-heavy post understates it. Treating the ~125-character Instagram fold or the ~160-character snippet budget as a hard cap turns a described display behavior into a rule the platform never published. And the national-language shift tables and carrier-side transcoding at the margins of SMS are not modeled; carriers can and do vary, so a large send's segment count is confirmed with the provider.

The honesty mechanism

A doctrine that separates modeled from described only means something if the modeled part is checked. The engine is a set of pure, typed functions with an automated test suite that pins the cases that have historically broken counters of this kind: GSM-7 alphabet membership (Ü and § are basic-set, curly quotes are not, form feed is extended), the one-segment/two-segment boundary at exactly 160 and 161 septets, extended characters priced at two septets, the encoding flip to UCS-2 on a single emoji with UTF-16-unit counting, astral emoji as one code point in limit checks, and Unicode word detection across scripts. A build cannot ship with a failing case. The worked figures on this page come from the same functions at build time, not from hand-typed numbers, so a change in the engine would change them too. Described rules are held to a weaker standard by design: stated as of mid-2026, labeled as described, and revised when a platform changes one. The platform limits guide keeps the current caps in one place. If you believe a count is wrong, the contact page explains what to include; a confirmed issue is fixed in the engine and locked in with a new test.

Frequently asked questions

What does "modeled" mean on this site?

A modeled rule runs in the engine against the text you type. Two things are modeled: code-point counts against each platform’s published cap, and the full GSM-7/UCS-2 segmentation for SMS. Both are pinned by the automated test suite.

What does "described, not modeled" mean?

A described rule is stated on the page but deliberately not simulated: the page explains the quirk and the direction it skews the count you see, while the meter keeps counting code points. As of mid-2026 the described rules are X’s weighted counting (every URL fixed at 23, emoji and CJK weighted 2), Instagram’s ~125-character feed fold, YouTube’s device-dependent title truncation, and Google’s pixel-width snippet cut.

Why not model X’s weighting if the rules are published?

Because the line is drawn where simulation would require guessing at server-side behavior that cannot be verified from outside. Platforms disagree on how to count joined emoji sequences, so a weighted total would be a guess on exactly the inputs where it matters. The page states the rule and its direction and leaves the final word to X’s own composer.

How should I read a count that is within a few characters of the limit?

Take the remaining figure as your margin in code points, then ask whether a described rule applies: URLs, emoji, or CJK text on X; the fold on Instagram; truncation on a YouTube title; the snippet cut on a meta description. If one does, the page tells you which direction it pushes, and the platform’s own composer is the final authority before a consequential send.

Why did my SMS switch to UCS-2 when I only pasted a sentence?

Usually because one character fell outside the GSM-7 alphabet, and a curly apostrophe pasted from a word processor is the classic cause. The first such character flips the whole message to UCS-2, where a single segment holds 70 UTF-16 units instead of 160 septets. The flip is modeled, and the counter shows the detected encoding so it is never silent.

What happens when a platform changes a limit?

Platform caps are product decisions, so each is stated as of mid-2026 and the page copy and preset are updated when a platform changes one. SMS arithmetic follows the GSM 03.38 standard and does not move. Until you have verified a consequential send against the platform itself, treat any third-party counter, including this one, as a drafting aid rather than a guarantee.

Which rules run in the engine and which are only stated is set out above and on the methodology page; counting is in Unicode code points. Nothing you type is transmitted or stored.