Counting characters in code points

Every limit checker on CharLimit.net reports one number first: how many characters you have used against a limit. This guide is about what that number is made of. It names the unit the engine counts in, shows what happens to accented letters, CJK text and emoji under that rule, sets out why the same glyph gets different counts on different platforms, states the engine's own ceiling in the engine's own words, and walks through how the remaining-characters bar on the limit checker is derived. Every figure below was computed by the site's engine at build time on the short sample strings shown, so what you read here is what the checkers do. It is one of the site's guides, and it expands on rules stated once, for every page, on the methodology page.

What a character is on this site

On this site a character is a Unicode code point: one numbered unit of the Unicode standard. The Characters card on every checker and the remaining-characters meter both count code points, which is the convention most platforms document for their caps. A code point is not a byte, not a unit of the UTF-16 representation JavaScript uses internally, and not a visible glyph. For plain English text those four things line up; they stop lining up as soon as accents, other scripts or emoji appear, which is where a draft that looks fine on screen comes back longer, or shorter, than expected. Where a platform's own arithmetic departs from a plain code-point count, the methodology page lists the departure as described rather than modeled, and the relevant checker page says which direction it skews.

Precomposed and combining forms

An accented letter can arrive in a text two ways: as one precomposed code point, or as a base letter followed by a combining accent mark. On screen the two are the same Γ©. The limit check counts the code points actually present, so to it the two forms differ in length. The word cafΓ© written with a precomposed Γ© is 4 code points, leaving 276 of 280 against the X preset; the same word written as e plus a combining acute accent is 5 code points and leaves 275.

The engine's letter breakdown works differently, and the difference is deliberate. Before it counts, letterCount normalizes the text to Unicode's composed form (NFC), so a decomposed accent is folded into the single letter Γ© first. Run on the precomposed word it reports 4 characters and 4 letters; run on the decomposed word it reports 4 characters and 4 letters β€” identical, which is the normalizing behavior the methodology page assigns to the letter breakdown, as distinct from the limit check. The limit meter does not normalize; it counts what the string contains. For the sample above, the decomposed form spends one more unit of the budget on the meter than the glyph count suggests, so in this case the meter reads higher, not lower. For a draft within a few characters of a limit that contains accented text of uncertain origin, the platform's own composer is the place to confirm.

CJK text

A CJK character is one code point, and the engine treats it as one. 東京タワー is 5 code points on the limit meter and 5 letters in the letter breakdown, one per visible character. Its UTF-16 representation is also 5 units, which is what the SMS counter reports for it, since text outside the GSM-7 alphabet is counted in UCS-2. For this script, code points and UTF-16 units agree; the only arithmetic that changes for CJK text here is the SMS budget. X's weighting of CJK characters at 2 (as of mid-2026, a rule the site describes but does not model) is a separate matter, taken up below.

Emoji and joined sequences

Emoji are where the ways of measuring text come apart. πŸŽ‰ is 1 code point, and the limit meter counts it as 1; the same emoji occupies 2 UTF-16 units, which is what the SMS counter charges for it once the message is UCS-2. Put it in a sentence and the gap carries through: Launch day πŸŽ‰ is 12 code points against a limit but 13 units in UCS-2 SMS arithmetic. The letter breakdown of that sentence reports 9 letters, 0 digits and 10 non-space characters, leaving 1 code point that is neither β€” the emoji.

Joined sequences go further. The family emoji πŸ‘¨β€πŸ‘©β€πŸ‘§ is a single visible glyph built from several code points joined by invisible joiners, and the counter reports the code points: 5 on the limit meter, 8 UTF-16 units in UCS-2. A skin-tone variant behaves the same way: πŸ‘ is 1 code point and πŸ‘πŸ½ is 2, because the tone travels as its own code point. None of this is a judgement about what a glyph ought to count as; it is what the string contains.

A mixed line shows the breakdown in full. Order #42 shipped πŸŽ‰ is 19 code points. The letter breakdown reports 12 letters, 2 digits and 16 characters without spaces, so 2 code points are neither letters nor digits β€” the hash sign and the emoji. That remainder is derived rather than reported; the engine's categories are letters, decimal digits, whitespace, and everything else.

Why platforms disagree on joined sequences

The methodology page says it plainly: platforms themselves disagree on such sequences, which is one reason the final authority on a borderline draft is always the platform's own composer. The site's response is to count code points consistently and to describe, rather than simulate, each platform's departure from that count. X's weighted scheme is the documented case: as of mid-2026, per X's developer documentation, every URL is fixed at 23 and emoji and CJK characters are weighted 2. So a draft that reads 1 here for a single πŸŽ‰ is weighted higher there, while a joined sequence that reads 5 here is whatever X's scheme makes of it β€” a server-side question this site declines to answer by guessing. The X counter checks against the 280-character standard-post limit as of mid-2026 (it was 140 until 2017; paid tiers now post longer) and explains which direction the weighting skews the count you see. The Instagram caption counter checks against the 2,200-character cap as of mid-2026 and describes, without modeling, the ~125-character feed fold (as of mid-2026). These caps are product decisions. When a platform changes one, the page copy and preset are updated; the site never extends a limit on its own.

How the remaining-characters bar is computed

The meter under each limit checker (the X, Instagram, YouTube and meta-description pages and the homepage) is a direct rendering of one engine call. limitCheck takes the text and the active limit and returns four values: the code-point count, the limit, the remaining budget (limit minus count, which goes negative past the line), and a flag that is true once the count exceeds the limit. The display reads "remaining" while the flag is off and "over the limit" once it is on; the bar fills to the fraction of the limit used and stops at full. Against the X preset, "Shipping the new release today. Notes in the thread below πŸ‘‡" is 59 code points, so the meter reads "221 of 280 remaining" and the bar fills to 21.1 percent. Repeat a short sentence 12 times and it tips over: the result is 312 code points, the remaining value is -32, the flag is on, the meter reads "32 over the 280 limit", and the bar sits at 100 percent. The Characters card beside the meter counts the same code points, so the two never disagree; the Words card has no bearing on the limit. The engine's limit check accepts a cap from 1 to 100,000; every preset on this site sits inside that range and there is no custom-limit field, so the guard matters only to the engine β€” a cap outside it would be refused with "Enter a valid character limit." instead of a count.

The two-million-character cap, as the engine states it

Every counting function in the engine checks its input before counting, and the check has a ceiling. A string of 2,000,000 characters passes and is counted in full; one character more is refused with the engine's own message: "Text longer than 2 million characters is not supported." The ceiling is applied to the string's length as JavaScript measures it β€” UTF-16 units β€” before any code points are counted. A text made only of πŸŽ‰ therefore reaches the ceiling at 1,000,000 code points, and one more emoji is refused with the same message: "Text longer than 2 million characters is not supported." No platform limit on this site comes within sight of that figure; the ceiling is a guard on the input. When the engine refuses, the meter is left unchanged rather than shown a wrong figure.

Reading a borderline draft

Together, the rules give a short procedure for a draft within a few characters of a limit. First, look at what the draft contains. Plain letters, digits, spaces, punctuation and newline characters count one each here; the uncertainty is concentrated in accents of unknown origin, emoji, joined sequences and, on X, URLs. Second, know the direction of each skew. A decomposed accent reads higher here than the glyph count. A single-code-point emoji reads lower here than X's weighting. A URL reads at face value here and at a fixed 23 on X. A joined sequence reads as its full code-point count here and as whatever the platform decides. Third, when any of those is in play inside the last few characters of headroom, confirm in the platform's own composer before a consequential send β€” the methodology page calls any third-party counter, including this one, a drafting aid rather than a guarantee.

The same discipline applies to SMS, with sharper consequences. There the question is not a code-point count against a cap at all but which encoding the message falls into: a curly apostrophe pasted from a word processor is the classic silent flip from the 160-septet GSM-7 budget to the 70-unit UCS-2 budget, and the SMS counter shows the detected encoding so the flip is never invisible. The segment arithmetic is worked through in full in SMS segments and GSM-7.

Frequently asked questions

Does an emoji count as one character or two?

On this site πŸŽ‰ is 1 code point, and the limit meter counts it as 1. The same emoji is 2 UTF-16 units, which is what the SMS counter charges once a message is UCS-2, and X weights emoji at 2 as of mid-2026 β€” a rule the site describes but does not model. The answer depends on which counter is asking.

Why does the family emoji count as more than one?

πŸ‘¨β€πŸ‘©β€πŸ‘§ is one visible glyph built from several code points joined by invisible joiners. The engine counts code points, so it reports 5; in UTF-16 units, the SMS measure, it is 8. Platforms disagree on such sequences, so for a draft that sits close to a limit and contains one, the platform's own composer is the final authority.

Is Γ© one character or two?

It depends on how it arrived. A precomposed Γ© is one code point, so cafΓ© is 4 on the limit meter. Written as e plus a combining accent the same word is 5, because the meter counts the code points actually present. The letter breakdown normalizes first and reports 4 letters for either form. For the decomposed form the meter reads higher, not lower, so it errs on the safe side.

Why does the SMS counter give a different number for the same text?

Because it runs a different algorithm. The limit checkers count code points against a cap; the SMS counter walks the message against the GSM-7 alphabet and, if any character falls outside it, counts the whole message in UTF-16 units instead. Launch day πŸŽ‰ is 12 code points but 13 UTF-16 units in UCS-2. The full arithmetic is in the SMS segments and GSM-7 guide.

How is the remaining figure on the meter calculated?

Remaining is the active limit minus the code-point count, from the engine's limitCheck. Past the limit it goes negative and the meter switches to "over the limit" wording; the bar fills to the fraction of the limit used and stops at full. Against the X preset, "Shipping the new release today. Notes in the thread below πŸ‘‡" leaves 221 of 280.

Is there a maximum amount of text the checker accepts?

Yes. The engine accepts up to 2,000,000 characters, measured in UTF-16 units before any counting, and refuses anything longer with the message "Text longer than 2 million characters is not supported." The engine's limit check itself accepts a cap from 1 to 100,000 β€” every preset on this site sits inside that range, and there is no custom-limit field; a cap outside it would be refused with "Enter a valid character limit."

Counting is in Unicode code points, as the methodology page states; every figure in this guide was computed by the site's engine at build time on the sample strings shown. Nothing you type is transmitted or stored.