Guide
SMS segments explained: GSM-7, UCS-2 and what they cost
A text isn't billed by the message or by the character. It's billed by the segment, and one curly apostrophe can double the count. Here's how segments work, with worked examples.
Key takeaways
- A single SMS holds 160 characters in GSM-7 or 70 in UCS-2. Both are the same 140 bytes.
- Longer texts are split into segments of 153 (GSM-7) or 67 (UCS-2) characters, because each part carries a small header that lets the phone reassemble it.
- One character outside the GSM-7 alphabet, such as an emoji or a curly apostrophe, switches the whole message to UCS-2.
^ { } \ [ ] ~ | €are in GSM-7 but take two slots each.- Twilio bills SMS per segment, so segments, not messages or characters, drive the cost.
Why SMS stops at 160 characters
The SMS standard gives each message 140 bytes of payload, which is 1,120 bits. How many characters fit into those bits depends on how each character is encoded:
- GSM-7 uses 7 bits per character: 1,120 ÷ 7 = 160 characters.
- UCS-2 uses 16 bits per character: 1,120 ÷ 16 = 70 characters.
You don't choose the encoding. The sending platform chooses it from the characters in your message. If every character is in the GSM-7 alphabet, the message goes as GSM-7. If even one isn't, the whole message goes as UCS-2. That single rule explains most surprise bills.
GSM-7, the default alphabet
GSM-7 comes from the GSM 03.38 standard. Its basic table has 128 characters: the English alphabet, digits, common punctuation, and a handful of accented letters and symbols used in Western European languages.
| Group | In the basic GSM-7 table |
|---|---|
| Letters and digits | A–Z, a–z, 0–9 |
| Punctuation | space . , ! ? ' " : ; ( ) - / @ # % & * + = < > _ |
| Accented letters | é è ù ì ò à É Ç Ä Ö Ü ä ö ü Ñ ñ Å å Ø ø Æ æ ß |
| Symbols | £ $ ¥ § ¿ ¡ ¤ and some Greek capitals such as Δ and Ω |
| Line break | A new line counts as one character |
Extension characters count twice
Nine characters live in an extension table rather than the basic one: ^ { }
\ [ ] ~ | and €. Each is sent as an escape
code followed by the character, so it takes two of the 160 slots. Extension characters don't change the
encoding; they only cost more room. A message with 158 ordinary characters and one euro sign fills exactly 160 slots and
still fits in one segment. A second euro sign pushes it into two.
UCS-2: when one character changes everything
When a message contains a character that isn't in GSM-7, every character in it is sent as 16 bits, including the plain letters. The limit for a single message drops from 160 to 70. The characters that cause this are rarely exotic. Most arrive through copy and paste or automatic formatting:
| Character | Where it usually comes from | GSM-7 alternative |
|---|---|---|
| ’ ‘ curly apostrophe and single quotes | Word processors, phone keyboards, email | ' straight apostrophe |
| “ ” curly double quotes | The same | " straight quotes |
| – — en and em dashes | Automatic formatting | - hyphen |
| … ellipsis as one character | Autocorrect | ... three periods |
| Non-breaking space (invisible) | Text copied from web pages and documents | An ordinary space |
| • bullet | Pasted lists | - or * |
| á í ó ú â ê ô ã õ | Names and words in Spanish, Portuguese and French | Often none; accept UCS-2 or rephrase |
| Emoji | Keyboards and templates | Words |
Emoji need a special mention. Most emoji sit outside the range that UCS-2 was designed for, so each one is sent as two 16-bit units. (Strictly speaking that's UTF-16, but the messaging industry still says UCS-2.) A single 📦 therefore counts as two characters toward the 70 or 67 limit. Emoji with skin tones, flags and combined figures are built from several code points and can take four units or more.
Long messages: 153 and 67
When a message doesn't fit in one SMS, it's split into several segments and the recipient's phone reassembles them. To make that possible, each segment carries a 6-byte header that says which message it belongs to and where it goes in the sequence. The header comes out of the same 140 bytes, which leaves room for 153 GSM-7 or 67 UCS-2 characters per segment.
So a 161-character message isn't 160 plus 1. It's two segments of up to 153 characters, and every segment in it holds 153, not 160.
| Segments | GSM-7, up to | UCS-2, up to |
|---|---|---|
| 1 | 160 characters | 70 characters |
| 2 | 306 characters | 134 characters |
| 3 | 459 characters | 201 characters |
| 4 | 612 characters | 268 characters |
| A full 1,600 characters | 11 segments | 24 segments |
For UCS-2, count 16-bit units rather than visible characters, so each emoji usually counts as two. Twilio accepts message bodies of up to 1,600 characters. Near a segment boundary, tools can disagree by one segment, because most senders avoid splitting a two-slot character across two segments. Leave a little headroom rather than writing to the exact limit.
Worked examples
Each count below follows the rules above. The captions show characters, encoding and segments.
1. A plain confirmation
Everything here is in the GSM-7 basic table, so it's one segment with 95 characters to spare. Most short transactional texts look like this.
2. One apostrophe
The words are identical. The second version has a curly apostrophe in you’re, which isn't in GSM-7, so the whole message switches to UCS-2. At 100 characters it no longer fits in 70, and it becomes 2 segments instead of 1 segment.
3. One emoji
The emoji version has only 70 visible characters, which looks like it should fit in 70. But the emoji takes two units, making 71, so it needs 2 segments.
4. Extension characters near the limit
Both messages have 159 characters and stay in GSM-7. In the first, the euro sign takes two slots, which makes exactly 160: 1 segment. The second swaps the parentheses for square brackets, which are extension characters too. That adds two more slots, 162 in all, and the message becomes 2 segments.
5. A longer message, with and without an emoji
The first version is 226 GSM-7 characters, so it's 2 segments of up to 153. Replacing the final period with a smiley switches it to UCS-2: 228 units at 67 per segment is 4 segments. One character doubled the cost of the message.
What segments cost
Twilio bills SMS per segment, and received texts are segmented and billed the same way. Prices depend on the destination country and the kind of number, and they change over time, so check Twilio's pricing for the countries you text rather than relying on a figure from an article.
The arithmetic is simple once you think in segments:
- Billable segments = recipients × segments per message.
- A two-segment message to 1,000 people is 2,000 segments. If an emoji pushes it to four segments, it's 4,000.
- MMS is priced per message rather than per segment, with its own size limits, so it follows different arithmetic.
Merge fields make the count personal. Hi {!Contact.FirstName|there} is shorter for Al than for
Maximiliano, so a template that fits in one segment for most people can take two for a few. A merged value can also
change the encoding: a first name with á or ó switches that person's message to UCS-2 even when every
other recipient gets GSM-7.
Tips to avoid surprise segments
- Type templates rather than pasting them. If you must paste, go through a plain-text editor first to strip curly quotes, long dashes and non-breaking spaces.
- Watch smart punctuation. Phones and word processors turn ' into ’ as you type. It's worth switching off where people write templates.
- Use emoji deliberately. If one belongs in the message, budget for 70 or 67 characters per segment.
- Prefer parentheses to brackets and braces. ( ) cost one slot each; [ ] and { } cost two.
- Keep fallbacks short, and test templates with the longest real values you have, such as the longest first name or product name.
- Count the opt-out line. Reply STOP to opt out. is 22 characters plus a line break.
- Keep links short. A long tracking URL can fill most of a segment on its own.
- Know what your Twilio settings do. A Twilio Messaging Service can use Smart Encoding, which swaps some Unicode lookalikes, such as curly quotes and long dashes, for GSM-7 equivalents before sending. If you turn it on, what Twilio bills can be lower than a count made before the swap.
How ConnectSMS shows encoding and segments
ConnectSMS is a Salesforce app that sends through your own Twilio account, so Twilio's per-segment billing applies directly. It counts segments with the rules in this guide wherever messages are written:
- On a record, the composer adds the length, segment count and encoding as soon as a message needs more than one segment, and names the characters that switched it to Unicode. The inbox reply box shows the count as you type.
- Preview shows the message with merge fields filled in, with its characters, segments and encoding.
- The template editor always shows a meter, how many characters are left in the segment, and a note when characters switch the text to Unicode.
- Outreach counts each message after merging and estimates the total number of segments on the review step, before anything is sent.
- Flow and Apex get the segment count back from Preview Text Message and Send Text Message.
Messages are limited to 1,600 characters, and ConnectSMS Setup shows usage as message and segment counts. See how the composer counts segments or how templates show the meter.
