Every character on screen is stored as a number. ASCII assigned numbers to 128 characters in 1963; Unicode now numbers more than 150,000, and UTF-8 is the byte encoding that carries them across the web. This converter shows those numbers for any text — decimal or hex code points, UTF-8 bytes, binary, HTML entities or JavaScript escapes — and works in reverse, turning a list of codes back into text. A per-character table labels invisible characters such as tabs, non-breaking spaces and zero-width joiners, which makes it a handy debugging tool when a string “looks right” but does not match.
How to use the text to ASCII converter
- Choose Text → codes or Codes → text.
- To encode, type or paste the text, choose an output format and a separator. The full output appears in a copyable box, and the table lists each character’s decimal code, U+ notation, UTF-8 bytes, binary, HTML entity and Unicode block.
- To decode, paste the codes separated by spaces, commas or new lines. Automatic mode understands decimal numbers, 0x hex, U+ notation, 8-bit binary bytes, HTML entities (H) and \u escapes; you can also force a specific format.
Characters, code points and bytes
- A code point is a character’s number in Unicode, written U+0041 for “A” (65 decimal). For codes 0–127 it is identical to the ASCII code.
- UTF-8 stores each code point in 1 to 4 bytes:
| Code point range | Bytes | Examples |
|---|---|---|
| U+0000 – U+007F | 1 | ASCII: A, z, 7, ? |
| U+0080 – U+07FF | 2 | é, ñ, Greek, Cyrillic |
| U+0800 – U+FFFF | 3 | €, most CJK characters |
| U+10000 – U+10FFFF | 4 | Most emoji, rare scripts |
The x positions hold the code point’s bits from left to right.
Worked example
Convert Hi, café to codes.
H = 72, i = 105, comma = 44, space = 32, c = 99, a = 97, f = 102 — all ASCII, one byte each.
é is U+00E9 = 233. It is outside ASCII, so UTF-8 needs two bytes: 233 = 000 1110 1001 in 11 bits → 11000011 10101001 = C3 A9.
So the eight characters are nine UTF-8 bytes: 48 69 2C 20 63 61 66 C3 A9.
Decoding 72 101 108 108 111 gives Hello: 72 = H, 101 = e, 108 = l, 111 = o.
Handy ASCII ranges
| Codes (decimal) | Hex | Characters |
|---|---|---|
| 0–31 | 00–1F | Control codes (TAB = 9, LF = 10, CR = 13) |
| 32 | 20 | Space |
| 48–57 | 30–39 | Digits 0–9 |
| 65–90 | 41–5A | Uppercase A–Z |
| 97–122 | 61–7A | Lowercase a–z |
| 127 | 7F | DEL |
Two useful patterns fall out of the table: a digit’s value is its code minus 48, and upper- and lowercase letters differ by exactly 32 — a single bit (0x20), which is why flipping that bit changes case.
Debugging invisible characters
Text copied from word processors, PDFs and chat apps often carries characters you cannot see: a no-break space (U+00A0) instead of a normal space, a zero-width space (U+200B), a byte-order mark (U+FEFF) at the start, or curly quotes (U+2019) instead of straight ones (U+0027). They break passwords, CSV imports, code and search matches. Paste the suspicious text here and the table labels each one by name.
To encode text for URLs or as Base64, use the URL encoder or Base64 encoder. To count characters, words and lines, try the character counter, and to convert a single number between bases, the number base converter.
Frequently asked questions
What is the difference between ASCII and Unicode?
ASCII defines 128 characters (codes 0–127): English letters, digits, punctuation and control codes. Unicode includes all of ASCII with the same numbers and extends it to over 150,000 characters from every writing system, plus symbols and emoji. So 'A' is 65 in both.
What is the ASCII code for a space?
32 in decimal, 0x20 in hex, 00100000 in binary. The digits 0–9 are codes 48–57, uppercase A–Z are 65–90 and lowercase a–z are 97–122.
Why does an emoji show one code but four bytes?
A code point is the character's number in Unicode; UTF-8 bytes are how it is stored. Code points above U+FFFF, which include most emoji, take four bytes in UTF-8 and two 16-bit units in JavaScript strings, which is why their .length is 2.
How do I convert binary back to text?
Switch to Codes → text and paste 8-bit groups such as 01001000 01101001. They are read as UTF-8 bytes, so the example gives Hi. Accented letters and emoji span two to four bytes, and the decoder joins them automatically.