Hex to Text Converter
To convert hex to text, split the hex into pairs of digits, turn each pair into a byte, and decode the bytes as ASCII or UTF-8. For example, 48 65 6C 6C 6F decodes to Hello. In UTF-8, characters outside ASCII use two to four bytes: D0 9F is П and F0 9F 98 80 is 😀.
5 bytes, 5 characters.
Text to hexCharacter breakdown
| Char | Code point | Bytes | Count |
|---|---|---|---|
| H | U+0048 | 48 | 1 |
| e | U+0065 | 65 | 1 |
| l | U+006C | 6C | 1 |
| l | U+006C | 6C | 1 |
| o | U+006F | 6F | 1 |
How to convert hex to ASCII or UTF-8 text
Hex text is almost always a byte dump: something that was text once, written out two digits per byte so it survives logs, URLs or a protocol trace. Decoding it means turning the pairs back into bytes and reading the bytes with the right character encoding.
- Remove separators such as spaces, commas, 0x or \x so only hex digits remain.
- Split the digits into pairs; each pair is one byte from 00 to FF.
- For ASCII, look each byte up in the 0–127 table.
- For UTF-8, read a lead byte to learn how many bytes the character uses, then combine them into one code point.
If you do not know the encoding, try UTF-8 first. It is what the web, JSON and most modern systems use, and for pure ASCII input the two give the same result. Choose ASCII when you want any byte above 7F flagged instead of interpreted: RFC 20 defines it as a 7-bit code, so nothing past 7F belongs to it.
Worked example: one, two and four bytes
ASCII letters take one byte, the Cyrillic Pe two, the emoji four
48 69 20 D0 9F 20 F0 9F 98 80 → Hi П 😀
| Char | Code point | Bytes | Count |
|---|---|---|---|
| H | U+0048 | 48 | 1 |
| i | U+0069 | 69 | 1 |
| space | U+0020 | 20 | 1 |
| П | U+041F | D0 9F | 2 |
| space | U+0020 | 20 | 1 |
| 😀 | U+1F600 | F0 9F 98 80 | 4 |
Reading UTF-8 lead bytes
UTF-8 is self-describing. The high bits of the first byte say how long the sequence is, and every following byte starts with 10, as RFC 3629, the definition of UTF-8, lays out. That is how a decoder, or you with this table, can tell that D0 opens a two-byte letter and F0 a four-byte one.
It also means you can start reading in the middle of a stream: skip any byte from 80 to BF and the next byte is the start of a character.
| First byte | Bit pattern | Bytes | Used for |
|---|---|---|---|
| 00–7F | 0xxxxxxx | 1 | ASCII itself |
| C2–DF | 110xxxxx | 2 | Latin, Greek, Cyrillic, Hebrew, Arabic |
| E0–EF | 1110xxxx | 3 | most other scripts, CJK, symbols |
| F0–F4 | 11110xxx | 4 | emoji and rare scripts |
| 80–BF | 10xxxxxx | — | continuation, never first |
When the bytes are not valid UTF-8
Real dumps contain mistakes: a string cut in half, a Latin-1 file read as UTF-8, binary data that was never text. In C3 28 41, C3 promises a second byte between 80 and BF, but 28 is an opening parenthesis. The decoder replaces the broken sequence with U+FFFD, the replacement character, and carries on.
The tool names the byte offset of each bad sequence, counting from 0, so you can find it in the original dump. A run of replacement characters usually means the data is in another encoding or is not text at all.
Not valid UTF-8 at byte 0 (C3); shown as �.
C3 28 41 → �(A
| Char | Code point | Bytes | Count |
|---|---|---|---|
| � | invalid | C3 | 1 |
| ( | U+0028 | 28 | 1 |
| A | U+0041 | 41 | 1 |
Separators and prefixes you can paste
Each tool writes bytes its own way. You do not need to clean the input first: spaces, line breaks, commas and the markers 0x, \x and % in front of a byte are skipped. What remains must be an even number of hex digits.
| Input | Typical source |
|---|---|
| 48656C6C6F | a plain hex dump |
| 48 65 6C 6C 6F | xxd, hexdump, Wireshark |
| 0x48, 0x65, 0x6C | a C or Rust byte array |
| \x48\x65\x6C | a Python or C string escape |
| %48%65%6C | URL percent-encoding |
Questions people ask
How do you convert hex to ASCII?
Take the hex two digits at a time and look up each pair as a character code: 48 is H, 65 is e, 6C is l and 6F is o, so 48656C6C6F reads Hello. ASCII covers 00 to 7F; bytes above 7F need an encoding such as UTF-8.
Why does the decoded text show a � character?
The replacement character U+FFFD marks bytes that are not valid UTF-8. For example, C3 must be followed by a byte from 80 to BF, so C3 28 decodes to � followed by (. The tool reports the byte offset of each bad sequence.
Do hex bytes need spaces between them?
No. 48656C6C6F and 48 65 6C 6C 6F decode the same way. Spaces, commas, line breaks and a 0x, \x or % before each byte are ignored, but the total number of hex digits must be even.
What is D0 9F in text?
П, the Cyrillic capital letter Pe (U+041F). In UTF-8, D0 is a lead byte that starts a two-byte sequence, and 9F carries the remaining six bits of the code point.
Related conversions
Last updated