Skip to content
Hexadecimal Converter

Text to Hex Converter

To convert text to hex, encode each character as bytes and write each byte as two hex digits. ASCII characters take one byte, so hello is 68 65 6C 6C 6F. In UTF-8, Cyrillic letters take two bytes each (Привет is D0 9F D1 80 D0 B8 D0 B2 D0 B5 D1 82) and most emoji take four.

Encoding
Hex format

5 bytes, 5 characters.

Hex to text
Character breakdown
Each character and its bytes
CharCode pointBytesCount
hU+0068681
eU+0065651
lU+006C6C1
lU+006C6C1
oU+006F6F1

How to turn text into hex bytes

Computers store characters as numbers, and hex is the usual way to look at those numbers directly. You need hex text when you build a test payload, compare what two systems actually sent, write a byte array into code or check that a string carries no hidden characters.

  1. Pick the encoding: ASCII for plain English text, UTF-8 for anything else.
  2. Find each character’s code point, its number in Unicode.
  3. Encode the code point as one to four bytes following the UTF-8 rules.
  4. Write every byte as two hex digits, with the separator your target expects.

A space is a byte too, 20, and so is a line break, 0A. The character list under the tool shows them by name so nothing invisible slips into the output unnoticed.

Worked example: Привет

Six Cyrillic letters, two bytes each

Привет → D0 9F D1 80 D0 B8 D0 B2 D0 B5 D1 82

Each character and its bytes
CharCode pointBytesCount
ПU+041FD0 9F2
рU+0440D1 802
иU+0438D0 B82
вU+0432D0 B22
еU+0435D0 B52
тU+0442D1 822

Choosing an output format

The bytes never change; only the punctuation around them does. Pick the format your destination parses, so the result can be pasted without editing. Lowercase digits are common in hashes and Unix tools; uppercase reads better in tables and documentation.

“Hi!” in each format
FormatOutputWhere it fits
Spaced48 69 21reading, hex editors, bug reports
Plain486921hashes, keys, compact storage
0x each0x48 0x69 0x21assembly listings, documentation
Escaped\x48\x69\x21string literals in C, Python, shells
Comma list0x48, 0x69, 0x21byte arrays in C, Rust, Java

Why an emoji takes four bytes

UTF-8 spends bytes according to the size of the code point. Up to U+007F fits in one byte, up to U+07FF in two, up to U+FFFF in three. Emoji were added above that line, so they need the four-byte form, which has room for 21 bits. The ranges and bit layout are set out in RFC 3629.

Encoding 😀 by hand: write U+1F600 in 21 bits, split them 3 + 6 + 6 + 6, and put 11110 in front of the first part and 10 in front of each of the others. Some emoji are several code points joined together, a family or a flag, so they can run to a dozen bytes or more.

U+1F600 packed into UTF-8
StageBitsHex
Code point1 1111 0110 0000 0000U+1F600
Split 3 + 6 + 6 + 6000 011111 011000 000000
With UTF-8 markers11110000 10011111 10011000 10000000F0 9F 98 80

Questions people ask

How do you say hello in hexadecimal?

In ASCII or UTF-8, hello is 68 65 6C 6C 6F: h is 68, e is 65, l is 6C and o is 6F. With a capital H, Hello starts with 48 instead.

Why is an emoji four bytes in UTF-8?

UTF-8 needs four bytes for code points above U+FFFF, and most emoji live there. 😀 is U+1F600, which UTF-8 writes as F0 9F 98 80. In UTF-16 the same emoji takes two 16-bit units, a surrogate pair.

Is ASCII hex the same as UTF-8 hex?

For the 128 ASCII characters, yes: UTF-8 was designed so that bytes 00 to 7F mean the same thing. They differ only for other characters, which ASCII cannot encode and UTF-8 writes as two to four bytes.

How many bytes does a Cyrillic letter take in UTF-8?

Two. Cyrillic letters sit between U+0400 and U+04FF, a range UTF-8 encodes in two bytes, so the six letters of Привет become twelve bytes.

Last updated