Text to Binary Converter_

Type anything and watch it become bits. Every character is turned into its numeric code, that number is written in base 2, and the result appears on the right as you type β€” Hi becomes 01001000 01101001. Pick the separator your worksheet or decoder expects: spaces, commas, one continuous stream, one byte per line, or one line per line of your input.

Two counters sit under the output and they rarely agree, which is the most useful thing this page shows you: characters and bytes. hello πŸ‘‹ is seven characters and ten bytes. The 7-bit ASCII mode is there for classroom worksheets, and it refuses to encode anything outside the ASCII range instead of quietly mangling it. Nothing is uploaded β€” the conversion runs in your browser.

toolkit.codes/text-to-binary
Binary_Output
Type or paste text to see it as binaryCharacters 0Bytes 0Groups 0
UTF-8
Ready
100% LOCAL
Input
Any text you can type or paste β€” Latin letters, accents, CJK, emoji, punctuation, line breaks and all.
Output
One group of bits per byte, at 8-bit UTF-8 or 7-bit ASCII, separated by whatever you find readable β€” with the character count and the byte count kept apart, because they are rarely the same number.
Processing
Encoded in this tab. Each character is resolved to its code point first, then to bytes, which is the step most converters skip and most confusion comes from.
Limits
7-bit mode covers code points 0 to 127 and refuses anything above, by name, rather than silently mangling it. 8-bit UTF-8 encodes everything Unicode defines.
Cost
One character is one byte only inside ASCII. An accented letter is two, most CJK is three, an emoji is four β€” and a flag or a skin-toned emoji is several characters joined invisibly, so it can run to eleven.

Turning writing into ones and zeros

Three steps, and everyone skips the middle one

Going from writing to bits happens in three moves. First each character is looked up and replaced by a number β€” its code point. H is 72, i is 105, a space is 32. Second, that number is stored as one or more bytes according to an encoding, which for the modern web means UTF-8. Third, each byte is written out in base 2, padded to a fixed width so the groups line up. Most explanations collapse the middle step, which is why the next part surprises people: character to byte is not always one to one, and that is where the interesting behaviour lives.

The encoding decision: 7-bit ASCII or 8-bit UTF-8

The original ASCII standard numbered 128 characters, which needs only seven bits. Computers settled on eight-bit bytes anyway, so ASCII is normally written with a leading zero wasted at the front. Both forms are in use: textbooks and puzzle worksheets often strip that zero and give you seven bits per letter, while anything a computer actually stores uses eight. This page does either, but the 7-bit mode has a hard limit β€” it can only reach characters 0 to 127. Ask it for Γ© and it stops and tells you which character it cannot represent, rather than dropping a bit and handing you binary that decodes to the wrong letter.

Why the output is longer than you counted

UTF-8 spends one byte on ASCII and more on everything else: two for most accented Latin and Greek, three for the majority of CJK, four for emoji and rarer scripts. So a message that looks short can produce a lot of bits. Worse, the counters can disagree three ways: a waving hand is one character, two UTF-16 code units, and four bytes. A woman-technologist emoji is three characters β€” a woman, an invisible joiner, and a laptop β€” glued into one picture, and eleven bytes. Nothing is broken when the numbers diverge; that divergence is what encoding actually is, and the counters above spell it out rather than making you guess.

What the output is not

Binary is a way of writing a number down, not a way of hiding one. Anyone can paste this output into any decoder and read it back in a second β€” there is no key and no secret. It is not compression either: writing a byte as eight visible characters makes the text roughly eight times bigger, not smaller. And it is not a file format; saving these digits gives you a text file full of the letters 0 and 1, not the original bytes. If you need secrecy use encryption, if you need to move bytes through a text channel look at Base64, and if you need a file smaller, compress it.

Type the text, pick a width, read both counts

  1. 01Type or paste into the left pane, or use Upload to read a text file from your device β€” it is decoded in the browser and never sent anywhere. The binary appears on the right immediately.
  2. 02Choose 8-bit UTF-8 for anything real, or 7-bit ASCII for classroom worksheets. In 7-bit mode a character outside the ASCII range stops the conversion and gets named, so you always know what could not be encoded.
  3. 03Pick the separator your target expects. Spaces are the readable default; None gives one unbroken stream; Comma suits a spreadsheet; One byte per line makes a vertical column; Line per input line keeps a pasted paragraph in its original shape.
  4. 04Read the counters. Characters and Bytes match only for plain ASCII β€” when they differ, the text contains something multi-byte. Copy_Result copies the whole output, Download saves it as a .txt file.

Four things worth encoding, and what they cost

A worksheet answer key

A homework sheet uses 7 bits per letter. Switch to 7-bit ASCII and the groups match the textbook exactly.

7-bit
Hi
Output
1001000 1101001

Checking what an encoder should emit

A unit test asserts some expected bytes. Paste the source string and compare the bits before trusting the implementation.

8-bit
Γ©
Output
11000011 10101001 β€” 1 character, 2 bytes

Writing a message in binary

A card, a tattoo, a puzzle for someone. One byte per line makes a column that is easy to transcribe by hand.

Separator
One byte per line
Output
a vertical stack, one 8-bit group per row

The real cost of an emoji

A field allows a fixed number of bytes, not characters. The counters show what a single emoji actually spends.

Input
πŸ‘©β€πŸ’»
Output
3 characters Β· 5 UTF-16 units Β· 11 bytes

Characters and their codes

CharacterDecimal8-bit binary7-bit binary
space32001000000100000
!33001000010100001
048001100000110000
957001110010111001
?63001111110111111
A65010000011000001
H72010010001001000
Z90010110101011010
a97011000011100001
i105011010011101001
z122011110101111010
~126011111101111110

The two binary columns hold the same value β€” the 8-bit column just carries a leading zero that the 7-bit column drops. Add 32 to a capital to get its lowercase partner, which in bits means flipping the third digit from the left: A is 01000001 and a is 01100001.

What multi-byte characters actually produce

CharacterCode point(s)BytesFull 8-bit output
A (ASCII)U+0041101000001
Γ© (accented Latin)U+00E9211000011 10101001
δΈ­ (CJK)U+4E2D311100100 10111000 10101101
πŸ‘‹ (emoji)U+1F44B411110000 10011111 10010001 10001011
πŸ‘©β€πŸ’» (emoji sequence)U+1F469 U+200D U+1F4BB1111110000 10011111 10010001 10101001 + 11100010 10000000 10001101 + 11110000 10011111 10010010 10111011

Read the last row as three separate characters that a font draws as one picture: a woman, an invisible zero-width joiner, and a laptop. Delete it with a single backspace and you may remove only the laptop β€” which is exactly why byte budgets and character limits disagree so often.

Habits that keep an encoding honest

  • Match the separator to whatever will read the output. A decoder written for spaces will choke on commas even though the bits are identical.
  • Stay on 8-bit unless a worksheet explicitly asks for seven. Everything a computer stores is eight bits wide, and 7-bit output has to be padded before most tools will touch it.
  • When Characters and Bytes disagree, that is your signal that the text is not plain ASCII β€” useful before pasting into a system with a byte limit.
  • Encode a known string first β€” the letter A should give 01000001. If that is right, the rest of the settings are too.
  • For a puzzle or a card, One byte per line is far easier to copy by hand than a long spaced line.
  • Paste the output straight into the Binary Translator to prove the round trip before you rely on it β€” emoji included.

Five ways the bits come out right and mean nothing

A 7-bit answer pasted into an 8-bit decoder

The two widths look almost identical on screen and produce completely different results. Feed 7-bit groups to something expecting bytes and the misalignment cascades: each letter gets assembled from the tail of one group and the head of the next, so you get plausible-looking nonsense rather than an error. If a decode comes out garbled, count the digits in one group before assuming the data is corrupt.

The output is longer than the text suggests

One character is one byte only within ASCII. Accented letters take two, most CJK takes three, emoji take four. A 20-character sentence with three emoji is 29 bytes, so anything sized from the character count will come up short. The counters above are separate for exactly this reason.

Separators are decoration, not data

The spaces, commas and line breaks between groups exist so a human can read the output. They are not encoded and they carry no information β€” the same bits with different separators are the same bits. What matters is the group width both sides agree on. A line break that is part of your input text is different: that gets encoded, as 00001010.

An emoji is often several characters

Flags, skin tones and people-with-objects are sequences joined by invisible zero-width joiners. The woman-technologist above is three code points and eleven bytes even though it draws as one glyph. Any count based on what you see on screen will be wrong, and cutting such a string in the middle can leave a stray joiner behind.

Binary is not encryption

This is the most common misconception about the format. Converting a message to ones and zeros hides nothing β€” every tool on this site and thousands of others will read it straight back, no key required. It is a change of notation, exactly like writing twelve as 1100. If you need something to stay private, encrypt it; binary makes text longer and no more secret.

Widths, code points, and what the encoder guarantees

File handling
Uploads are read inside the page with the browser File API and are never transmitted; Download writes out what is already in the tab.
Encoding
8-bit mode runs the text through UTF-8 (TextEncoder) and writes each byte as eight digits, zero-padded
7-bit mode
ASCII only, one 7-digit group per character; the first character above code point 127 stops the conversion and is named with its code point and position, never truncated or substituted
Separators
Space, none, comma, one byte per line, or one output line per input line β€” a display choice that never alters the bits
Counts
Characters are Unicode code points, UTF-16 units appear only when surrogate pairs are present, and bytes are what UTF-8 actually produces
Round trip
Every output decodes back to the original text through the Binary Translator; asserted in the test suite across all separators and both widths
Large inputs
Binary is eight times the size of its input, so the display stops at 300k characters and says so. Copy and Download are unaffected and always carry the complete result.
Network
None from tool code. A test sweep calls every function this page uses with fetch and XMLHttpRequest replaced by stubs that throw, so a stray request fails the build instead of shipping. Disconnect from the network and the page still works.

Questions about encoding text as binary

How do I convert text to binary?

Look up each character's numeric code, then write that number in base 2 and pad it to a fixed width. The letter H is 72, and 72 in base 2 is 1001000, which padded to a byte is 01001000. Typing into the box above does this for whole paragraphs at once, including characters that need more than one byte.

What is the difference between 7-bit and 8-bit binary?

They are the same numbers written with different amounts of padding. ASCII only defined 128 characters, which fits in seven bits, but computers work in eight-bit bytes so the eighth is normally a leading zero. Worksheets often use seven; real storage uses eight. Only 8-bit can represent anything beyond ASCII.

Why is my binary output longer than my text?

Because UTF-8 spends more than one byte on anything outside ASCII: two for accented Latin letters, three for most Chinese, Japanese and Korean characters, four for emoji. Compare the Characters and Bytes counters above β€” when they differ, that gap is what made the output longer.

How do I convert binary back to text?

Split it into groups of the same width it was written with, read each group as a number, and look the number up. Paste the output into the Binary Translator and it will do exactly that, telling you if the grouping does not divide evenly instead of guessing.

Is text written in binary secure?

No. Binary is a notation, not a cipher β€” there is no key, and any decoder reads it instantly. It also makes the text about eight times longer. Use encryption for privacy; use this for learning, debugging and puzzles.

How do I write my name in binary?

Type it above and read the result. For a card or a tattoo, choose One byte per line so each letter sits on its own row, and check the first letter against the character table before committing to anything permanent.