Comparisons
9 min read·Published

Base64 vs Base32 vs Base58 vs Base85: which encoding, and why

Every binary-to-text encoding is a trade between size and the places the output has to survive. We ran the same 3,000 bytes through each: Base85 came out 25% larger, Base64 33%, Base58 37% and Base32 60%. Here is what each one buys you for that cost, the gotcha in each, and a decision guide for picking one.

By offlineutils.com

Binary-to-text encodings exist because a lot of systems that carry text will mangle arbitrary bytes. Every one of them pays for safety with size, and they differ in what they are optimising for: the smallest output, surviving a case-insensitive system, being read aloud, being pasted into source code. Picking the wrong one usually works — until the output passes through the one system it was not designed for.

The short answer

If the output has to…Use
go anywhere that handles text, with nothing special about itBase64
sit in a URL, filename, cookie or JWTBase64url
survive case-insensitive storage, DNS, or being read aloudBase32
be copied and typed by people, and is shortBase58
be as small as possible, or already is Ascii85 or Z85Base85

What each one costs, measured

We ran the same 3,000 random bytes through each encoder. Three thousand divides evenly by three, four and five, so padding does not distort the ratio:

EncodingCharacters outOverheadPacking
Base85 (Ascii85 or Z85)3,750+25.0%4 bytes → 5 characters
Base644,000+33.3%3 bytes → 4 characters
Base584,097+36.6%none — whole-input base conversion
Base324,800+60.0%5 bytes → 8 characters
Hex (Base16)6,000+100.0%1 byte → 2 characters

The Base58 figure matches the theory: each character carries log₂58 ≈ 5.858 bits, so a byte needs 8 ÷ 5.858 ≈ 1.366 characters. It is the only row without a fixed packing ratio, and that turns out to matter for more than size.

Base64: the default, for good reason

Base64 is the densest encoding that is safe almost everywhere text goes, which is why email attachments, data URIs, PEM certificates and most APIs use it. Its 64 characters are the upper and lower case letters, the digits, + and /, with = as padding.

The gotcha is those last two characters. In a query string + means a space, and / is a path separator. Any Base64 that ends up in a URL, a filename or a cookie should be the URL-safe variant from RFC 4648 §5, which swaps them for - and _ and usually drops the padding. JWTs use exactly that. More on the format in the Base64 glossary entry.

Base32: when case or humans get involved

Base32 uses only uppercase letters and the digits 2 to 7. That makes it 60% larger than the input instead of 33%, and it buys robustness: Base32 survives case-insensitive storage, DNS labels, case-folding filesystems and being dictated down a phone line. That is why TOTP authenticator secrets in otpauth:// URIs are Base32, and why DNSSEC uses it for NSEC3 hashes.

There are four alphabets in real use, and the same string decodes differently under each: the RFC 4648 standard alphabet; RFC 4648's extended-hex alphabet, whose digits-first ordering means encoded strings sort in the same order as the bytes; Douglas Crockford's, which drops I, L, O and U and, when decoding, reads I and L as 1 and O as 0 so a misread character still lands on the right value; and z-base-32.

The gotcha is the final character. It only carries part of a byte, and RFC 4648 says the unused low bits must be zero. Most decoders do not check: we tried MZXW6YTBOJ====== — one letter off the canonical MZXW6YTBOI======— and both Python's b32decode and Go's base32.StdEncoding returned foobar without complaint. So two different strings decode to the same bytes, which quietly breaks anything that compares encoded values — a cache key, a deduplication check, a signature over the encoded form. It also usually means the value was truncated or hand-edited. Our Base32 tool decodes such strings and says so, rather than silently discarding the bits.

Base58: built for people copying addresses

Base58 exists for one job: values that people copy, paste and occasionally retype, such as Bitcoin addresses and IPFS content identifiers. The rationale is spelled out in a comment in Bitcoin's source:

Don't want 0OIl characters that look the same in some fonts and could be used to create visually identical looking data. A string with non-alphanumeric characters is not as easily accepted as input. E-mail usually won't line-break if there's no punctuation to break at. Double-clicking selects the whole string as one word if it's all alphanumeric.

Every one of those is a reason Base64's +, / and = were unacceptable. It is also not a bit-packing scheme at all: 58 is not a power of two, so the whole input is treated as one big integer and repeatedly divided by 58. That has two consequences people miss.

First, leading zero bytes vanish unless handled separately. A zero byte at the front of a number adds nothing to its value, so a naive implementation drops it. Base58 encodes each one as a leading 1 instead — and getting that wrong corrupts exactly the values that matter, because a standard Bitcoin address begins with a 0x00 version byte. That is where the leading 1 in those addresses comes from.

Second, it is quadratic. Dividing an ever-shorter big integer by 58 once per output character is O(n²). We measured our own implementation:

InputBase58 encodeBase64 encode
256 bytes0.12 ms0.01 ms
1 KB1.9 ms0.03 ms
4 KB29 ms0.08 ms
16 KB461 ms0.36 ms

The absolute numbers depend on the machine; the shape does not. Each fourfold increase in input multiplies Base58's time by about sixteen, while Base64 grows in a straight line. At 16 KB Base58 is roughly a thousand times slower. That is irrelevant for a 25-byte address and disqualifying for a file — which is why nobody uses Base58 for files.

Bitcoin addresses add a further layer, Base58Check: a four-byte checksum from a double SHA-256, so a mistyped address fails verification instead of decoding to plausible garbage. Our Base58 tool verifies that checksum and names the version byte.

Base85: the smallest, and two dialects

Base85 turns four bytes into one 32-bit integer and writes it as five base-85 digits — 25% overhead, the best of the common encodings. You meet it in two forms. Ascii85, from Adobe, uses the 85 printable characters from ! to u, collapses an all-zero group to a single z, and appears inside PDF and PostScript streams. Z85, from ZeroMQ, uses a different 85 characters chosen to avoid quotes, backslashes and commas, so its output can be pasted into source code, JSON or a shell command without escaping.

The gotcha is arithmetic. Five base-85 digits can express values up to 85⁵, about 4.44 billion, but four bytes only reach 4,294,967,295. So some five-character groups — anything above s8W-! — are not valid at all. A decoder that ignores this wraps the value and returns four wrong bytes with no complaint. Z85 adds a second trap: it is defined only for whole four-byte words and carries no length field, so padding a ragged input is something you have to own rather than something the format records.

A decision guide

  • No special constraints? Base64. It is everywhere, it is fast, and every language has it built in.
  • Going into a URL, filename, cookie or token? Base64url.
  • Case might be changed, or a human will read it out? Base32 — Crockford's alphabet if people will type it back in.
  • Short identifier that people copy and paste? Base58, with a checksum if a typo would be expensive.
  • Every byte counts, or the data is already Ascii85 or Z85? Base85 — Z85 if it is going into source code.
  • Debugging, and you want to see the bytes? Hex. Twice the size, but every byte is two characters you can read directly; our number base converter helps with individual values.

One thing they all share

None of these is encryption. Each is a reversible change of alphabet that anyone can undo, so a secret encoded in Base64 is exactly as exposed as the secret itself. That also means it is safe to decode a token or a key to inspect it — as long as the decoder runs in your browser rather than on someone else's server. Every encoder linked above does; here is how to check that for yourself.

Frequently asked questions

Which binary-to-text encoding is the most compact?

Base85, at 25% overhead: four bytes become five characters. Base64 is next at 33%, Base58 at about 37%, Base32 at 60% and hex at 100%. We confirmed those figures by encoding the same 3,000 random bytes with each — the outputs were 3,750, 4,000, 4,097, 4,800 and 6,000 characters. Size is rarely the deciding factor, though; where the output has to survive usually is.

When should I use Base32 instead of Base64?

When the output has to survive something that ignores or changes letter case, or has to be read by a person. Base32 uses only uppercase letters and the digits 2 to 7, so it passes through case-insensitive systems, DNS labels and case-folding filesystems intact, and it can be dictated over the phone. That is why TOTP authenticator secrets use it. The price is size: 60% overhead against 33% for Base64.

Why does Bitcoin use Base58 instead of Base64?

Because addresses are copied and typed by people. Bitcoin's source code gives four reasons: Base58 drops 0, O, I and l so nothing can be misread; strings with punctuation are less readily accepted as input; email clients will not line-break a string with no punctuation; and double-clicking selects the whole string when it is all alphanumeric. Base64's + and / break all of the last three.

Is Base58 slower than Base64?

Yes, dramatically so for anything long, because it is a big-integer base conversion rather than a bit-packing scheme. In our measurements, quadrupling the input multiplied Base58's encode time by about sixteen — quadratic growth — and 16 KB took around 460 ms against well under a millisecond for Base64. That is irrelevant for a 25-byte address and disqualifying for a file, which is why Base58 is only ever used for short values.

What is the difference between Ascii85 and Z85?

The alphabet and the strictness. Ascii85, from Adobe, uses the printable ASCII characters from ! to u, including quotes and backslashes, which is fine inside a PDF stream and awkward inside a string literal. ZeroMQ's Z85 uses a different 85 characters chosen to avoid those, so its output can be pasted into source code, JSON or a shell command unescaped. Z85 also only accepts input that is a whole number of four-byte words.

What is Base64url and when do I need it?

It is Base64 with - and _ in place of + and /, defined in RFC 4648 section 5, usually with the = padding dropped. Standard Base64's + becomes a space in a query string and / is a path separator, so any Base64 that ends up in a URL, a filename or a cookie should be Base64url. JWTs use it for exactly that reason.

Tools mentioned in this post

Base64 Encoder / Decoder

Encode text to Base64 and decode it back, all offline.

Base32 Encoder / Decoder

Encode and decode Base32 — RFC 4648, extended hex, Crockford and z-base-32.

Base58 Encoder / Decoder

Encode, decode and checksum-verify Base58 and Base58Check.

Base85 Encoder / Decoder

Ascii85 and Z85 encoding — denser than Base64, with strict validation.

Number Base Converter

Convert numbers between binary, octal, decimal, hex and any base.

Image to Base64

Encode an image as a Base64 data URI (and decode back), in your browser.

Concepts in this post

Related posts