Engineering
11 min read·Published

Most punycode converters disagree with your browser. We measured it.

Punycode is only the last step of turning a Unicode domain into an xn-- label, and tools that stop there return a different domain than your address bar does. We tested a bare RFC 3492 implementation against the browser's own IDNA on seventeen realistic names: six came back wrong. Then we measured our own against ICU across 1,112,042 code points.

By offlineutils.com

If you paste a Unicode domain into an online punycode converter and into your browser's address bar, you can get two different answers. Not an error message — a different, well-formed xn-- label that points at a domain someone else could register. We measured how often that happens, and then measured our own implementation against the library browsers actually use.

The short answer

Punycode is only the last step of turning a Unicode domain into its ASCII form. Browsers run a preprocessing pass first, defined by Unicode Technical Standard #46, and a tool that calls a bare RFC 3492 library skips all of it. On seventeen realistic domain names, a bare implementation disagreed with the browser on six.

What the preprocessing pass does

UTS #46 specifies four steps: map every character, normalise the result to NFC, split it into labels, then validate and encode each label. The mapping step folds case, collapses compatibility characters to their plain equivalents, and discards a handful of characters entirely. The normalisation step is the one that causes the most damage when skipped.

“café” can be stored two ways: four code points ending in a precomposed é (U+00E9), or five ending in a plain e followed by a combining acute accent (U+0301). They render identically, and you cannot tell them apart by looking. Both turn up in real text: HFS+, the Mac file system until 2017, stored every filename decomposed, and its successor APFS preserves whatever form a name was created in — so decomposed names copied off older Macs and external drives are still common. Without the mandatory NFC pass, those two spellings encode to two different domains.

The measurement

We ran seventeen names through a bare RFC 3492 implementation and through ICU — the library browsers use — reachable from JavaScript as new URL('http://' + host + '/').hostname. Six disagreed:

InputBare punycodeBrowser (ICU)
cafe + combining acutexn--cafe-yvc.frxn--caf-dma.fr
example.com (fullwidth)xn--mi7chab1aes7c.comexample.com
ⓐⓑ.com (circled letters)xn--vuhc.comab.com
①.com (circled digit)xn--orh.com1.com
file.com (fi ligature)xn--le-1b1n.comfile.com
℀.comxn--z1g.comrejected

Every one of those bare-punycode answers is a syntactically valid xn-- label. Nothing about the output looks wrong. That is what makes the failure mode expensive: you copy the label into a registrar form or a TLS certificate request and it is accepted, because it is a perfectly good domain name. Just not yours.

Why the fullwidth row is a security problem

Look at row two again. Fullwidth example.com maps down to plain example.com. Now imagine a service that accepts a hostname, punycodes it without the mapping pass, and checks the result against an allowlist. The check sees xn--mi7chab1aes7c.com, which is not on the list, so the name passes as “some other domain” — or fails closed, depending on the logic. The browser, or any HTTP client using a conforming URL parser, resolves it to example.com.

Any time a validator and a fetcher disagree about what a hostname means, you have a request-forgery primitive. The mapping pass is not a cosmetic detail; it is the thing that makes both halves agree.

Testing our own against ICU

Having been rude about other people's implementations, we owed a number for ours. ICU is a usable oracle here, so we pushed every Unicode code point through both and compared. Each code point was wrapped in a label — a<char>b.test — encoded by both, and the results diffed.

The result: 1,112,042 code points compared, 99.9862% agreement, 154 differences. The path there is more informative than the final number:

StageDifferences
Mapping, NFC and validity only2,973
After adding CheckBidi and CheckJoiners183
After adding the URL Standard's forbidden characters154

The first drop is the two contextual rules almost no converter implements. CheckBidi requires a label containing right-to-left text to be consistently ordered, which is why an Arabic or Hebrew label with a stray Latin letter is invalid even though both characters are fine individually. CheckJoiners allows the invisible joiner characters only where a script needs them — a zero-width joiner directly after a virama, a zero-width non-joiner after a virama or between two Arabic joining letters — because anywhere else they do nothing visible and exist only to make two different names look identical.

The bug the sweep caught

The second drop, 183 to 154, came from a bug we would not have found by clicking around. UTS #46's table marks the space character valid. Separately, the URL Standard forbids it in a host. If you implement only the first specification, you are correct by that document and still wrong in practice — and it is not hypothetical, because U+FE70 ARABIC FATHATAN ISOLATED FORM maps to a space followed by a mark. We were returning labels with a literal space inside them.

Worse, the first version of our sweep hid this. It skipped any comparison where our own output contained URL syntax, on the theory that such cases could not be compared fairly. That reasoning quietly excluded exactly the code points we were getting wrong. The lesson generalises: never skip a test case based on your own output. Producing malformed output is a finding, not a reason to look away.

What the remaining 154 are

All 154 are cases where we reject a name ICU accepts, and all of them involve characters Unicode assigned in version 14.0 or later — U+061D, the Arabic Extended-B block, the Garay script. We did not want to assume that was a version skew, so we checked it two ways.

First, behaviour: ICU accepts a؝b and rejects ؝א. If it treated ؝ as an Arabic letter, those results would be the other way round — so ICU is not classifying it as right-to-left at all. Second, distribution: bucketing every divergence by the Unicode version that assigned the character put all 154 in the 14.0-or-later bucket and none anywhere else. A real bug in our rule would not correlate that cleanly with a character's age. ICU's bidi property data is simply older than the Unicode 17 table we generate from.

The direction matters more than the count. Nothing remains in the opposite direction: there is no name we accept that ICU refuses. That asymmetry is the one worth protecting, because telling someone a domain works when their browser will reject it is the expensive kind of wrong.

How the tables are built

Shipping UTS #46 to a browser means shipping its data. The raw mapping table is 9,262 rows, and the naive encoding of it is around 50 kB. Two observations cut that to 23.5 kB including the bidi and joining-type data:

  • Status runs tile the whole code space, so merging adjacent runs of the same status collapses 9,262 rows to 2,921 runs.
  • The mapping column is essentially NFKC plus case folding, which JavaScript can already compute. 5,956 of 6,381 mapped code points are reproducible that way, so only the 425 divergences ship explicitly.

Deriving data rather than shipping it is normally a trade of accuracy for size, and here it is not, for two reasons. The generator verifies the derivation against the official file for all 1,114,112 code pointsand refuses to write anything out on a single mismatch, so every divergence is captured by construction. And leaning on the platform's Unicode data costs no extra risk, because UTS #46 mandates an NFC pass anyway — normalize()is already load-bearing — while Unicode's normalization stability policy guarantees that a character's decomposition never changes once assigned.

What to do with this

If you convert domain names in code, use a UTS #46 implementation rather than a punycode library — and check that the one you have really is one, because two of the obvious choices are not:

  • Python: the idna package, called as idna.encode(name, uts46=True). Without that flag it skips the mapping and rejects anything with a capital letter, and the standard library's encodings.idna codec implements only IDNA2003.
  • Java: ICU4J's IDNA.getUTS46Instance(…). The built-in java.net.IDN implements RFC 3490 — IDNA2003, with Nameprep — so it gives the old answer for names like straße.de.
  • JavaScript: tr46, which is the reference implementation behind the WHATWG URL parser, or simply new URL() when you only need the ASCII form.

Reaching for a punycode library directly opts you out of the mapping and normalisation passes, and so does a default that silently means IDNA2003. Either way the output still looks like a perfectly good domain.

One related mix-up worth clearing up: punycode applies only to the host. The path and query string of the same URL use percent-encoding of UTF-8 bytes instead — a completely different mechanism, which is why münchen.de/straße becomes xn--mnchen-3ya.de/stra%C3%9Fe. Our URL parser splits a URL into those parts so you can see which rule applies where, and the URL encoder handles the percent-encoding half.

If you are checking a name by hand, our Punycode / IDN convertershows both forms at once, accounts for every character it changed, and cross-checks its own answer against your browser's IDNA so you can see if they ever disagree. The concepts have plain-English write-ups under Punycode and IDN.

The security consequences of two spellings that render identically get their own post: IDN homograph attacks, and why your browser sometimes shows you xn--.

Frequently asked questions

Why does my punycode converter give a different result than my browser?

Because punycode is only the final step. Browsers first run the UTS #46 preprocessing pass — case folding, compatibility mapping, and NFC normalisation — and a tool that calls a bare RFC 3492 library skips all of it. The clearest case is normalisation: café can be stored as four code points with a precomposed é or five with a combining accent. They look identical, and without the required NFC pass they encode to xn--caf-dma and xn--cafe-yvc respectively — two different domains.

Is punycode the same thing as IDNA?

No. IDNA is the whole process of converting an internationalised domain name to the ASCII form DNS carries; punycode is one step inside it. The full pipeline defined by UTS #46 is: map every character, normalise to NFC, split into labels, validate each label, then punycode-encode the ones that still need it. Bugs cluster in the steps around punycode rather than in punycode itself, which is well specified and easy to get right.

How can I tell if a punycode converter is correct?

Give it cafe with a combining acute accent (U+0065 U+0301) rather than a precomposed é. A correct implementation returns xn--caf-dma, the same as your browser; one that skips normalisation returns xn--cafe-yvc. A second test is fullwidth example.com, which should map all the way down to plain example.com. A third is straße.de, which should give xn--strae-oqa.de under the rules browsers use today.

What does the browser's address bar actually use to convert domains?

ICU, the International Components for Unicode library, applying UTS #46 with the options the WHATWG URL Standard specifies: non-transitional processing, CheckBidi and CheckJoiners on, STD3 ASCII rules off. You can reach the same code path from JavaScript with new URL('http://' + host + '/').hostname, which makes it a usable reference implementation for testing your own.

Can a Unicode character be valid in UTS #46 but still be rejected in a domain?

Yes, and this is a gap worth knowing about. UTS #46's own table marks the space character and the C0 controls as valid. The URL Standard separately forbids them in a host. The two specifications together are what a correct converter has to apply — U+FE70 maps to a space followed by a mark, so a tool implementing only UTS #46 will hand back a label with a literal space inside it.

How many differences are acceptable between an IDNA implementation and ICU?

Direction matters more than count. Differences where you reject a name ICU accepts are usually a Unicode version skew and are survivable. Differences where you accept a name ICU rejects are the dangerous kind, because the tool tells someone a domain works when their browser will refuse to load it. In our own sweep of 1,112,042 code points we ended with 154 differences, all in the first category and all traceable to ICU's bidi data being older than Unicode 17.

Tools mentioned in this post

Punycode / IDN Converter

Convert between Unicode domain names and xn-- Punycode, with full UTS #46 and homograph checks.

URL Parser

Break any URL into its protocol, host, path and query parameters.

URL Encoder / Decoder

Percent-encode and decode URL components instantly.

Concepts in this post

Related posts