IDN homograph attacks: how a domain can look exactly like another
A homograph attack registers a domain that renders identically to a real one by swapping in letters from another script — Cyrillic а for Latin a. In 2017 an all-Cyrillic xn--80ak6aa92e.com displayed as apple.com in Chrome, Firefox and Opera. Here is how the attack works, why browsers sometimes show you xn-- instead, and the defences that actually hold.
In April 2017 a developer named Xudong Zheng registered xn--80ak6aa92e.com and put up a page that said, in effect, “this is not Apple.” In Chrome 57, Firefox 52 and Opera, the address bar displayed it as apple.com — with a padlock. Nothing about the name looked wrong, because to a human reader nothing was. That is a homograph attack, and it is worth understanding properly, because the defences that work are not the ones most people reach for.
The short answer
A homograph attack registers a domain that renders identically to a real one by using characters from another script. Cyrillic а and Latin a are different characters with different code points, so аpple.com and apple.com are different domains that look the same. Browsers defend by showing you the raw xn-- encoding when a name looks suspicious, and the most reliable personal defence is a password manager, because it matches the real name rather than the rendering.
How the attack works
Domain names were originally restricted to ASCII letters, digits and the hyphen. Internationalised domain names let people register names in their own scripts — bücher.de, 日本語.jp, россия.рф — which DNS still carries as ASCII by encoding each label with Punycode and prefixing it with xn--. The browser decodes the label and shows you the readable version.
The problem is that Unicode contains many characters that look like Latin letters without being Latin letters. By Unicode's own confusables data, Cyrillic alone has domain-legal lookalikes for seventeen of the twenty-six: a, c, d, e, h, i, j, l, o, p, q, r, s, v, w, x and y. Greek, Armenian, Cherokee and others add more. Combine them carefully and you can spell a well-known brand name in characters that are, as far as DNS is concerned, completely unrelated to it.
A short history
- 2002. Evgeniy Gabrilovich and Alex Gontmakher describe the attack in “The Homograph Attack” in Communications of the ACM, demonstrating it by registering a
microsoft.comspelled with Russian с and о. - 2005. A Cyrillic-а spelling of
paypal.comcirculates, and ICANN issues a warning about homograph attacks. Browser vendors start restricting when they will display a domain in Unicode — at first by allowlisting TLDs, later by detecting labels that mix scripts. - 2017. Zheng's
xn--80ak6aa92e.combypasses those mixed-script checks entirely, for a reason worth understanding — below. Chrome fixes it in version 58.
Why the 2017 version got through
By 2017 every major browser refused to render a label that mixed Latin and Cyrillic, so аpple.com — one Cyrillic letter among four Latin ones — showed up as xn--pple-43d.com. The check worked. Zheng's insight was that it only fired on mixing.
His label was аррӏе: Cyrillic а, two Cyrillic р, Cyrillic palochka ӏ, and Cyrillic е. Every character came from one script, so there was nothing mixed to detect. Every character also happened to have a Latin twin, so the whole label read as “apple.” This is called a whole-script confusable, and it defeats any defence that only asks “are these letters from the same alphabet?”
| Label | Scripts | Mixed-script check | Reads as |
|---|---|---|---|
аpple | Cyrillic + Latin | Caught | apple |
аррӏе | Cyrillic only | Not caught | apple |
россия | Cyrillic only | Not caught | — (и and я have no Latin twin) |
The third row is the reason the problem is hard. A rule that rejected every all-Cyrillic label would break the internet for a few hundred million Russian, Ukrainian, Bulgarian and Serbian speakers. The fix has to distinguish a Russian word from an English word spelled in Cyrillic.
How browsers decide what to show you
Chrome's policy, which it documents publicly, runs a series of checks on each label separately and falls back to the xn-- form if any of them trips. Among them: Latin, Cyrillic and Greek may not be mixed in one label; ASCII Latin may only be mixed with Chinese, Japanese or Korean; and — the 2017 fix — a label made entirely of letters from a whole-script-confusable set is shown as punycode unless the top-level domain is in that same script, or is one known to host many such domains.
That last clause is what keeps россия.рф readable while sending аррӏе.com to punycode. Under .рф, an all-Cyrillic label is the expected case and there is no ASCII brand for it to impersonate. Under .com, the same letters spell a name someone could register in plain ASCII.
The useful question is not “is this name foreign?” It is “does this name have an ASCII twin, and is that twin a different domain?” If the answer to both is yes, the name is a spoof regardless of which scripts it uses.
The defences that hold up
Use a password manager
This is the strongest personal defence and it works for a reason that is easy to miss. A password manager fills credentials by exact domain match, against the real underlying name, not against how the name renders. On аррӏе.com it looks for a saved login for xn--80ak6aa92e.com, finds nothing, and fills nothing. Chrome's IDN documentation lists it alongside Safe Browsing as a protection that does not depend on the display policy at all: password managers “won't automatically fill a password into a domain that is not the exactly correct one.” If yours suddenly offers nothing on a site you use daily, stop and look at the address.
Check the encoded form when it matters
Before entering anything sensitive on a domain you reached through a link, look at its xn-- form. A name that should be plain ASCII and has an xn-- form at all contains non-ASCII characters, full stop. The certificate details in your browser show the encoded name, and a punycode converter will show it too — ours reports which ASCII domain a lookalike impersonates and names each impostor character, so xn--80ak6aa92e.com comes back as “looks exactly like apple.com — but it is not.” It runs entirely in your browser, so a domain lifted from a phishing email never leaves your machine; if you would rather confirm that than take it on trust, here is how to check a web tool is not uploading your data.
Turn off Unicode display entirely, if you can
Firefox exposes this directly: set network.IDN_show_punycode to true in about:config and every internationalised domain is shown in its encoded form. It eliminates the attack, and it makes legitimate non-ASCII domains unreadable. If you only visit ASCII sites, that is a good trade.
Do not trust the padlock
The padlock says the connection to this domain is encrypted. It says nothing about whether this domain is the one you meant. Certificates for homograph domains are issued exactly like any other, because a homograph domain is a perfectly legitimate domain that belongs to someone else.
If you build software that handles names
The same attack works anywhere two strings can render alike: usernames, package names, display names, email addresses. The standard tooling is Unicode Technical Standard #39, which defines a skeleton for any string — every character replaced by its prototype lookalike — so that two names that would be confused share a skeleton. Comparing skeletons rather than raw strings catches аррӏе colliding with apple.
Compare hosts only after parsing them out of the URL with a real parser — https://apple.com@аррӏе.com/ is a URL whose host is the part after the @, which naive string matching gets backwards. Our URL parser shows how a conforming parser splits one.
Two details trip people up. First, apply the full normalisation pipeline before comparing anything; a decomposed accent and a precomposed one are different strings that render identically, which is its own variant of the problem — we measured how often domain converters get that wrong in most punycode converters disagree with your browser. Second, decide what to do about whole-script confusables deliberately, because rejecting every Cyrillic name is not an answer for a global user base.
A related trap: same name, different site
Homographs are two different names that look the same. There is also the opposite problem: one typed name that two browsers resolve differently. Four characters — German ß, Greek final ς, and the invisible zero-width joiner and non-joiner — were handled differently by the old and new IDNA standards. Firefox and Safari moved to the new behaviour in 2016; Chrome kept the old one until Chrome 110 in 2023. For seven years, typing faß.de opened one site in Chrome and a different one in Firefox. The full story is in the IDN glossary entry.
Takeaways
- A domain can be visually identical to another and still be a different domain owned by someone else.
- Mixed-script checks catch
аpple.combut notаррӏе.com; the question that matters is whether a name has an ASCII twin. - A password manager is the most dependable defence, because it matches the real name, not the rendering.
- The
xn--form never lies. When it matters, look at it.
Frequently asked questions
What is an IDN homograph attack?
It is registering a domain name that renders identically, or nearly so, to a legitimate one by using characters from another writing system. Cyrillic а, е, о, р, с and х are indistinguishable from Latin a, e, o, p, c and x in most fonts, so аpple.com written with a Cyrillic а is an entirely different domain that looks like apple.com. The attack was described by Evgeniy Gabrilovich and Alex Gontmakher in Communications of the ACM in 2002.
Why does my browser show xn-- instead of the website's name?
Because it judged the Unicode form unsafe to display. Every internationalised domain is really stored as an ASCII xn-- label, and browsers decide label by label whether to show you the readable Unicode version or the raw encoding. When a label mixes scripts that should not be mixed, or consists entirely of letters that impersonate Latin ones, the browser shows the xn-- form so you can see the name is not what it appears to be.
How can I tell if a domain is fake when it looks identical?
You often cannot tell by looking, which is the point of the attack. Check the encoded form instead: paste the domain into a punycode converter, or look at the certificate details, which show the xn-- label. If a name that should be plain ASCII has an xn-- form at all, it contains non-ASCII characters. A tool that reports which ASCII domain the name impersonates — аррӏе.com reads as apple.com — answers the question directly.
Do browsers protect against homograph attacks?
Yes, through display policies rather than blocking. Chrome, Firefox and Safari each show the xn-- form for labels that look suspicious — mixed Latin and Cyrillic, for instance, or a label made entirely of Cyrillic letters that resemble Latin ones under a TLD that is not Cyrillic. Those policies have been tightened over time after real bypasses; the all-Cyrillic apple.com lookalike rendered as apple.com in Chrome 57 and was fixed in Chrome 58 in April 2017.
Can a password manager protect against homograph phishing?
Yes, and it is one of the strongest defences available. A password manager fills credentials by exact domain match, and it compares the real underlying name rather than how it renders. On аррӏе.com it finds no saved login for apple.com and fills nothing. If your password manager unexpectedly offers nothing on a site you use every day, treat that as a warning rather than a glitch.
How do I make Firefox always show punycode?
Open about:config, search for network.IDN_show_punycode, and set it to true. Firefox will then display every internationalised domain in its xn-- form. That removes the attack entirely at the cost of making legitimate non-ASCII domains unreadable, which is a reasonable trade if you only ever visit ASCII sites.
Can anyone register a homograph domain?
It depends on the top-level domain. Many registries restrict which characters may be used and which scripts may be combined in a single label, which blocks the crudest lookalikes. Many others do not, and the rules differ from one TLD to the next. That patchwork is exactly why browsers maintain their own display policies instead of relying on registries to have prevented the registration.
Tools mentioned in this post
Concepts in this post
IDN
An IDN (Internationalized Domain Name) is a domain name containing characters outside ASCII, which is carried through DNS as an ASCII xn-- form produced by the IDNA processing defined in UTS #46.
Punycode
Punycode is the encoding defined by RFC 3492 that represents a Unicode string using only the letters, digits and hyphen that DNS allows, producing the xn-- labels that carry internationalised domain names.
Related posts
Most punycode converters disagree with your browser. We measured it.
Punycode is only the last step of turning a Unicode domain into an xn-- label, and tools that stop there return a different domain than your address bar does. We tested a bare RFC 3492 implementation against the browser's own IDNA on seventeen realistic names: six came back wrong. Then we measured our own against ICU across 1,112,042 code points.
Is offlineutils.com safe? An honest, verifiable answer
Yes — offlineutils.com is safe to use. Every tool runs entirely inside your browser, so the files and text you work with are never uploaded to a server, and there are no accounts, no cookies, and no tracking scripts. Here's exactly why it's safe, what data is and isn't collected, and how to verify every claim yourself in under a minute.