What is Punycode?
The DNS was designed for host names made of ASCII letters, digits and hyphens. To allow domain names in other scripts — internationalized domain names (IDNs) — each non-ASCII label is converted to an ASCII string with the Punycode algorithm (RFC 3492) and prefixed with xn--. Browsers show the Unicode form in the address bar, but DNS lookups, TLS certificates and server configs use the xn-- form.
| Unicode domain | Punycode (ASCII) |
|---|---|
| münchen.de | xn--mnchen-3ya.de |
| bücher.example | xn--bcher-kva.example |
| españa.com | xn--espaa-rta.com |
| café.fr | xn--caf-dma.fr |
| ελλάδα.gr | xn--hxakic4aa.gr |
| пример.рф | xn--e1afmkfd.xn--p1ai |
How the conversion works
A domain is split at the dots and each label is handled separately; pure-ASCII labels such as de or com stay as they are. Within a label, Punycode first copies all the ASCII characters (münchen → mnchen), adds a hyphen, and then appends a compact code (3ya) that records which non-ASCII characters to insert and where. Finally xn-- marks the label as encoded: xn--mnchen-3ya.
When you need the xn-- form
- DNS records (A, AAAA, CNAME, MX, TXT) at your registrar or DNS provider
- TLS/SSL certificate requests (the CSR's Common Name and Subject Alternative Names)
server_namein nginx orServerNamein Apache- Mail server settings, allow lists and application config files
- Older software and APIs that reject non-ASCII host names
Convert in code
| Language | To ASCII | To Unicode |
|---|---|---|
| JavaScript (browser) | new URL('https://münchen.de').hostname | — |
| Node.js | url.domainToASCII('münchen.de') | url.domainToUnicode(s) |
| Python | 'münchen.de'.encode('idna') | b'xn--mnchen-3ya.de'.decode('idna') |
| Go | idna.ToASCII(s) | idna.ToUnicode(s) |
Python's built-in idna codec implements the older IDNA2003 rules; the third-party idna package implements IDNA2008. Go's package is golang.org/x/net/idna.
Frequently asked questions
What happens to uppercase and full-width characters?
Domain names are case-insensitive, so each label is lowercased before conversion (MÜNCHEN.DE → xn--mnchen-3ya.de). Full-width letters and digits are normalized to their ASCII forms (Unicode NFKC) first.
Is the path of a URL converted too?
No. Only the host name is converted to Punycode. Non-ASCII characters in the path or query string are percent-encoded instead — for example ü becomes %C3%BC. Use the URL encoder for that part.
What is an IDN homograph attack?
Many characters in different scripts look identical — Cyrillic а and Latin a, for example. Attackers register look-alike domains to imitate real sites in phishing links. The Punycode form exposes the trick: аррӏе.com spelled with Cyrillic letters is really xn--80ak6aa92e.com. Modern browsers display the xn-- form for domains they consider confusable, and converting any suspicious link here shows whether it contains non-ASCII characters.
Why does straße.de give a different result elsewhere?
This tool, like Node.js's url.domainToASCII, follows IDNA2008 and keeps ß: straße.de → xn--strae-oqa.de. Older IDNA2003 implementations (including Python's built-in codec) map ß to ss and produce strasse.de.