When str.lower() is a security vulnerability in Python

SethMLarson1 pts0 comments

When str.lower() is a security vulnerability in Python — Seth Larson

Blog :<br>About :<br>RSS :<br>Blogroll

When str.lower() is a security vulnerability in Python

Seth Larson @ 2026-08-18

Some internet standards only support ASCII characters, but the world uses<br>much more than the Latin alphabet. Thus, a mapping from Unicode to<br>ASCII for use in domain names is required.

NamePrep was part of that solution,<br>defined in RFC 3491 as a profile of StringPrep,<br>and is crucially a component of Internationalizing Domain Names in Applications<br>(IDNA), also known as “IDNA 2003”. The StringPrep algorithm is defined in<br>RFC 3454. IDNA 2003<br>has been obsoleted by IDNA 2008 defined in RFC<br>5890,<br>5891,<br>5892, and<br>5893.

Python supports IDNA 2003 through the idna codec (str.encode('idna')) and<br>IDNA 2008 is supported by the idna package on the Python package Index.<br>Python's implementation of StringPrep is implemented in the stringprep module<br>in the standard library. In general, you should be using the idna package (IDNA 2008)<br>and not .encode("idna") (IDNA 2003), but sometimes you<br>do need the older behavior.

StringPrep defines the “case folding” step (case folding is approximately “how to lowercase/uppercase a codepoint”) in Section 3.2, enabling case-insensitive<br>comparisons of strings, by mapping all characters through mapping tables B.2 and B.3.<br>B.2 is effectively str.lower(), lowercasing all characters according to<br>Unicode rules and B.3 contains the exceptions. The Python code implementing<br>this (and assuming B.3 table is captured correctly) is the following code below:

def map_table_b3(code):<br>r = b3_exceptions.get(ord(code))<br>if r is not None: return r<br>return code.lower()

And that might seem fine... and the title probably gave it away already.<br>The str.lower() call in this function is a vulnerability!

Why? Because str uses whatever Unicode data that the particular Python<br>interpreter is shipped with, you can figure out what Unicode version your<br>Python interpreter uses by accessing unicodedata.unidata_version:

>>> import unicodedata<br>>>> unicodedata.unidata_version<br>'17.0.0'

There's also a database of Unicode 3.2.0 data available on every version<br>of Python (unicodedata.ucd_3_2_0) specifically for the StringPrep and IDNA algorithms:

$ grep -I "ucd_3_2_0" -R Lib/<br>Lib/stringprep.py:from unicodedata import ucd_3_2_0 as unicodedata<br>Lib/encodings/idna.py:from unicodedata import ucd_3_2_0 as unicodedata

This is important! StringPrep depends on this specific version of Unicode<br>to operate consistently, the B.2 and B.3 tables in RFC 3454 are essentially<br>Unicode 3.2.0 case-folding rules encoded into<br>a table. So we need to use Unicode 3.2.0 case-folding rules, not newer<br>Unicode case-folding rules. This is why calling str.lower() represents<br>a difference in the implementation and the specification,<br>and therefore a vulnerability:

# RFC 3454 compliant value ('Ꭰ' is U+13A0)<br>>>> "ᎠᎠ".encode("idna")<br>'xn--58da'

# Value if using Unicode 17.0.0 case-folding<br>>>> "ᎠᎠ".encode("idna")<br>'xn--kz9aa'

The fix was to create new exceptions so that str.lower() would behave<br>as if it was using Unicode 3.2.0 for only particular function. So, we<br>go through each Unicode codepoint and record when the behavior of<br>str.lower() is different when comparing the Unicode version shipped with Python and Unicode 3.2.0.<br>And that's all, now IDNA 2003 is consistent with the specification.

Thanks to Bitshift for reporting the vulnerability, Stan Ulbrych for co-developing<br>the remediation, and Marc-Andre Lemburg and Petr Viktorin for reviewing the<br>remediation. See CVE-2026-17084 for more details.

My work as the Security Developer-in-Residence at the Python Software Foundation is sponsored<br>by Alpha-Omega. Thanks to Alpha-Omega for supporting<br>security in the Python ecosystem.

Wow, you made it to the end!

Share your thoughts with me on Mastodon, email, or Bluesky.

Browse this blog’s archive of 193 entries.

Check out this list of cool stuff I found on the internet.

Follow this blog on RSS or the email newsletter.

Go outside (best option)

idna unicode python lower stringprep unicodedata

Related Articles