Python's str.lower() Can Be a Security Vulnerability
When str.lower() is a security vulnerability in Python – Seth Larson

Seth Larson explains how Python's str.lower() can introduce security vulnerabilities in IDNA 2003 implementations. The issue arises because str.lower() uses the Unicode version shipped with the interpreter, while the StringPrep specification (RFC 3454) requires Unicode 3.2.0 case-folding rules. This discrepancy can lead to incorrect domain name processing, as demonstrated with Cherokee characters. The fix involves adding exceptions to mimic Unicode 3.2.0 behavior. The vulnerability is tracked as CVE-2026-17084.
The str.lower() call in this function is a vulnerability!
- echoangle
> This is why calling str.lower() represents a difference in the implementation and the specification, and therefore a vulnerability:
I wish there was some explanation how this is a vulnerability and not just a bug generating erroneous data.
Vulnerability for me sounds like there’s a reasonable way to create an exploit from the bug, and I don’t see one here as someone who’s not very familiar with the topic.
- tialaramex
This idiocy is a big part of why it was so important to get Python people working on TLS implementations to understand that the defined mechanism for SANs (no the "alternative" in Subject Alternative Name doesn't mean in the sense of more than one, X.509 is originally for the X.500 system and the Internet repurposed X.509 so these are alternative names from the Internet) says that these are DNS names, they specifically are not to be understood as some sort of human readable text, and thus "decoding" them to Unicode is definitely nonsense even though Python really wanted to do that and I think used to do it or at least proposed to.
The rule for how SAN DnsNames match againt like names, from the DNS is very, very simple so that you don't screw it up. You handle a single wildcard (ASCII * code 42 matches any single DNS label) and beyond that it's literally byte comparison. You don't care what these bytes mean, either the bytes are all identical or that's not a match and we're done.
- ummonk
> The fix was to create new exceptions so that str.lower() would behave as if it was using Unicode 3.2.0 for only particular function. So, we go through each Unicode codepoint and record when the behavior of str.lower() is different when comparing the Unicode version shipped with Python and Unicode 3.2.0
This sounds like a really hacky solution compared to implementing a separate frozen Unicode 3.2.0 lower.
- jooon
Reminds me of an old security incident at Spotify https://engineering.atspotify.com/2013/06/creative-usernames
- ike_sh
Hit this with the Kelvin sign once. Took embarrassingly long to track down.