UTF-8000 Extends UTF-8 to Encode Any Codepoint, No Matter How Large

UTF-8000: Unlimited UTF-8

UTF-8000 Extends UTF-8 to Encode Any Codepoint, No Matter How Large

A fun proposal called UTF-8000 stretches UTF-8 to support arbitrarily large codepoints by striping start bits across continuation bytes. It keeps ASCII as a subset, preserves self-synchronization and self-punctuation, and introduces no special cases beyond those inherited from UTF-8. A reference implementation is available via pipx install UTF-8000. Not affiliated with the Unicode Consortium.

The main contribution of UTF-8000's specification is clarity on splitting the highest bits of the first byte of UTF-8 code units into self-synchronization bits and start bits, and then making it clear how to stripe the start bits across the continuation bytes if needed, to achieve arbitrarily large code units.
  1. 2shortplanks

    On a practical matter, it seems like a bad idea to have codepoints that can take up to an arbitrary number of bytes - this just screams buffer overflow problems.

    So in practicality, you’re going to want an arbitrary limit on this (the article suggests as much). But if you place a limit on it then you’ve got one implementation of the standard that can decode certain characters and another that can’t. Better to have one standard that puts a hard limit on the number of bytes and another standard that uses more bytes and so on.

  2. sph

    > UTF-8000 is in no way endorsed by or representative of the Unicode Consortium.

    Not until they decide to expand the emoji range, allocate space for all past and future fictional languages, as well as birdsong and dog barks.

    Someone at the consortium is rubbing their hands with glee with all the newfound space.

    But honestly, cool hack! If you invent a method to encode large numbers into bytes, why limit yourself to 24-bit numbers?

  3. Sharlin

    UTF-8 originally supported up to six-byte encodings (see eg. RFC 2279), but it was restricted to four bytes in 2003 in order to match UTF-16 constraints :(

  4. stbenjam

    > No special cases introduced. All properties preserved.

    I don’t actually know if this is LLM-generated, but phrasing like this is weirdly triggering to me now

  5. zahlman

    Since the Unicode Consortium isn't going to actually assign those code points, this is functionally just a scheme for encoding variable-length integers designed as an extension of UTF-8 more or less arbitrarily.

    There's a long history of designs for these (https://en.wikipedia.org/wiki/Variable-length_integer) that the author might be interested in. I used to think about these things myself, including the "zigzag encoding" for signed values (not a difficult idea; this "marvelous bijective mapping" is the standard one used in math class to demonstrate that the integers are countable, and the nice implementation properties are a consequence of the choice to "zig" from 0 to -1 first combined with how two's-complement works).

  6. yyyk

    Just limit it to 8 bytes at which point you always do 'know the number of follow on bytes' from the first byte.

    Nobody needs more than 4.47 trillion characters. (famous last words)

  7. bastawhiz

    At some point it just collapses into a sort of Huffman coding of every possible 4096 bit embedding vector.

  8. achille

    > Ken Thompson: "...i really dont think it is useful. it is like replacing ipv6 with ipv50"

More from this day

2026-09-20