unicode/norm: don't truncate runes in the recomposition map key combine built its map key as uint32(uint16(a))<<16 + uint32(uint16(b)), truncating both runes to 16 bits. Supplementary-plane runes therefore aliased BMP entries in both directions: entries whose decomposition uses a supplementary rune were stored under a truncated key, and lookups with a supplementary rune found unrelated BMP entries. The result is a composition that is not canonically equivalent to its input, reachable from plain BMP text: NFC(U+05D2 U+0307) = U+105C9 // Hebrew gimel + dot above NFC(U+0300 U+11F41) = U+1F43 // Kawi rune destroyed This is not a Unicode 17 regression; it affects the Unicode 15 tables too, at 15056 two-rune inputs, and 15379 under Unicode 17. Pack each recompMapPacked entry as a big-endian uint64 holding three 21-bit fields: the two decomposition runes and the composed rune. That is wide enough for any rune and keeps the entries at 8 bytes each, so the table size is unchanged. An exhaustive scan of every rune followed by every combining or backward-combining rune now reports no output that fails to be canonically equivalent to its input, and NFC is idempotent throughout. Change-Id: I2795f7d129d7a1adb40e5666f2cd33888935e9fb Reviewed-on: https://go-review.googlesource.com/c/text/+/822082 LUCI-TryBot-Result: golang-scoped@luci-project-accounts.iam.gserviceaccount.com <golang-scoped@luci-project-accounts.iam.gserviceaccount.com> Reviewed-by: Damien Neil <dneil@google.com> Reviewed-by: David Chase <drchase@google.com>
This repository holds supplementary Go packages for text processing, many involving Unicode.
It is important that the Unicode version used in x/text matches the one used by your Go compiler. The x/text repository supports multiple versions of Unicode and will match the version of Unicode to that of the Go compiler. At the moment this is supported for Go compilers from version 1.7.
To submit changes to this repository, see http://go.dev/doc/contribute.
The git repository is https://go.googlesource.com/text.
To generate the tables in this repository (except for the encoding tables), run go generate from this directory. By default tables are generated for the Unicode version in core and the CLDR version defined in golang.org/x/text/unicode/cldr.
Running go generate will as a side effect create a DATA subdirectory in this directory, which holds all files that are used as a source for generating the tables. This directory will also serve as a cache.
Run
go test ./...
from this directory to run all tests. Add the “-tags icu” flag to also run ICU conformance tests (if available). This requires that you have the correct ICU version installed on your system.
TODO:
To generate the tables in this repository (except for the encoding tables), run go generate from this directory. By default tables are generated for the Unicode version in core and the CLDR version defined in golang.org/x/text/unicode/cldr.
Running go generate will as a side effect create a DATA subdirectory in this directory which holds all files that are used as a source for generating the tables. This directory will also serve as a cache.
To update a Unicode version run
UNICODE_VERSION=x.x.x go generate
where x.x.x must correspond to a directory in https://www.unicode.org/Public/. If this version is newer than the version in core it will also update the relevant packages there. The idna package in x/net will always be updated.
To update a CLDR version run
CLDR_VERSION=version go generate
where version must correspond to a directory in https://www.unicode.org/Public/cldr/.
Note that the code gets adapted over time to changes in the data and that backwards compatibility is not maintained. So updating to a different version may not work.
The files in DATA/{iana|icu|w3|whatwg} are currently not versioned.
The main issue tracker for the text repository is located at https://go.dev/issues. Prefix your issue with “x/text:” in the subject line, so it is easy to find.