Releases: cometkim/unicode-segmenter
Release list
unicode-segmenter@0.17.3
Patch Changes
- 308be8c: Fixed G9Bc edge case. The
0x0A4D(Gurmukhi Sign Virama) was incorrectly treated as InCB=Linker.
unicode-segmenter@0.17.2
Patch Changes
-
7374446: Removed pinned
graphemeSegments()in the module scope to make all APIs able to be three-shaken properly.It was introduced when they all use
graphemeSegments()as the core. But now they are all have their own loop. -
957898b: Optimize the hot loop based on a deep analysis of the V8 optimization chain.
As the result, the bundle size, speed, and memory usage. All three axes are improved. See PR #144 for detailed explanation.
- Bundle: −4.2% min+gzip, −3.9% min+brotli on
unicode-segmenter/grapheme(2,453 → 2,351 gzip); −3.4% / −2.8% on the full entry - Hermes bytecode: −20.6% (20,015 → 15,892 bytes), −18.6% gzipped
- Runtime (Node.js/V8, per benchmark case)
splitGraphemes()1.5–2.2x,countGraphemes()1.20–1.43x,graphemeSegments()1.05–1.21x,collectGraphemes()1.01–1.19x.- Bun/JSC gains are larger, and the interpreter tiers (Hermes, QuickJS) improve 5–23%
- Memory: lookup tables 20.6 kB → 19.1 kB, retained heap 228 kB → 218 kB, module init 1.7 ms → 1.5 ms
The state compaction strategy is the major part. It is valid across all optimization tiers of the V8 runtime (Jitless, Maglev, TurboFan) and has been consistently improved across all other engines.
Another noticeable change is
splitGraphemes(), it now owns its loop, just likecountGraphemes().
It produces a 30-60% performance improvement. The size increase is roughly free after compression, since the fourth byte-aligned copy of the loop back-references the other three. And the uncompressed size is amortized by other improvements.All the analysis have done by Claude Opus 5, well-done!
- Bundle: −4.2% min+gzip, −3.9% min+brotli on
unicode-segmenter@0.17.1
Patch Changes
-
e94d203: Add
collectGraphemes(input)API that collects grapheme clusters into an array directly.This is a fast version of
[...splitGraphemes(input)], which is acually the most used pattern in practice.It's 2-4x faster than the iterator-based approach. However, it collects all grapheme clusters at once, so it's not good for large text or streamed input.
-
5fcaa4f: Promote
countGraphemes()API to be a fast path.It avoids generator overhead and only counts boundaries without allocating segments.
When only counting, it achieves ~10x faster runtime performance and no GC pressure.
unicode-segmenter@0.17.0
Minor Changes
-
483ff75: Rewrite the grapheme segmenter with flat lookup tables and pair-rule dispatch.
- Bundle: −31% minified, −34% min+gzip, −26% min+brotli, −10% Hermes bytecode
- Runtime: fastest in the ecosystem on every benchmark case, ~1.1x faster on Hermes
- Memory: −44% retained heap, no retained per-range JS objects
Also fixes several segmentation bugs where results diverged from
Intl.Segmenter:- GB11: ZWJ joined non-pictographic characters (e.g.
'👍a'), double ZWJ, and SpacingMark-interruptedExtend*sequences; ExtPic after a Prepend never armed the rule - GB9c:
InCB=Nonecharacters (matras, ZWNJ) did not reset the conjunct sequence; consonants after a Prepend never started one - Unassigned
U+D7FC..U+D7FFwere treated as HangulT
The public API is unchanged. The internal
_incb_datamodule is removed
(its data is folded into the grapheme table), and all modules now share a
single flat-table binary search —general/emojipredicates get 27–36%
faster with 65% less retained heap.
unicode-segmenter@0.16.0
Minor Changes
- 9ee7b6d: Remove pre-built bundled entrypoints from the package
Patch Changes
-
12dc573: Fix GB9c rule; reset internal "InCB=Consonant" state properly.
So giving the following input:
# Malayalam KA + Virama + SPACE + VA "क् क"Will now produces three sperated segments correctly.
Thanks to @spaceemotion for reporting this issue.
-
f5d3453: Fix G9Bc rule;
ZWNJ(InCB=None) handling was missing. Thanks to @spaceemotion for reporting this. -
8ec376e: Reset
InCB=Linkertracking state for a new boundary. -
877b76c: Fix
Extend + Extended_Pictographiccluster break
unicode-segmenter@0.15.0
Minor Changes
-
97a871e: Update to Unicode® 17.0.0
Unicode® Standard Annex #29 - Revision 47
Tested with Node.js v25.5.0 (icu 78.2)
Patch Changes
-
38a37f2: Fix TypeScript Node16 module resolution for CommonJS modules.
More specifically, the "Masquerading as CJS" issue has been fixed by including re-export declaration files.
Due to the library continues to support CommonJS (at least up to v1), this change is necessary and slightly increases the size of node_modules.
Also, pre-bundled files (
unicode-segmenter/bundle/*) are included for browsers and miniprograms. They were missing in previous versions due to a path typo in the build script.
unicode-segmenter@0.14.5
unicode-segmenter@0.14.4
Patch Changes
- 41a7920: Inlining more Hangul ranges (Hangul Jamo Extended-B) to reduce index memory usage (8.5KB -> 7.4KB)
Slightly improved the bundle size as well.
unicode-segmenter@0.14.3
Patch Changes
-
65c38ce: Move GB9c rule checking to be after the main boundary checking.
To try to avoid unnecessary work as much as possible.No noticeable changes, but perf seems to be improved by ~2% for most cases.
-
8b23df9: Two further optimizations:
- Remove inlined ranges from the data file.
- Add inlined range: 0xAC00-0xD7A3 (Hangul syllables) can easily be inlined.
The 1 is something I forgot in #104 task, but it was a slight chance.
Btw, the number 2 is a huge finding. It is a pretty extensive range to be newly inlined.
Applying both optimizations significantly reduced the bundle size and memory footprint.- Size(min): 12,549 bytes -> 6,846 bytes (-45.5%)
- Size(min+gz): 5,314 bytes -> 3,449 bytes (-35.1%)
- Index memory usage: 14,272 bytes -> 8,686 bytes (-39.2%)
Of course, without perf regression.
unicode-segmenter@0.14.2
Patch Changes
-
b7a6e12: Optimizing grapheme break category lookup for better runtime trade-offs.
See issue for the explanation.
With this change, the library's constant memory footprint is reduced from 64 KB to 14 KB without performance regressions.
However, the code size increases slightly due to inlining. It's still relatively small.