Skip to content

Releases: cometkim/unicode-segmenter

unicode-segmenter@0.17.3

Choose a tag to compare

@github-actions github-actions released this 29 Jul 21:26
d9d3c4e

Patch Changes

  • 308be8c: Fixed G9Bc edge case. The 0x0A4D (Gurmukhi Sign Virama) was incorrectly treated as InCB=Linker.

unicode-segmenter@0.17.2

Choose a tag to compare

@github-actions github-actions released this 26 Jul 19:58
dd517de

Patch Changes

  • 7374446: Removed pinned graphemeSegments() in the module scope to make all APIs able to be three-shaken properly.

    It was introduced when they all use graphemeSegments() as the core. But now they are all have their own loop.

  • 957898b: Optimize the hot loop based on a deep analysis of the V8 optimization chain.

    As the result, the bundle size, speed, and memory usage. All three axes are improved. See PR #144 for detailed explanation.

    • Bundle: −4.2% min+gzip, −3.9% min+brotli on unicode-segmenter/grapheme (2,453 → 2,351 gzip); −3.4% / −2.8% on the full entry
    • Hermes bytecode: −20.6% (20,015 → 15,892 bytes), −18.6% gzipped
    • Runtime (Node.js/V8, per benchmark case)
      • splitGraphemes() 1.5–2.2x, countGraphemes() 1.20–1.43x, graphemeSegments() 1.05–1.21x, collectGraphemes() 1.01–1.19x.
      • Bun/JSC gains are larger, and the interpreter tiers (Hermes, QuickJS) improve 5–23%
    • Memory: lookup tables 20.6 kB → 19.1 kB, retained heap 228 kB → 218 kB, module init 1.7 ms → 1.5 ms

    The state compaction strategy is the major part. It is valid across all optimization tiers of the V8 runtime (Jitless, Maglev, TurboFan) and has been consistently improved across all other engines.

    Another noticeable change is splitGraphemes(), it now owns its loop, just like countGraphemes().
    It produces a 30-60% performance improvement. The size increase is roughly free after compression, since the fourth byte-aligned copy of the loop back-references the other three. And the uncompressed size is amortized by other improvements.

    All the analysis have done by Claude Opus 5, well-done!

unicode-segmenter@0.17.1

Choose a tag to compare

@github-actions github-actions released this 23 Jul 07:28
5b34af9

Patch Changes

  • e94d203: Add collectGraphemes(input) API that collects grapheme clusters into an array directly.

    This is a fast version of [...splitGraphemes(input)], which is acually the most used pattern in practice.

    It's 2-4x faster than the iterator-based approach. However, it collects all grapheme clusters at once, so it's not good for large text or streamed input.

  • 5fcaa4f: Promote countGraphemes() API to be a fast path.

    It avoids generator overhead and only counts boundaries without allocating segments.
    When only counting, it achieves ~10x faster runtime performance and no GC pressure.

unicode-segmenter@0.17.0

Choose a tag to compare

@github-actions github-actions released this 06 Jul 23:16
f4abccc

Minor Changes

  • 483ff75: Rewrite the grapheme segmenter with flat lookup tables and pair-rule dispatch.

    • Bundle: −31% minified, −34% min+gzip, −26% min+brotli, −10% Hermes bytecode
    • Runtime: fastest in the ecosystem on every benchmark case, ~1.1x faster on Hermes
    • Memory: −44% retained heap, no retained per-range JS objects

    Also fixes several segmentation bugs where results diverged from Intl.Segmenter:

    • GB11: ZWJ joined non-pictographic characters (e.g. '👍‍a'), double ZWJ, and SpacingMark-interrupted Extend* sequences; ExtPic after a Prepend never armed the rule
    • GB9c: InCB=None characters (matras, ZWNJ) did not reset the conjunct sequence; consonants after a Prepend never started one
    • Unassigned U+D7FC..U+D7FF were treated as Hangul T

    The public API is unchanged. The internal _incb_data module is removed
    (its data is folded into the grapheme table), and all modules now share a
    single flat-table binary search — general/emoji predicates get 27–36%
    faster with 65% less retained heap.

unicode-segmenter@0.16.0

Choose a tag to compare

@github-actions github-actions released this 23 Apr 18:23
61a4701

Minor Changes

  • 9ee7b6d: Remove pre-built bundled entrypoints from the package

Patch Changes

  • 12dc573: Fix GB9c rule; reset internal "InCB=Consonant" state properly.

    So giving the following input:

    # Malayalam KA + Virama + SPACE + VA
    "क्‌ क"
    

    Will now produces three sperated segments correctly.

    Thanks to @spaceemotion for reporting this issue.

  • f5d3453: Fix G9Bc rule; ZWNJ(InCB=None) handling was missing. Thanks to @spaceemotion for reporting this.

  • 8ec376e: Reset InCB=Linker tracking state for a new boundary.

  • 877b76c: Fix Extend + Extended_Pictographic cluster break

unicode-segmenter@0.15.0

Choose a tag to compare

@github-actions github-actions released this 28 Jan 19:46
fa2356d

Minor Changes

  • 97a871e: Update to Unicode® 17.0.0

    Unicode® Standard Annex #29 - Revision 47

    Tested with Node.js v25.5.0 (icu 78.2)

Patch Changes

  • 38a37f2: Fix TypeScript Node16 module resolution for CommonJS modules.

    More specifically, the "Masquerading as CJS" issue has been fixed by including re-export declaration files.

    Due to the library continues to support CommonJS (at least up to v1), this change is necessary and slightly increases the size of node_modules.

    Also, pre-bundled files (unicode-segmenter/bundle/*) are included for browsers and miniprograms. They were missing in previous versions due to a path typo in the build script.

unicode-segmenter@0.14.5

Choose a tag to compare

@github-actions github-actions released this 29 Dec 00:02
a47e430

Patch Changes

  • 9d482aa: Inlined the grapheme boundary checking
    to avoid unnecessary function calls in the hotpath and consolidating internal state.

    This achieved the runtime perf by 2% and a slight bundle size reduction.

  • d737dfe: Inlined the InCB=Linker checking for Indic scripts

unicode-segmenter@0.14.4

Choose a tag to compare

@github-actions github-actions released this 14 Dec 23:25
e50d821

Patch Changes

  • 41a7920: Inlining more Hangul ranges (Hangul Jamo Extended-B) to reduce index memory usage (8.5KB -> 7.4KB)
    Slightly improved the bundle size as well.

unicode-segmenter@0.14.3

Choose a tag to compare

@github-actions github-actions released this 14 Dec 22:18
6c8b4b4

Patch Changes

  • 65c38ce: Move GB9c rule checking to be after the main boundary checking.
    To try to avoid unnecessary work as much as possible.

    No noticeable changes, but perf seems to be improved by ~2% for most cases.

  • 8b23df9: Two further optimizations:

    1. Remove inlined ranges from the data file.
    2. Add inlined range: 0xAC00-0xD7A3 (Hangul syllables) can easily be inlined.

    The 1 is something I forgot in #104 task, but it was a slight chance.

    Btw, the number 2 is a huge finding. It is a pretty extensive range to be newly inlined.
    Applying both optimizations significantly reduced the bundle size and memory footprint.

    • Size(min): 12,549 bytes -> 6,846 bytes (-45.5%)
    • Size(min+gz): 5,314 bytes -> 3,449 bytes (-35.1%)
    • Index memory usage: 14,272 bytes -> 8,686 bytes (-39.2%)

    Of course, without perf regression.

unicode-segmenter@0.14.2

Choose a tag to compare

@github-actions github-actions released this 12 Dec 20:59
c06976d

Patch Changes

  • b7a6e12: Optimizing grapheme break category lookup for better runtime trade-offs.

    See issue for the explanation.

    With this change, the library's constant memory footprint is reduced from 64 KB to 14 KB without performance regressions.
    However, the code size increases slightly due to inlining. It's still relatively small.