Skip to content

Bump the nuget group with 1 update - #171

Closed
dependabot[bot] wants to merge 1 commit into
masterfrom
dependabot/nuget/src/Benchmarks/nuget-ddb496c2ed
Closed

dependabot[bot] wants to merge 1 commit into
masterfrom
dependabot/nuget/src/Benchmarks/nuget-ddb496c2ed

Conversation

@dependabot

@dependabot dependabot Bot commented on behalf of github Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

Updated FastBertTokenizer from 1.0.28 to 1.2.6.

Release notes

Sourced from FastBertTokenizer's releases.

1.2.6

Highlights

  • New net10.0 target (#​134). The package now ships net10.0, net8.0 and netstandard2.0. On .NET 9+ the span-based vocabulary lookup uses FrozenDictionary.GetAlternateLookup instead of the unsafe char* key trick, so that target is compiled without unsafe code.
  • added_tokens support (#​100). LoadTokenizerJson now honours the added_tokens section of tokenizer.json; if several added tokens match at a position, the longest one wins, like in Hugging Face (#​101, thanks @​kudima03).
  • strip_accents support. tokenizer.json files that set normalizer.strip_accents explicitly (including false) now load. If it is unset, it follows lowercase, as in the original BERT.
  • cleanup_tokenization_spaces from the configuration (#​105, #​107, thanks @​kudima03). Decode now uses the tokenizer's decoder.cleanup setting.
  • Faster pre-tokenization (#​151). ASCII characters are classified with a lookup table, which cuts single-threaded encode time by roughly 6–9%.
  • Bounded worst case for very long words (#​156). WordPiece matching now starts at the length of the longest vocabulary entry, not the whole word. Encoding a single 10 000-char word took 20.6 s on net8.0 before and takes about 0.35 ms now. This matters if you tokenize untrusted input. Output is unchanged.
  • LoadFromHuggingFaceAsync on netstandard2.0 (#​159), so it is also available on .NET Framework.

Deprecations

  • Decode(ReadOnlySpan<long> tokenIds, bool cleanupTokenizationSpaces) is obsolete (#​107, #​159). Use Decode(ReadOnlySpan<long> tokenIds), which takes the value from decoder.cleanup in tokenizer.json (default true; always true for vocabularies loaded with LoadVocabulary).
    • Compiled binaries keep working unchanged. When you recompile, a call like Decode(ids) switches to the new overload. Output only differs if your tokenizer.json sets "cleanup": false.

Behaviour changes

  • tokenizer.json files with added_tokens tokenize differently. Earlier versions ignored added_tokens; the resulting token ids now match Hugging Face. Configurations with an added token that has single_word: true are rejected with an ArgumentException.
  • Decode of a sequence starting with a suffix token now puts the configured continuing-subword prefix (e.g. ##) in front of the result ("##m ipsum" instead of "m ipsum").
  • The net6.0 target was removed. .NET 6 and 7 apps now use the netstandard2.0 build. Its public API is identical, except for the [Experimental] CreateAsyncBatchEnumerator(ChannelReader<…>, …) overload, which needs net8.0 or later.
  • On .NET Framework, LoadFromHuggingFaceAsync needs TLS 1.2. Apps targeting .NET Framework below 4.7 may have to set ServicePointManager.SecurityProtocol themselves.

Other

  • Verifiable releases (#​160). The package is built and published to nuget.org by GitHub Actions, using trusted publishing. Each release comes with a build provenance attestation for the .nupkg and for the FastBertTokenizer.dll files inside it. nuget.org adds its own signature to the package, which changes the .nupkg's hash. To verify a package from nuget.org, either:

    • use NuGet Provenance Verifier, which removes the nuget.org signature and checks the package against GitHub's attestations. It runs in your browser and takes a package id and version or a .nupkg file, or
    • extract an assembly from the package and run
      gh attestation verify FastBertTokenizer.dll -R georg-jung/FastBertTokenizer

    GitHub releases are immutable from this version on; gh release verify v1.2.6 -R georg-jung/FastBertTokenizer checks the release itself.

  • API docs are published at https://fastberttokenizer.gjung.com/, which now also links to these release notes.

  • Package versions follow SemVer 2. The package is validated against 1.0.28 to catch accidental API breaks.

  • Dependency floors for netstandard2.0: System.Memory 4.6.3 (was 4.5.5), System.Text.Json 8.0.6 (was 8.0.3).

  • Many new parity tests against Hugging Face tokenizers. Known remaining divergences are documented as skipped tests (#​151).

  • Reworked benchmarks (.NET 10, fixed baselines, current competitors) with fresh results in the README (#​123, #​158).

Full Changelog: v1.0.28...v1.2.6

1.1.30-alpha

  • Add ## (or what is configured in tokenizer.json) to the beginning of the decode result if the first encoded token id is a suffix token.

1.1.21-alpha

Full Changelog: v1.0.28...v1.1.21-alpha

  • Add support for added_tokens, #​100
  • Add proper support for strip_accents
    • todo: add tests for this

Commits viewable in compare view.

Dependabot compatibility score

Dependabot will resolve any conflicts with this PR as long as you don't alter it yourself. You can also trigger a rebase manually by commenting @dependabot rebase.


Dependabot commands and options

You can trigger Dependabot actions by commenting on this PR:

  • @dependabot rebase will rebase this PR
  • @dependabot recreate will recreate this PR, overwriting any edits that have been made to it
  • @dependabot show <dependency name> ignore conditions will show all of the ignore conditions of the specified dependency
  • @dependabot ignore <dependency name> major version will close this group update PR and stop Dependabot creating any more for the specific dependency's major version (unless you unignore this specific dependency's major version or upgrade to it yourself)
  • @dependabot ignore <dependency name> minor version will close this group update PR and stop Dependabot creating any more for the specific dependency's minor version (unless you unignore this specific dependency's minor version or upgrade to it yourself)
  • @dependabot ignore <dependency name> will close this group update PR and stop Dependabot creating any more for the specific dependency (unless you unignore this specific dependency or upgrade to it yourself)
  • @dependabot unignore <dependency name> will remove all of the ignore conditions of the specified dependency
  • @dependabot unignore <dependency name> <ignore condition> will remove the ignore condition of the specified dependency and ignore conditions

Bumps FastBertTokenizer from 1.0.28 to 1.2.6

---
updated-dependencies:
- dependency-name: FastBertTokenizer
  dependency-version: 1.2.6
  dependency-type: direct:production
  update-type: version-update:semver-minor
  dependency-group: nuget
...

Signed-off-by: dependabot[bot] <support@github.com>
@dependabot dependabot Bot added .NET Pull requests that update .net code dependencies Pull requests that update a dependency file labels Oct 2, 2026
@dependabot @github

dependabot Bot commented on behalf of github Oct 5, 2026

Copy link
Copy Markdown
Contributor Author

Looks like FastBertTokenizer is updatable in another way, so this is no longer needed.

@dependabot dependabot Bot closed this Oct 5, 2026
@dependabot
dependabot Bot deleted the dependabot/nuget/src/Benchmarks/nuget-ddb496c2ed branch October 5, 2026 18:05
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

dependencies Pull requests that update a dependency file .NET Pull requests that update .net code

Projects

None yet

Development

Successfully merging this pull request may close these issues.

0 participants