A comprehensive, Pydantic-native tokenization toolkit for building, analyzing, and managing tokenizers. Designed as a control surface for training frameworks - tokenizers are inspectable, mutable, schedulable, and debuggable.
This module provides utilities for:
- Custom tokenizer creation - Build character or word-level tokenizers
- Token statistics & analysis - Frequency counts, coverage, compression ratios
- Batch processing - Efficient encoding/decoding with padding and chunking
- Special token handling - BOS, EOS, PAD, UNK token management
- Vocabulary management - Merge, filter, validate vocabularies
- Debugging & visualization - Token inspection and comparison tools
- Validation - Roundtrip testing and integrity checks
- Format conversion - Export/import JSON, TSV, CSV, HuggingFace format
| Submodule | Purpose |
|---|---|
analyze/ |
Coverage, entropy, fit scoring, efficiency metrics, vocab induction |
curriculum/ |
Token-length buckets, reasoning density scoring |
preprocessing/ |
Pre-tokenization transforms, profiles, byte fallback |
runtime/ |
Special token registry, dynamic vocab, semantics mapping, chat templates |
training/ |
Sequence packing, throughput profiling |
regression/ |
Token regression test framework |
research/ |
Soft tokens, token morphing, embedding analysis |
instrumentation/ |
Token histograms, OOV analysis, waste metrics, vocab comparison |
backends/ |
Pluggable backends (HuggingFace + fast MLX CharTrie) + benchmarking |
fingerprint.py |
Tokenizer fingerprinting for compatibility verification |
- Pydantic-native: All data structures use Pydantic BaseModel for validation
- No magic strings: Enums and constants for type safety
- Protocol-based: TokenizerProtocol allows any compatible tokenizer
- Composable: Functions work independently or together
- Training-aware: Tokenizers as a control surface, not just preprocessing
from chuk_lazarus.data.tokenizers import CustomTokenizer
from chuk_lazarus.data.tokenizers.batch_processing import create_batch, PaddingSide
from chuk_lazarus.data.tokenizers.token_stats import get_top_tokens
# Create a tokenizer
tokenizer = CustomTokenizer()
tokenizer.build_vocab("The quick brown fox jumps over the lazy dog.", min_freq=1)
# Encode text
tokens = tokenizer.encode("The fox is quick.")
print(f"Tokens: {tokens}")
# Batch processing
batch = create_batch(
["Hello world", "How are you?"],
tokenizer,
padding=True,
padding_side=PaddingSide.RIGHT,
)
print(f"Batch shape: {len(batch.input_ids)} x {len(batch.input_ids[0])}")Lightweight character-level tokenizer for sequence encoding. Useful for:
- Character-level language models
- Testing and debugging pipelines
- Domains where subword tokenization is overkill
from chuk_lazarus.data.tokenizers import CharacterTokenizer, CharacterTokenizerConfig
# Create from corpus (learns vocabulary)
tokenizer = CharacterTokenizer.from_corpus(["hello", "world"])
print(f"Vocab size: {tokenizer.vocab_size}") # 11 (7 chars + 4 special)
# Factory methods
ascii_tokenizer = CharacterTokenizer.from_ascii()
lowercase_tokenizer = CharacterTokenizer.from_ascii_lowercase()
digits_tokenizer = CharacterTokenizer.from_digits()
# Encode/decode
tokens = tokenizer.encode("hello", add_special_tokens=True)
text = tokenizer.decode(tokens, skip_special_tokens=True)
# Batch encoding with padding
batch = tokenizer.encode_batch(
["hi", "hello"],
max_length=10,
padding=True,
)Config options:
config = CharacterTokenizerConfig(
pad_token_id=0, # Padding token ID
unk_token_id=1, # Unknown token ID
bos_token_id=2, # Beginning of sequence
eos_token_id=3, # End of sequence
lowercase=True, # Lowercase input before tokenizing
)
tokenizer = CharacterTokenizer.from_corpus(texts, config)Bag-of-Words character tokenizer for classification tasks. Converts text into fixed-size vectors of character frequencies (not token IDs).
from chuk_lazarus.data.tokenizers import BoWCharacterTokenizer, BoWTokenizerConfig
# Create from corpus
tokenizer = BoWCharacterTokenizer.from_corpus(["cat", "dog"])
print(f"Vocab size: {tokenizer.vocab_size}") # 6 (unique chars)
# Encode returns float vector (not token IDs)
vec = tokenizer.encode("cat") # [0.33, 0.0, 0.33, ...]
print(f"Vector length: {len(vec)}") # Same as vocab_size
print(f"Sum: {sum(vec)}") # ~1.0 (normalized by default)
# Batch encoding
vecs = tokenizer.encode_batch(["cat", "dog"])
# Factory methods
ascii_tokenizer = BoWCharacterTokenizer.from_ascii()
lowercase_tokenizer = BoWCharacterTokenizer.from_ascii_lowercase()Config options:
config = BoWTokenizerConfig(
lowercase=True, # Lowercase input (default: True)
normalize=True, # L1-normalize vectors (default: True)
)
tokenizer = BoWCharacterTokenizer.from_corpus(texts, config)
# Without normalization (raw counts)
config = BoWTokenizerConfig(normalize=False)
tokenizer = BoWCharacterTokenizer("abc", config)
vec = tokenizer.encode("aab") # [2.0, 1.0, 0.0] (raw counts)Use case - Sentiment Classification:
from chuk_lazarus.data import ClassificationDataset
from chuk_lazarus.data.tokenizers import BoWCharacterTokenizer
# Load data
dataset = ClassificationDataset.from_jsonl("train.jsonl")
# Build tokenizer from corpus
tokenizer = BoWCharacterTokenizer.from_corpus(dataset.texts)
# Encode for classification
import mlx.core as mx
inputs = mx.stack([mx.array(tokenizer.encode(s.text)) for s in dataset])Custom tokenizer implementation extending HuggingFace's PreTrainedTokenizer.
from chuk_lazarus.data.tokenizers import CustomTokenizer
tokenizer = CustomTokenizer()
tokenizer.build_vocab(text, min_freq=2)
tokenizer.save("./my_tokenizer")
loaded = CustomTokenizer.load("./my_tokenizer")Statistics and analysis for tokenization.
from chuk_lazarus.data.tokenizers.token_stats import (
get_token_frequencies,
get_vocabulary_coverage,
calculate_compression_ratio,
get_top_tokens,
get_rare_tokens,
)
# Analyze token frequency
frequencies = get_token_frequencies(texts, tokenizer)
# Check vocabulary coverage
coverage = get_vocabulary_coverage(text, tokenizer)
print(f"Coverage: {coverage.coverage_ratio:.2%}")
# Compression analysis
stats = calculate_compression_ratio(text, tokenizer)
print(f"Chars per token: {stats.chars_per_token:.2f}")Efficient batch encoding and padding.
from chuk_lazarus.data.tokenizers.batch_processing import (
create_batch,
pad_batch,
chunk_text,
PaddingSide,
ChunkConfig,
)
# Create padded batch
batch = create_batch(
texts,
tokenizer,
max_length=512,
padding=True,
truncation=True,
padding_side=PaddingSide.LEFT,
)
# Chunk long text
config = ChunkConfig(chunk_size=128, overlap=16)
chunks = chunk_text(long_text, tokenizer, config)Special token utilities.
from chuk_lazarus.data.tokenizers.special_tokens import (
SpecialTokenConfig,
SpecialTokenType,
get_special_token_ids,
strip_special_tokens,
add_bos_token,
add_eos_token,
)
# Get config from tokenizer
config = SpecialTokenConfig.from_tokenizer(tokenizer)
all_special = config.all_special_ids()
# Manipulate sequences
tokens = add_bos_token(tokens, tokenizer.bos_token_id)
tokens = strip_special_tokens(tokens, tokenizer)Vocabulary manipulation utilities.
from chuk_lazarus.data.tokenizers.vocab_manager import (
merge_vocabularies,
filter_vocabulary,
validate_vocabulary,
get_vocabulary_diff,
ConflictResolution,
)
# Merge two vocabularies
merged = merge_vocabularies(
vocab1, vocab2,
conflict_resolution=ConflictResolution.FIRST
)
# Validate integrity
issues = validate_vocabulary(vocab)
if issues.has_issues():
print(f"Found {len(issues.duplicate_ids)} duplicate IDs")Debugging and visualization tools.
from chuk_lazarus.data.tokenizers.token_debug import (
get_token_info,
compare_tokenizations,
highlight_tokens,
format_token_table,
)
# Inspect a token
info = get_token_info(token_id, tokenizer)
print(f"Token: {info.token_str}, Bytes: {info.byte_repr}")
# Compare tokenizers
comparison = compare_tokenizations(text, tokenizer1, tokenizer2)
print(f"Tokenizer 1: {comparison.tokenizer1_count} tokens")
# Visualize token boundaries
highlighted = highlight_tokens("Hello world", tokenizer, separator="|")Tokenizer validation and testing.
from chuk_lazarus.data.tokenizers.validation import (
check_roundtrip,
validate_special_tokens,
create_validation_report,
assert_valid_tokenizer,
)
# Test roundtrip encoding
result = check_roundtrip("Hello world", tokenizer)
if not result.is_lossless:
print(f"Lost characters: {result.diff_chars}")
# Full validation
report = create_validation_report(tokenizer, "MyTokenizer")
print(f"Valid: {report.is_valid}")Format conversion utilities.
from chuk_lazarus.data.tokenizers.conversion import (
export_vocabulary,
import_vocabulary,
save_huggingface_format,
ExportFormat,
)
# Export vocabulary
export = export_vocabulary(vocab, format=ExportFormat.JSON)
# Save as HuggingFace format
result = save_huggingface_format(vocab, "./hf_tokenizer")Coverage, entropy, fit scoring, efficiency metrics, and retokenization comparison.
from chuk_lazarus.data.tokenizers.analyze import (
analyze_coverage,
analyze_entropy,
analyze_efficiency,
calculate_fit_score,
analyze_vocab_induction,
)
# Coverage analysis
coverage = analyze_coverage(texts, tokenizer)
print(f"UNK rate: {coverage.unk_rate:.2%}")
print(f"Tokens/word: {coverage.tokens_per_word:.2f}")
# Entropy analysis
entropy = analyze_entropy(texts, tokenizer)
print(f"Entropy: {entropy.entropy:.2f} bits")
# Tokenizer-dataset fit score
fit = calculate_fit_score(texts, tokenizer)
print(f"Fit score: {fit.overall_score:.2f}")
# Vocabulary induction - find high-impact tokens to add
report = analyze_vocab_induction(texts, tokenizer)
print(f"Potential savings: {report.total_potential_savings:,} tokens")Token-length buckets and reasoning density scoring.
from chuk_lazarus.data.tokenizers.curriculum import (
create_length_buckets,
get_curriculum_schedule,
score_reasoning_density,
)
# Create length-based curriculum buckets
buckets = create_length_buckets(texts, tokenizer, config)
for bucket in buckets:
print(f"Bucket {bucket.bucket_id}: {bucket.sample_count} samples")
# Get curriculum schedule (easy -> hard)
schedule = get_curriculum_schedule(texts, tokenizer)
# Score reasoning density
score = score_reasoning_density(text, idx, tokenizer)
print(f"Reasoning density: {score.overall_score:.2f}")Special token registry, dynamic vocabulary, semantics mapping, and chat templates.
from chuk_lazarus.data.tokenizers.runtime import (
SpecialTokenRegistry,
TokenCategory,
create_standard_registry,
ChatTemplateRegistry,
validate_chat_template,
patch_chat_template,
)
# Special token registry with collision detection
registry = create_standard_registry(vocab_size=50000)
registry.register(
token_str="<TOOL_CALL>",
token_id=50001,
category=TokenCategory.TOOL_CALL,
)
# Chat template management
result = validate_chat_template(tokenizer)
print(f"Format: {result.format.value}, Valid: {result.is_valid}")
# Patch a missing template
patch_chat_template(tokenizer, "chatml") # or llama, phi, gemma, etc.Sequence packing and throughput profiling.
from chuk_lazarus.data.tokenizers.training import (
create_packed_batch,
calculate_packing_efficiency,
profile_tokenization,
)
# Smart sequence packing (20-40% throughput improvement)
packed = create_packed_batch(texts, tokenizer, max_seq_length=512)
print(f"Packing ratio: {packed.packing_ratio:.2f}x")
# Throughput profiling
metrics = profile_tokenization(texts, tokenizer)
print(f"Tokens/second: {metrics.tokens_per_second:.0f}")Pre-tokenization hooks, profiles, and byte fallback.
from chuk_lazarus.data.tokenizers.preprocessing import (
normalize_numbers,
inject_structure_tokens,
HookPipeline,
HookedTokenizer,
create_training_profile,
wrap_with_fallback,
)
# Numeric normalization
text = "Pi is 3.14159 and e is 2.71828"
encoding = normalize_numbers(text)
print(encoding.encoded_text) # "Pi is <NUM_0> and e is <NUM_1>"
# Hook pipeline
pipeline = create_standard_pipeline(numeric=True, structure=True)
hooked = HookedTokenizer(tokenizer, pipeline)
tokens = hooked.encode(text) # Transforms applied automatically
# Byte fallback for robustness
wrapper = wrap_with_fallback(tokenizer)
tokens = wrapper.encode("Emoji: 123") # No UNK tokensToken regression test framework for CI/CD pipelines.
from chuk_lazarus.data.tokenizers.regression import (
TokenTestSuite,
run_token_tests,
load_tests_from_yaml,
)
# Load from YAML (for CI/CD)
suite = load_tests_from_yaml("tests/tokenizer_regression.yaml")
result = run_token_tests(suite, tokenizer)
print(f"Pass rate: {result.pass_rate:.2%}")YAML format:
name: My Tokenizer Tests
tests:
- name: math_symbols
text: "x^2 + y^2 = z^2"
assertion: max_tokens
expected: 8
- name: roundtrip
text: "Hello world"
assertion: roundtrip_losslessResearch playground for soft tokens, embedding manipulation, and analysis.
from chuk_lazarus.data.tokenizers.research import (
SoftTokenBank,
create_prompt_tuning_bank,
morph_token,
find_nearest_neighbors,
cluster_tokens,
analyze_embeddings,
)
# Create soft prompt tokens for prompt tuning
bank = create_prompt_tuning_bank(
num_tokens=10,
embedding_dim=768,
prefix="task",
)
# Analyze embedding space
neighbors = find_nearest_neighbors(
embeddings[0], embeddings, token_ids, token_strs, k=10
)
# Comprehensive embedding analysis
analysis = analyze_embeddings(embeddings)
print(f"Isotropy: {analysis.isotropy_score:.2f}")Pluggable backend architecture with multiple implementations.
from chuk_lazarus.data.tokenizers.backends import (
HuggingFaceBackend,
FastBackend,
get_best_backend,
is_fast_backend_available,
benchmark_tokenizer,
compare_backends,
)
# Create backends
backend = HuggingFaceBackend(hf_tokenizer)
result = backend.encode("Hello, world!")
# Fast backend (MLX CharTrie, optional)
if is_fast_backend_available():
fast = FastBackend.from_tokenizer(hf_tokenizer)
batch_result = fast.encode_batch(texts, num_workers=4)
# Benchmark and compare
comparison = compare_backends(hf_tokenizer, corpus, num_workers=4)
print(comparison.summary()) # Shows speedup ratioGenerate stable hashes for compatibility verification.
from chuk_lazarus.data.tokenizers.fingerprint import (
compute_fingerprint,
verify_fingerprint,
save_fingerprint,
load_fingerprint,
)
# Compute fingerprint
fingerprint = compute_fingerprint(tokenizer)
print(f"Fingerprint: {fingerprint.fingerprint}")
# Save/load
save_fingerprint(fingerprint, "tokenizer_fingerprint.json")
loaded = load_fingerprint("tokenizer_fingerprint.json")
# Verify tokenizer matches
mismatch = verify_fingerprint(tokenizer, fingerprint)
if mismatch is None:
print("Tokenizer matches!")Pure observability tools for analyzing tokenization behavior.
from chuk_lazarus.data.tokenizers.instrumentation import (
compute_length_histogram,
format_histogram_ascii,
analyze_oov,
analyze_waste,
compare_vocab_impact,
)
# Token length histogram
histogram = compute_length_histogram(texts, tokenizer, num_bins=20)
print(format_histogram_ascii(histogram))
# OOV analysis
oov_report = analyze_oov(texts, tokenizer)
print(f"UNK rate: {oov_report.unk_rate:.2%}")
# Waste analysis
waste = analyze_waste(texts, tokenizer, max_length=512)
print(f"Padding rate: {waste.padding.padding_rate:.1%}")
print(f"Truncation rate: {waste.truncation.truncation_rate:.1%}")Support for OpenAI's tokenizers via tiktoken.
pip install 'chuk-lazarus[openai]'from chuk_lazarus.data.tokenizers.tiktoken_wrapper import TiktokenWrapper
# Load by model name
tokenizer = TiktokenWrapper.from_model("gpt-4")
tokenizer = TiktokenWrapper.from_model("gpt-4o")
# Use like any other tokenizer
tokens = tokenizer.encode("Hello, world!")| Model | Encoding | Vocab Size |
|---|---|---|
| gpt-4, gpt-4-turbo, gpt-3.5-turbo | cl100k_base | 100,277 |
| gpt-4o, gpt-4o-mini, o1, o1-mini | o200k_base | 200,019 |
All utilities accept any tokenizer implementing this protocol:
class TokenizerProtocol(Protocol):
def encode(self, text: str, add_special_tokens: bool = True) -> list[int]: ...
def decode(self, ids: list[int]) -> str: ...
def get_vocab(self) -> dict[str, int]: ...
pad_token_id: int
unk_token_id: int | None
bos_token_id: int | None
eos_token_id: int | NoneReserve token ID ranges for different purposes:
[0-49999] -> language tokens
[50000-50999] -> tool/action tokens
[51000-51999] -> memory/paging tokens
[52000-52999] -> solver/operator tokens
[53000+] -> experimental
Use runtime.SpecialTokenRegistry to enforce these boundaries and detect collisions.
pytest tests/data/tokenizers/ -v --cov=src/chuk_lazarus/data/tokenizers