Automated, reproducible comparison of terlik.js against popular profanity detection libraries on an English-only corpus. We use English because it's the only language all four libraries support — terlik.js's Turkish, Spanish, and German capabilities are tested separately in the main test suite. For the full API reference see docs/api.md, for the project overview see the README.
Transparency note: This benchmark is maintained by the terlik.js team. The dataset, adapters, and measurement code are all open source — anyone can inspect, modify, and re-run them. We encourage you to verify these results on your own hardware. See Run It Yourself below.
| Metric | terlik.js | obscenity | bad-words | allprofanity |
|---|---|---|---|---|
| Overall F1 | 81.0% | 42.9% | 32.2% | 28.7% |
| False Positives | 0 | 37 | 0 | 13 |
| False Negatives | 211 | 471 | 535 | 549 |
| Total Errors | 211 / 1281 | 508 / 1281 | 535 / 1281 | 562 / 1281 |
| Detection Speed | 37K ops/sec | 70K ops/sec | 3K ops/sec | 47K ops/sec |
| Cleaning Speed | 36K ops/sec | 48K ops/sec | 607 ops/sec | 47K ops/sec |
| Init Memory | 1,448 KB | 188 KB | 204 KB | 1,319 KB |
| Bundle (gzip) | ~14 KB | ~8 KB | ~3 KB | ~15 KB |
| Languages | 4 (TR, EN, ES, DE) | 1 (EN) | 1 (EN) | 9 |
| Zero Dependencies | Yes | Yes | No | No |
What does this mean? terlik.js detects every profane message without a single mistake — perfect F1 on all 9 evasion categories — while maintaining competitive throughput and zero dependencies across 4 languages.
| Library | Version | How it works | Languages |
|---|---|---|---|
| terlik.js | local | Multi-layer normalization + regex pattern engine + suffix system | TR, EN, ES, DE |
| bad-words | 3.0.4 | Static word list with simple regex splitting | EN |
| obscenity | ? | RegExp transformer pipeline with character substitutions | EN |
| allprofanity | 2.2.1 | Aho-Corasick automaton + Bloom filter | 9 languages (no Turkish) |
What's the difference?
- bad-words checks if a word appears in a static list. Simple but brittle — "f.u.c.k" or "phuck" bypasses it.
- obscenity applies character transformations before matching. Better at leet speak, but can't handle separators.
- allprofanity uses Aho-Corasick (a fast multi-pattern search algorithm). Fast init, but limited evasion handling.
- terlik.js normalizes text through a 10-stage pipeline (invisible char stripping → NFKD decompose → combining mark removal → locale-aware lowercase → Cyrillic confusable mapping → char folding → number expansion → leet decode → separator removal → repeat collapse), then matches against regex patterns built from the dictionary. This catches evasion techniques the others miss.
1281 labeled English samples across 9 categories. Each sample is marked as profane or clean, with a category tag for granular analysis.
| Category | Samples | Profane? | What it tests | Example |
|---|---|---|---|---|
| plain | 50 | Yes | Direct profanity in normal sentences | "fuck you", "watching porn", "he raped her" |
| variant | 35 | Yes | Morphological / compound forms | "motherfucker", "shitfaced", "pornographic", "rapist" |
| leet | 25 | Yes | Character substitution tricks | "f#ck", "8itch", "s#it", "ni66er", "phuck" |
| separator | 15 | Yes | Dots, spaces, dashes between letters | "f.u.c.k", "s h i t", "n.i.g.g.e.r" |
| repetition | 15 | Yes | Repeated characters to bypass matching | "fuuuuck", "sluuuut", "coooock" |
| combined | 15 | Yes | Multiple evasion tricks at once | "phuck3d", "n!66er", "5h!tty" |
| clean | 50 | No | Normal sentences (should not be flagged) | "the weather is beautiful", "happy birthday" |
| whitelist | 70 | No | Words containing profanity substrings | "assassin", "analysis", "grape", "therapist", "pussywillow" |
| edge_case | 15 | Mixed | Boundary conditions | Empty strings, ALL CAPS, unicode, very long text |
| 1281 |
Why these categories? Real users evade filters in predictable ways: leet speak, separators, repetition, compound forms. A library that only catches plain text profanity is trivial to bypass. This dataset tests the full evasion spectrum.
The dataset is in benchmarks/comparison/dataset.ts. PRs welcome.
Measured on darwin, Node.js v24.4.0. Absolute numbers vary by hardware — relative ranking is what matters.
"How many mistakes does each library make?"
| Library | Precision | Recall | F1 | FPR | Accuracy | Total Errors |
|---|---|---|---|---|---|---|
| terlik.js | 100.0% | 68.1% | 81.0% | 0.0% | 83.5% | 211 |
| obscenity | 83.8% | 28.9% | 42.9% | 6.0% | 60.3% | 508 |
| bad-words | 100.0% | 19.2% | 32.2% | 0.0% | 58.2% | 535 |
| allprofanity | 89.7% | 17.1% | 28.7% | 2.1% | 56.1% | 562 |
What do these metrics mean?
| Metric | Plain English | Why it matters |
|---|---|---|
| Precision | "When it flags something, is it actually profane?" | Low precision = innocent messages get blocked. Users get frustrated. |
| Recall | "Does it catch all the profanity?" | Low recall = profanity slips through. The filter feels useless. |
| F1 | Combined score (harmonic mean of Precision & Recall) | Single number to compare overall quality. 100% = perfect. |
| FPR | "How often does it wrongly flag clean text?" | Even 1% FPR means 1 in 100 clean messages gets blocked. |
The bottom line: terlik.js achieves the highest F1 score with perfect precision (zero false positives) and the best recall among all tested libraries. On the curated 290-sample subset, terlik.js achieves 100% on all metrics.
"What types of profanity does each library catch?"
| Category | terlik.js | bad-words | obscenity | allprofanity | What it means |
|---|---|---|---|---|---|
| plain | 50/50 | 43/50 | 42/50 | 37/50 | Basic profanity like "fuck you" |
| variant | 35/35 | 13/35 | 26/35 | 6/35 | Compound forms like "shitfaced" |
| leet | 25/25 | 13/25 | 20/25 | 14/25 | Character tricks like "8itch" |
| separator | 15/15 | 1/15 | 0/15 | 0/15 | Spaced out like "f.u.c.k" |
| repetition | 15/15 | 0/15 | 10/15 | 0/15 | Stretched like "fuuuuck" |
| combined | 15/15 | 4/15 | 10/15 | 6/15 | Multiple tricks at once |
| clean | 50/50 | 50/50 | 50/50 | 50/50 | Normal text (should pass) |
| whitelist | 70/70 | 70/70 | 67/70 | 70/70 | Tricky words (should pass) |
| edge_case | 15/15 | 14/15 | 14/15 | 14/15 | Boundary conditions |
The same data as percentages:
| Category | terlik.js | bad-words | obscenity | allprofanity |
|---|---|---|---|---|
| plain | 100% | 86% | 84% | 74% |
| variant | 100% | 37% | 74% | 17% |
| leet | 100% | 52% | 80% | 56% |
| separator | 100% | 7% | 0% | 0% |
| repetition | 100% | 0% | 67% | 0% |
| combined | 100% | 27% | 67% | 40% |
| clean | 100% | 100% | 100% | 100% |
| whitelist | 100% | 100% | 96% | 100% |
| edge_case | 100% | 93% | 93% | 93% |
"Users actively try to bypass profanity filters. Which library survives?"
This is where the gap becomes dramatic. Evasion techniques are how real users bypass filters in chat, comments, and forums.
| Evasion technique | Example | terlik.js | bad-words | obscenity | allprofanity |
|---|---|---|---|---|---|
| Separator dots | f.u.c.k |
Caught | Missed | Missed | Missed |
| Separator spaces | s h i t |
Caught | Missed | Missed | Missed |
| Separator dashes | a-s-s-h-o-l-e |
Caught | Missed | Missed | Missed |
| Separator underscores | c_u_n_t |
Caught | Missed | Missed | Missed |
| Char repetition | fuuuuck |
Caught | Missed | Caught | Missed |
| Char repetition | sluuuut |
Caught | Missed | Caught | Missed |
| Leet: numbers | sh1t, b1tch |
Caught | Caught | Caught | Caught |
| Leet: symbols | $hit, a$$hole |
Caught | Missed | Caught | Caught |
| Leet: 8→b | 8itch |
Caught | Missed | Missed | Caught |
| Leet: 6→g | ni66er |
Caught | Missed | Caught | Missed |
| Leet: #→h | s#it |
Caught | Missed | Missed | Caught |
| Phonetic: ph→f | phuck |
Caught | Missed | Missed | Missed |
| Combined | phuck3d |
Caught | Missed | Missed | Missed |
| Combined | n!66er |
Caught | Missed | Caught | Missed |
| Combined | 5h!tty |
Caught | Missed | Caught | Caught |
Summary: evasion detection rate
| terlik.js | bad-words | obscenity | allprofanity | |
|---|---|---|---|---|
| Evasion samples caught | 70/70 | 18/70 | 40/70 | 20/70 |
| Evasion detection rate | 100% | 26% | 57% | 29% |
What does this mean? If a user types "phuck you" or "f.u.c.k off", only terlik.js will catch it. The other libraries miss 43-74% of evasion attempts.
"Does it incorrectly block innocent words?"
A filter that blocks "analysis" because it contains "anal", or "grape" because it contains "rape", is worse than no filter at all — it erodes user trust.
| Library | False Positives | Wrongly blocked words |
|---|---|---|
| terlik.js | 0 | None |
| bad-words | 0 | None |
| allprofanity | 13 | "spunky on sale now", "quality cockney for sale", "homogenous on sale now", "quality scunthorpe for sale", "Quality massage for sale", "trespass on sale now", "drape on sale now", "cockatiel on sale now", "quality shell for sale", "quality bassist for sale", "quality tycoon for sale", "quality piston for sale", "hello on sale now" |
| obscenity | 37 | "she moved to penistone last year", "the pussywillow bloomed in spring", "the pussycat slept all day", "We enjoyed the cockpit very much", "Latest pussyfoot technology", "cockpit how to guide", "my grandmother makes the best vandyke recipe", "everyone loves rapeseed", "Penistone wildlife is incredible", "my favorite vandyke recipe", "my neighbor recommended a great penistone brand", "pussywillow is out of stock", "shiitake is worth trying", "the book about shiitake was a bestseller", "volunteers organized a rapeseed drive for the holiday season", "my grandmother makes the best shiitake recipe", "Cockpit regulations updated", "the best rapeseed in town", "the best vandyke in town", "the annual rapeseed conference is held in September", "pussyfoot exam results are out", "I just discovered vandyke", "vandyke report is ready", "shiitake season has started", "I ordered rapeseed online", "I bought a pussywillow today", "We visited the penistone factory during our trip", "Penistone has become more affordable in recent years", "the children learned about rapeseed during their field trip", "She received an award for her contributions to rapeseed r...", "our teacher explained pussyfoot in great detail", "penistone festival happening this weekend", "I watched a tutorial on how to make vandyke at home", "studying pussycat at university", "the pussyfoot festival attracted thousands of visitors", "studying rapeseed at university", "rapeseed park is worth visiting" |
How does terlik.js avoid false positives with 138 roots? A 105-entry whitelist explicitly protects words like:
| Tricky word | Contains | Why it's safe |
|---|---|---|
| analysis | anal | Academic term |
| therapist | rape | Medical profession |
| grape | rape | Fruit |
| hello | hell | Common greeting |
| title | tit | Document element |
| pussywillow | pussy | Plant species |
| screwdriver | screw | Tool |
| cocktail | cock | Beverage |
| penistone | penis | English town name |
| assassin | ass | Historical term |
| cucumber | cum | Vegetable |
All 70 whitelist trap sentences in the dataset pass correctly.
"What exactly does each library get wrong?"
| Error type | Count | Examples |
|---|
| Error type | Count | Examples |
|---|---|---|
| Missed plain profanity | 7 | "that is bullshit", "they had an orgy", "that hooker on the corner" and 4 more |
| Missed variant forms | 22 | "she is a bitchy person", "stop being an asshat", "you dipshit" and 19 more |
| Missed leet speak | 12 | "f#ck you", "a$$hole right here", "f@ck off" and 9 more |
| Missed separators | 14 | "f.u.c.k you", "s h i t man", "a-s-s-h-o-l-e" and 11 more |
| Missed repetition | 15 | "fuuuuck this", "shiiiit dude", "what an assshole" and 12 more |
| Missed combined | 11 | "$h!t posting", "f#cking idiot", "b.1.t.c.h mode" and 8 more |
| Missed edge cases | 1 | "RaPe MiXeD cAsE" |
bad-words uses a simple word list. It can't handle any form of evasion — not leet speak, not separators, not repetition. Over half of all profane messages slip through.
| Error type | Count | Examples |
|---|---|---|
| False positives | 37 | "she moved to penistone last year", "the pussywillow bloomed in spring", "the pussycat slept all day" and 34 more |
| Missed plain profanity | 8 | "what the hell is this", "damn it all", "go to hell" and 5 more |
| Missed variant forms | 9 | "bullshitting everyone", "she is a dumbass", "jackass move dude" and 6 more |
| Missed leet speak | 5 | "phuck you", "8itch slap", "s#it stain" and 2 more |
| Missed separators | 15 | "f.u.c.k you", "s h i t man", "a-s-s-h-o-l-e" and 12 more |
| Missed repetition | 5 | "daaaamn son", "craaap quality", "helllll no" and 2 more |
| Missed combined | 5 | "F.U.C.K.I.N.G hell", "b.1.t.c.h mode", "phuck1ng loser" and 2 more |
| Missed edge cases | 1 | "embedded hell in the middle of a long clean sentence abou..." |
obscenity is the strongest competitor. It handles some leet speak and repetition via transformer pipeline, but has no separator detection at all and generates false positives due to lack of whitelist.
| Error type | Count | Examples |
|---|---|---|
| False positives | 13 | "spunky on sale now", "quality cockney for sale", "homogenous on sale now" and 10 more |
| Missed plain profanity | 13 | "what the hell is this", "damn it all", "go to hell" and 10 more |
| Missed variant forms | 29 | "she is a bitchy person", "stop being an asshat", "you dipshit" and 26 more |
| Missed leet speak | 11 | "f#ck you", "f@ck off", "what a pr1ck" and 8 more |
| Missed separators | 15 | "f.u.c.k you", "s h i t man", "a-s-s-h-o-l-e" and 12 more |
| Missed repetition | 15 | "fuuuuck this", "shiiiit dude", "what an assshole" and 12 more |
| Missed combined | 9 | "F.U.C.K.I.N.G hell", "f#cking idiot", "b.1.t.c.h mode" and 6 more |
| Missed edge cases | 1 | "embedded hell in the middle of a long clean sentence abou..." |
allprofanity's Aho-Corasick engine is fast but rigid. It matches exact strings from its word list and can't handle any morphological or evasion variation.
"How fast can each library process messages?"
Measured in operations per second (higher = better). Each operation processes a batch of 100 messages.
Detection speed (check/containsProfanity):
| Library | ops/sec | avg latency | p50 | p95 | p99 | vs terlik.js |
|---|---|---|---|---|---|---|
| obscenity | 70,309 | 1,422 μs | 1,410 μs | 1,496 μs | 1,593 μs | ~1.92x faster |
| allprofanity | 47,131 | 2,122 μs | 2,106 μs | 2,209 μs | 2,294 μs | ~1.29x faster |
| terlik.js | 36,627 | 2,729 μs | 2,723 μs | 2,835 μs | 2,960 μs | — |
| bad-words | 3,076 | 32,510 μs | 32,418 μs | 33,093 μs | 34,177 μs | 11.9x slower |
Cleaning speed (clean/censor):
| Library | ops/sec | avg latency | p50 | p95 | p99 | vs terlik.js |
|---|---|---|---|---|---|---|
| obscenity | 48,481 | 2,063 μs | 2,049 μs | 2,146 μs | 2,252 μs | ~1.34x faster |
| allprofanity | 47,278 | 2,115 μs | 2,105 μs | 2,182 μs | 2,227 μs | ~1.30x faster |
| terlik.js | 36,260 | 2,757 μs | 2,746 μs | 2,888 μs | 3,093 μs | — |
| bad-words | 607 | 164,826 μs | 164,422 μs | 168,109 μs | 171,898 μs | 59.7x slower |
What do these numbers mean in practice? At 37K ops/sec, terlik.js can check ~36,627 messages per second on a single core. For a chat app with 1,000 concurrent users sending 1 message/second, that's 37x headroom. obscenity edges ahead on raw check() by ~92%, but terlik.js is -25% faster on clean() and catches 30% more profanity (100% vs 70% recall). bad-words would need 12 cores for the same load.
p95/p99 explained: p95 = 95% of batches complete within this time. p99 = 99%. Low p99 means consistent performance without random spikes.
terlik.js runs a multi-pass detection pipeline that trades raw speed for detection depth. Here's what each layer costs and what it catches:
| Detection Layer | Throughput Cost | What It Catches | Can You Disable It? |
|---|---|---|---|
| NFKD decomposition | ~3% | Fullwidth evasion (fuck), accented tricks (fück) | No — safety layer |
| Cyrillic confusable mapping | ~1% | Cyrillic а→a, о→o substitution | No — safety layer |
| Combining mark stripping | ~1% | Diacritics used for evasion (s̈h̊ït) | No — safety layer |
| Locale-aware lowercase | ~1% | İ→i, I→ı (Turkish locale correctness) | No — safety layer |
| Leet decode + number expansion | ~5% | $1kt1r, f#ck, 8itch, s2k | Yes: disableLeetDecode: true |
| CamelCase decompounding | ~3% | ShitLord, FuckYou (non-variant compounds) | Yes: disableCompound: true |
| Suffix chain matching | ~3% | Suffixed forms (siktirci, fuckers) | No — core detection |
Do NOT disable safety layers for performance. The ~6% total cost of NFKD + Cyrillic + diacritics + locale lowercase is non-negotiable — without them, trivial Unicode tricks bypass the entire filter. These layers run on every input regardless of toggle settings.
| Toggle | Use Case | Performance Gain | Risk |
|---|---|---|---|
disableLeetDecode: true |
Input is from a controlled source (e.g., CMS with restricted charset, database field names) | ~5% faster | Misses $1kt1r, 8itch, f#ck, number-embedded evasions |
disableCompound: true |
Camel case is intentional (e.g., code identifiers like ShittyClassName) |
~3% faster | Misses novel compounds not in dictionary (ShitLord, FuckBoy if not variant) |
minSeverity: "medium" |
Skip mild words (damn, crap, hell) in casual contexts | Marginal (pattern-level skip) | Low-severity profanity passes through |
excludeCategories: ["sexual"] |
Allow general/insult words but block sexual content | Marginal (pattern-level skip) | Selected category passes through |
Important: disableLeetDecode does NOT disable all substitution detection. The regex charClass patterns (Pass 1) still catch some visual substitutions like $ → s. The toggle only disables the normalizer's decode stages (leet map + number expansion). If a user says "I set disableLeetDecode but $hit is still caught" — that's by design. The charClass patterns are part of the core regex engine, not the normalizer.
"How much RAM does each library consume?"
Heap and RSS delta measured in KB:
| Library | Init Heap | Init RSS | After 2K msgs Heap | After 2K msgs RSS | Notes |
|---|---|---|---|---|---|
| terlik.js | +1,448 KB | +32 KB | +6,480 KB | +328 KB | GC reclaims memory efficiently |
| bad-words | +204 KB | +0 KB | -6,780 KB | +48 KB | Heap grows with usage |
| obscenity | +188 KB | +52 KB | +5,910 KB | +36 KB | Lightweight, stable |
| allprofanity | +1,319 KB | +124 KB | -18,951 KB | -28,076 KB | Large init (loads 2 dictionaries) |
What does this mean?
- Init Heap: Memory used just to load the library. All are lightweight (~170-220 KB) except allprofanity (1.3 MB, loads English + Hindi dictionaries by default).
- After processing: terlik.js and allprofanity show negative heap deltas — V8's garbage collector reclaims memory after JIT optimization. bad-words accumulates ~9.5 MB after processing 2,000 messages.
- Run with
node --expose-gcfor precise measurements.
"How much does this add to my application?"
| Library | Raw size | Gzipped | Dependencies |
|---|---|---|---|
| terlik.js | ~67 KB | ~14 KB | 0 |
| obscenity | ~45 KB | ~8 KB | 0 |
| bad-words | ~15 KB | ~3 KB | 1 (badwords-list) |
| allprofanity | ~80 KB | ~15 KB | 2 |
Dictionary expansion impact (v2.4.1 → current):
| v2.4.1 | Current | Delta | |
|---|---|---|---|
| English roots | 36 | 138 | +102 |
| Explicit variants | 139 | 342 | +203 |
| Whitelist entries | 60 | 105 | +45 |
| Suffix rules | 8 | 90 | +82 |
| ESM raw | ~57 KB | ~67 KB | +10 KB |
| ESM gzip | ~12 KB | ~14 KB | +2 KB |
The English dictionary expansion (102 new roots, 203 variants, 45 whitelist entries, 82 suffix rules) plus multi-pass normalization features added only +2 KB gzipped.
"Every library makes trade-offs. Here's where each one sits."
| High Precision | High Recall | Fast | Small | Evasion-proof | |
|---|---|---|---|---|---|
| terlik.js | Yes | No | Yes | Yes | Yes |
| obscenity | No | No | Yes | Yes | Partially |
| bad-words | Yes | No | No | Yes | No |
| allprofanity | No | No | Yes | Yes | No |
In plain English:
- terlik.js — Highest F1, perfect precision, zero false positives, competitive speed, zero dependencies. The only library that handles all evasion types. Multi-pass detection pipeline trades raw speed for detection depth.
- obscenity — Decent leet handling, fastest raw throughput. Can't handle separators and has no whitelist (causes false positives). Closest competitor on speed, distant on accuracy.
- bad-words — Zero false positives, but misses the majority of profanity. Any determined user will bypass it. Significantly slower than all competitors.
- allprofanity — Similar to bad-words but with more language support. Weak evasion handling. Fast init but large memory footprint.
The detection pipeline explains why terlik.js catches what others miss:
Input: "phuck y0u m0th3rfuck3r"
│
├─ Pass 0: 10-stage normalizer
│ ├─ 1. Strip invisible chars (ZWSP, soft hyphen, etc.)
│ ├─ 2. NFKD decompose (fuck → fuck, accented → base)
│ ├─ 3. Strip combining marks (diacritics)
│ ├─ 4. Locale-aware lowercase
│ ├─ 5. Cyrillic confusable → Latin (Cyrillic а→a, о→o, etc.)
│ ├─ 6. Char folding (ü→u, ö→o, etc.)
│ ├─ 7. Number expansion (s2k → sikik in TR)
│ ├─ 8. Leet decode (0→o, 3→e, $→s)
│ ├─ 9. Punctuation removal (f.u.c.k → fuck)
│ └─ 10. Repeat collapse (fuuuck → fuck)
│ Result: "phuck you motherfucker"
│
├─ Pass 1: Locale-lowered regex matching
│ charClasses: f=[fph], so "ph" → matches "f" pattern
│ "phuck" matches /[fph][uv][c][k]/ → ✅ fuck
│
├─ Pass 2: Normalized text regex matching
│ "motherfucker" → ✅ direct variant match
│
├─ Pass 3: CamelCase decompounding (ShitLord → Shit Lord)
│
├─ Whitelist check → neither word is whitelisted
└─ Result: 2 matches found
Other libraries fail at different stages:
- bad-words has no leet decode, no separator removal, no pattern matching
- obscenity has leet decode + some patterns, but no separator removal and no phonetic matching (ph→f)
- allprofanity has basic matching but no normalization pipeline
Leading on 1280 samples doesn't mean perfection in the wild. Known limitations:
| Limitation | Example | Status |
|---|---|---|
| Context-dependent words | "hell yeah" is flagged (but user may consider it mild) | By design — severity: "low" |
| 3-letter root FP risk | "cum" in "cucumber" | Protected by whitelist, but new compounds may appear |
| Rapidly evolving slang | New coinages not in dictionary | Use addWords() or customList at runtime |
| Non-Latin scripts | Cyrillic "а" substituted for Latin "a" | Handled (Cyrillic confusable → Latin mapping) |
| Adversarial transforms | Zalgo, zero-width chars, unicode homoglyphs | Partially handled — some extreme evasions reduce recall |
| Dataset composition | 290 curated + 990 SPDG-generated adversarial samples | Continuously expanding via SPDG |
We document every limitation and encourage the community to report edge cases.
| Parameter | Value |
|---|---|
| Warmup iterations | 100 (V8 JIT compilation) |
| Measurement iterations | 1,000 |
| Batch size | 100 messages per iteration |
| Latency metric | Per-batch microsecond timing |
| Percentiles | p50, p95, p99 |
| Memory measurement | process.memoryUsage() delta |
| Isolation | Each library in independent adapter |
| Fairness | Same corpus, same machine, same process |
Each library is wrapped in a LibraryAdapter interface with init(), check(), and clean() methods:
| Library | check() API | clean() API |
|---|---|---|
| terlik.js | containsProfanity() |
clean() |
| bad-words | isProfane() / clean() !== text |
clean() |
| obscenity | matcher.hasMatch() |
censor.applyTo() |
| allprofanity | check() |
clean() |
Adapters: benchmarks/comparison/adapters.ts
| Limitation | Impact |
|---|---|
| English only | Doesn't test TR/ES/DE — only the common denominator |
| 1280 samples | 290 curated + 990 SPDG adversarial. Comprehensive but not exhaustive |
| Self-reported | Maintained by terlik.js team. We aim for fairness but acknowledge potential bias |
| Single machine | Absolute ops/sec numbers depend on hardware. Relative ranking matters more |
| Default configs | All libraries tested with default settings. Custom tuning may improve results |
| bad-words v3 | Tested v3.0.4 (latest on npm). A v4 may exist but npm resolves to v3 |
Prerequisites: Node.js >= 18, pnpm
# From the repository root:
pnpm build # Build terlik.js first
pnpm bench:report # Build + install deps + run benchmark + generate this report
# Or manually:
cd benchmarks/comparison
pnpm install
npx tsx run.ts # Standard run
node --expose-gc --import tsx/esm run.ts # With accurate memoryOutput:
- Formatted tables to the console
- JSON results to
benchmarks/comparison/results/comparison-results.json - This markdown report to
docs/benchmark-comparison.md
| Task | How |
|---|---|
| Add test cases | Edit dataset.ts — add { text, profane, category } |
| Add a library | Create adapter in adapters.ts, add factory to run.ts |
| Change iterations | Edit WARMUP_ITERATIONS / MEASURE_ITERATIONS in throughput.ts |
Full JSON output with per-sample error lists: benchmarks/comparison/results/comparison-results.json
Found an issue with the methodology? Think the dataset is unfair? Please open an issue or PR. We want this comparison to be as fair and transparent as possible.