Skip to content

fix(core): reset client-supplied is_builtin on custom model registration - #5530

Merged
qinxuye merged 1 commit into
xorbitsai:mainfrom
kah-ja:fix/register-model-is-builtin-reset
Sep 12, 2026
Merged

qinxuye merged 1 commit into
xorbitsai:mainfrom
kah-ja:fix/register-model-is-builtin-reset

Conversation

@kah-ja

@kah-ja kah-ja commented Sep 12, 2026

Copy link
Copy Markdown
Contributor

What

Supervisor.register_model (xinference/core/supervisor.py:2200) and Worker.register_model (xinference/core/worker.py:2357) build the custom model spec with model_spec_cls.parse_raw(model), applying the caller's JSON as is. CustomLLMFamilyV2 (xinference/model/llm/llm_family.py:202), CustomEmbeddingModelFamilyV2 (xinference/model/embedding/core.py:105), the rerank equivalent (xinference/model/rerank/core.py:87) and CustomImageModelFamilyV2 (xinference/model/image/core.py:58) declare is_builtin: bool = False as a plain field with no write protection, so a registration request body containing "is_builtin": true produces a spec with that field set regardless of the Config.extra setting.

Why

allow_trust_remote_code() (xinference/model/utils.py:3261) returns true when XINFERENCE_TRUST_REMOTE_CODE is set or when getattr(model_family, "is_builtin", False) is true, and loaders such as xinference/model/embedding/sentence_transformers/core.py:303 pass its result straight into trust_remote_code. is_builtin marks models the project itself loaded from the bundled registry: xinference/model/llm/__init__.py:308 and xinference/model/utils.py:3234 set family.is_builtin = True right after loading a real built-in family, never from external input. register_model is the one path where a caller-controlled value reaches the same field unchecked.

How

Both register_model methods reset model_spec.is_builtin = False right after parse_raw(), before the spec reaches register_fn or the cache, guarded by hasattr so model types without an is_builtin field (audio, flexible) stay untouched; video and world are not registrable through this code path at all. The reset is duplicated in worker.py because Supervisor.register_model forwards the raw, unmodified request body to every worker, both through _sync_register_model's cluster-wide fan-out and through an explicit worker_ip target, without constructing a reset spec on the worker side, so the same gate is needed there independently. Nothing else in either function changed. Not covered by this change: register_custom_model() in each xinference/model/<type>/__init__.py reloads persisted registration JSON from XINFERENCE_MODEL_DIR at process start via parse_raw and calls the register function directly, so a file written before this fix keeps its is_builtin value on restart.

Testing

Environment: fresh clone at origin/main 99868ea7, editable install with pip install -e ".[dev]" (CPU-only torch wheel).

  • pytest -vv xinference/core/tests/test_register_model_is_builtin.py (new test): 2 passed with the patch applied.
  • Same command against the unpatched base (git checkout origin/main -- xinference/core/supervisor.py xinference/core/worker.py): 2 failed, both asserting is_builtin is False where the actual value was True, confirming the reset is what the test exercises.
  • pre-commit run --files xinference/core/supervisor.py xinference/core/worker.py xinference/core/tests/test_register_model_is_builtin.py: black, end-of-file-fixer, trailing-whitespace, ruff check, isort, mypy, codespell all passed.
  • pytest -vv xinference/model/llm/tests/test_llm_family.py: 69 passed, 2 skipped, 1 failed (test_query_engine_general, missing llama.cpp engine entry). Same failure reproduces on the unpatched base with the identical command, so it predates this change (no llama-cpp engine installed in this environment).
  • Not run: the full non-GPU CI test command from AGENTS.md and xinference/core/tests/test_restful_api.py, which need a running supervisor/worker cluster and heavier optional dependencies than this environment has.

References

GHSA-v9h5-42jh-j8h5 (found during penetration test by turingpoint, reported privately, unpublished at time of writing)

CustomLLMFamilyV2, CustomEmbeddingModelFamilyV2 and CustomRerankModelFamilyV2
declare is_builtin with Config.extra = "allow", so parse_raw() in
Supervisor.register_model and Worker.register_model applies a caller-supplied
is_builtin value verbatim. allow_trust_remote_code() treats is_builtin as
proof that a model family was vetted and loaded from the bundled registry,
so a client that sets is_builtin=true on its own custom registration gets
the same trust_remote_code grant as a real built-in model.

Both register_model methods now reset is_builtin to False right after
parse_raw(), before the spec is handed to register_fn or cached. This
mirrors the existing family.is_builtin = True assignment done for real
built-ins in model/utils.py and llm/__init__.py. Model types without an
is_builtin field are unaffected.
@XprobeBot XprobeBot added bug Something isn't working gpu labels Sep 12, 2026
@XprobeBot XprobeBot added this to the v3.x milestone Sep 12, 2026

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request ensures that client-submitted model registrations cannot falsely claim to be built-in by resetting the 'is_builtin' attribute to 'False' during parsing in both the supervisor and worker. A security review comment suggests re-serializing the sanitized 'model_spec' back to the 'model' string in the supervisor to prevent forwarding the unsanitized payload to workers.

Comment thread xinference/core/supervisor.py

@qinxuye qinxuye left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM. Reviewed 89b9dbe; 132 focused tests passed locally. Additionally verified targeted-worker forwarding, the remote-code trust gate, and sanitized persistence serialization. The defense-in-depth review thread has been addressed and resolved. Existing persisted registrations are outside the stated scope of this patch.

@qinxuye
qinxuye merged commit 2ab499e into xorbitsai:main Sep 12, 2026
14 of 15 checks passed
@Su1ph3r

Su1ph3r commented Oct 6, 2026 •

Copy link
Copy Markdown

Hi, I'm the reporter of the issue this PR fixes, filed privately as
GHSA-jv65-rccw-j2j6 on 9 July. Thanks for shipping the fix in v3.5.0.

The advisory is still in triage and unpublished, so there's no public record telling
operators on 3.4.0 and earlier that they need to upgrade. Two things only maintainers can
do, and both would help:

  1. Publish the advisory. Nothing in it is sensitive any more. The fix is in v3.5.0 and
    the diff here is public.
  2. Hit "Request CVE" on it. The Request CVE action is maintainer and owner only, so I
    can't trigger it. GitHub has issued CVEs for this repo before (CVE-2026-61539 was
    assigned by GitHub as CNA), so publishing with that action is the cleanest path to an
    ID. Failing that I'd fall back to a third-party MITRE request, as I noted on the
    advisory in August, but an ID attached to your own advisory is better for your users.

Correcting my earlier note on versions. I have now tested this against running
servers rather than reasoning from the source, and I had two things tangled. Each
result below is a pair of requests that are identical except for the single
is_builtin field, so the control shows what the gate does when it is absent.

Unauthenticated by default: <= 2.12.0. v3.0.0 added the advanced auth system and
is_auth_advanced() defaults XINFERENCE_AUTH_ADVANCED to true. On a stock 3.4.0
server, GET /v1/models, POST /v1/model_registrations/embedding and POST /v1/models
all return 401 and nothing executes. So my earlier suggestion to widen the range to
<= 3.4.0 was wrong, and the draft's existing <= 2.12.0 is right for the
unauthenticated framing.

The is_builtin forgery is present and exploitable through 3.4.0, fixed by this PR.
Where it fires, with no credentials:

  • 2.12.0 default config: control 400 "Loading this model executes code shipped in the
    model repository..."
    , forged 200 and the repo's scripts/qwen3_vl_embedding.py
    executes in the model process.
  • 2.12.0 with xinference-supervisor + xinference-worker: same, and the payload line
    appears in the worker log, not the supervisor log.
  • 2.12.0 with model_uri unset and model_id set: the files are fetched via
    snapshot_download and executed, so this does not need any pre-existing path on the
    host. To be precise about method, I pointed HF_ENDPOINT at a local
    HuggingFace-compatible endpoint serving the repository rather than publishing anything
    to huggingface.co. The calls your code made were GET /api/models/<repo>/revision/main,
    then HEAD and GET on resolve/<sha>/config.json and
    resolve/<sha>/scripts/qwen3_vl_embedding.py.
  • 3.4.0 with XINFERENCE_AUTH_ADVANCED=0, which the comment beside that flag describes as
    running "with no authentication at all": same result.
  • 3.4.0 default config with an admin token: same result, so it is also an authenticated
    privilege issue, not only an unauthenticated one.

I also verified your fix holds. On 3.5.0 the forged request is refused with the same
gate error and the replica fails to launch, so model_spec.is_builtin = False closes it.

One note for the advisory text if it is useful: on xprobe/xinference:v2.12.0-aarch64
there is no USER directive, the process runs as uid=0, and it binds 0.0.0.0:9997,
so on that image the executed code runs as root.

Two limits on my own testing, so you can weigh it. First, I fired the
Qwen3-VL-Embedding* exec_module branch, not the sentence_transformers auto_map
one. Second, the control and exploit pairs above were run on pip installs of each version.
On xprobe/xinference:v2.12.0-aarch64 the forged request did execute, but I could not get
a clean paired control there, because that image's GPU probe hangs the launch on my
GPU-less test host, so treat the official image as confirming the root and bind behaviour
and that execution occurs, rather than as supplying the differential.

On timing, for the record rather than as news: I gave notice of a disclosure date on the
advisory back in August, and that date was 7 October, tomorrow. I am moving it to
Friday 9 October 2026 so there is room for you to publish the advisory and request the
CVE first, which I would much rather do than publish instead of it. The write-up covers root cause and impact
only. I won't publish anything that isn't already visible in this PR's diff, and I'm
holding back the turnkey exploit, as I said I would.

Happy to draft the advisory text or fill in the CVE fields if that saves you time.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bug Something isn't working gpu

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants