Your RSA-2048 keys break in 2030. Find every one of them before attackers do.
🐍
🐍 PyPI
Not in CISA KEV
CRITICAL severity

GHSA-x2rj-828p-hx9m — xinference

CRITICALFix: xorbitsai/inference#4786

GHSA-x2rj-828p-hx9m is a critical-severity (CVSS 10) CWE-95 vulnerability in xinference. A fix is available for xinference — see the affected versions and patch details below.

Xinference vulnerable to remote code execution via unsafe `eval()` in Llama3 tool-call parsing

Also known asCVE-2026-61539PYSEC-2026-3946
Published
Updated
Affected
1 pkg
Patched
1 / 1
Exploits
None indexed
Exploitation data as of Oct 4, 2026 · OSV.dev, NVD, FIRST.org (EPSS)

Exploitation Status

No confirmed exploitation observed yet

  • CISA assesses this as automatable — exploitation doesn’t require manual, per-target effort, which raises the odds of mass scanning and opportunistic attacks.
  • A successful exploit gives an attacker total control of the affected component, not partial access.
  • CISA’s own triage has not observed active exploitation or public proof-of-concept code for this CVE as of its last assessment.

Exploitation and automatability from CISA’s SSVC triage for GHSA-x2rj-828p-hx9m.

EPSS Exploitation Probability

via FIRST.org ↗
1.2%probability of exploitation in next 30 days
Lower Risk0.00%
Lower risk than most CVEs68th percentile — riskier than 68% of all scored CVEsHighest risk
0.16%0.68%1.20%1.73%0.7%1.2%1.2%Sep 26Oct 26Oct 26

Probability of exploitation in the next 30 days, from FIRST.org EPSS.

How urgent is this, really

GHSA-x2rj-828p-hx9m by exploitation likelihood (EPSS) against impact (CVSS). Outside the shaded patch-first corner.

Where this sits among everything scored

Of 382,795 CVEs with a current EPSS score, this one falls in the < 10% band (highlighted). Counts from FIRST.org, log-scaled.

Real-World Exposure

1 pkg affected
🐍xinference

Real-time download stats are indexed for npm and PyPI packages. This vulnerability affects PyPI packages — download data is not available via public APIs for these ecosystems.

Description

Summary

Xinference used Python's unsafe eval() function when parsing Llama3 tool-call output generated by a large language model. Because the model output can be influenced by attacker-controlled prompts sent to the chat completion API, a remote attacker can craft prompts that cause the model to return a Python expression. Xinference then evaluates that expression on the server while post-processing the tool-call result. In the tested default deployment, authentication was not enabled, so the vulnerability was exploitable by an unauthenticated remote attacker through the /v1/chat/completions endpoint.

Details

Users can interact with deployed models through Xinference's OpenAI-compatible /v1/chat/completions API. The request entry point is implemented in xinference/api/restful_api.py; non-streaming requests call the model instance's chat() method and return the inference result.

When the Transformers backend is used, inference results flow through the batching logic in xinference/model/llm/transformers/core.py. Non-streaming chat results are handled by handle_chat_result_non_streaming(). If the request contains a tools field, Xinference calls _post_process_completion() to parse tool-call output from the model response.

The Llama3 tool-call parser is implemented in xinference/model/llm/tool_parsers/llama3_tool_parser.py. In affected versions, extract_tool_calls() parsed model output with eval():

def extract_tool_calls(
    self, model_output: str
) -> List[Tuple[Optional[str], Optional[str], Optional[Dict[str, Any]]]]:
    try:
        data = eval(model_output, {}, {})
        return [(None, data["name"], data["parameters"])]
    except Exception:
        return [(model_output, None, None)]

The intended behavior was to convert a Python dictionary-like string generated by the model into a dictionary object. However, eval() executes the input as a Python expression, and eval(model_output, {}, {}) is not a security sandbox. If an attacker can influence the model output through prompt injection or direct chat input, the attacker can cause the model to return an expression such as:

__import__('os').system('touch /tmp/hacked')

When the expression reaches eval(), it is executed in the Xinference server process context. The harmless touch /tmp/hacked command can be replaced with other payloads, such as a reverse shell, malware download, sensitive file read, or lateral-movement payload.

Score

Severity: Critical

CVSS v3.1: 10.0

Vector: CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:C/C:H/I:H/A:H

Rationale:

  • AV:N: the vulnerable API is remotely reachable over the network;
  • AC:L: exploitation only requires a crafted chat-completion request and tool-call parameter;
  • PR:N: the tested default configuration did not require authentication;
  • UI:N: no user interaction is required;
  • S:C: command execution can affect resources beyond the Xinference application boundary;
  • C:H/I:H/A:H: remote code execution can fully compromise confidentiality, integrity, and availability.

Credit

This vulnerability was discovered by:

Affected Packages

1 total 1 fixed
EcosystemPackageVulnerable rangeFix
🐍PyPIxinferenceall versions2.7.0pip install --upgrade 'xinference==2.7.0'

Detection & mitigation playbook

Open-source dependency
  1. Detect

    Scan your dependency tree (package-lock.json, pnpm-lock.yaml, requirements.txt, go.sum, etc.) for xinference, including transitive dependencies — a direct dependency you never call can still pull in a vulnerable version.

  2. Fix

    Update xinference to 2.7.0 or later, then make sure no transitive (indirect) dependency still pins the vulnerable range — O3 confirms GHSA-x2rj-828p-hx9m is resolved across your whole dependency graph.

  3. Workarounds

    If you can't upgrade right away: gate or disable the affected feature, validate untrusted input at the boundary, and avoid passing attacker-controlled data into the vulnerable path. O3's runtime protection blocks exploitation in production as an interim safeguard until the upgrade lands.

Frequently Asked Questions

### Summary Xinference used Python's unsafe `eval()` function when parsing Llama3 tool-call output generated by a large language model. Because the model output can be influenced by attacker-controlled prompts sent to the chat completion API, a remote attacker can craft prompts that cause the model to return a Python expression. Xinference then evaluates that expression on the server while post-processing the tool-call result. In the tested default deployment, authentication was not enabled, so the vulnerability was exploitable by an unauthenticated remote attacker through the `/v1/chat/compl
O3 Security · Impact-Aware SCA

Is GHSA-x2rj-828p-hx9m in your dependencies?

Find it across PyPI, including transitive dependencies.

GHSA-x2rj-828p-hx9m: RCE — Fixed in 2.7.0 | O3 Security