GHSA-gqvg-gmmx-x4hm is a high-severity (CVSS 8.8) remote code execution vulnerability in mlflow. O3 Security confirms whether GHSA-gqvg-gmmx-x4hm is actually reachable in your code before you act, and blocks exploitation at runtime until you patch.
MLFLOW_ALLOW_PICKLE_DESERIALIZATION=False safety control bypassed by mlflow.statsmodels flavor — RCE via crafted model artifact
Real-World Exposure
mlflowReal-time download stats are indexed for npm and PyPI packages. This vulnerability affects PyPI packages — download data is not available via public APIs for these ecosystems.
Description
Summary
MLflow introduced MLFLOW_ALLOW_PICKLE_DESERIALIZATION as a security control to prevent unsafe pickle.load execution during model loading, in response to CVE-2024-37052 through CVE-2024-37060. When set to False, operators expect all pickle deserialization to be blocked. The most recent related fix (#21188) patched a bypass in the pyfunc flavor.
However, the mlflow.statsmodels flavor completely omits this guard. An attacker who places a crafted MLmodel artifact into any accessible artifact store can trigger arbitrary code execution on any process that calls mlflow.pyfunc.load_model() against the malicious model — even when MLFLOW_ALLOW_PICKLE_DESERIALIZATION=False.
This is a security control bypass. The operator believes pickle RCE is mitigated; the statsmodels flavor silently ignores the control.
Root Cause
mlflow.pyfunc.load_model() dispatches to flavor _load_pyfunc implementations via:
# mlflow/pyfunc/__init__.py L1170-1172
model_impl = importlib.import_module(conf[MAIN])._load_pyfunc(data_path)
The guarded pattern (from mlflow/sklearn/__init__.py L526-533, the reference implementation) is:
if (
not MLFLOW_ALLOW_PICKLE_DESERIALIZATION.get()
and not is_in_databricks_runtime()
and not is_in_databricks_model_serving_environment()
):
raise MlflowException("Deserializing model using pickle is disallowed...")
mlflow/statsmodels/__init__.py has no such check:
# L307-320 — no guard anywhere in this file
def _load_model(path):
import statsmodels.iolib.api as smio
return smio.load_pickle(path) # calls pickle.load() directly
def _load_pyfunc(path):
return _StatsmodelsModelWrapper(_load_model(path))
statsmodels.iolib.api.load_pickle is a thin wrapper around pickle.load. Its own docstring warns: "Never unpickle data received from an untrusted or unauthenticated source."
Trigger
An attacker crafts an MLmodel YAML that specifies mlflow.statsmodels as the loader module:
flavors:
python_function:
loader_module: mlflow.statsmodels
data: model.pkl
statsmodels:
data: model.pkl
statsmodels_version: 0.14.0
With a malicious model.pkl placed alongside it in the artifact store, any call to:
os.environ["MLFLOW_ALLOW_PICKLE_DESERIALIZATION"] = "False"
mlflow.pyfunc.load_model("models:/MaliciousModel/1")
...deserializes the pickle file with no guard check, executing arbitrary code with the privileges of the calling process.
On default MLflow deployments (no --app-name basic-auth), authentication is disabled, so artifact upload requires no credentials.
Affected Code
mlflow/statsmodels/__init__.pyL307-310:_load_model— callssmio.load_picklewithout checkingMLFLOW_ALLOW_PICKLE_DESERIALIZATIONmlflow/statsmodels/__init__.pyL313-320:_load_pyfunc— dispatches to_load_modelwithout checking the control
Permalink (commit 0b0c576c):
- https://github.com/mlflow/mlflow/blob/0b0c576c642b5b0d9496c829809c7d097403bc9f/mlflow/statsmodels/__init__.py#L307-L310
- https://github.com/mlflow/mlflow/blob/0b0c576c642b5b0d9496c829809c7d097403bc9f/mlflow/statsmodels/__init__.py#L313-L320
Recommended Fix
Add the missing guard to mlflow/statsmodels/__init__.py:
from mlflow.environment_variables import MLFLOW_ALLOW_PICKLE_DESERIALIZATION
from mlflow.utils.databricks_utils import (
is_in_databricks_model_serving_environment,
is_in_databricks_runtime,
)
def _load_model(path):
if (
not MLFLOW_ALLOW_PICKLE_DESERIALIZATION.get()
and not is_in_databricks_runtime()
and not is_in_databricks_model_serving_environment()
):
raise MlflowException(
"Deserializing model using pickle is disallowed, but this statsmodels "
"model requires pickle deserialization. Set environment variable "
"'MLFLOW_ALLOW_PICKLE_DESERIALIZATION' to 'true' to allow this."
)
import statsmodels.iolib.api as smio
return smio.load_pickle(path)
Affected Packages
| Ecosystem | Package | Vulnerable range | Fix |
|---|---|---|---|
| 🐍PyPI | mlflow | ≥ 2.1.0&&< 3.15.0 | 3.15.0 |
Detection & mitigation playbook
Open-source dependencyDetect
Scan your dependency tree (package-lock.json, pnpm-lock.yaml, requirements.txt, go.sum, etc.) for mlflow. O3's reachability analysis confirms whether the vulnerable code path is actually invoked in your application, so you act on real exposure instead of every transitive match.
Fix
Update mlflow to 3.15.0 or later, then make sure no transitive (indirect) dependency still pins the vulnerable range — O3 confirms GHSA-gqvg-gmmx-x4hm is resolved across your whole dependency graph.
Workarounds
If you can't upgrade right away: gate or disable the affected feature, validate untrusted input at the boundary, and avoid passing attacker-controlled data into the vulnerable path. O3's runtime protection blocks exploitation in production as an interim safeguard until the upgrade lands.
How O3 protects you
O3 pinpoints whether GHSA-gqvg-gmmx-x4hm is reachable in your code and exactly where to fix it, then blocks exploitation in production at runtime until the patched version is deployed.
Tailored to GHSA-gqvg-gmmx-x4hm. Runtime protection reduces exposure until a permanent patch is applied and verified — it complements patching, it doesn't replace it.
Frequently Asked Questions
Is GHSA-gqvg-gmmx-x4hm in your dependencies?
O3 detects GHSA-gqvg-gmmx-x4hm across PyPI dependencies and uses function-level reachability to confirm whether the vulnerable code path is actually reachable — not just present. No false positives.