LightLLM Mass Disclosure — 2 CVSS 9.8 Unauthenticated RCE in LLM Serving Framework
Two unauthenticated CVSS 9.8 RCEs in LightLLM — the LLM serving framework used with LLaMA, Mistral, and Qwen deployments. No patch confirmed. All versions through 1.2.0 are affected. The CVEs CVE CVSS Type Vector CVE-2026-103040 9.8 Pickle deserialization RCE Router profiler RPyC service CVE-2026-103041 9.8 Pickle deserialization RCE Embed cache RPyC service CVE-2026-103042 7.5 Memory exhaustion…
Two unpatched critical vulnerabilities have been discovered in LightLLM, an LLM serving framework used by popular language models like LLaMA, Mistral, and Qwen. Both vulnerabilities are due to unauthenticated Remote Procedure Call (RPC) services passing untrusted input directly to Python's pickle.loads() function, which executes arbitrary code upon deserialization.
The first vulnerability, CVE-2026-103040, is a Pickle deserialization Remote Code Execution (RCE) flaw. It triggers only when the --enable_profiling flag is enabled. If this flag is unnecessary for production use, it should be avoided. The second vulnerability, CVE-2026-103041, affects multimodal deployments and is also a Pickle deserialization RCE issue. The exposed embed cache RPyC service is accessible on all interfaces by default, posing a significant security risk.
The third vulnerability, CVE-2026-103042, is a Memory Exhaustion Denial of Service (DoS) flaw. It allows unauthenticated attackers to call the exposed_set_value function on the NCCL control channel without any size limits. This rapidly consumes the memory of the KV-transfer worker, ultimately crashing the node.
The impact of these vulnerabilities is severe, as a compromised LightLLM node exposes every user prompt, model weight, and API key processed by the service. These components are crucial to production AI stacks. Due to the well-known nature of the pickle deserialization pattern, it should not be present in frameworks deployed on such a large scale.
To mitigate the risks, users should avoid enabling the --enable_profiling flag in production, firewall the RPyC ports as the services should never be exposed to the internet, and restrict the --pd_trans_mode nccl parameter to trusted nodes only. Additionally, applying memory limits through cgroups or container resource quotas can help prevent exhaustion impact. However, no patch has been confirmed for LightLLM version 1.2.0 at this time.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.