LMCache ships a critical, unpatched RCE reachable with a single ZeroMQ message on vLLM servers
On October 7, 2026, JFrog disclosed CVE-2026-105192, an unauthenticated remote code execution flaw rated 9.8 in LMCache, the cache layer that speeds up LLM servers such as vLLM, with no fix available. If your LMCache cache listens on a routable address, isolate it now instead of waiting for a patch.
October 7, 2026. JFrog published a flaw it rates 9.8 out of 10, the critical band. CVE-2026-105192 is an unauthenticated remote code execution bug in LMCache, the open-source cache that accelerates LLM servers such as vLLM. There is no fixed version. Why it matters: a single network message sent to the cache is enough to run commands as the process user — root on the official images — and the project’s own Kubernetes example exposes that port on every interface.
One ZeroMQ message is enough to run code
The bug lives in LMCache’s multiprocess mode. In that mode, the cache no longer runs inside the vLLM process: it becomes a standalone server that workers reach over the ZeroMQ messaging library. That server is where the vulnerability sits.
The ZeroMQ socket the server opens, so workers can register and share cached data, has no authentication. One message type is deserialized with pickle, Python’s native serialization format. But pickle is not a data format — it is an execution format. It can carry arbitrary code and run it the moment it is decoded. The server unpacks the message before checking its type, while it is still reading the message’s arguments. The result: a crafted message runs the sender’s code with the privileges of the LMCache process.
On the official container images, that process runs as root, JFrog notes. The full chain is a single step: reach the socket, send a malicious pickle message, get a root shell on the machine hosting the cache.
The real exposure comes down to one setting
One nuance changes the practical exposure. By default, the multiprocess server listens only on the loopback address (127.0.0.1), so another host cannot reach it. It becomes reachable as soon as an operator starts it on a routable address, which is exactly what multi-node deployments do when several machines share one cache.
This is where the documentation works against the user. The Kubernetes example LMCache itself ships starts the server listening on all network interfaces. An operator who follows the reference manifest without adapting it exposes the port to the whole cluster — and, depending on the topology, to anything that can reach that cluster.
The affected range is broad: CVE-2026-105192 affects LMCache from 0.3.9 (October 2025) through 0.5.5, the latest stable release, and is still present in the 0.5.6 release candidates and the development branch. No fixed version exists at publication time.
pickle, once again
The root mistake — handing data from an unauthenticated socket to pickle — is hardly new. It is the same mechanism researchers flagged in November 2025 across other inference frameworks, in a group of flaws they called ShadowMQ. The lesson is old and unanimous: never deserialize with pickle data that can come from outside, whatever the transport layer.
What makes LMCache dangerous in practice is less the technicality of the flaw than its invisibility. LMCache has published no security advisory. JFrog’s advisory gives operators no way to determine whether a server has already been attacked: no indicator of compromise, no network signature, no log to inspect. A team that learns of the flaw today cannot tell whether its cache has already served as an entry point.
An ecosystem in motion around it
The flaw does not stand alone. On October 6, 2026, the day before the disclosure, a GitHub account filed six additional reports against LMCache. They allege unauthenticated access to cached data belonging to other tenants, plus exposure of several network services that run commands without a login. These reports rest on proof-of-concept claims, with no CVE, no maintainer confirmation, and no fix. One points to a default that has since changed: the admin HTTP server that listened on all interfaces in 0.5.5 now listens only on the loopback in the 0.5.6 release candidates.
A neighboring flaw in vLLM is already closed. Before version 0.30.0 (released September 22, 2026), a single request carrying a malformed cache_salt value could crash the engine on deployments using the LMCache multiprocess connector. That denial of service, CVE-2026-105756, is rated 6.5 and does not allow code execution. It still shows how fast the attack surface moves around inference acceleration layers.
What to do right now
With no patch available, the protection is architectural, not software. First, never bind the multiprocess server to a routable address: keep it on 127.0.0.1 or, at worst, on a trusted cluster network. Second, audit your manifests: if you copied LMCache’s Kubernetes example verbatim, your server is probably listening on 0.0.0.0. Third, apply a firewall that restricts which hosts can reach the port — while keeping in mind that the filter reduces the risk without removing it, since any host still able to open a connection can run code.
For a typical single-machine deployment — LMCache running inside the vLLM process, cache not exposed — the port is never opened and the exposure is zero. The risk concentrates on multi-node deployments that share a cache across machines, and on the official images that run as root.
Why the inference layer is becoming a target
The cache is the last link anyone thinks about. LMCache is not the model — it is the plumbing that feeds it, and that is exactly what makes it dangerous. Inference acceleration components — KV caches, request routers, schedulers — get deployed in a hurry, with far less security review than the firewall or the web control plane. Their job is performance, and performance is often won by removing the steps that slow things down — authentication first among them.
LMCache’s multiprocess mode is the exact illustration. Splitting the cache out of the vLLM process speeds up inference by sharing a cache across workers, but it also creates a network service — a ZeroMQ socket — that did not exist before. A network service with no authentication is a target; a network service that runs code through pickle is a backdoor. The two together make this plumbing as serious an entry point as an exposed web server, but without the usual visibility: no WAF, no application logs, no dedicated SIEM alert.
What worsens it is the concentration of privilege. The official images run LMCache as root, which means a crafted message does not just gain access to the cache — it gains the whole machine, with the ability to pivot toward the models, the data flowing through the inference cluster, and the secrets workers share with the cache.
Verdict
If your LMCache cache listens on a routable address, treat the machine as potentially compromised: isolate it, close the port, and plan a re-assessment once a fix ships, because JFrog offers no way to prove it was not already attacked. If you run LMCache single-process inside vLLM, you are not exposed to CVE-2026-105192 — but watch the announcements, since the six reports from October 6 suggest other attack surfaces remain to be clarified. In every case, ban pickle from any path that touches an unauthenticated socket: applied upstream, that rule would have made this flaw harmless.