EN
en direct
Élevée CVSS 7.5

CVE-2026-54234

Vllm

vLLM is a high-throughput and memory-efficient inference and serving engine for LLMs. Prior to 0.24.0, a frontend-legal multi-request speculative decoding workload can cause the rejection sampler to produce a recovered token equal to the model vocabulary size boundary value, which is then converted to negative one when the engine selects the next live token for a request and is written back into the drafter's input ids; that out-of-vocabulary value is later consumed by the model's embedding and attention path and crashes the engine worker with a GPU device-side assertion. The same triggering request sequence is reachable through the public gRPC Generate and Abort endpoints, so a remote client that can send generation requests can crash the shared engine worker, aborting concurrent requests and causing a service-wide denial of service for other clients of the deployment until the worker is restarted. This issue is fixed in version 0.24.0.

Ce que ça veut dire

Exposition
Exploitable à distance depuis le réseau, sans authentification et sans action de la victime.
Impact
Un attaquant peut mettre le service hors ligne.
Faiblesse
L’application accepte une entrée sans la contrôler, ce qui ouvre la porte à des comportements non prévus.
Probabilité
Le score EPSS reste bas : rien n’annonce une exploitation imminente, ce qui ne dispense pas de corriger.

À faireÀ intégrer au prochain cycle de correctifs. Commencer par les instances exposées à Internet.

Lecture automatique du vecteur CVSS, du type de faiblesse (CWE) et du score EPSS. La description technique ci-dessus reste celle publiée par le NIST, en anglais.

Publié
6 juillet 2026
CVSS
7.5 (v3.1) CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H
EPSS
0,36 % probabilité d'exploitation sous 30 jours · au-dessus de 29 % des CVE
Faiblesse
CWE-20CWE-1284
Sources
nvd
Références

Tapez au moins deux caractères.

naviguer ouvrir esc fermer