A critical vulnerability has been discovered in vLLM's weight caching feature, specifically in version 0.30.0. This bug allows a running model daemon to serve weights from a different checkpoint than the one requested by an engine, provided the tensor layouts match. This can lead to the engine producing plausible but incorrect outputs, as the system incorrectly assumes the cached weights are for the requested model. The issue arises because the weight cache's fingerprinting mechanism only hashes metadata like tensor names and shapes, not the actual tensor values, making it susceptible to mismatches when only the weights differ. AI
IMPACT This vulnerability could lead to incorrect model outputs in production systems using vLLM's weight caching, potentially impacting applications that rely on accurate AI responses.
RANK_REASON Discovery of a bug in a specific version of an open-source inference engine.
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →