Study Finds Open-Weight LLMs Broadly Vulnerable to Prefill Attacks
The paper presents the largest systematic study to date of 'prefill attacks,' where an attacker predefines initial response tokens before generation to bypass safety training. Testing over 20 strategies across multiple open-weight model families, the authors find these attacks consistently succeed against essentially all major contemporary open-weight models, including reasoning models that show only partial robustness to generic prefilling but remain vulnerable to model-specific tailored strategies. The authors argue this exposes a critical, underexplored vulnerability demanding urgent attention from model developers.
Discussion: 2 tweets from 2 authors · @yugen_matuni, @Iamwonderingif