DeepSeek Adopts YOCO 'Decoder-Decoder' Architecture for KV Cache Efficiency
YOCO (You Only Cache Once) is a decoder-decoder LLM architecture that stores key-value pairs just once: a self-decoder efficiently builds a global KV cache, which a cross-decoder then reuses via cross-attention, mimicking a standard decoder-only Transformer's behavior while cutting memory use. This design lets prefill exit early without altering outputs, and the authors report order-of-magnitude gains in memory, prefill latency, and throughput, plus near-perfect retrieval at 1M-token context. Twitter commentary noted that DeepSeek's newly released V4.1 Flash model has reportedly adopted YOCO as a core architectural choice, seen as validation of its prefill/KV-caching advantages carrying over into production-scale systems.