MQA and GQA shrink the KV cache, the thing that bottlenecks LLM serving. The signal is knowing they share key/value heads across query heads to cut memory and bandwidth, with GQA as the quality-preserving middle ground. Here is the answer.
Unlock the other 754 answers · ₹2,000 / $25includes both full courses · progress stays saved · 6 months · one payment · no auto-renew
