I'm always wondering why new models don't always adopt deepseek's KV tweaks. It's insanely valuable to have such powerful prefix caching and so cheap at inference time.Are there drawbacks to this?
Ey7NFZ3P0nzAe · · focus · HN ↗
Are there drawbacks to this?