‹ BackHN Continuity

Thread

Aleph Alpha Kolibri: How the sovereign German LLM works

419 points · 12 comments · tejaskumar__

  1. Ey7NFZ3P0nzAe · · focus · HN ↗
    I'm always wondering why new models don't always adopt deepseek's KV tweaks. It's insanely valuable to have such powerful prefix caching and so cheap at inference time.

    Are there drawbacks to this?

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.