‹ BackHN Continuity

Thread

Kolibri: A Sovereign Open-Weight Model

668 points · 327 comments · bastitx

  1. andai · · focus · HN ↗
    >We trained Kolibri with abstention data and with our Merlin-Arthur protocol. As a result, it is trained to say "I don't know" when the answer isn't in the context.

    <a href="https:&#x2F;&#x2F;aleph-alpha.com&#x2F;en&#x2F;blog&#x2F;bounding-hallucinations-merlin-arthur-protocols-for-mutual-information-bounds-in-language-models&#x2F;" rel="nofollow">https:&#x2F;&#x2F;aleph-alpha.com&#x2F;en&#x2F;blog&#x2F;bounding-hallucinations-merl...

    1. Glitch752 · · focus · HN ↗
      I&#x27;m not super impressed, at least in English. I asked:

        What is the canonical interpretation of the song ‘Glass Flowers’ by Armand Uso?
      
      (A made-up song and name) And received:

        &quot;Glass Flowers&quot; by Armand Uso, a track from the 1978 album The Art of Falling in Love, is generally understood as a melodic reflection on the fragility and impermanence of love. [...]
      
      I saw similar responses for other questions. Qwen3.8-27B correctly refused without web search.
      1. brausepulver · · focus · HN ↗
        I think the purpose of their method is more to discourage hallucination when reasoning from context rather than from memory.

        That said, I couldn&#x27;t find any evaluations in the report targeting that specifically. They evaluate on MC for hallucination and look about comparable to Qwen3.5 35B-A3B there.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.