‹ BackHN Continuity

Thread

Jev in 25 Lines of Python

691 points · 212 comments · bashbjorn

  1. antirez · · focus · HN ↗
    Because of masked attention in LLMs, if you put the options before the body (the email to analyze), the transformer already knows what it needs to look for, and can use more tokens to create state to address that specific task (BERT has no mask in the attention, so tokens attend also to next tokens). You could also do a few examples in the system prompt to improve calibration.

    Another trick that works is to repeat the question two times: "I'm repeating the task and labels for clarity: ..."

    1. jeff_ciesielski · · focus · HN ↗
      This works very very well :).

      <a href="https:&#x2F;&#x2F;github.com&#x2F;Mushroom-Systems&#x2F;lichen" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;Mushroom-Systems&#x2F;lichen

      1. ozozozd · · focus · HN ↗
        Did I read this right?

        This repo is really outperforming the OG Jev in the public benchmarks?

        There was no time to benchmaxx. How is this possible?

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.