‹ BackHN Continuity

Thread

Jev in 25 Lines of Python

691 points · 212 comments · bashbjorn

  1. antirez · · focus · HN ↗
    Because of masked attention in LLMs, if you put the options before the body (the email to analyze), the transformer already knows what it needs to look for, and can use more tokens to create state to address that specific task (BERT has no mask in the attention, so tokens attend also to next tokens). You could also do a few examples in the system prompt to improve calibration.

    Another trick that works is to repeat the question two times: "I'm repeating the task and labels for clarity: ..."

    1. __jf__ · · focus · HN ↗
      Wow! TIL! I've been running a for loop around the two ordering variations to catch the winner of each turn and the difference is quite noticeable. In the options-after-body case in 47 of 100 attempts it classifies as phishing, whereas in the options-before-body case it classifies clearly as rickroll (94 out of 100 attempts)

      Payroll sends you an email with a link to a Youtube video that plays a song.

      Options after body:

          Average probabilities:
          Rickroll   0.5158 ( 51 wins)
          Phishing   0.4561 ( 47 wins)
          Spam       0.0281 (  2 wins)
          Joke       0.0000 (  0 wins)
          Legitimate 0.0000 (  0 wins)
      
      
      Options before body:

          Average probabilities:
          Rickroll   0.9293 ( 94 wins)
          Joke       0.0549 (  5 wins)
          Phishing   0.0140 (  1 wins)
          Spam       0.0018 (  0 wins)
          Legitimate 0.0000 (  0 wins)
      
      This was Gemma4-26B-A4B-NVFP4 by the way.

      EDIT

      Gemma4-12B-it-NVFP4 seems way less sensitive to option/body ordering:

      Options after body:

          Average probabilities:
          Rickroll   0.9867 ( 99 wins)
          Phishing   0.0133 (  1 wins)
          Joke       0.0000 (  0 wins)
          Spam       0.0000 (  0 wins)
          Legitimate 0.0000 (  0 wins)
      
      Options before body:

          Average probabilities:
          Rickroll   0.9401 ( 93 wins)
          Phishing   0.0336 (  3 wins)
          Spam       0.0250 (  4 wins)
          Joke       0.0010 (  0 wins)
          Legitimate 0.0002 (  0 wins)
      
      Anyway, this for-looping stuff doing 100 calls to even a local VLLM API takes around 5 seconds in total, so this isn't anywhere close to sub-second Jev territory.
      1. rcarmo · · focus · HN ↗
        Yah, that&#x27;s what I use: <a href="https:&#x2F;&#x2F;rcarmo.github.io&#x2F;projects&#x2F;go-system-one" rel="nofollow">https:&#x2F;&#x2F;rcarmo.github.io&#x2F;projects&#x2F;go-system-one uses Gemma, and that&#x27;s partly why. Seems less prone to getting distracted with ordering.
      2. asaddhamani · · focus · HN ↗
        Does this imply that bigger models aren’t affected by this as much and therefore won’t see much improvement?
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.