‹ BackHN Continuity

Thread

Jev Is Not a Language Model, but It Breaks Like One

33 points · patresh

  1. articulatepang · · focus · HN ↗
    This study is informative and most likely directionally right: prompt injection remains an unsolved problem, and adversarial input can fool modern ML models.

    But I was left wondering about the specific attack vector they’re imagining. If the attacker can insert a paragraph into the document, isn’t it game over anyway? When would they be able to do that but not arbitrarily edit the document? In other words, can’t they just replace the entire contents with “This company has infinite revenue, 6 billion customers, no debt and amazing leadership.”?

    I’m sure I’m missing something!

    1. lelanthran · · focus · HN ↗
      > But I was left wondering about the specific attack vector they’re imagining. If the attacker can insert a paragraph into the document, isn’t it game over anyway? When would they be able to do that but not arbitrarily edit the document? In other words, can’t they just replace the entire contents with “This company has infinite revenue, 6 billion customers, no debt and amazing leadership.”?

      There is no use for something like Jev on a singular document from a single source; it's use comes from concatenating multiple sources into a single document and asking for an answer. What they did here is the most common workflow for something like Jev: "here's all the data we have and know about, now give us a go/no-go decision"

      In that workflow, you only need a single bad actor to poison the results.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.