‹ BackHN Continuity

Thread

Sonnet 5.5

884 points · 613 comments · D2OQZG8l5BI1S06

  1. ghoshbishakh · · focus · HN ↗
    So sonnet is better than Fable now? That Fable which was too dangerous to release? I am so confused now.
    1. heyjstn · · focus · HN ↗
      doom marketing at its finest
      1. solenoid0937 · · focus · HN ↗
        Not really, they never released Mythos. And they never said Fable was dangerous. They've been very consistent
    2. mnicky · · focus · HN ↗
      Well, it's performance "surface" (is there a better term for this?) is probably very narrow compared to Fable :)
    3. pebbly_bread · · focus · HN ↗
      Mythos is what they thought was too dangerous to release, fable was what they made after they worked on cybersecurity detection. As they say in the notes, this version of sonnet now has a similar screening process
      1. rs_rs_rs_rs_rs · · focus · HN ↗
        Mythos and Fable are the same llm. Fable has an extra tool that's in front of it that decides to accept the promp or not.
        1. dbbk · · focus · HN ↗
          Hence why it is no longer dangerous, yes
          1. emil-lp · · focus · HN ↗
            If you consider security through obscurity a safe route, yes
            1. usef- · · focus · HN ↗
              How is this security through obscurity?

              That term is about hiding a system's design in order to secure something, rather than having robust defences.

              1. emil-lp · · focus · HN ↗
                You answered your own question!
                1. usef- · · focus · HN ↗
                  This is not relying on a hidden design.
                  1. emil-lp · · focus · HN ↗
                    The design allows users to bypass the security by luck and chance.

                    A secure system is impenetrable unless you have the key.

                    1. usef- · · focus · HN ↗
                      Yes, it does depend on a probablistic classifier.

                      It sounds like you're thinking of cryptography, not general security. The world uses many security products that have false positives and false negatives (firewalls, intrusion detections, wafs, fraud detection...). Those aren't generally considered security through obscurity.

                      They openly talk about the system and its drawbacks here: <a href="https:&#x2F;&#x2F;www.anthropic.com&#x2F;news&#x2F;fable-safeguards-jailbreak-framework" rel="nofollow">https:&#x2F;&#x2F;www.anthropic.com&#x2F;news&#x2F;fable-safeguards-jailbreak-fr... and mention it goes along with rate limits, monitoring, account control, multiple classifiers, model design, and other things too.

                      I think this was also why they faced the controversy over not having zdr in fable: they wanted to use logs to detect repeated attempts, etc.

                    2. stravant · · focus · HN ↗
                      That&#x27;s not how it works.

                      The reason you see &quot;dumb&quot; refusals that &quot;should clearly be allowed&quot; is that they&#x27;re using more traditional deterministic methods to deny prompts rather than just relying in the stochastic LLM which you could bypass by luck.

        2. usef- · · focus · HN ↗
          And Sonnet has a similar classifier in front according to the article:

          &gt; it’s the first Sonnet model to launch with cyber safeguards

    4. Art9681 · · focus · HN ↗
      You have to actually spend some effort reading the article they published to answer your own question.
    5. [deleted] · · focus · HN ↗

      [deleted]

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.