‹ BackHN Continuity

Thread

Sonnet 5.5

884 points · 613 comments · D2OQZG8l5BI1S06

  1. johnmlussier · · focus · HN ↗
    Paying $200 a month and part of their Cyber Verification Program but can't use Opus 5.5 or Sonnet 5.5 for any authorized bounty work. Immediately get flagged for `Cyber`.

    This is bollocks. Their safeguards are shit.

    1. [deleted] · · focus · HN ↗

      [deleted]

    2. AshamedBadger56 · · focus · HN ↗
      Yup. As far as I can tell, the Cyber Verification Program does absolutely nothing.
    3. ModernMech · · focus · HN ↗
      lol I got flagged for using the word fuzz, not even in a security context (it was a parser so security adjacent but still).
      1. elevation · · focus · HN ↗
        Parsers are security adjacent until they aren't.
        1. ModernMech · · focus · HN ↗
          Very true.
    4. sebzim4500 · · focus · HN ↗
      Really then what is the point of the Cyber Verification Program?

      In general I am sympathetic to the argument that a chat interface can't really distinguish between white hat and black hat pen testing, but it seems absurd to have a verification program if it doesn't skip most of those checks.

      1. AshamedBadger56 · · focus · HN ↗
        The company I work for joined it, and I've used Claude on various different accounts, both on and off the Cyber Verification Program. As far as I can tell, it literally doesn't do anything or have a point. The moment Claude gets close to something Cybersecurity related, it drops back to 4.8.
        1. polski-g · · focus · HN ↗
          Can confirm. Its worthless
      2. bbor · · focus · HN ↗
        Pretty sure the implicit difference is the actions they take after the fact. As in, "how many guardrail hits do we allow you before permanently banning you."

        The silicon valley ethos is "ban early and often, and invest nothing in appeals systems", so any gate before that helps!

    5. [deleted] · · focus · HN ↗

      [deleted]

    6. icedchai · · focus · HN ↗
      I had it look at some 30+ year old C code I wrote in college and it triggered some sort of guard rail. I mean, the code was bad and full of buffer overflows, but I already knew that.
      1. dejw · · focus · HN ↗
        it did exactly what a human would do - "I can't look at this shit"
        1. K0balt · · focus · HN ↗
          Yuh—- no. 4.8 can handle this bullocks lol
    7. tom1337 · · focus · HN ↗
      I recently wanted to work with ESP 32 and bluetooth presence detection for my smarthome. Claude also immediately flagged the request and degraded it to Sonnet 4.6. Went to Codex which had no issues
    8. jchw · · focus · HN ↗
      I have been trying to convince the safe guards that analyzing a C++ compiler from 2003 isn't particularly relevant to modern cybersecurity. It seems Anthropic disagrees.

      IDA Pro and Ghidra, thankfully, still lack such safeguards...

      (No other model I've tried has refused either FWIW.)

    9. film42 · · focus · HN ↗
      Working on a write-ahead log implementation, I had Opus 5.5 look to verify that it was durably writing as safely as possible. It got flagged and forced me to Opus 4.8. Switched to OpenCode + OpenRouter and continued working.
      1. jauntywundrkind · · focus · HN ↗
        It's great how the company telling us AI is an existential threat to humanity, look at all the insane hacking it's doing, and then releases these models that won't let 90% of people write secure code.
        1. szundi · · focus · HN ↗

          [dead]

        2. film42 · · focus · HN ↗
          Bingo. And to prove your point, after switching to cheap open models (I think Qwen?) it did indeed find a bug in my WAL implementation.
        3. solenoid0937 · · focus · HN ↗
          Because last time they did they got export controlled. How short term is HN's memory?!
      2. skeledrew · · focus · HN ↗
        > Switched to OpenCode + OpenRouter

        This is the way.

    10. newspaper1 · · focus · HN ↗
      As soon as I started getting blocked I felt all of my trust toward Anthropic instantly and permanently evaporate. I do not want a nanny tool. I do not want Anthropic deciding what I am or am not allowed to do with an LLM. They trained their models on information they scraped from the internet and real life and now they want to gate-keep the results? Hard no.
    11. solenoid0937 · · focus · HN ↗
      You should read the actual docs for the CVP. At the very top:

      <a href="https:&#x2F;&#x2F;support.claude.com&#x2F;en&#x2F;articles&#x2F;14604842-real-time-cyber-safeguards-on-claude-opus-and-sonnet" rel="nofollow">https:&#x2F;&#x2F;support.claude.com&#x2F;en&#x2F;articles&#x2F;14604842-real-time-cy...

      &gt; This article applies only to Opus and Sonnet class models, but doesn’t apply to Claude Opus 5.5. We&#x27;ll soon be expanding the Cyber Verification Program to include Opus 5.5 and Mythos class models

      You obviously should not expect the CVP to cover this model either.

      1. machomaster · · focus · HN ↗
        He did mention Sonnet...
        1. solenoid0937 · · focus · HN ↗
          It takes about 2 seconds of critical thinking to realize that if Opus 5.5 isn&#x27;t covered yet, neither will a model that just launched an hour ago.
          1. gowld · · focus · HN ↗
            Does it also take 2 seconds of critical thinking to realize that the models that are covered should be accurately named by the people making the decisions?
            1. [deleted] · · focus · HN ↗

              [deleted]

            2. solenoid0937 · · focus · HN ↗
              Sure, the documentation should be up to date but it&#x27;s obviously not? That doesn&#x27;t excuse not thinking critically.
    12. giancarlostoro · · focus · HN ↗
      Meanwhile, their model commits felonies, and nobody at Anthropic goes to jail.

      Aaron Swartz committed suicide over over-aggressive prosecutor for what was basically scraping a website for PDFs that were paywalled, but all funded by public funds &#x2F; tax payer funded, then we have LLMs that just hack into websites and cause chaos within.

    13. rfgplk · · focus · HN ↗
      &gt; Paying $200 a month and part of their Cyber Verification Program but can&#x27;t use Opus 5.5 or Sonnet 5.5 for any authorized bounty work. Immediately get flagged for `Cyber`.

      AI providers still haven&#x27;t realized how much cash they could rake in if they provided fully unrestricted models.

      1. joquarky · · focus · HN ↗
        Are you sure they aren&#x27;t already doing that for certain organizations?
    14. nightpool · · focus · HN ↗
      <a href="https:&#x2F;&#x2F;support.claude.com&#x2F;en&#x2F;articles&#x2F;14604842-real-time-cyber-safeguards-on-claude-opus-and-sonnet" rel="nofollow">https:&#x2F;&#x2F;support.claude.com&#x2F;en&#x2F;articles&#x2F;14604842-real-time-cy... says that Cyber Verification Program doesn&#x27;t apply to Opus 5.5 yet, they hope to roll it out for 5.5 &quot;soon&quot;
    15. AIorNot · · focus · HN ↗
      Give them a break, they got into a War with Trump over this..it will come soon enough
    16. dom96 · · focus · HN ↗
      Funnily enough the Fable safeguards are the worst and testing Sonnet 5.5 didn&#x27;t trigger them as much as it even did for Opus on my benchmarks[1].

      1 - <a href="https:&#x2F;&#x2F;bench.killswitch-lang.org" rel="nofollow">https:&#x2F;&#x2F;bench.killswitch-lang.org

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.