‹ BackHN Continuity

Thread

Once Claude can measure something, it can make it faster

231 points · 153 comments · matthieu_bl

  1. hungryhobbit · · focus · HN ↗
    How about you make Opus 5.5 actually work?

    I had it try to prepare a code review for me. Not only did it refuse, it refused to even tell me what the prompt (written by another Claude!) was. Why?

    When I had another model read the session (all of the "stupider" models handled it just fine) it explained that it had the word "reasoning" in it

    That's the entirety of Anthropic's billions of dollars of research: any prompt with the word "reasoning" is trying to hack Claude to figure out how it reasons!

    A model like that should never have gotten out of QA, let alone been released.

    1. railgunmerlin · · focus · HN ↗
      seems a bit weird to complain about the model issues in a post about the harness/sites?
    2. Marciplan · · focus · HN ↗

      [dead]

      1. cyanydeez · · focus · HN ↗
        the skill issue is "having to use a cloud model to do work of any value"

        might as well offer your life to a king to work in their fields.

      2. hungryhobbit · · focus · HN ↗
        Again, I used a slightly older model in the same series, Opus 4.6. It read the same prompt without any problem whatsoever. Also, a (non-Opus 5.5) Claude wrote the prompt in the first place.

        Opus 5.5 literally refused to work OR EVEN TELL ME WHAT I'D "SAID" when it read that prompt.

        Nothing to do with skill or the user at all: same exact prompt, three different models ... two worked, one didn't.

    3. bitpush · · focus · HN ↗
      I understand the frustration but shows a lack of critical thinking. Esp when you start with 'How about ..'.

      This blogpost is about frontend performance. It'll be akin to you commenting on a swift blogpost saying 'How about Airpods noise cancellation'. Sure both are Apple, but they are wildly different teams.

      1. hungryhobbit · · focus · HN ↗
        This is a techie discussion forum.

        In such a forum, it seems to me like it's fair game to point out that the company patting itself on the back about how great they are at programming (as evidenced in the article above about their 3x speed improvement) ...

        ... can't even make their latest model handle basic English without refusing to work.

        1. jfidjcjwjcjwjd · · focus · HN ↗
          Not that I’m defending Anthropic but OP is right, you’re complaining about the taste of a pear on a blog post about roses with the excuse that you’re in a biology forum. Sure, they both come from the same family, and sure, they both fall under the purview of biology, but they’re not the same.
          1. s3p · · focus · HN ↗
            I struggle to understand this argument. It is very likely that the user who is complaining about Opus 5.5 would be unable to do the very tasks anthropic is mentioning in their blog post, because of the refusals.

            In a blog post about Claude, i find it strange you get upset when people talk about Claude

        2. [deleted] · · focus · HN ↗

          [deleted]

      2. s3p · · focus · HN ↗
        Through critical thinking I believe this would actually be akin to them commenting that AirPods firmware, written in swift, is bad because of swift limitations.

        In both instances, it's a side point that is actually tangentially related to the first. Not completely unrelated as you are implying

    4. frumplestlatz · · focus · HN ↗
      I’ve had the same thing occur five or six times over the past week; they seem to be attempting to prevent anything resembling chain of thought extraction.

      Every single time it triggered, it was due to a prompt written by their own model in a dynamic workflow. The self-serving nanny oversight has to go.

      The fact that they label model distillation as an “attack” is genuinely hilarious after they “distilled“ their models from all of our work, and continue to do so.

      I believe AI is here to stay and an incredibly powerful tool, but these companies, and especially Dario and Altman, are the very last people I want to see in charge of it.

      1. copperx · · focus · HN ↗
        > “distilled“ their models from all of our work

        They distilled all digitized human knowledge and artifacts and they're now complaining about someone copying their outputs saying it's a "national security concern."

        I'm not sure about how to classify that. Hilarious? Pathetic? Sad? Hypocritical? Hyperdramatic? All of the above?

    5. vikramkr · · focus · HN ↗
      Probably it thinks you're doing some sort of system prompt exfiltration/distillation attack. Also what even is the workflow you're trying to have it do? It's doing code review but you're having it read some other AI models prompt/session history? Are you doing code review or like session history retrospectives?
    6. post-it · · focus · HN ↗
      > When I had another model read the session (all of the "stupider" models handled it just fine) it explained that it had the word "reasoning" in it

      Did it explain it did it hallucinate?

      1. hungryhobbit · · focus · HN ↗
        This happened at the classifier level, there was no "Claude thought X about it (hallucinating or otherwise)": this was a glorified regex deciding Claude couldn't work on a prompt (a code review prep) because it contained a string ("reasoning") it didn't like.

        It's more or less the same mistake we've seen Anthropic make repeatedly with it's brain-dead regex-based Fable/Mythos gates.

        1. post-it · · focus · HN ↗
          > because it contained a string ("reasoning") it didn't like.

          How do you know? How would the stupider model know?

          1. nfcampos · · focus · HN ↗
            Because it’s in a page accessible from its search tool presumably <a href="https:&#x2F;&#x2F;claude.dev&#x2F;blog&#x2F;getting-the-most-out-of-opus-5-5&#x2F;" rel="nofollow">https:&#x2F;&#x2F;claude.dev&#x2F;blog&#x2F;getting-the-most-out-of-opus-5-5&#x2F;
    7. dolmen · · focus · HN ↗
      This issue is mentioned on the Opus 5.5 post [1] from Anthropic (no idea if it has been added after your rant):

        &gt; Don’t ask it to show its reasoning in the reply
        &gt;
        &gt; What to do. Remove requests to reproduce its internal reasoning in the reply from your prompts and instructions.
        &gt;
        &gt; Why it matters on Opus 5.5. A request to reproduce its internal reasoning in the reply can be declined. It’s one of the flag categories.
        &gt;
        &gt; How. Ask Claude for what you need instead, for example, “Explain why you chose this approach in three sentences.”
      
      [1]: <a href="https:&#x2F;&#x2F;claude.dev&#x2F;blog&#x2F;getting-the-most-out-of-opus-5-5&#x2F;" rel="nofollow">https:&#x2F;&#x2F;claude.dev&#x2F;blog&#x2F;getting-the-most-out-of-opus-5-5&#x2F;
      1. ronsor · · focus · HN ↗
        Anthropic is trying so hard to &quot;crack down&quot; on distillation that they&#x27;re ruining their own product. I do not know why.
        1. PunchyHamster · · focus · HN ↗
          AI companies are seeing the endgame, more and more tasks can be done just fine by cheaper model, and there will be not enough money to be made to support model that is only needed for top 1% of the tasks.

          And the goal is to slow competition down enough before IPO

      2. crooked-v · · focus · HN ↗
        That&#x27;s extremely stupid, but also it&#x27;s a completely different thing from literally just having the word &quot;reasoning&quot; in the text.
        1. rendaw · · focus · HN ↗
          If only we had some sort of artificial intelligence-like system that could decide not only based on the word but also surrounding context, and maybe even in cases when that word is not specifically used.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.