‹ BackHN Continuity

Thread

Using Opus 5.5 to discover a new eyewitness record of the dodo

226 points · 79 comments · benbreen

  1. nl · · focus · HN ↗
    > Epistemological weirdness

    > They are also notably bad at judging the historical significance of what they find.

    I use LLMs for some things that are outside the more common use-cases (in my case 3D design for 3D printing) and one thing I've noticed is that the errors it makes are so completely unlike human errors that they are hard to anticipate.

    It will do things like build perfect snap catches but put them so the the pieces they are connecting are rotated 90 degrees from how they should be. It's "dumb" error, but hard to say the model itself if dumb because it does other very hard things so perfectly.

    > seven chord groups

    This sounds a lot more like Opus 5.0 than Opus 5.5 TBH. I wonder if that was an earlier investigation because 5.5 has improved that kind of language a lot.

    1. morpheos137 · · focus · HN ↗
      In general llms are weak with spatial reasoning. This seems to be an unsolved problem. Probably because human language is generally imprecise spatially and humans think about spatial problems in visual terms. I wonder if having an llm make a 3d design in a format an image model could check would result in a better outcome?
      1. NiloCK · · focus · HN ↗
        I think that this is an unsolved problem in the same way that mangled fingers in image generation was an unsolved problem.

        Through at least Opus 4, LLMs were practically useless for authoring any sort of coherent procedural closed-curve geometry (I know this with strong confidence because of the little animated guys at <a href="https:&#x2F;&#x2F;letterspractice.com" rel="nofollow">https:&#x2F;&#x2F;letterspractice.com).

        Opus 5.5 can bang it all out. Possibly a deliberate RL sort of thing or maybe another surprise emergent capability.

      2. Aerroon · · focus · HN ↗
        Is it that LLMs are weak with spatial reasoning (and memory) or is it that we are unusually good at it?

        When I need to use a program I seldomly use I&#x27;m far more likely to remember where I need to click to open it than the word I need to search for to open it.

        1. kfarr · · focus · HN ↗
          Yes I like to think of humans with built-in accelerators for certain tasks -- our visual and spatial reasoning is off the charts presumably because it&#x27;s a life or death skill!
        2. treis · · focus · HN ↗
          I made a building and had astra fill out the interior of the bathroom with toilets. It put 6 of them in two rows back to back with no way to reach the second row. Other than climbing over the stalls of the first row I suppose.

          So yeah, they are very weak at spatial reasoning.

      3. nl · · focus · HN ↗
        Having used them heavily for 3D model understanding since February I can say it&#x27;s nuanced.

        Opus 4.6 and 4.7 were bad, but GPT 5.2 and above were very usable. Opus 4.8 was usable, but the GPT 5.x series was better.

        Fable is great.

        Opus 5.0 was interesting. It could solve some problems that Sol 5.x couldn&#x27;t solve (applying a G2 curve on a 3 way corner where one face was a Bezier curve) but you had to be super prescriptive (&quot;only answer the question&quot;&#x2F;&quot;only do what I tell you and stop when done&quot;) or it would go on a hugely involved validation journey that didn&#x27;t really achieve a lot.

        Opus 5.5 is better than that was in that respect.

        Sol 6.x is great, and my daily driver for this (I use Opus for coding though)

        Astra can solve problems that Sol can&#x27;t but for some reason on easy stuff makes uglier solutions.

        For all models it&#x27;s very interactive though - we aren&#x27;t at the &quot;agentic design&quot; phase for most things yet.

        Here&#x27;s a sample of what I&#x27;ve been able to get them to design with me: <a href="https:&#x2F;&#x2F;x.com&#x2F;nlothian&#x2F;status&#x2F;2099023496794018067" rel="nofollow">https:&#x2F;&#x2F;x.com&#x2F;nlothian&#x2F;status&#x2F;2099023496794018067

      4. komodo99 · · focus · HN ↗
        There&#x27;s an interesting comparison in the creative&#x2F;literary end of llm output too, they&#x27;re in my experience, dreadful with anatomy. Like, it knows humans have hands, heads, etc, but often times a seemingly limited concept of how anything is connected, or degrees of freedom. (e.g., Why yes, certainly there are many examples of humans rotating their torsos 180º at the hip, seems perfectly cromulent)

        I honestly don&#x27;t know if an image model would help, or if it might analyze the output and go &quot;13 fingers? ship it!&quot; anyway.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.