‹ BackHN Continuity

Thread

I asked Meta’s Muse for its filesystem and it sent me 6.8GB

356 points · 170 comments · Aeroi

  1. tolugenius · · focus · HN ↗
    > About 20 Markdown files described browser use, connectors, payments, credentials, data handling, generated files, voice, goals, and scheduling.

    This the state of software engineering in 2026.

    Edit: clarified engineering to software engineering, which is more correct

    1. s08148692 · · focus · HN ↗
      To be fair there's probably a considerable amount of engineering that went into evaluating those markdown files so the agent behaviour is statistically reliable. The markdown is the product, not the process
      1. xnickb · · focus · HN ↗
        which part of that is engineering exactly?

        Not trying to be snarky. I genuinely don't get it

        1. thornewolf · · focus · HN ↗
          Write a prompt, evaluate the prompt, understand that is succeeds 95% of the time.

          Write a new prompt, evaluate, it now succeeds 99% of the time. Measure what changes between prompt #1 and prompt #2, understand what contributed to the performance jump.

          Write a third prompt, this one succeeds 100% of the time. Increase the size of your evaluation set, find a 1/5000 error-class and a 1/10000 error-class, add some explicit code to correct for this cases.

          Roll out to production, collecting usage metrics. You make some tweaks to your harness, your prompts. Eventually you have confidence that your system has fewer mistakes than 1 in 100k.

          Now, multiply this iteration across all your different prompts and different ways that they might interact with one another.

          1. kevin_thibedeau · · focus · HN ↗
            95% is shit tier engineering. Would you be satisfied if your keyboard randomly failed 5% of the time.
            1. Anon1096 · · focus · HN ↗
              Things like Voice to Text and biometric unlocks (fingerprint scanners, face ID) have worse success rates and they're used every day by billions of people.
              1. TheOtherHobbes · · focus · HN ↗
                Voice to text and biometrics are noisy sources, so a big part of the problem is dealing with that noise.

                Typing is not a noisy source. It should be reliable and deterministic.

                Protecting an agent from fairly obvious attacks should also be deterministic.

                1. eliaspro · · focus · HN ↗
                  The fundamental issue is, that "we" somehow decided it would be a good idea to throw all the fundamental ideas of computing (determinism, context, separation between data and execution,...) away and try to solve the issues by running a probabilistic/stochastic word generator on top of deterministic circuits instead at 10 magnitude worse efficiency.
                  1. TeMPOraL · · focus · HN ↗
                    It's not a fundamental issue. Determinism and "separation between data and execution" are artificial constructs, make-believe universe in which we design classical code, and a whole lot of hardware engineering goes into allowing us to briefly forget it's all fake.

                    Real world is probabilistic in practical / metrological, if not fundamental sense, and separation between data and execution does not exist. Our reality does not support such separation.

                    > a probabilistic/stochastic word generator on top of deterministic circuits instead at 10 magnitude worse efficiency

                    It's 10 magnitude better efficiency end-to-end, if you factor in design time you'd have to spend to get your "deterministic circuits" (which really aren't, we just paper over it) into shape so they deterministically solve a specific problem, for each problem you want to solve - where with the "stochastic word generator", you just need to change the text prompt.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.