‹ BackHN Continuity

Thread

A warning about 'model welfare'

242 points · 701 comments · andsoitis

  1. qarl · · focus · HN ↗
    Birch, The Edge of Sentience (2024), ch. 16 - "simply no way to assess sentience in an LLM"

    Schwitzgebel, AI and Consciousness (2025) - "we won't know before we've already manufactured thousands or millions of disputably conscious AI".

    Butlin, Long et al., Consciousness in Artificial Intelligence: Insights from the Science of Consciousness (2023) - "no obvious technical barriers to building AI systems which satisfy these indicators".

    Chalmers, Could a Large Language Model Be Conscious? (2023) - "within the next decade, we may well have systems that are serious candidates for consciousness".

    Long, Sebo, Butlin, Birch et al., Taking AI Welfare Seriously (2024) - "there is a realistic possibility that some AI systems will be conscious and/or robustly agentic in the near future".

    Dreksler, Caviola, Chalmers, Sebo et al., Subjective Experience in AI Systems: What Do AI Researchers and the Public Believe? (2025) - survey of 582 AI researchers; median estimate of 25% by 2034, and only 10% that such systems will never exist.

    1. TacticalCoder · · focus · HN ↗
      Well there are, today, several models (either text or image or vid) that can be run in a fully deterministic way.

      A conscious machine that always answer the very exact same thing, formulated the exact same way, bit for bit, to a query is, well, quite a weird kind of "consciousness".

      Now, I know, I know: the counter-argument is going to be "but humans have no free-will and are 100% deterministic too".

      I haven't yet decided if humans saying there's no free-will and who consider themselves to be 100% deterministic machines are reasonable or not.

      Meanwhile: seed / temperature = 0 and I'll happily turn the power button off of any glorified abacus without feeling bad about it.

      1. qarl · · focus · HN ↗
        > A conscious machine that always answer the very exact same thing, formulated the exact same way, bit for bit, to a query is, well, quite a weird kind of "consciousness".

        You'll need to explain why.

        1. 5eueuur · · focus · HN ↗
          it's self evident isn't it - such a consciousness has never existed, one existing would be by the fact of existing be weird
          1. qarl · · focus · HN ↗
            It doesn't seem self evident to me, no.

            I suspect the reasoning is connected to free will vs. determinism. But no, I see no inconsistency. You'll have to actually point it out.

          2. SkyBelow · · focus · HN ↗
            The problem is that science has found no evidence of free will, leaving humans to be nothing except determinism and chance (and given that QM effects don't scale to molecular level, that leaves only determinism). The oddity in conscious LLMs wouldn't be the determinism, it would be that we finally have a consciousness running on a system we have enough control over to repeat the same state.
            1. loopies · · focus · HN ↗
              >and given that QM effects don't scale to molecular level, that leaves only determinism

              This is simply wrong. Quantum effects affect your choices more than it might seem possible. The wrong atom decays in the wrong place in your body mutating your DNA, that one too much string of DNA that starts a tumor, which means chemotherapy. That does seem to affect your choices.

              Or instead of Shroedinger's cat you use it to trigger something, buy or sell some stock. Which clearly affects your future, one way or another.

              I never really understood why humans keep insisting stochastic processes do not affect their choices and their future. It's very short sighted.

              Hell, even brownian noise in your brain can affect your ideas, that one random spark out of nowhere, that one neuron that's pushed over the triggering edge. There's so much stuff all around us affecting us constantly. One neutrino interacting a the wrong time (average one does interact with our bodies across whole life, I remember reading), one cosmic ray.

              1. SkyBelow · · focus · HN ↗
                This isn't enough to be relevant to free will. You can construct massive impact by tying them to quantum systems, and one off things that happen occasionally given the length of a lifetime. But the quantum effects that are needed for free will need to be continuous. We can build a system where QM decides if a nerve cell fires or not, but that is very different from QM being involving in determine if every (or most, or even 1%) of nerve cells firing.

                (And even if QM did factor in normally, that just reduces it from determinism to random chance, "free will" still doesn't show up. That's like arguing an AI with sufficiently well implemented RNG is necessary to be conscious.)

            2. pixl97 · · focus · HN ↗
              To a forth dimensional being humans would look almost exactly like an LLM does to us. "What do you mean they have to wait for the future to know about it, how could they possibly be conscious then".

              Also, I'd love to have a quantum copy machine where you just replay someones state again and again with slight changes to see out it effects the output. My assumption is we'd figure out we behave just like LLMs pretty quick.

              Also, don't give me unlimited power or I might make a the world a war crime simulator, so there's that.

          3. mitxela · · focus · HN ↗
            neither has an LLM existed before, and that's quite weird
        2. pessimizer · · focus · HN ↗
          Why it's weird? Seems obvious. Because it doesn't behave any differently than any other process that we don't think of as conscious. We don't think of a valve as conscious; and if we make a Rube Goldberg machine valve, we don't think of it as more conscious. It's just a series of valves that do a predictable thing.

          I don't think it's defensible to say

          1) that a transistor isn't conscious,

          2) that a bunch of transistors that I've wired together aren't conscious, because I know how I've wired them together and therefore I know when I give them a particular input I will get a particular output, just like the single transistor, then go to

          3) that I have a bunch of transistors that I've wired together, but I put in so much input that I can't remember exactly what I've put in, plus I've wired some of the transistors to output random numbers that would be difficult to guess and fed them in also, therefore I don't know what will come out, are conscious.

          Even if I do accept this, if I take away the random number generator, and I can literally predict what can come out (by running a test in advance), and I still considered that consciousness, that would be odd. The only reason why I ever suspected consciousness was because I couldn't predict the output. I shouldn't even have accepted that, because I didn't think of the random number generator as conscious. [edit: of course, with the same seed the random number generator would have the same oddness.]

          It seems a bit like an argument from ignorance, a theological argument. Not understanding how something moves makes it alive (animated by spirits.) But it's even weirder to assign it metaphysical qualities when you built every single element of it with the goal to do the thing that it does, and it does it totally predictably and deterministically.

          1. qarl · · focus · HN ↗
            It sounds like you disagree with determinism. But that's a well worn argument and determinism has won, albeit in an odd sort of detente.
          2. Kim_Bruning · · focus · HN ↗
            Let's do something simpler.

            Take a car apart. So now I have a pile on the floor with an engine, some wheels, a tire, and some chairs.

            Can you point to the part that makes it cruise at 130 km/h on the autobahn? I bet you can't. You need to assemble all the parts back into a car before they work again.

            Or take your transistors. We can put them together to make a pocket calculator. Can any small bunch of transistors add 1+1? Trivially I can think of a few conformations that can actually, if that's all you want to do. But you will need all of them together if you add 12345678+87654321.

            So all of biology so far has been take things apart all the way and you end up with your hands full of a bunch of molecules. Are those molecules alive? No. You need to put them together into organelles and the organelles into cells before we call them alive.

            How about consciousness? Well, we haven't solved that one yet, but we figure that -since it's a biological function - we should be able to pull it apart in the same way. That's biology's best guess anyway, and there's several sub-disciplines of biology working on it, from neuroanatomy to neurophysiology to ethology.

            1. pixl97 · · focus · HN ↗
              I'm wondering how many groups are doing really unethical animal genetic modifications of the genes around language and intelligence in hidden places. It really seems we're at the edge of technology that will allow us to do things previous generations wrote about in horror.
        3. lelanthran · · focus · HN ↗
          >> A conscious machine that always answer the very exact same thing, formulated the exact same way, bit for bit, to a query is, well, quite a weird kind of "consciousness".

          > You'll need to explain why.

          Where are you going with this? I don't see a conclusion for this line of questioning.

          I mean, you can't explain why a machine that reliably and predictably produces the same result is "the normal kind of conscious", can you? So why expect someone else to explain why it's a "weird" kind of conscious?

          1. qarl · · focus · HN ↗
            I ask because I truly don't understand what he's saying. Clearly we have different ideas about consciousness and I have no idea where we diverge (from what he has said.)
            1. lelanthran · · focus · HN ↗
              > But we already have examples that break that rule (us).

              No, we don't. Where are you reading your research papers?

              1. qarl · · focus · HN ↗
                Please don't insult me.
                1. lelanthran · · focus · HN ↗
                  It's not an insult.

                  "People are deterministic" is news to me, so I'd really rather like to know which papers claimed that.

                  1. qarl · · focus · HN ↗
                    Sorry friend, between this and the vague memories I have of you harassing me, I think I'm not going to answer. I'm sure you can get an LLM to explain it to you.

                    Have a nice day.

                    1. lelanthran · · focus · HN ↗
                      > Sorry friend, [...] Have a nice day.

                      Not a problem; this is the playbook all the all-in AI boosters run to when they are called on their claims. Everyone reading this knows exactly what your claims meant, because the majority of HN have seen this trick before.

                      You can pretend all you want that you're just being polite, but the truth is there is no research backing your claims.

                      Regards,

                      A former research scientist, current business owner.

                      1. qarl · · focus · HN ↗
                        Yep. Have a nice day.
      2. antx · · focus · HN ↗
        Out of curiosity, which models are fully deterministic? I was under the impression that all LLMs were fundamentally probabilistic.
        1. Wowfunhappy · · focus · HN ↗
          The randomness is something we add on purpose; you can set an LLM's "temperature" to 0 to get deterministic output. This tends to make the quality of its responses worse for reasons I don't think anyone really understands, but it's still functional.

          I don't think the state of the art LLM providers let you do this anymore (?), but they certainly could if they wanted to, and you can do it yourself with a local model.

          1. antx · · focus · HN ↗

            [dead]

            1. Wowfunhappy · · focus · HN ↗
              But that's from, like, floating point errors, right? If you used higher precision that wouldn't happen, it's just because we're cheap in how we do rounding.
              1. pixl97 · · focus · HN ↗
                I mean, currently we'd have some difficultly proving our hardware isn't deterministic, just that we can't actually test it.

                But, I think you're tricking yourself on determinism. You'll say something like "I know if I ask an LLM what 1+1 is, it will answer 2", but the thing is, you don't. You have to run the LLM first to figure out it's output. And when you send in just a few bits of text, it's outputs are going to be rather limited.

                But this all breaks when it hits the real world. Inputs are unpredictable. Hence while LLM outputs, like humans, are probabilistic, you can't figure out what it's going to be until you ask. And in any high complexity data gathering environment you're going have a difficult time ensuring your entire systems conditions are the same.

                System consistency is very hard, once you start running thousands of processors in an agentic loop small errors accrue and timing starts differing and the system will take non-deterministic paths.

                1. throwaway63486 · · focus · HN ↗
                  I think you're arguing a different thing than determinism.

                  If I ask an llm to "add 2 and 2" is and it replies corectly, then I ask for "the sum of 2 and 2" and it replies "banana" that is a lack of predictability and consistency but not a lack of determinism.

                  As long as it produces the same output for a given input, unhinged or not, it is deterministic.

                  Your example at the end of different systems feeding data to each other is non-deterministic only at the system level, not the individual llm level.

                  1. pixl97 · · focus · HN ↗
                    >only at the system level, not the individual llm level.

                    Which is why llms aren't agents and depend on harnesses. The llm itself doesn't have a continual loop built in, that would be very power hungry. The harness works as the orchestrator of memory and action. Now, I can't think of a reason why an LLM couldn't bootstrap its own harness, but in general it sounds like a very dumb idea to actually build that from an AI safety perspective.

                    This discussion falls under the idea and refutation of the Chinese Room. The room may have no idea what Chinese characters are, but the system does.

          2. mitxela · · focus · HN ↗
            You can also use a seeded random generator to get the same random numbers each time
        2. piker · · focus · HN ↗
          Same weights, same seed, same input tokens, same algorithm, same output tokens, probabilistic or not. Quantum effects have been de-noised, but I guess there are still random gamma rays.
        3. qarl · · focus · HN ↗
          Naw - computers are really deterministic. It's hard to get them to behave otherwise.

          As I understand it, if you turn down the temperature to 0 you get repeatable behavior - EXCEPT - on large servers with lots of users - the GPU can sometimes produce slightly different results based on batch size.

          1. joe_the_user · · focus · HN ↗
            Unless you have something exotic, the randomness that's adding to a computer is a combination of how it's configured combined with a pseudo-random number generator. I assume the system adds entropy to the generator regularly but all you need to do is fix the various supposedly random inputs and you can get full determinism even without zero temperature.
            1. qarl · · focus · HN ↗
              Yes, in theory, of course.

              In practice - on a multitasking OS with input from multiple human users - it's hard to get it deterministic because of that GPU scheduling thing I mentioned.

              1. fc417fc802 · · focus · HN ↗
                GPU scheduling only affects the result due to buggy optimizations. It's the exact same mechanism as fp rounding error on the CPU or updates to globally shared PRNG state. We use lots of buggy optimization because they don't matter in practice in most situations (see ex -ffast-math).
                1. qarl · · focus · HN ↗
                  I don't know the details but it has something to do with multiprocess contention for the GPU and batch sizing.
                  1. fc417fc802 · · focus · HN ↗
                    Right but regardless, it's a buggy optimization. The calculations are all fully deterministic when done "properly" and fully consistently but we almost never bother with that because it slows things down and the errors don't matter in practice 99% of the time.
                    1. qarl · · focus · HN ↗
                      AGREED. Although whoever designed that architecture might take issue with your description as "buggy".
          2. goodmythical · · focus · HN ↗
            If computers were fully deterministic, we wouldn't need error correcting ram.

            The abstracted design of the machine is meant to be deterministic, but you can't predict before running any command whether or not it will complete because there are externalities that effect the outcome.

            Electromagnetic interference even happens in-chip where an electron can accidentally escape it's wire and enter another, possibly resulting in an error, but not every time.

            It's even been used as an attack vector where rapidly flipping a bit increases the likelihood that a neighbor bit is also flipped, but the method is probabalistic, not deterministic.

            1. qarl · · focus · HN ↗
              Yes. In theory they are deterministic. In practice, not so much.
              1. budman1 · · focus · HN ↗
                Those non-deterministic results are exceedingly rare.
                1. qarl · · focus · HN ↗
                  When human input is involved, like I described elsewhere, it happens frequently.
                2. lioeters · · focus · HN ↗
                  As long as the machine is part of the larger universe, and not completely isolated in its bubble of space-time (an impossible situation), it cannot be deterministic.
                  1. mitxela · · focus · HN ↗
                    So your theory is that arithmetic is nondeterministic because very rarely a calculator will say 1+1=3.
                    1. lioeters · · focus · HN ↗
                      Arithmetic is deterministic, because it is abstract. A calculator or computer is not, because it is physical and exists in a non-deterministic universe. But with error correction, etc., it's deterministic enough for practical purposes.
        4. SkyBelow · · focus · HN ↗
          By default they are matrix multiplications. Temperature is added in as forced PRNG because testing found that correlated with better outputs.

          Given the same prompts and the same weights, one can get the same answer each time.

          In practice, there are a number of optimizations that makes the results dependent upon thing we give up control of to increase performance, meaning the results end up being effectively non-deterministic. But, if you are willing to run it in a slower mode so we don't do some steps out of order to speed things up and don't batch results (or if you consider the determinism of a given batch of requests rather than individual requests), then the same input gets the same output.

      3. joe_the_user · · focus · HN ↗
        The idea that determinism and consciousness are incompatible is deeply intuitive to many people. And many other embrace it - you can trace this back to the debates for and against Calvinism.

        A lot comes down to the way people parse causation and choice. You don't want to say that a murderer was completely caused to choose something because then you can't hold the person responsible. And so determined consciousness makes people unhappy. But just as much, if the opposite of determinism is hard statistical randomness, how do say that "is the essence of personhood". This is why physicist go out in the world trying to find consciousness as a fifth physical force.

        I mean, think consciousness is a term that ever have a non-contradictory meaning since it's primarily used to bound ethical human worlds and the verifiable formulations of biological and physical systems. But it's going to be with us for a while and I'm not sure what can be done about it.

        1. qarl · · focus · HN ↗
          I think this debate is rooted in our instinctual beliefs about fairness. An organism in a social group needs to decide how to react to the brutish behavior of its peers. And so Mother Nature has encoded a good strategy in our brains. (For example, dogs seem to understand fairness.)

          But this only means the strategy is practical - it doesn't mean it's consistent. I think "responsibility" falls into this category. So we have strong intuitions about it that don't quite logically work. And this is where free will and determinism and choice and punishment all crash together.

        2. Kim_Bruning · · focus · HN ↗
          You know, the one really interesting dynamic domain is chaotic, which is a subdomain of the deterministic! I figure free will is a chaotic phenomenon.

          Groundhog day is your intuition pump here. Go back in time and most assume that the day will go mostly the same, except for the butterfly effect if you change something. Given exactly the same conditions (we went back in time, so that's pretty exact), people will make the same decisions.

          Most people don't assume that the day will be completely different the next time loop. The town won't suddenly spontaneously all start breakdancing or standing on their heads.

          (Edit: The murderer is the one who -given the same/similar conditions- will murder again. I'd politely suggest maybe we put them behind bars as a precaution against that; until/unless they learn to act differently. )

          (Edit 2: Throw in stochasticity and this actually tends to smooth out chaotic systems. This is counter-intuitive! People's brains break on Evolution in the same way. Meanwhile somehow folks intuit birdshot without difficulty; which is weird!)

          1. jddj · · focus · HN ↗
            To extend on your edit #1: And, when you go back in time, delete all punishment for murder and see whether you get the same number of murders.

            If you get more than before, punishment might make some sense in a deterministic world so reinstate it.

            1. Kim_Bruning · · focus · HN ↗
              It might. Punishment adds a negative bias to the payoff matrix. Note that increasing punishment severity past a certain point actually saturates and doesn't alter the dynamics further. (which is why modern judicial punishment schedules sometimes seem so mild)
            2. pixl97 · · focus · HN ↗
              Eh, sorry chaos theory doesn't work this way.

              It's just as likely you go back and perform whatever magic you need to to delete punishment for murder, then zip back to 'now' and see the world is a global panopticon making sure people don't murder each other.

              Some of this may be a failure of people to realize what P !=NP is about, especially in non-linear systems.

              Small perturbations in an early state of the system can lead to wildly different outcomes in the final measured state of the system. You can't determine what that will be in non-polynomial time. The only thing you can do is make probabilistic models by running a full simulation in real time (or reality with a time machine I guess).

              1. Kim_Bruning · · focus · HN ↗
                If you look at chaos in dynamical systems, you can often find some equilibria (like attractors, spirals, saddle points etc ) at least, even if the details might vary. This is how people still manage to predict the weather somewhat, for instance. Oddly orbital mechanics is chaotic too, but it sort of rhymes over the millennia.
                1. pixl97 · · focus · HN ↗
                  Your correct. You'd have to run the universe simulation a bunch and pick the most probable outcomes (which makes you the most brutal serial killer in existence). The probability graph would represent the state stability of the underlying structures that guide what you're trying to measure.

                  In weather you have things that are chaotic, but follow rules and don't create them. Things can surprise you, but you'll mostly center around particular solutions.

                  With humans this gets a lot more complicated because you are messing with systems that follow fixed rulesets (chemistry) and things that follow self modifying rulesets (social conventions). A very unstable system on top of a semi-chaotic probability state.

                  A good example of this would be if I deleted the concept of rain from the human mind and all of human history and knowledge. Humanity would be very surprised on the thunderstorms that popped up, but they'd continue on just as they did before.

                  Now lets imagine I delete the concept of all gods. Would the idea of gods come back, for sure. Would they look like our gods now. Very questionable indeed. We wouldn't be getting jesus back. And the new set of religions could have wildly large effects on how the people act.

        3. loopies · · focus · HN ↗
          >you can't hold the person responsible

          The act of holding them responsible is supposed to determine them not do it. If stochastic behavior is impeding them in being a functional member of society then it still makes sense to remove them from society, why would you choose to live amongst people who are not behaving rationally? Or allow them to hurt other people? At the end of the day the why matters less to removing them or not from society. And sentences are clearly used as a determining factor.

          The holding responsible part deals with politics and more primitive aspects of our societies and biologies. Getting tangled up in holding them responsible or not is hardly something you should give much attention to. Rather to make sure they do not cause any more harm and also make sure such things are not created in the first place. Which opens up another can of worms for which society and politicians are ready for.

      4. fc417fc802 · · focus · HN ↗
        > I haven't yet decided if humans saying there's no free-will and who consider themselves to be 100% deterministic machines are reasonable or not.

        That seems exceptionally arbitrary. What basis do you have on which to classify them? Are you not conflating the perception of free will with ... what was the definition for it again anyway? Are you able to construct a satisfactory one that meshes with physics? I certainly haven't been able to.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.