‹ BackHN Continuity

Thread

OpenJev

722 points · 296 comments · ilreb

  1. kul_ · · focus · HN ↗
    Is it only me or do others also find LLM generated websites so off-putting?
    1. djaro · · focus · HN ↗
      Unsolvable problem.

      Why was the aesthetic standard to be pale when workers worked the fields and royals were inside, but tan when workers moved into factories and only the rich could afford to go on a beach vacation?

      Aesthetic standards are formed by association. Its why sites that are "well designed" but obviously just use a squarespace or wix template feel so cheap. Why millenial flannel went from hip to standard to outdated. Why purple was the color of royalty before we could synthesize the pigment.

      Having good design is about associations. Whatever design LLMs will default to, it will always feel cheap because we will learn over time that that design means cheap. Having good taste is about being ahead of the curve. An LLM cant be ahead of the curve because then that becomes the standard, and theres a new ahead.

      You can use LLMs to make novel looking websites by carefully telling it to add certain details, use certain elementd, etc. At that point youve looped back to being a graphic designer.

      1. miki123211 · · focus · HN ↗
        This is, once again, about diversity and the lack thereof (and I don't mean diversity in a political sense).

        LLMs seem fundamentally incapable of producing truly diverse outputs, truly creative and different responses to the same prompts in different runs. Because you and me use the same Claude, if you want a website and I want a website, we'll get (almost) the same website. This is not some BS about "the average of its training data", most of the LLM style (both in design and in text) comes from reinforcement learning. You could RL Claude to produce a very different style, but you couldn't RL it to produce a different style for me than it does for you.

        I think this is also where a lot of the complaints about "Claude writing" come from.

        1. CuriouslyC · · focus · HN ↗
          RL causes distributional collapse, it's how the models get consistent. Anyone who generated images with early gen (SD1.5-2) models will remember the wild variance between seeds, which newer models have mostly lost, and similarly GPT3.5/4 could produce weirder, more original outputs even if they were less consistently "good" in some sense.

          It's worth mentioning that they do RL for aesthetics to some degree based on human expert feedback, but whatever the model tends to produce quickly becomes debased by its ubiquity. They could RL for output diversity, but it's less well studied and likely to cause minor regressions in coding performance, at least until the algorithms are dialed in.

          1. miki123211 · · focus · HN ↗
            Yes!

            There was a great paper at Neurips 2025, where they showed that all "capabilities" that models get via RLVR were in fact already there, if you did pass@k with a sufficiently high k. RLVR makes the models consistently use tricks that tend to work and bring high rewards, which makes them better when k is low, but it suppresses the unusual, which is actually worse when k is high.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.