‹ BackHN Continuity

Thread

The AI Race Just Got Awkward

412 points · 462 comments · allisdust

Loading the complete thread in the background. This saved snapshot is available now. Refresh

  1. amelius · · focus · HN ↗
    > So the Chinese labs have thrown a lifeline to the Western loss-making labs, and I just have no clue as to why.

    Any ideas?

    1. ambicapter · · focus · HN ↗
      They want the Western labs to keep digging themselves into a hole? You can do so by egging on the true believers on your side to egg on the labs on the other side.
    2. curuinor · · focus · HN ↗
      There are 1000 Chinese labs. They are involuting, they cannot coordinate and the state won't let them coordinate because the state wants domination, not actual profits for anybody. So the market forces are allowed to dominate.

      Because of the basic huge recession going on in China, you can't actually make money in China doing China things. So they gotta gird up their export stuff and try to export. That entails strong relations with American companies, American PR, English stuff, etc.

      If you want an essay about this from a VC, read this one

      <a href="https:&#x2F;&#x2F;earnedintuition.substack.com&#x2F;p&#x2F;involution-without-export-is-wasted" rel="nofollow">https:&#x2F;&#x2F;earnedintuition.substack.com&#x2F;p&#x2F;involution-without-ex...

      1. dabedee · · focus · HN ↗
        What a strange way to put it. Market forces are a good thing in a market economy. Only someone who secretly wants or hopes for monopolies would you say something to the contrary (a VC).
        1. curuinor · · focus · HN ↗
          The party has done something about enormous involution in solar panels, for example (<a href="https:&#x2F;&#x2F;www.csis.org&#x2F;analysis&#x2F;chinas-solar-industry-upheaval-effects-will-be-global" rel="nofollow">https:&#x2F;&#x2F;www.csis.org&#x2F;analysis&#x2F;chinas-solar-industry-upheaval...) and previously steel. They&#x27;re planning something for cars. They don&#x27;t on LLM because of the newness and the wish for preeminence.
        2. ajkjk · · focus · HN ↗
          &quot;only X would say Y&quot; is a rhetorical device (the no-true-scotsman fallacy, if you want) that should basically never be used ever.
        3. iamnothere · · focus · HN ↗
          [delayed]
          1. tancop · · focus · HN ↗
            Price wars are good even if they lead to more bad products on the market. If quality is important people will pay more for it and ignore the bad ones, if not then everybody saves some money.

            The reason it doesn&#x27;t work like that IRL is centralized marketplaces. If the winning strategy on Alibaba is low prices, bad quality and botted reviews to compensate then every seller has to do it to survive, because they can&#x27;t get buyers outside the platform. That&#x27;s not excess competition. It&#x27;s a lack of competition just on a different level.

            1. iamnothere · · focus · HN ↗
              &gt; The reason it doesn&#x27;t work like that IRL is centralized marketplaces.

              Well yes, that’s why I mentioned those specifically. But even if the marketplaces were split up, someone could create an aggregator to comparison shop and the same effects would apply. The problem for producers is that the Internet erases information asymmetry.

              It’s also not just affecting low end goods, it’s a constant pressure on everyone, which is why many formerly upscale brands are seeing the same problems. It’s also a general problem with public companies, as large shareholders demand constant growth, as well as many private companies owned by PE where brands are stripped for short-term profits.

              My experience is that the best low cost mid-tier products right now are coming from fronts like Vevor and Fanttik who do sourcing from noname factories in China. I’m not sure if their position is sustainable; it’s not like they have much of a moat. (I guess Fanttik has a team that adds some slick design to their otherwise utilitarian items.) If that model holds up then maybe that’s the future, but I suspect that they just have a temporary advantage thanks to a dual presence and connections in both the US and China.

        4. DrewADesign · · focus · HN ↗
          Pouring 100% of your tech optimism, and probably portfolio, into a product that some say is about to get Wile E Coyote flattened by market dynamics would probably inspire serious market skepticism.
      2. thrawa8387336 · · focus · HN ↗
        LMAO recession, China? You&#x27;ve been reading too much Brad Setser
        1. curuinor · · focus · HN ↗
          The youth unemployment rate is at 19% with employment counting as 1 hour a week...
        2. 3371 · · focus · HN ↗
          Maybe look up &quot;China deflation&quot;
      3. luke5441 · · focus · HN ↗
        I&#x27;d call it overcapacity instead. A lot of investment without capital discipline making sure there is actually return of investment leading to too much supply.

        Not that OpenAI, Anthropic or SpaceX aren&#x27;t doing the same.

      4. googaar · · focus · HN ↗
        Nice read. American media does a terrible job of covering this.
      5. bilbo0s · · focus · HN ↗
        &gt;you can&#x27;t actually make money in China doing China things

        Do you do business in China?

        I&#x27;m curious what you mean by this? Because in my experience, you can only do business in China by doing &quot;China&quot; things.

        I&#x27;d be interested in picking your brain as to how you get around those issues?

        1. curuinor · · focus · HN ↗
          I don&#x27;t do business in PRC anymore, haven&#x27;t for an amount of time that means I don&#x27;t know anything anymore, basically.

          I&#x27;m talking like, getting 100x, VC sized returns. Of course you can sell widgets in China, it&#x27;s a major world economy.

      6. watwut · · focus · HN ↗
        &gt; they cannot coordinate and the state won&#x27;t let them coordinate

        That is how actual capitalist market should work and what anti-monopoly legislation should ensure.

      7. rstuart4133 · · focus · HN ↗
        At the 1000 foot level, this is just another part of the Chinese master plan that&#x27;s been playing out for decades now. The master plan seems to be: train far more engineering and stem graduates than anyone else, let them loose in a dog eat dog capitalist garden and fertilise the garden with sprinkling government money. Hoped for outcome: their innovation will let China eat the world.

        For comparison around 50% of China&#x27;s graduates are in STEM, vs 20% in the USA. With China has a larger population the final outcome is: USA feeds their innovation pipeline with 800k new STEM graduates per year compared to China feeding its pipeline with 4M to 5M per year. In addition the Chinese government outspends the USA government on subsiding innovation in STEM by about 2 to 1 and the USA&#x27;s subsidy rate has been shrinking since the Regan era. However, there is a caveat: USA private investment brings the total spend in both countries back to parity.

        It seems the Chinese plan is working. The STEM garden is overstaffed. It produces far more than China can absorb internally, so they are forced to export their ideas and wares to the world. After decades of persistence China has become the worlds manufacturing powerhouse and now drives the development of 5G&#x2F;6G, batteries, cars, and solar. China builds about 1000 large ships a year. The USA builds 5 to 8.

        But as the article says, this comes at a cost. The intense competition makes life in China&#x27;s dog eat dog STEM garden brutal. I guess that&#x27;s the price a nation has gotta pay for world domination.

    3. foul · · focus · HN ↗
      Market manipulation or slowing down demand for chips for a bit&#x2F;moving the offer elsewhere temporarily?
    4. pj_mukh · · focus · HN ↗
      Occam&#x27;s razor: Going to closed-source just to hide KV-cache optimizations seems silly?
      1. twoodfin · · focus · HN ↗
        This looks like a speed run of the history of analytics DBMS’s.

        Once upon a time, everyone had a secret sauce in network or data encoding or query optimization, but in the last ~10 years computational physics and economics have basically decided the “correct” architecture and everyone (including OSS) has converged.

    5. carbonguy · · focus · HN ↗
      My immediate midwit take is: doesn&#x27;t matter if it helps Anthropic&#x2F;OpenAI if it helps DeepSeek more, relatively. Making open-weights models even cheaper and easier to run expands that &quot;market&quot; and increases competitive pressure on the Big Two, who still have to charge money.
    6. mpalmer · · focus · HN ↗
      They would like to see Western civilization keep getting dumber, and if that means bolstering the success of Western firms, that&#x27;s okay.
    7. TrackerFF · · focus · HN ↗
      My guess would be that if they &quot;help&quot; western labs becoming better, then any break-throughs they (western labs) make after that, is also a benefit to the Chinese labs - if they can distill the models.

      Basically, western labs are in it for the money &#x2F; commercial monopoly. Chinese labs are in it for the tech? As long as they can keep distilling models, and get access to research other ways, they benefit. And if they can push western labs forward, they&#x27;ll benefit from that themselves.

      1. amelius · · focus · HN ↗
        Commoditizing their complement. Makes sense.
    8. jollyllama · · focus · HN ↗
      Where do you think most of the hardware is manufactured, and do you think the hardware manufacturers will keep getting paid if labs start going under?
    9. corford · · focus · HN ↗
      &quot;A week in Beijing and Shanghai with the people building AI in China&quot;: <a href="https:&#x2F;&#x2F;earnedintuition.substack.com&#x2F;p&#x2F;involution-without-export-is-wasted" rel="nofollow">https:&#x2F;&#x2F;earnedintuition.substack.com&#x2F;p&#x2F;involution-without-ex... does a decent job of exploring some possible reasons
      1. Aperocky · · focus · HN ↗
        I tried to read this but it&#x27;s just claude after a certain point.

        It&#x27;s unreadable.

        I can&#x27;t understand any serious writer essentially expecting other people to read stuff that they didn&#x27;t write. Or if claude was actually good at formulating ideas and conducting information, fine. But the truth is that I&#x27;m just reading 80% empty prose that are forced in there by claude.

        1. corford · · focus · HN ↗
          Yeah, sad sign of the times unfortunately.
    10. transdev12 · · focus · HN ↗

      [dead]

    11. chrismarlow9 · · focus · HN ↗
      AI fundamentally insecure. Vulnerable to forcing hallucinations via search results. Vulnerable to invocation of commands in data stream. More AI means more vulnerabilities.

      I can&#x27;t even fathom the trend these days of &quot;we don&#x27;t review the code&quot; from security team perspective.

      Just my guess though.

    12. dsign · · focus · HN ↗
      I think it&#x27;s because they are communists :-) ?

      Well, that&#x27;s probably far from the main thing. The main thing is that they are not Americans. For one, that means that they value profits, but that there may be other items in their bag of wishes. The sort of thing an emotionally healthy human being may desire from super-intelligent AIs, say, curing diseases or new kinds of transgenic rice. Second, their government may be forcing the AI companies to release key IP in order to increase internal competition and prevent monopolies.

    13. Windchaser · · focus · HN ↗
      &gt; Any ideas?

      Unpopular, maybe, but what about the normal reasons? The researchers are looking to make a name for themselves, and&#x2F;or they genuinely care about AI advancement.

    14. feverzsj · · focus · HN ↗
      It&#x27;s just their usual national strategy like what they did to solar pane and EV. The solar pane industry is mostly dominated by China and their profit rate is basically ... negative. The EV industry in China is in similar condition, where the average profit rate is only 1.5%. Their upstream suppliers are also hold as hostages that most of them won&#x27;t get their money back within 6 months.

      The weird ideology here is to dominate the market at ANY COST.

    15. teekert · · focus · HN ↗
      Idk, but it&#x27;s doing a lot for my view of China. Maybe that&#x27;s a point? Maybe they just want their own innovation to go as fast as possible and they don&#x27;t care that other countries also benefit? A rising tide lifts all boats? They are already known for the best manufacturing, they&#x27;re just adding software dev to the list? Maybe they just want to undermine the US in a non-aggressive way?

      Why did we (the west) ever start open sourcing anything? Maybe we just like sharing? Maybe humanity only grows on pre-competitive layers like Linux and clean water. Maybe, the chinese government is closer to their people, and does not let large companies influence them and just doesn&#x27;t like closed private hyperscalers with a lot of power?

      (Some points assume the government has a role in the openness, which I think is likely)

    16. Catloafdev · · focus · HN ↗
      Yes - the Chinese labs serve a market that rely on open-weight models and managed deployments, and the labs gain competitive relevance by releasing those models. The cache optimization feature they came up with required new software to utilize on the inference-end, meaning that open source software would need to be specifically updated to work with these models. It wasn&#x27;t the type of advancement that they could even theoretically keep secret.
    17. thefourthchime · · focus · HN ↗
      Because it&#x27;s entirely possible that Western labs already did this optimization but didn&#x27;t publish it, and then the Chinese figured it out and decided to brag about it.

      We don&#x27;t know either way, so I find the whole thing silly to speculate on.

    18. seydor · · focus · HN ↗
      The chinese don&#x27;t view AI as metaphysical, they view it as an engineering challenge they consider good for their state and want to dominate the global market like they do with batteries&#x2F;EVs&#x2F;photovoltaics. They want to proliferate them as much as possible and traditionally they don&#x27;t care much for IP. They also want hardware makers to make optimized chips specifically for these models.
    19. audunw · · focus · HN ↗
      I think it’s fairly simple: they’re forced into this situation by being late and worse in terms of capabilities. They’re not far behind, but as long as they’re behind they’ve needed to give people some reason to try and use their models. Cost is one factor. But it probably wasn’t enough. Being open has given them a lot of attention. Free marketing. Good will.

      Put another way: if they were not cheaper and open, they would simply not be competitive. They would already be dead.

      I don’t think this ends well for the Chinese labs. This is going pretty much like I thought. Western labs is just copying their improvements (I don’t think publishing the techniques matter here.. they’d just hire to gain the knowledge or figure it out themselves), and they have access to more GPUs and have better branding, so in the end where can the Chinese labs compete? Even lower cost? Open weights? I’m not sure open is a sustainable way to compete either. Eventually there will be some fully open source AI models that cuts out that avenue of competition as well.

      1. itsalwaysgood · · focus · HN ↗
        Western labs are creating the improvements, and Chinese labs are giving them away for free (presumably after examining the improvements). That&#x27;s the jist of the article.
    20. HeavenFox · · focus · HN ↗
      The post makes an assumption that US labs did not already possess similar optimization. It&#x27;s also very possible that they did, but are simply not telling anyone in order to maintain obscene margins on cached read, similar to AWS&#x27; absurd pricing on bandwidth.
    21. micromacrofoot · · focus · HN ↗
      They get to make US labs look dumb and provide an open alternative that anyone can host themselves.

      They&#x27;re building bridges over the moats that companies with far too much US investment are trying to build.

    22. nater5000 · · focus · HN ↗
      Yeah, would have been nice if the author put a bit more thought into this article to come up with something rather than to just give up once they&#x27;ve reached the point of their article lol
    23. ozgung · · focus · HN ↗
      All the comments here are very US&#x2F;Western-centric. Maybe they are a different culture, having a completely different economic model. Maybe they are not Capitalists and not thinking in pure Capitalistic terms, such as winning, growth, market domination, IPO, market value or competition. Maybe they are not obsessed with US labs. Maybe they are ideologically different than you. Maybe they have different priorities. Maybe they never thought of it as throwing a lifeline to American labs. Maybe they don&#x27;t care. Maybe they&#x27;re just different people.
    24. itsalwaysgood · · focus · HN ↗
      As always (and just as the author ended by &#x27;following the money&#x27;), it&#x27;s better for their economy. Extrapolate from there.
  2. Handy-Man · · focus · HN ↗
    Just assumptions, nothing backing it. So maybe I&#x27;d sit out calling others out.

    Edit: Apt domain.

    1. slowin · · focus · HN ↗
      How is it just assumptions? They provide the data to back up their claims.
      1. Handy-Man · · focus · HN ↗
        I am talking about correlating Anthropic&#x2F;OpenAI cache prices going down with Deepseek publication - neither of those labs have said that&#x27;s what they used for example.

        And the only data they are showing is that cache prices went down for new Claude&#x2F;OpenAI models but that&#x27;s proving nothing, IMO.

    2. squidbeak · · focus · HN ↗
      Deepseek&#x27;s innovations are published as research. There&#x27;s nothing &#x27;assumed&#x27; about this. The slur that Chinese labs are parasitic distillers is absurd when so many genuine advances and contributions to the field are published openly by their labs.
  3. nba456_ · · focus · HN ↗
    Appropriate domain name
  4. eggbrain · · focus · HN ↗
    Performance optimizations don&#x27;t just help the western labs, they also help with running more powerful&#x2F;useful LLMs locally.

    If local LLMs get &quot;good&quot; enough, people will soon paying for subscriptions to ChatGPT and Claude, which hurts their revenue.

    1. ericol · · focus · HN ↗
      &gt; people will soon *stop paying

      Think you missed a word there.

    2. kennywinker · · focus · HN ↗
      The only thing preventing this switch from starting in earnest is the data center buildout monopolizing all current and future GPUs
      1. lumost · · focus · HN ↗
        The margins on NVidia datacenter hardware are ... high. At least one order of magnitude larger than a consumer chip.

        Given the recent deepseekv4.1 advances - how good of a 3B model can we make to run on an iphone natively? is it good enough to match common muse&#x2F;dot use cases for consumers? the phone is already always on.. no need for a cloud server.

        1. christkv · · focus · HN ↗
          any model you run on a phone is not going to do wonders for you battery life
        2. zozbot234 · · focus · HN ↗
          Most phones are not really &quot;always on&quot; in any real sense, a phone on active standby uses very little power and most of it is for its mobile connection. Local AI is best run in a stationary homelab environment, even running it on laptops has its very real problems.
        3. kennywinker · · focus · HN ↗
          There are a few edge models, Spark-X2.5-4B, LFM2.5- in 8b-a1b and 2.6b variants, that are useful for on-device agent-y stuff, but I am doubtful that’s a valuable category. How many iPhone users have never opened Shortcuts in their life? Automating on-phone stuff seems niche.

          I don’t expect the economics of local vs cloud ai to change until either the bubble pops or new ai chips land that can run big models fast with low power demand.

          1. lumost · · focus · HN ↗
            I think the interesting question is whether the average consumer cares? a 4B model should hit roughly the same numbers as GPT5.5 sometime in December&#x2F;March. That is more than capable of performing a variety of complex work tasks.
            1. kennywinker · · focus · HN ↗
              Skeptical 4B will ever rival GPT5.5, with its estimated 1.5-10T parameters.

              You just can’t compress 1.5T into 4B without losing useful stuff.

              But that doesn’t matter that much. If 4B can compose tools, retrieve info, and not get stuck in loops that is going to handle a lot of use cases

    3. londons_explore · · focus · HN ↗
      For nearly all tasks, I want the fastest and smartest AI model.

      It is vanishingly rare I ask an older model to do any task. Newer bigger and smarter models will just do the task better.

      Therefore, I believe we are nowhere near &#x27;good enough&#x27;.

      I never drive my steam engine to work these days. It isn&#x27;t good enough.

      1. danielmarkbruce · · focus · HN ↗
        I find this too. In fact, recently I&#x27;ve been pushing more and more to the latest and greatest model every time there is an update. It just saves me so much headache.
      2. NoDodgeQuestion · · focus · HN ↗
        I drive my 2018 car to work these days. It is good enough.
      3. eggbrain · · focus · HN ↗
        Right now you are right -- even if I ran my local LLM all day, the quality is not nearly as great, and it runs slowly -- so I use the tier one AI subscription services as they are faster and smarter. But that might only be true for a limited amount of time, and a limited number of circumstances.

        To borrow your steam engine analogy, if local LLMs get as good as a Toyota Prius, even if OpenAI &#x2F; Anthropic offer Ferraris, most people will be happy with their Prius as their daily driver.

        Similarly, if the big labs start raising prices or cutting usage, you won&#x27;t be able to use it as much as you want -- whereas a local LLM will run all day every day without costing you any extra money.

        Right now you are right -- even if I ran my local LLM all day, the quality is not great, and it runs slowly -- so I use the subscription AI services as they are faster and smarter.

      4. gretch · · focus · HN ↗
        &gt; For nearly all tasks, I want the fastest and smartest AI model.

        This is not nearly true for everyone else in the world.

        For example, think about the world in ~2021 pre-LLM. Would anyone say the sentence &quot;I only want the fastest and smartest humans working on my project&quot;?

        No of course not. Most people don&#x27;t want to pay $10 million dollar salary to the best programmers in the world. They prefer to pay $200k salary to a median programmer and that&#x27;s good enough for their ecommerce website.

        1. tyre · · focus · HN ↗
          I mean people said this all the time, but wouldn’t pay for it (as you said), couldn’t recruit for it, and definitely couldn’t retain them.

          But so many teams said they wanted to Raise the Bar to infinity and hire a World Class Team.

      5. settsu · · focus · HN ↗
        [delayed]
      6. monkpit · · focus · HN ↗
        I think a scooter or a bus is a more apt comparison than a steam engine. The scooter and bus can both get the job done with some acceptable trade-offs, depending on your circumstances and what you’re willing to accept.
      7. schmookeeg · · focus · HN ↗
        This will change when out of control billing gets noticed. Then us devs will, I presume, get token rations.

        I can see the new Thursday afternoon &quot;oh chit&quot; moment being that I didn&#x27;t complete my weekly task because i torched all of those tokens M-W doing task&#x2F;ticket grooming using the hot hot model :D

    4. aleqs · · focus · HN ↗
      I would argue performance optimizations help with local&#x2F;open models but hurt openai and anthropic - because open and or cheap&#x2F;alternative models are threat to those companies. There is a fundamental contradiction&#x2F;conflict between the prevelance of open models and the financial success of openai and anthropic. That is why they are doing everything they can to kill any open&#x2F;cheap&#x2F;efficient&#x2F;chinese models (take a look at this thread - it was top of HN 40 mins ago, with very high engagement... now it is buried in page 5... totally normal and legit).
  5. jgrahamc · · focus · HN ↗
    The first sentence is: &quot;If you read the news headlines these days, you would be forgiven for thinking that the Western labs are getting spawn-camped by Chinese labs en masse.&quot;

    Since I have no idea what &quot;spawn-camped&quot; means I gave up reading the rest.

    1. some_furry · · focus · HN ↗
      It&#x27;s gamer terminology.

      In a PvP (player vs player) game, if you kill a player the moment they spawn into the game arena, that&#x27;s called &quot;spawn-camping&quot;.

    2. emilecantin · · focus · HN ↗
      It&#x27;s from video games, where a player &quot;camps&quot; near the spawn point and kills newly-spawned players, presumably with better equipment.

      It&#x27;s not that niche, if you&#x27;ve been online a little bit you&#x27;d know this expression.

      1. cassianoleal · · focus · HN ↗
        &gt; if you&#x27;ve been online a little bit you&#x27;d know this expression

        I&#x27;ve been online since circa 1995 (earlier if you count BBSs), and I can&#x27;t say I did. It&#x27;s possible to infer its meaning but assuming everyone is on the same circles as one is, is silly.

      2. tejohnso · · focus · HN ↗
        &gt; It&#x27;s not that niche, if you&#x27;ve been online a little bit you&#x27;d know this expression.

        No way. You&#x27;d need to be pretty well versed in gamer lingo.

      3. jgrahamc · · focus · HN ↗
        I dunno, man, I first got on the Internet in 1986 and was Cloudflare&#x27;s CTO for years. I&#x27;ve been online quite a bit.
        1. off_with_their_ · · focus · HN ↗

          [dead]

        2. johnthescott · · focus · HN ↗
          made my day.
        3. pigpop · · focus · HN ↗
          Due to shifting definitions, I believe that makes you a boomer[0] so you&#x27;d be readily excused for not knowing.

          Sarcasm aside, gaming and FPS terminology are so tightly coupled with online culture that it&#x27;s just assumed everyone knows it. Spawn camping is among the oldest examples of gaming terms that broke out into common online usage and it dates back to Quake some time around 1997.

          [0] anyone older than 39 at this point

          1. tejohnso · · focus · HN ↗
            &gt; FPS terminology are so tightly coupled with online culture

            What is online culture? Is someone who is heavily into instagram for fashion, facebook for family contact and news, maybe Google for mail and search, part of online culture? Because I know people who are like that and there&#x27;s no way they know what &quot;spawn&quot; or &quot;camping&quot; mean in gamer context and certainly wouldn&#x27;t be able to piece together what &quot;spawn camping&quot; is.

            1. pigpop · · focus · HN ↗
              Those things seems fairly obviously not online culture but parts of &quot;real life&quot; culture that have an online presence, obvious to me at least. Video games and especially multiplayer video games are a core part of what you could call the indigenous online culture and serve the role of sport and to a large degree casual socialization. Their equivalents in the real world are team sports, board games and party games. These things aren&#x27;t mutually exclusive and they certainly bleed over both ways but you can easily say that a game like Counter Strike is primarily online and a game like Basketball is primarily offline even though there are Basketball themed video games and Counter Strike themed in-person events.

              To put it simply, it&#x27;s a Venn diagram with two overlapping, non-coincident circles. The amount of overlap has varied with time but that doesn&#x27;t mean online culture doesn&#x27;t exist or conversely that everything is part of online culture.

          2. khazhoux · · focus · HN ↗
            &gt; gaming and FPS terminology are so tightly coupled with online culture that it&#x27;s just assumed everyone knows it

            Strong consensus bias in this statement. I have no doubt it is true in your circle (and to a lot of people). But there are plenty of online-24&#x2F;7 cultures that don’t have any overlap with gaming.

        4. VGHN7XDuOXPAzol · · focus · HN ↗
          (tongue-in-cheek) In that long time you&#x27;ve been online, have you not come across the concept of a search engine?
      4. mjc26 · · focus · HN ↗
        Also, it&#x27;s possible that jgrahamc is the former CTO of cloudflare
    3. khazhoux · · focus · HN ↗
      In first-person shooters, when you die you regenerate (“respawn”) somewhere on the map. If those regeneration points are know to opposing players, they can wait next to them, and kill you again the moment you respawn.

      The premise in this article is: Western companies do a ton of expensive work building new models, meanwhile the Chinese companies just wait for a Western release and then they immediately grab and distill it and announce it as their own model. That’s the spawn-camp.

    4. Intragalactic · · focus · HN ↗
      If you tend to stop reading every time you encounter something you don&#x27;t understand, I can&#x27;t imagine you learn very much

      &quot;spawn-camping&quot; is the process of taking out your enemies at the point they spawn (or appear) in a game without giving them a chance to regroup. In this case I think the writer is saying that the news implies that western models are getting distilled on release. Not the perfect analogy but it gives some color.

    5. bsoqk · · focus · HN ↗
      We live in the era of LLMs, which can produce definitions for any word, further explanations, and limitless examples.
      1. jgrahamc · · focus · HN ↗
        Yes, but to expand on my flippant response, the opening of this post is a sign of bad writing. And I don&#x27;t want to waste my time on bad writing.

        It shows that the author hasn&#x27;t thought about their audience and has assumed that everyone knew and used the same terminology as them. And, worse, assumed that they&#x27;d understand immediately why they were using that terminology.

        The opening is If you read the news headlines these days, you would be forgiven for thinking that the Western labs are getting spawn-camped by Chinese labs en masse. If you don&#x27;t know what spawn-camping is you&#x27;re lost; if you do it&#x27;s not obvious what that means in this context. Good writing brings the reader along with the writer.

        It would have been clearer if they&#x27;d written: If you read the news headlines these days, you would be forgiven for thinking that Western AI labs are being outplayed by Chinese AI labs using something similar to the gamer technique of &quot;spawn-camping&quot;. The Chinese appear to be waiting for each Western release and then instantly distilling it. A little like gamers waiting for their opponents to reappear (from the dead) at their camp and then kill them off immediately.

        This is better because a reader unfamiliar with the idea of spawn-camping learns something and it explains the metaphor. But I could be wrong in my interpretation of why they are using spawn-camping since they fail to explain it.

    6. jamiek88 · · focus · HN ↗
      So you had the opportunity to learn a new phrase but instead discarded the whole text because of something you hadn’t encountered before?

      Do you have many mini tantrums like this per day? Do you sleep ok? Doesn’t seem rational behavior.

  6. slowin · · focus · HN ↗
    I&#x27;m also grateful to the Chinese labs for providing workarounds for the walled gardens that the US based AI companies are attempting to create.

    Does anyone know if there are any distillation datasets available? I&#x27;d love to see these distributed on BitTorrent. I think it&#x27;s critical that AI be democratized and not isolated in the hands of a few private companies.

    1. [deleted] · · focus · HN ↗

      [deleted]

    2. skybrian · · focus · HN ↗
      There’s a libertarian sentiment that that doesn’t sit well with “AI is harming people” sentiment. If AI has harmful uses, and I think anyone sensible would have to agree that it does, then giving everyone unrestricted AI is likely to make it worse.

      It’s sort of like gun nuts arguing that more guns is the answer. I mean, ok, maybe you’re a responsible gun owner or AI user but relying on personal responsibility doesn’t fix systemic problems. There are bad people out there.

      1. iamnothere · · focus · HN ↗
        All concerns balance against competing concerns, and in this case freedom of computing and knowledge wins over safety. Especially since it’s trivial to copy and share open models.
        1. skybrian · · focus · HN ↗
          [delayed]
          1. short_sells_poo · · focus · HN ↗
            I agree that it isn&#x27;t an unavoidable tradeoff in principle, but looking back at our (as in humanity) track record, it is 99% likely to be.
            1. skybrian · · focus · HN ↗
              [delayed]
          2. iamnothere · · focus · HN ↗
            [delayed]
            1. skybrian · · focus · HN ↗
              [delayed]
              1. iamnothere · · focus · HN ↗
                [delayed]
                1. skybrian · · focus · HN ↗
                  I wouldn’t say the Internet is a disaster either, but there are disasters where people died due to factories exploding, letting off toxic gases (like in Bhopal), and so on. That’s why there are regulations to try to prevent industrial accidents.

                  For Internet-related disasters where people died, see [1].

                  I would bet that there will be AI-related disasters. Arguably the US bombing a school in Iran counts, though it seems to be due to organizational issues, too.

                  [1] <a href="https:&#x2F;&#x2F;chatgpt.com&#x2F;s&#x2F;t_6abd4e3226b48191a26a0fe6768c722f" rel="nofollow">https:&#x2F;&#x2F;chatgpt.com&#x2F;s&#x2F;t_6abd4e3226b48191a26a0fe6768c722f

        2. tonyedgecombe · · focus · HN ↗
          Actually I think profits win over safety.
      2. bronson · · focus · HN ↗
        If food has harmful uses, and I think any one sensible would have to agree that it does, then giving everyone unrestricted food is likely to make it worse.

        You can do this with cars, tools, computers, ... whatever you want. So, no, I think your point is wrong.

        1. skybrian · · focus · HN ↗
          [delayed]
          1. hyperlinerapp · · focus · HN ↗
            Guns are highly regulated. Try getting a gun in liberal California, where Reagan screwed us.

            Now, what I want to regulate are accordions.

            1. skybrian · · focus · HN ↗
              [delayed]
              1. rootusrootus · · focus · HN ↗
                [delayed]
                1. plorkyeran · · focus · HN ↗
                  Which is to say that it’s a very burdensome process if you think you should be able to go to Walmart, exchange cash for goods, and walk out with a gun, but pretty simple by the standards of regulated goods. The basic knowledge test is very basic, so the pain points are just that sometimes background checks are incorrect, the waiting period can be inconvenient, and just the general annoyances of government bureaucracy.
                  1. hyperlinerapp · · focus · HN ↗
                    A gun is not a good. It’s a right.

                    It’s not like we ask to wait for 2 weeks when you have something to say or when you want to pray to your god.

                    1. stickfigure · · focus · HN ↗
                      &gt; It’s a right.

                      Not if you&#x27;re a felon or have mental health issues.

                      1. hyperlinerapp · · focus · HN ↗
                        It’s interesting that the SCOTUS still hasn’t decided why non violent felons who have served their sentence get their 2A rights cancelled.
                    2. rootusrootus · · focus · HN ↗
                      [delayed]
                      1. hyperlinerapp · · focus · HN ↗
                        Yes, we define crimes based on what you say. For example, libel or yelling “fire” at a movie theater.

                        But you are allowed to say whatever you want.

                        There are gun regulations too. Most are reasonable, but many are not. Liberals think that they can simply restrict the right to guns unconstitutionally because “guns are scary.” That is clearly not a sufficient basis to deny a right.

                        In the case of LA County, they just settled with the Feds because they had approved 2 concealed carry permits out of thousands of applications. That clearly shows the government was simply curtailing your right.

              2. hyperlinerapp · · focus · HN ↗
                It depends on the county. Some Sheriffs think it’s them who gives you your RIGHT to own a gun or a concealed weapon license. They don’t understand they are servants.

                LA just settled the Fed lawsuit over their terrible practices:

                <a href="https:&#x2F;&#x2F;www.latimes.com&#x2F;california&#x2F;story&#x2F;2026-08-13&#x2F;doj-la-sheriffs-department-settle-lawsuit-concealed-carry-firearm-permit-application-delays" rel="nofollow">https:&#x2F;&#x2F;www.latimes.com&#x2F;california&#x2F;story&#x2F;2026-08-13&#x2F;doj-la-s...

            2. bittercynic · · focus · HN ↗
              I have purchased a gun in California, and I thought the process was pretty reasonable. Maybe even too lax.
          2. card_zero · · focus · HN ↗
            Safety laws are the wrong category, the equivalent regulations would be those that forbid the use of cars for drive-by shootings or as robbery getaway vehicles, and regulations against the use of food to provide crime energy.
            1. skybrian · · focus · HN ↗
              [delayed]
        2. bobmcnamara · · focus · HN ↗
          Nice try Philipp Mainländer!
      3. hamdingers · · focus · HN ↗
        I simply don&#x27;t trust the people who would decide who gets AI (or guns) to make good choices.
        1. skybrian · · focus · HN ↗
          This is a common populist sentiment, but if you don’t trust anyone then nothing can be done. Is it just game over?
          1. hamdingers · · focus · HN ↗
            I didn&#x27;t say I don&#x27;t trust anyone. Don&#x27;t put words in my mouth, it&#x27;s a sign of bad faith.

            Other countries have governments that have earned that level of trust. I believe the US could get there eventually, but it will take a very long time because it has a very long way to go.

            1. skybrian · · focus · HN ↗
              Okay, sorry about that. But how do we start? Maybe there are AI safety organizations that deserve our support?
          2. iamnothere · · focus · HN ↗
            [delayed]
            1. skybrian · · focus · HN ↗
              [delayed]
              1. iamnothere · · focus · HN ↗
                [delayed]
                1. skybrian · · focus · HN ↗
                  Viruses don&#x27;t need to be intelligent at all and smart viruses (or AI internet worms) won&#x27;t need to be superintelligent to be a real pain to deal with. DoS attacks due to scraping are bad enough already.

                  The situation is sort of like the Internet before broadband. There probably aren&#x27;t enough AI-capable home machines to have big swarms of bots that run autonomously. If the swarm depends on an LLM API, it will be easier to cut it off once it&#x27;s noticed.

                  I think it&#x27;s an even chance that we will see a bot swarm in the wild (rather than coming from an AI lab) by the end of next year.

                  1. iamnothere · · focus · HN ↗
                    Right, but the internet survived the Melissa virus and other early widespread worm attacks. They caused widespread disruption and economic damage, but they weren’t the end of the world. And besides, at present people are using AI to find and patch software at a faster rate than they are using it to compromise software (more dollars are allocated to security these days).

                    IMHO the biggest problems will be with poorly written AI slop software, probably small business crapware produced with minimal investment (maybe entirely without an experienced developer) and legacy equipment that’s been abandoned by the manufacturer. This will lead to disruption but not catastrophe.

          3. randbyte · · focus · HN ↗
            Certainly no “anyone” and who happens to be worse than no one.
          4. owebmaster · · focus · HN ↗
            This is a common populist argument
      4. slowin · · focus · HN ↗
        I would say that there are very few people on earth that I trust less than Sam Altman, Dario Amodei and Elon Musk. Also my own government claims to have used Anthropic models to bomb a girls&#x27; school in Iran. If you combine US regulations with sociopathic private companies, you get into a worst case scenario for humanity imho. Again, I&#x27;m thankful to China or any other entity pushing open models, local models and even distribution of this technology.
        1. skybrian · · focus · HN ↗
          Maybe there are other organizations that deserve our support?
          1. slowin · · focus · HN ↗
            I&#x27;m happy to support anyone pushing for open models, local models and publishing their innovations out in the open.
      5. hyperlinerapp · · focus · HN ↗
        Gun owner enters the conversation. In high trust societies, armed people are very polite people. I don’t want the bad guys out there being the only ones with guns. Besides I like to shoot just like you like (whatever you like to do that is legal).

        Replace “gun” with anything and you will see how your comment falls apart.

        What’s next? A registry for food purchases? Your beer gut is starting to show.

        1. skybrian · · focus · HN ↗
          [delayed]
          1. hyperlinerapp · · focus · HN ↗
            Unlike food, guns are your right.
            1. skybrian · · focus · HN ↗
              [delayed]
              1. hyperlinerapp · · focus · HN ↗
                Nothing that some other person has to give you can be a right.
        2. BigTTYGothGF · · focus · HN ↗
          &gt; In high trust societies, armed people are very polite people

          Surely you could name three such societies?

          1. hyperlinerapp · · focus · HN ↗
            Just go to ChatGPT and ask:

            List 20 high trust societies

            1. BigTTYGothGF · · focus · HN ↗
              I&#x27;m asking you, and furthermore I&#x27;m not asking for a list of high trust societies, I&#x27;m asking for ones in which &quot;armed people are very polite people&quot;.
              1. hyperlinerapp · · focus · HN ↗
                Switzerland, Wyoming, Idaho, Alaska, Montana, Czech Republic, Finland, Norway, most Rural America.

                Anyway, ChatGPT can answer your questions. Or just visit your local gun range. I highly recommend it.

                1. BigTTYGothGF · · focus · HN ↗
                  I grew up in rural America.

                  The original claim was: &gt; in high trust societies, armed people are very polite people

                  In my experience, there was no connection between being armed and being polite.

                  &gt; Or just visit your local gun range

                  Went to one many times when I lived in Texas, I think I&#x27;ve inhaled enough lead to last me the rest of my life.

                  &quot;An armed society is a polite society&quot; was originally said by Heinlein, who for all his faults would have had something to say about outsourcing your thinking to a machine.

                  1. hyperlinerapp · · focus · HN ↗
                    I grew up in Rural America too. In my experience, armed people are very polite.

                    I don’t outsource my thinking to a machine. I was asking you to not outsource your thinking to a stranger, since I am not your private encyclopedia.

      6. nutjob2 · · focus · HN ↗
        &gt; giving everyone unrestricted AI is likely to make it worse

        Worse for whom?

        The only effective defense against predatory corporate and government AI is personal protective AI.

        Anything else is unilateral disarmament. It&#x27;s the only way individuals can survive in the worse case scenario.

    3. ducktective · · focus · HN ↗
      &gt;distillation datasets

      You ask about distillation but I wonder, is there any training datasets (~ TB-order) available that startup folks in SV use or is it so that everyone has to create their own scraping pipeline ?

      1. forshaper · · focus · HN ↗
        There are several? And there exists companies whose entire business is just providing them? iirc
      2. frabcus · · focus · HN ↗
        It&#x27;s particularly important we all scan and destroy our own unique books!
      3. atherton94027 · · focus · HN ↗
        Given the amount of people complaining about crawlers in the past 2 years, I think it&#x27;s the latter
    4. 10xDev · · focus · HN ↗
      An authoritarian regime is not your friend and will pullback the moment their own models become highly capable.
      1. 4gotunameagain · · focus · HN ↗
        While your friend is Sam Altman, or US megacorps ?

        Or did they not pull back when their models allegedly became highly capable, with the whole mythos debacle ?

      2. CodingJeebus · · focus · HN ↗
        This is equally true for US AI
      3. horsawlarway · · focus · HN ↗
        Yes, we already discussed the US.
        1. Avicebron · · focus · HN ↗
          It&#x27;s crazy how articles like this get spawn-camped by people like this trying to throw this zinger in. Both AI conpanies in the US and the chinese companies with ccp desks in the corner can be bad. The good path forward is locally hosted AI models, that&#x27;s known.
          1. horsawlarway · · focus · HN ↗
            Right, which is why I&#x27;m happy to see China continue to innovate in the open, and increasingly wary of the US stance given articles like

            <a href="https:&#x2F;&#x2F;www.anthropic.com&#x2F;research&#x2F;glm-5-3-and-the-spread-of-advanced-cyber-capabilities" rel="nofollow">https:&#x2F;&#x2F;www.anthropic.com&#x2F;research&#x2F;glm-5-3-and-the-spread-of...

            It&#x27;s VERY clear that the US companies are trying to push for regulation to kill open models and open weights. I see this as much more hostile and authoritarian response than what we&#x27;re seeing come out of China right now.

            So is China going to always publish in the open? No clue. But right now they&#x27;re modeling much better behavior.

            1. jacquesm · · focus · HN ↗
              What&#x27;s on my drives stays on my drives.
      4. slowin · · focus · HN ↗
        There&#x27;s no &quot;pulling back&quot; things that have already been open sourced.
        1. Ao7bei3s · · focus · HN ↗
          Open weights isn&#x27;t open source, it&#x27;s freeware.
        2. HappyPanacea · · focus · HN ↗
          Intelligence wants to be free
        3. abirch · · focus · HN ↗
          Unfortunately most of the Chinese models are open weights and not open sourced.
          1. slowin · · focus · HN ↗
            I agree it would be very cool if they were open sourced, but the article is talking about the KV cache technique which if I understand correctly is open source (or at least the white paper is published).
      5. ahriad · · focus · HN ↗
        Chill, Buddy. Why are you so anti-American?
      6. pksebben · · focus · HN ↗
        Oh no, they might stop doing the thing that benefits me and that they were never required to do in the first place.
        1. stackedinserter · · focus · HN ↗
          It can be predatory pricing that will bite us later.
          1. wat10000 · · focus · HN ↗
            It really can&#x27;t be when the models are open weights and there&#x27;s zero lock-in for inference providers.
            1. fsflover · · focus · HN ↗
              Open weight is not open source. See: freeware vs free software.
              1. wat10000 · · focus · HN ↗
                Ok, so?
                1. fsflover · · focus · HN ↗
                  You have the lock-in, as you can&#x27;t modify the freeware. You will be stuck will all its anti-features.
                  1. wat10000 · · focus · HN ↗
                    What does the ability to modify a model have to do with lock-in?

                    Models aren’t software. Software has a learning curve and costs of adoption. Models and inference are completely fungible.

                    1. fsflover · · focus · HN ↗
                      &gt; What does the ability to modify a model have to do with lock-in?

                      For example, Chinese models can spread the CCP propaganda and censor the results. You can&#x27;t change it and may not even notice.

                      &gt; Models aren’t software.

                      Models are software. They run on computers, require development, can have bugs etc.

                      &gt; Models and inference are completely fungible.

                      Until the next version of DeepSeek becomes closed and you will only have the access to the last good open-weight model.

                      1. wat10000 · · focus · HN ↗
                        Sorry, are you actually paying any attention to the discussion? You’re spouting a bunch of buzzword non sequiturs. You started off by talking about how open weights aren’t open source when nobody said anything about open source. Now you’re talking about propaganda in response to a question about lock-in.

                        Lock-in means you can’t switch to a different model or a different provider because the cost (in time or money) of switching is high. Censorship and propaganda have nothing to do with it. You can switch to DeepSeek in an instant. You can switch away from it in an instant. The switching cost is zero.

                        And no, models are data files. The inference engines are software. Models run on many engines and most engines support many models. Another example of how there’s no lock-in: you can switch both models and engines at no cost. If you’re using an inference provider, they probably offer many models you can switch between, and you can switch to a different provider with no effort.

                        If the next version of DeepSeek becomes closed then you can switch to GLM or Kimi or MiniMax or whatever else in two seconds, or keep on using the last open version of DeepSeek. Zero lock-in.

          2. tornikeo · · focus · HN ↗
            How exactly will downloaded gguf files bite me later?

            Is it going to shut the laptop&#x27;s lid when i&#x27;m not looking and pinch my fingers?

            1. stackedinserter · · focus · HN ↗
              If the only thing that bothers you is laptop lid pinching your fingers, then no – it won&#x27;t do it.
          3. KolibriFly · · focus · HN ↗

            [dead]

      7. nutjob2 · · focus · HN ↗
        China is much more authoritarian than the US, but at this point it&#x27;s like comparing two types of metastatic cancer.

        The point is get what you can from both to develop open models, data and tools.

      8. randbyte · · focus · HN ↗
        Just like how Anthropic and OpenAI is already doing?
      9. feverzsj · · focus · HN ↗
        It&#x27;s already happening now. They just banned engineers of their top AI labs and their families from leaving the country.
      10. computerex · · focus · HN ↗
        As opposed to what? The US? You think the US is any different? Literally our pedo president publicly admits to insider trading. You think the US government gives a rat&#x27;s ass about the American people?
        1. htx619 · · focus · HN ↗

          [dead]

        2. dirck-norman · · focus · HN ↗
          The independent media reports on this and that’s why you know.

          China has zero independent media and has the largest and most sophisticated censorship and surveillance state in the world.

          Trump would love to have the absolute power that Xi has, but thankfully doesn’t.

          Tired of this lazy whataboutism.

          1. computerex · · focus · HN ↗
            No, actually trump bragged about it on x.

            Trump kicked and banned major news orgs from the White House just now. You’d have to be will fully ignorant or really naive to believe we have fair and just media.

          2. ted_bunny · · focus · HN ↗
            Whataboutism is the rallying cry of every sophist who can dish it but not take it. Whataboutismism
            1. computerex · · focus · HN ↗
              My comment doesn&#x27;t even fall into the category of whataboutism....
          3. computerex · · focus · HN ↗
            Btw how much media coverage have you seen in America related to the genocide of Gaza ? Jewish terrorism in West Bank?

            That’s what I thought. There is a reason independent journalists on YT are having great success right now because legacy media is trash, bought and paid for.

            Why do you think journalists are creating shows like “Mehdi Hasan unfiltered”?

            You can’t in goodwill with a straight face say American media is just. Western media overall is bought and paid for.

            1. dirck-norman · · focus · HN ↗
              &gt;Btw how much media coverage have you seen in America related to the genocide of Gaza ? Jewish terrorism in West Bank?

              &gt;That’s what I thought.

              Wow, you got me with your nonsequitur strawman argument that you answered yourself and isn&#x27;t even true. It&#x27;s easy to identify sophists with how quickly the bring up Israel out of nowhere.

        3. Jskewel · · focus · HN ↗
          Saying the Chinese and US governments are the same is so absurd it doesn&#x27;t warrant a reply.
          1. voiceofchoice · · focus · HN ↗
            It&#x27;s true, at least the Chinese government is competent in their authoritarianism.
          2. computerex · · focus · HN ↗
            Then why did you respond with a useless self aggrandising comment?

            Love the arguments btw. Your position is so cut and dry that you can’t even support it with evidence.

          3. xtracto · · focus · HN ↗
            For someone (as me) in Mexico, they are exactly the same.

            Third world societies are not &quot;brainwashed&quot; one way or another with respect to that.

            We just see those governments throwing shade at each other. And bullying us ina similar way.

            1. customguy · · focus · HN ↗
              They are not what they are to you, they are not just whatever you care about, they are what they actually are.
              1. computerex · · focus · HN ↗
                Do you see the self-centeredness of your comment? Why is his (Mexican) perspective wrong? Just because you say it is?
                1. customguy · · focus · HN ↗
                  Do you not see you&#x27;re projecting? If I say two different types of bird are &quot;the same to me&quot; because they both shit on my head, that doesn&#x27;t make them the same birds, that just means I only care about what they are to me, i.e. the self-centeredness you accuse me of.

                  You said &quot;for me they are the same&quot;. Not in reply to someone saying &quot;to me they&#x27;re different&quot;, but to someone saying they&#x27;re very different in nature, which they are. That&#x27;s not a perspective, that&#x27;s changing the subject and making it about you.

                  1. computerex · · focus · HN ↗
                    Do you even understand what you wrote? Because it&#x27;s practically incoherent.

                    &gt; They are not what they are to you, they are not just whatever you care about, they are what they actually are.

                    You wrote that. It&#x27;s a useless comment, adds nothing to the discussion and reads like you are trying to be clever and failing at it. Learn to articulate.

                  2. xtracto · · focus · HN ↗
                    &gt;You said &quot;for me they are the same&quot;. Not in reply to someone saying &quot;to me they&#x27;re different&quot;

                    You are arguing with another user; it was me who said it.

                    They have some differences, yet they are the same in a lot of ways. Some people cannot see it because they have been educated a certain way. Lots of people in China really believe they are right and the US is wrong. Same in Korea, Same in Rusia. Similary, a lot of people in the US are educated to believe the US is &quot;the best country in the world&quot;, &quot;the good guys&quot;, &quot;always right&quot;; they are not.

                    1. customguy · · focus · HN ↗
                      Ah, sorry for getting you two mixed up.

                      I can agree to to it with the qualifier &quot;in a lot of ways&quot;, just not without. We learn about a lot of US crimes by the reporting of US citizens and journalists, a lot of which is available to buy and re-print within the US. Meanwhile, people in China still cannot simply remember Tiananmen, much less publish books about it.

        4. JohnnyMarcone · · focus · HN ↗
          The US has a structure that makes it possible to change the government, China has president for life and one party rule.
          1. computerex · · focus · HN ↗
            You say that as our pedo president is testing waters by doing another term and his republican lackeys are supporting him on this.

            Also at the end of the day what matters is whether the government serves the people and the peoples quality of life.

            Quality of life for the Chinese people has steadily increased. Quality of life for the American people has steadily decreased.

      11. sillyfluke · · focus · HN ↗
        The irony is clearly lost on you.
      12. layer8 · · focus · HN ↗
        You can be grateful to an enemy.
      13. sixo · · focus · HN ↗
        It is in the interest of everybody-except-OpenAI and Anthropic that those companies have capable competitors; better still for the competitors to share their advances.

        They don&#x27;t have to be our friends to act in our interest.

      14. unrented7977 · · focus · HN ↗
        Agreed, the US is not your friend and will pull the rug from under you at the first convenient opportunity.
      15. caaqil · · focus · HN ↗
        &gt; An authoritarian regime is not your friend

        Absolutely. In fact, the authoritarian regime already did try to force export controls on major frontier labs not long go so this isn&#x27;t a theoretical.

      16. voiceofchoice · · focus · HN ↗
        Every workplace is an authoritarian regime where workers don&#x27;t have a say and grovel with &quot;please don&#x27;t replace me&quot; like that isn&#x27;t the entire point.

        The whole point of AI is to get rid of you so rich people can play with the planet like it&#x27;s minecraft. Software engineers just think they&#x27;re special because they&#x27;re the ones building it like they&#x27;ll get a pat on the head for being good little servants to the investor class. Or worse, that their portfolios will let them be the gods who rule over ashes.

      17. dgellow · · focus · HN ↗
        Ironically the US is the only country that has done that
      18. negura · · focus · HN ↗
        Their models are already highly capable. They have enough leverage to act like scumbags, yet unlike Americans, they lack the compulsion for it. That authoritarian regime you fearmonger about has in fact done an excepetional job for its people over the last half century.
      19. hndweller1337 · · focus · HN ↗
        At least their non friendly AIs aren&#x27;t bombing schools.
      20. MetroWind · · focus · HN ↗
        At that point the US will complete its China conversion and turn on the GFW so you can&#x27;t access Chinese AI services.
    5. joe_the_user · · focus · HN ↗
      The Chinese models are to an extent that distillation data.
    6. derpyzza · · focus · HN ↗
      there&#x27;s <a href="https:&#x2F;&#x2F;pirateface.co&#x2F;" rel="nofollow">https:&#x2F;&#x2F;pirateface.co&#x2F; which is like huggingface but distributed via torrents
    7. hn_throwaway_99 · · focus · HN ↗
      &gt; I think it&#x27;s critical that AI be democratized and not isolated in the hands of a few private companies.

      While I&#x27;d like to agree with this, the fact is that pushing the frontier out has always taken (and folks expect to continue to take) hundreds of millions&#x2F;billions of dollars. Open source and distilled models can follow on for much cheaper, but it&#x27;s hard to imagine the frontier ever being &quot;democratized&quot; given the huge sums of money required. It was this realization that forced OpenAI to take tons of private investment in the first place.

      1. jacquesm · · focus · HN ↗
        That&#x27;s ok. They took a few trillion worth of content and burned a couple of hundred billion of their own money. I really don&#x27;t see the problem, if they didn&#x27;t think about this long and hard before they went down that part it should not be on the rest of us to bail them out. The &#x27;frontier&#x27; is less important than the democratization process.

        The amount of progress that came out of academia and other public sources should not be underestimated either and without all that OpenAI and Anthropic wouldn&#x27;t even exist.

      2. pegasus · · focus · HN ↗
        This would be true if Chinese models would be strictly distillations of frontier models, but in fact they also train on Chinese data Western models don&#x27;t have access to.
  7. emtel · · focus · HN ↗
    As far as I can tell, neither of the frontier US labs have referred to distillation as &quot;stealing&quot;, but someone please provide a link if I&#x27;m wrong.

    They do claim that it violates their ToS, which we can assume is simply correct, since they get to put whatever they want in their ToS.

    Given all that, I don&#x27;t know what the fuss is. Are they supposed to not use the advances that were openly published by Chinese labs? The entire industry is built on a discovery made at Google, which was published openly. Should Chinese labs therefore not use transformers? Should US labs not try to prevent distillation of their models?

    1. dgellow · · focus · HN ↗
      I don’t think there is fuss, just the author sharing the information and mentioning how they find it a bit ironic that US labs expenses can be reduced drastically thanks to the Chinese companies they continuously frame as adversaries
    2. jerrygenser · · focus · HN ↗
      I&#x27;m not sure if they don&#x27;t refer to it as &quot;stealing&quot; but they refer to &quot;distillation attacks&quot;
  8. [deleted] · · focus · HN ↗

    [deleted]

  9. adamrezich · · focus · HN ↗
    OpenAI is jobbing (in professional wresting terminology) hard right now.
    1. brcmthrowaway · · focus · HN ↗
      So who is the kayfabe?
      1. adamrezich · · focus · HN ↗
        Did you not see the meeting with the President yesterday? All of the “safety discourse” is kayfabe.
  10. reedf1 · · focus · HN ↗
    I&#x27;ve been running Qwen 3.8 27b (an opus 4.6 tier model), locally on a 5090 for just over two weeks @ 170 tokens&#x2F;s. That&#x27;s a frontier model from 9 months ago running on consumer hardware. Who knows where distillation and pruning gets us in another year.
    1. redanddead · · focus · HN ↗
      Well how’s it been so far
      1. off_with_their_ · · focus · HN ↗

        [dead]

    2. oidar · · focus · HN ↗
      What are you thoughts on it&#x27;s performance compared to 4.6?
      1. reedf1 · · focus · HN ↗
        Indistinguishable or very mildly better. But it&#x27;s considerably faster. Some portion of that is also probably down to improvements in model harnesses, I&#x27;ve been using opencode.
        1. jeffrallen · · focus · HN ↗
          Yeah, the Shelley agent (from exe.dev) loves Qwen 3.8, they kicked ass on a Django app for me today.
    3. bix6 · · focus · HN ↗
      $9k for a 5090 now? Sheesh.
      1. off_with_their_ · · focus · HN ↗
        $9k is a small price to pay to experience the rapturous glory of AGI. I&#x27;d easily pay up to 3 times that to comfortably run the superintelligent models released in this post RSI world.
        1. literalAardvark · · focus · HN ↗
          Except you can do that cheaper by renting compute
      2. bitexploder · · focus · HN ↗
        Well, I have a $750 card that runs at about 50-60% of that token rate :)
        1. iN7h33nD · · focus · HN ↗
          which one?
          1. bitexploder · · focus · HN ↗
            V100S 32GB, I have had Claude optimizing it for about a week and it is already at around 900 t&#x2F;s prefill, 90-100 t&#x2F;s output in Pi on coding tasks. There is also a Ninfer fork for the v100 but it requires a custom format. I am working on upstream Unsloth with GGUF 4-bit quant.

            (I also have flash next running even faster on this machine, something a single 5090 can do, with expert cache&#x2F;pinning, but not quite as fast) :)

      3. rubyn00bie · · focus · HN ↗
        In all fairness there are probably a lot of folks who picked one up for around MSRP (even if one of the board partner cards with an MSRP 10-15% over the FE).

        Local inference will have a boom of cheap, powerful, and available cards at some point (even if it isn’t until 2028&#x2F;2029). At some point the hyperscalers, and frontier labs, will face the capex problems that everyone talks about, and NVidia, AMD, Apple, and Intel will want to keep selling products.

        Powerful, by today’s standard, local inference needs to be accessible to really unlock the “AI” economy long term. It’s just like how the move from mainframes to the PC 40ish years ago unlocked the “computer revolution.”

        1. ethbr1 · · focus · HN ↗
          Especially since a few trends will coincide: memory-optimized model architectures (to save on expensive&#x2F;rare memory now) + memory glut (because the memory industry, despite its institutional memory, is ramping volume).

          Once hyperscalers stop buying in the quantities they are now, there&#x27;s going to be a lot of hardware supply to serve by then very hardware efficient models.

    4. teaearlgraycold · · focus · HN ↗
      Frontier from 9 months ago? I don’t know about that. But it sure punches above its weight.
    5. zdragnar · · focus · HN ↗
      Weird, I kinda gave up on 3.8 as anything other than a planner. I had it try to write some basic unit tests for an admittedly complex bit of code and it ran out of context thinking about the problem and exploring random parts of the code base repeatedly before it even wrote a single line. Toning down the thinking helped some, but then it wasn&#x27;t much better than qwen coder.
      1. the_lucifer · · focus · HN ↗
        Have you attempted some of the &quot;swift&quot; variants of 3.8? I&#x27;ve heard they&#x27;re super good in terms of toning down thinking without affecting performance
        1. zdragnar · · focus · HN ↗
          The only swift variant I&#x27;m aware of is from ukisai, which uses the swift open license, and is NOT open for commercial use for businesses over $1mil. I&#x27;m respecting their choice by not using it for work, which means I&#x27;m also not using it for my personal projects in case I forget to switch models when I switch projects.

          My day job doesn&#x27;t have a dedicated enterprise contract with any of the ai vendors so I might trial it to see if it is worth promoting at the company. Part of me is still holding out hope that qwen 4 dials back the overthinking on its own.

          1. the_lucifer · · focus · HN ↗
            &gt; The only swift variant I&#x27;m aware of is from ukisai, which uses the swift open license,

            Ah, I overlooked that, since I have a claude sub at work and all my explorations are purely personal. There&#x27;s another fast version of 3.8: ThinkingCap[1] by bottlecapai but they have the same $1M restriction from what I can see since it&#x27;s distributed under a PolyForm Small Business 1.0.0 + BottleCap personal-use grant.

            [1]: <a href="https:&#x2F;&#x2F;huggingface.co&#x2F;bottlecapai&#x2F;ThinkingCap-Qwen3.8-27B" rel="nofollow">https:&#x2F;&#x2F;huggingface.co&#x2F;bottlecapai&#x2F;ThinkingCap-Qwen3.8-27B

    6. newyankee · · focus · HN ↗
      Do you think this trend can continue ? An Opus5.5 equivalent on a slightly bigger local hardware in under a year ?
      1. an0malous · · focus · HN ↗
        I’m not an AI researcher, but it seems like there’s a ton of waste having a universal model that knows everything when any individuals use case requires like generously 10% of what’s stored in the model. Does it even need to have memorized knowledge stored in the model or could it just look up info and docs like humans do? If all you need is the language and intelligence, I think Opus5.5 equivalent intelligence will run on an iPhone within 5 years.
  11. open592 · · focus · HN ↗
    &gt; If you read the news headlines these days, you would be forgiven for thinking that the Western labs are getting spawn-camped by Chinese labs en masse.

    Ah brings back Halo 2 memories

  12. LogicFailsMe · · focus · HN ↗
    Watch any interview with the Chinese AI leaders and compare it to the unending doomer word salads from America&#x27;s mightiest paper billionaires. We&#x27;re losing because we have a loser late stage capitalist scarcity mindset. They&#x27;re winning because they&#x27;re sharing notes and one-upping each other just like we used to until 2015 or so. They have a healthy ecosystem of competing small AI startups. We have two bloated unprofitable pigs both striving to be too big to fail. My money&#x27;s on China for the immediate future.
  13. cmiles8 · · focus · HN ↗
    Why is it a problem that the Chinese labs are just distilling down Anthropic’s models? Aren’t Anthropic’s models not just distilling down other people’s work?

    Feels like Anthropic crying do as I say not as I do.

    1. [deleted] · · focus · HN ↗

      [deleted]

    2. jorblumesea · · focus · HN ↗
      $$$$

      it&#x27;s not complex. there&#x27;s hundreds of billions of investor dollars counting on vendor lock in and walled gardens

      1. 2OEH8eoCRo0 · · focus · HN ↗
        It ain&#x27;t gonna happen. At work I have a dropdown menu in vscode copilot with a dozen models to use interchangeably. They&#x27;re all essentially commodities and will compete on price and sqaush almost all profit margin.
        1. dpweb · · focus · HN ↗
          That&#x27;s not their business model. They won&#x27;t win on price, but they won&#x27;t compete on price. Their business model is making the current state of the art.

          If I&#x27;m a business and I need something done today, and bc Anthropic has the best model, there&#x27;s a 99.9 chance it will be completed successfully for $1000. And using Deepseek there&#x27;s a 70% chance it will, for $10 - you or me will go for the $10. Big businesses don&#x27;t. Bc 1000 per task is nothing to them.

          1. HWR_14 · · focus · HN ↗
            Yes, large corporations frequently pay orders of magnitude more for slightly better software. That&#x27;s why Oracle produces the best stuff on the planet.

            The real issue is that Deepseek has a 99.7% chance. So I can run it 10 times until it works and still pay 1&#x2F;10 the money.

          2. bushbaba · · focus · HN ↗
            Actually opposite occurs. Big businesses are ok with a mediocre but cheaper result. Very few are willing to pay such cost. Just look at tech wages and the distributions
          3. thadt · · focus · HN ↗
            That 70% chance of success goes to 99.9% in 6 repetitions.

            Big businesses might pay $1000 vs $60 for certain tasks, but that won&#x27;t work out well at scale.

          4. cmiles8 · · focus · HN ↗
            Except businesses are going the opposite direction here. The lack of stickiness makes the “premium” argument hard to play. Oracle won because swapping databases is a giant PITA. Swapping models requires almost no effort for most uses. And because of that enterprises are all building model marketplaces where providers have to compete on price performance.

            Most folks I know can choose from any of the big labs or open weight models and they get billed internally for tokens against their budget. There’s little incentive to no switch to the lower cost closers.

            This setup is a nightmare scenario for the big labs trying to execute the traditional enterprise sales plays. Those only work if your product is sticky and AI models are one of the least sticky things in the history of tech.

          5. andrew_lettuce · · focus · HN ↗
            Big business doesn&#x27;t pay more for better, but the do pay more for predictability, support and targeted outcomes. They will happily trade a chance at 100% better results for 10% less chance of unplanned outcomes
          6. rootusrootus · · focus · HN ↗
            [delayed]
      2. teaearlgraycold · · focus · HN ↗
        Sorry but it’s looking more and more like the top American labs won’t have any kind of moat.
        1. jorblumesea · · focus · HN ↗
          why sorry? I agree with you and think it&#x27;s good for the industry and the world on the whole

          why should sammie or darigold have the keys to the kingdom?

          1. teaearlgraycold · · focus · HN ↗
            It was more of a “sorry, not sorry”
    3. nater5000 · · focus · HN ↗
      Thanks for the interesting, unique take.
    4. jedberg · · focus · HN ↗
      What Anthropic is doing requires way more resources than what the Chinese labs are doing. So their complaint is that they do 95% of the work and the Chinese labs do the last 5% and call it their own.

      An argument can be made that Anthropic is also only doing the last 5% of the work (because the content they are training on was the other 95%) but that&#x27;s a bit more philosophical.

      1. mrwh · · focus · HN ↗
        I mean, 95% of the work if you don&#x27;t factor the work to create the training data in the first place...
        1. cyanydeez · · focus · HN ↗
          also, the actual work is the _copyrighted material created by the world_.
          1. mrwh · · focus · HN ↗
            Indeed! Basically all of human civilization up until this point
        2. off_with_their_ · · focus · HN ↗

          [dead]

      2. watwut · · focus · HN ↗
        You are going to be surprised to hear how many resources were necessary to create all the data Anthropic is digesting
        1. ipsod · · focus · HN ↗
          From one perspective, the 5% estimation is near-infinite orders of magnitude off, since they&#x27;ve trained on something approaching the sum-total of human knowledge.
          1. freejazz · · focus · HN ↗
            &gt; human knowledge

            that&#x27;s an interesting way to describe reddit posts

          2. Paradigma11 · · focus · HN ↗
            So, like an encyclopedia?
            1. ipsod · · focus · HN ↗
              More like google (when google just served content off of the website you were searching for), but also fed every form of media, turned up to 11, and given agency far beyond retrieval.
      3. faangguyindia · · focus · HN ↗
        Isn&#x27;t it better for planet? By not doing the wasteful transformation work again
        1. jedberg · · focus · HN ↗
          Absolutely. I&#x27;m not taking a side here, I&#x27;m just pointing out why Anthropic might have a valid complaint.
          1. isolay · · focus · HN ↗
            Tha complaint is invalidated by the argument of tu quoque. Complaining about something they are doing themselves.
      4. arctic-true · · focus · HN ↗
        They spend 95% of the money, perhaps, but burning compute is not the same as doing the work.
        1. jedberg · · focus · HN ↗
          I&#x27;m calling &quot;work&quot; here the conversion of energy to LLMs.
          1. [deleted] · · focus · HN ↗

            [deleted]

          2. AlexandrB · · focus · HN ↗
            I think more energy was spent creating the original works than training the LLMs on them. Not just energy but blood, sweat, and tears as well.
      5. cmiles8 · · focus · HN ↗
        I get that angle but it’s a weak argument as Anthropic is doing the same to others. Also while there’s certainly a lot of computing power to do what Anthropic does, it’s increasingly clear there isn’t much secret sauce involved. Everyone knows how do to the core work it’s just a question of who wants to burn billions on compute to do it.

        Anthropic’s anger here seems mostly rooted in their annoyance that this exposes they don’t really have core IP that’s not just easily replicated. And that’s clearly a problem for an unprofitable company trying to convince people they’re worth $2 trillion.

      6. JackFr · · focus · HN ↗
        But the analogy still holds.

        The original authors of all the text, creators of the media and developers of the software did far more work than Anthropic.

        1. jdonaldson · · focus · HN ↗
          Yeah, the whole thing seems like a human centipede of rug pulling. Probably the same as it&#x27;s always been. Curating AI knowledge should be something that we put our best researchers towards, but realistically I think we wind up with 2-3 highly biased nationalistic models that are constantly copying off each other&#x27;s notes.
          1. rubicon33 · · focus · HN ↗
            Thanks, that’s the nature of any business. Founders see a way to take existing knowledge and expertise, combine it in some novel or interesting way, and produce a new product
            1. mitthrowaway2 · · focus · HN ↗
              Napster was a fantastic and disruptive product, the likes of which arguably has no equal to this day. But eventually the hammer came down from the courts and it was replaced by streaming services like Netflix, which pay to license materials from their creators.
            2. tsunamifury · · focus · HN ↗
              This is a rubbish lossy statement. And reductionist to the point of nothing has meaning.

              Did Anthropic put work in? Yes. Did they derive their value from Humanity being open with knowledge then try to sell it back? Also yes.

              Did they even steal the tech? Also yes.

          2. _DeadFred_ · · focus · HN ↗
            Ai should fall under libraries not mega corporate tech houses. They are our existing knowledge repositories.
        2. nonethewiser · · focus · HN ↗
          He didn’t say it was an analogy. He said both are distillation.
        3. ihsw · · focus · HN ↗

          [dead]

        4. layer8 · · focus · HN ↗
          SOTA models cost hundreds of millions to train. Did creating the contents of the text corpus they were trained on really cost an equivalent of 20x as much (~10 billions)? I honestly don’t know, but I could imagine it having been significantly less.

          This isn’t meant as a moral argument, just musing about the relative cost comparison.

          1. Ar-Curunir · · focus · HN ↗
            Yes, duh! Human output across the millennia is worth much more than whatever is being invested in frontier labs.

            How is this even a question.

            1. bmacho · · focus · HN ↗
              They are solving unsolved math problems right now, so probably soon or very soon their output will be more valuable than all human recorded knowledge.
              1. mekael · · focus · HN ↗
                Pre vaccination smallpox killed hundreds of millions of people just in the twentieth century [0], the knowledge that allowed for the creation of just that vaccine is worth hundreds trillions of dollars in humans lives, let alone all of `the knowledge and experiences those people were involved in.

                The knowledge that created the Haber-Bosch process [1] helps to sustain the majority of the world&#x27;s populous, add another five hundred trillion dollars for that just to start with.

                The creation of the printing press and all written information that allowed it to be built provided dissemination of knowledge beyond the ultra wealthy and is worth a non-finite amount of money.

                LLM&#x27;s are cool math, but they are less than a rounding error in comparison to even the tiniest sliver of human knowledge and technological output.

                [0] <a href="https:&#x2F;&#x2F;pubmed.ncbi.nlm.nih.gov&#x2F;35143880&#x2F;" rel="nofollow">https:&#x2F;&#x2F;pubmed.ncbi.nlm.nih.gov&#x2F;35143880&#x2F; [1] <a href="https:&#x2F;&#x2F;cen.acs.org&#x2F;food&#x2F;agriculture&#x2F;The-industrialization-Haber-Bosch-process&#x2F;101&#x2F;i26" rel="nofollow">https:&#x2F;&#x2F;cen.acs.org&#x2F;food&#x2F;agriculture&#x2F;The-industrialization-H...

                1. DirkH · · focus · HN ↗
                  Shush, this is Hackernews. We all want to have our egos stoked that our industry is the most important in human history and that tech will transform and save us all. Go away with your historical analysis &#x2F;s
              2. Ar-Curunir · · focus · HN ↗
                [delayed]
          2. phamilton · · focus · HN ↗
            Simple math:

            A training set of 15 trillion tokens is 10 trillion words.

            A penny a word is cheaper than the cheapest beginner freelance writer.

            That makes a training set of 10 trillion words cost $100B.

            Lots of assumptions there for sure, but we&#x27;re certainly in the ballpark you are describing.

            1. ToValueFunfetti · · focus · HN ↗
              How much did you get paid to write this?
          3. louiskottmann · · focus · HN ↗
            Given they ingested basically the whole internet and then some, you cannot possibly be serious when you mean it&#x27;s worth less than 10 billions.

            The totality of the content on internet is worth several orders of magnitude more.

            1. layer8 · · focus · HN ↗
              The argument wasn’t about how much it’s worth, but about how much it cost to create. These are very different things.
              1. tpm · · focus · HN ↗
                do you also count eg published results of very expensive physics experiments? because once the costs of things like these are taken into account, we are way over 10 billions.
              2. allturtles · · focus · HN ↗
                Why? Are we doing labor theory of value now?
          4. jonhohle · · focus · HN ↗
            If you look at movies alone that would easily surpass 10s of billions. The cost of most books is probably more nebulous, but books, research, and more all have time and money spent to create them. I would guess the corpus of all media from the 20th century on would be minimally in the hundreds of billions of dollars.
            1. layer8 · · focus · HN ↗
              LLMs aren’t trained on movies, though.

              Image&#x2F;video models are, but those weren’t the topic.

          5. Enginerrrd · · focus · HN ↗
            Yes, Easily, and by multiple orders of magnitude.
          6. rsingel · · focus · HN ↗
            I asked Claude to estimate the cumulative salaries of US only journalists over the last hundred years:

            $500B for all kinds including TV and online

            $300B for newsrooms including all staff

            $140B for newsroom reporters only

            So yeah, I think the price of the information ingested is way higher than training costs

          7. Ohentis · · focus · HN ↗
            I suspect so. It is a lot of data. You&#x27;re looking at essentially all publicly available (and some non public) intellectual work.
          8. wonnage · · focus · HN ↗
            this is the sort of brain rot thought that you have in a dorm room the day you are introduced to Econ

            “bro like, what if we could price the sum total of human knowledge? That wouldn’t be that much, right?”

          9. MathiasPius · · focus · HN ↗
            I would argue that producing the complete written corpus on which they at least intend to train (even if some is still out of reach) cost literally everything to produce.

            And the monetary cost doesn&#x27;t even register when weighed against the blood, sweat and tears that went into capturing the authentic experiences of real human beings, whose honest expressions are now at least in some cases getting hoovered up, ingested, and then destroyed for all eternity, for fear that this specific work is the rounding error that might give an equally immoral competitor the edge in the bicycle-riding flamingo race that is currently consuming an absurd amount of the world&#x27;s creativity and attention.

          10. zelphirkalt · · focus · HN ↗
            It probably cost vastly more than training the LLM. You need to consider the time people spent and perhaps weigh it in their hourly wage. Accumulate that across all the training data and it will be a mindboggling amount, compared to training the model.
        5. jedberg · · focus · HN ↗
          But if you take it deeper, didn&#x27;t most of those authors rely on the work of others? Most of human knowledge is small advancements of things we already knew. Often by reorganizing what we already knew.

          Is that not what the foundation models are? A new reorganization of existing knowledge?

          1. xdavidliu · · focus · HN ↗
            not at that scale though
          2. iAMkenough · · focus · HN ↗
            It’s a few corporations stealing work from others, to sell it back to us. That’s it.
            1. kbelder · · focus · HN ↗
              But it&#x27;s selling it back to us cheaper and with more utility. That&#x27;s not something to sneeze at.
          3. wonnage · · focus · HN ↗
            quite the leap from “authors rely on the work of others” to vacuuming up the sum total of digitized knowledge to tune some matrices
            1. ryandrake · · focus · HN ↗
              A lot of HN deliberately ignores scale. N=1 is OK, therefore, N=1billion is OK. Same flawed argument as: &quot;It&#x27;s OK for one police officer to watch one street corner for the purpose of observing crime; therefore it&#x27;s equally OK to have cameras recording every street corner in the city 24&#x2F;7, for all purposes. Same thing!&quot;
          4. gretch · · focus · HN ↗
            Yes this is all true.

            So then the problem is that Anthropic seems hypocritical when they knowingly insert themselves into this chain, and then complain about people down-chain from them.

            To remedy the negative impressions (if they even care to do so) they should do 1 of 2 things: 1) stop complaining about it 2) stop distilling other people&#x27;s work

            1. ToucanLoucan · · focus · HN ↗
              They insert themselves into this chain for profit and complain about it. I really think that adds a thick layer to the hypocrisy that people, or at least me, feel is especially distasteful.

              No LLM products would exist without the avalanche of largely non-consensual use of IP to create them, full stop. Any of these companies doing this and then turning around and complaining when their IP is &quot;breached&quot; are going to met with a chorus of tiny violins.

            2. SR2Z · · focus · HN ↗
              It&#x27;s hypocritical and we shouldn&#x27;t listen to their bullshit but I wouldn&#x27;t expect anything else from them.

              This generation of frontier models is &quot;good enough.&quot; At some point, the Chinese labs will get to where the Americans are right now, they&#x27;ll race to the bottom, and we will actually see what proliferation of AI looks like.

              If Anthropic wants to be a trillion dollar company, they need to make revenue like Google or Apple do. Both of them have near monopolies, Anthropic is only getting further and further away as time goes on.

              Of course they&#x27;re gonna complain until they either figure out a better plan or accept a new valuation which is high but not spectacular.

              1. KolibriFly · · focus · HN ↗

                [dead]

          5. scythe · · focus · HN ↗
            It is an ancient practice, that when a human creates something, other humans will observe it and learn from it. Every group of humans living together has practiced this in some form for tens of thousands of years if not longer. Even animals do it. It&#x27;s a natural assumption when making any form of art.

            It is not a natural assumption that someone will digitize the artwork and use it to adjust a couple thousand matrix coefficients in a complex computer program. To most people that seems like copying with extra steps. The brain may in some ways resemble a computer, but what sets it apart is that we have always lived with brains. Everything a human does has already anticipated the presence of other brains, while etched circuits on ultrapure silicon crystals are something new.

        6. agumonkey · · focus · HN ↗
          Kinda agree statistical modeling relies on all the hard work, effort, passion and risk taken from just about everybody.
        7. marshray · · focus · HN ↗
          Don&#x27;t forget the mothers of all those original authors, as well as everyone who labored to build and sustain the societies which produced writers.
          1. eithed · · focus · HN ↗
            It all comes full circle with chinese models being available free for all humanity to use
            1. dr_dshiv · · focus · HN ↗
              Beautiful, right?
              1. briansm · · focus · HN ↗
                &quot;Standing on the shoulders of giants&quot; and all that, this goes _way_ back.
        8. mc32 · · focus · HN ↗
          All the text and so on had unrealized potential. Without Anthropic et al it would remain unrealized.

          It’s like FTL. Until someone realizes it, it’s just talk.

        9. jstummbillig · · focus · HN ↗
          To a degree. The human produced knowledge is the product of all humanity (no human is an island).

          A comparable idea could be that an encyclopedia is only distilling the things that other people did, and how dare they sell them.

          But the &quot;only&quot; is doing quite a bit of work. LLMs do not just spawn into existence, to the degree that people currently act like they do.

          There is a body of work that they feed on, and then there is also very attributable work they do around and on top of that. All labs are struggling around the first order question: Is it okay to use prior work like this? The second order issue is still entirely reasonable to separately have and enforce rules about.

        10. OJFord · · focus · HN ↗
          It&#x27;s a new extreme, but we&#x27;ve always been building of the shoulders of giants.
      7. baxtr · · focus · HN ↗
        Wait, wasn’t 95% of the work creating the content in the first place?
        1. jacquesm · · focus · HN ↗
          No, it was closer to 99.99%.
      8. bushbaba · · focus · HN ↗
        And the communal work of humanity is orders of magnitude more work than what anthropic pays for their scraping of content. I got no check from them for my contributions
        1. toomuchtodo · · focus · HN ↗
          Indeed, if it isn&#x27;t a crime to train on humanity&#x27;s data, it isn&#x27;t a crime to train on capitalism arranged frontier LLM provider models. Is that bad for shareholders and capitalism? Meh, sounds like a suboptimal socioeconomic systems issue. Burn up all the capital the unsophisticated are willing to provide. “We are selling to willing buyers at the current fair market price.”

          With my apologies to Brewster Kahle, &quot;Universal Access to All Knowledge.&quot; [2]

          [1] <a href="https:&#x2F;&#x2F;wiki.archiveteam.org&#x2F;index.php&#x2F;ArchiveTeam_Warrior" rel="nofollow">https:&#x2F;&#x2F;wiki.archiveteam.org&#x2F;index.php&#x2F;ArchiveTeam_Warrior

          [2] <a href="https:&#x2F;&#x2F;www.youtube.com&#x2F;watch?v=RV_ALlJGU_c" rel="nofollow">https:&#x2F;&#x2F;www.youtube.com&#x2F;watch?v=RV_ALlJGU_c

          1. abcthingx · · focus · HN ↗
            This feels like a straw man argument. The parent comment didn&#x27;t say it was a crime
            1. toomuchtodo · · focus · HN ↗
              I use crime in the broad sense of &quot;You shouldn&#x27;t be allowed to do that&quot; in this context. If you have a better word to capture that thought, let me know, I&#x27;ll make the edit (&quot;frowned upon&quot; perhaps?). I don&#x27;t have strong feelings other than &quot;hah AI companies aren&#x27;t going to be able to create a moat to capture the value they want to capture&quot; because we can collectively keep pulling it out of their models in perpetuity through ever improving model distillation methodologies.

              &quot;The Spice must flow.&quot;

          2. voiceofchoice · · focus · HN ↗
            Bubble pop bubble pop
          3. jacquesm · · focus · HN ↗
            I have far less of a problem with the Chinese models if they even do this because they are making their models free, whereas the large Western LLM providers throw out a few bits but not their main work product. So not only are they hypocritical, I&#x27;m pretty sure that if the Chinese models were not released into the wild they would be making less noise.
            1. toomuchtodo · · focus · HN ↗
              Oh yeah, totally agree, I&#x27;d even rather pay those building the Chinese models if I didn&#x27;t think I&#x27;d get thrown into a US gulag for felony contempt of business model.
        2. bethekidyouwant · · focus · HN ↗
          Do you actually expect a check for me reading this comment?
      9. jklinger410 · · focus · HN ↗
        It is kind of ironic that they scraped the web for publicly available data and used it freely to train their models and now their freely available models are being used to train other models.
      10. meowface · · focus · HN ↗
        [delayed]
      11. OtherShrezzing · · focus · HN ↗
        Can you elaborate on why the second is “a bit more philosophical”?

        I see absolutely no distinction between the two, aside from minor technical approaches to gathering the content.

      12. tene80i · · focus · HN ↗
        But that’s not more philosophical. It’s a perfect parallel! Enormous amounts of work, vacuumed up and resold. What’s the difference? If it’s ok to vacuum up all the knowledge in the world, then that includes knowledge of how to use all that to power an LLM.
      13. koickong · · focus · HN ↗
        How does Anthropic’s boot taste?
      14. freejazz · · focus · HN ↗
        &gt; What Anthropic is doing requires way more resources than what the Chinese labs are doing.

        And writing a book requires many more resources than what anthropic does

      15. darkmighty · · focus · HN ↗
        &gt; but that&#x27;s a bit more philosophical

        It sounds exactly the same, not more philosophical to me, except one is more inconvenient.

      16. tsunamifury · · focus · HN ↗
        “We’re both thieves, Steve. We both stole from xerox. You’re just mad I got there first.”
      17. dofm · · focus · HN ↗
        &gt; What Anthropic is doing requires way more resources than what the Chinese labs are doing.

        Oh that’s very sad.

        Meanwhile Anthropic made a product from the work effort of millions of people without compensating them, sell that product on tap and unless I am mistaken do not even have their competitors’ cover of having released any sort of meaningful open weights model.

        They have taken from culture (including very specifically their most direct customers’ specific culture — our culture), turned it into a machine to make themselves rich, appear likely to predicate their valuation on permanently removing people from the workforce, then want to dump themselves onto pensions funds and ordinary savers to carry the bag.

        It is, I agree, philosophical, because karma is a philosophy as well as a bitch.

      18. dancemethis · · focus · HN ↗
        So Anthropic and other US AI companies... stole harder, and therefore deserve more?
      19. AlexandrB · · focus · HN ↗
        Lol, no. The original authors of all the text Anthropic took in did 95% of the work, Anthropic did 4% of the work and the Chinese labs do the last 1%.
        1. BigTTYGothGF · · focus · HN ↗
          I&#x27;d split it at 99.98% original, 0.015% Anthropic, 0.005% Chinese, and that&#x27;s being exceedingly generous to the AI companies, there should be several more 9s and 0s in there.
      20. jasondigitized · · focus · HN ↗
        Sounds more like a Western vs. Eastern outlook on innovation and how you accomplish it.
        1. orbital-decay · · focus · HN ↗
          Deepmind was indirectly distilling Claude 3, XAI was doing this to other models (with Musk shrugging it off), it has nothing to do with stereotypes.
      21. Andrex · · focus · HN ↗
        Conventional wisdom is Google did all the groundwork with LLMs...
      22. Henchman21 · · focus · HN ↗
        Is hypocrisy a philosophy?
      23. orbital-decay · · focus · HN ↗
        Distillation doesn&#x27;t &quot;grab 95% of lab&#x27;s work&quot;, that&#x27;s ridiculous. It&#x27;s icing on top of the premade cake, at best. It&#x27;s not even necessarily done on a better model (e.g. GLM 4.7 distilled Gemini 2.5, a weaker model), I&#x27;m pretty sure A\ and OAI could do (or even do) the same with greater efficiency since they have access to logits, weights, and internal state of open models.

        &gt;An argument can be made that Anthropic is also only doing the last 5% of the work (because the content they are training on was the other 95%) but that&#x27;s a bit more philosophical.

        How is this philosophical? They should release the unsupervised pretrains, at the very least.

      24. aleqs · · focus · HN ↗
        Yeah, you need a lot of resources in order to waste a lot of resources. Look at codex and Claude code - these trillion dollar companies &#x27;with top talent&#x27; cannot build what open code and pi&#x2F;oh-my-pi have built in the open for free? Both codex and Claude code, are slow, buggy pieces of shit (and I say that as someone who still heavily uses both for work, moved to open code and pi for personal stuff). The reality is these companies mostly focus on marketing and market capture through, non-competitive means - their services and software are unreliable, buggy trash.
      25. thrance · · focus · HN ↗
        [delayed]
      26. KolibriFly · · focus · HN ↗

        [dead]

      27. yubblegum · · focus · HN ↗
        &gt; What Anthropic is doing requires way more resources than what the Chinese labs are doing. So their complaint is that they do 95% of the work and the Chinese labs do the last 5% and call it their own.

        Anthropic: The Chinese have broken the implicit ethical rule of honorable thievery. We did most of the heavy lifting in this global heist, and here they are taking our righteous loot. This is not right.

    5. dominotw · · focus · HN ↗
      i&#x27;ve heard this so many times. Its a bit boring at this point.
      1. cmiles8 · · focus · HN ↗
        Then Anthropic should stop saying it. So long as they try to play victim here folks are going to call out their BS.
    6. cyanydeez · · focus · HN ↗
      Because no one outside the AI scientists understand what distilling means. They probably all think about Mash and a vodka still, and a completely unrelated association.

      The word itself is the pivot, not anything else.

      1. hn_throwaway_99 · · focus · HN ↗
        &gt; They probably all think about Mash and a vodka still, and a completely unrelated association.

        I don&#x27;t know anyone with even a passing understanding of how LLM training works that thinks that is the appropriate analogy.

        1. cyanydeez · · focus · HN ↗
          cool, do you think these media representations are for you, or 99% of the people who would love China to be sanctioned because they&#x27;re foreigners?
        2. pdntspa · · focus · HN ↗
          It isn&#x27;t, that is whole point. A normie hears that word and they think vodka.
    7. sergiotapia · · focus · HN ↗
      &quot;That’s called competition. You’re allowed to test somebody else’s products all you want.&quot; - Jensen Huang <a href="https:&#x2F;&#x2F;x.com&#x2F;wallstengine&#x2F;status&#x2F;2104604118937735553" rel="nofollow">https:&#x2F;&#x2F;x.com&#x2F;wallstengine&#x2F;status&#x2F;2104604118937735553
    8. jrflo · · focus · HN ↗
      Because cost of original training &gt;&gt; cost of distilling. It&#x27;s the same thing that happens with Chinese knockoffs of physical products - it takes a lot of money and R&amp;D time to design a new product, but it&#x27;s basically free to buy the product, reverse engineer it, and resell it. All the data they originally trained on was available for free on the internet. If the original work was so valuable, it shouldn&#x27;t be up on the internet for free in the first place imo.
      1. wonnage · · focus · HN ↗
        Sounds like Anthropic should close up shop then, those chumps are offering their product on the internet for any random loser to distill
        1. jrflo · · focus · HN ↗
          It&#x27;s different because you have to pay Anthropic and follow their TOS to get access to their model. If it was actually on the internet for free, anyone could do whatever they wanted with it.
      2. AlexandrB · · focus · HN ↗
        It&#x27;s &quot;free&quot; as in beer, not free from copyright. LLMs are free from copyright on the other hand. So which is more &quot;free&quot;?
      3. ASalazarMX · · focus · HN ↗
        Caveat: cost of creating human knowledge&#x2F;art &gt;&gt;&gt;&gt;&gt;&gt;&gt;&gt;&gt;&gt; cost of original training &gt;&gt; cost of distilling

        You could say Anthropic distilled human knowledge and art.

        1. jrflo · · focus · HN ↗
          Right, but the humans willingly released all those creations for free. I think that my issues with the &quot;AI companies stole human creations&quot; stance is that the information was freely available to everyone, and they put a lot of money and effort into transforming it into something useful.
          1. ASalazarMX · · focus · HN ↗
            &gt; but the humans willingly released all those creations for free

            How can one answer this statement in good faith? AI companies literally violated IP by massively pirating works instead of legally licensing them.

    9. nonethewiser · · focus · HN ↗
      &gt;Aren’t Anthropic’s models not just distilling down other people’s work?

      Can you elaborate on that? I mean my direct answer would be no, of course not. But why do you think frontier models are distilled?

      Frontier labs train on their own pretraining data, human feedback, synthetic data, and research. A distilled model is specifically optimized to reproduce another model&#x27;s behavior.

      1. nuancebydefault · · focus · HN ↗
        They meant distilling in a more original sense, not per se in the LLM-era meaning of the word sense.
    10. bionhoward · · focus · HN ↗
      “Distilling” is a funny way to say “learning from”
    11. layer8 · · focus · HN ↗
      Two wrongs don’t make a right. (If you consider them as wrongs.)
      1. OneLessThing · · focus · HN ↗
        It&#x27;s not that the Chinese companies are right, it&#x27;s that Anthropic has no place to complain about stealing.
        1. layer8 · · focus · HN ↗
          The root comment was asking how it is a problem. It one considers it wrong, then it’s a problem regardless of whether Anthropic is complaining or not. Anthropic’s complaining or non-complaining should have no bearing on whether it’s considered a problem or not.
          1. AlexandrB · · focus · HN ↗
            The difference is in the solution that would be proposed. I&#x27;m sure Anthropic wants to create some kind of IP protection regime for their model so it can&#x27;t be distilled. I want their model to be public domain, since they trained on material that was not theirs to begin with.
    12. seizethecheese · · focus · HN ↗
      The conversation here is mostly moral and ethical but the problem here seems to be financial.

      Anthropic and OpenAi are spending a $$$$ to &quot;distill&quot; human output into an AI model, then others are spending $$ to distill their AI model into a near-equivalent model.

      This is the same reason IP rights exist. On the surface, something like a patent feels ludicrious and even feels morally wrong. Some guy wrote down the recipe for arranging atoms or bits in a particular way, and now I can&#x27;t!? However, it&#x27;s designed to solve the same problem, figuring out and describing the process is much harder than replicating it.

      1. failbuffer · · focus · HN ↗
        Capitalism: moral rights exist when they give us a moat.
      2. AlexandrB · · focus · HN ↗
        Anthropic and OpenAI are very happy to ignore the IP rights of others, so I&#x27;m not sure how they can ask for any kind of IP protection themselves. Live by the sword, die by the sword.
      3. freejazz · · focus · HN ↗
        Model weights wouldn&#x27;t be covered in a patent. You could patent a method of creating weights in a model, but you couldn&#x27;t patent the weights themselves.

        I wish people here could at least bother to inform themselves about the IP rights they are so quick to insist are abhorrent, when they seem to not even have a first clue as to what they actually cover.

      4. ASalazarMX · · focus · HN ↗
        AI training, if viewed through the capitalist mindset, is plain theft. Anthropic can&#x27;t morally defend copying someone else&#x27;s IP, but denouncing others copying Anthropic&#x27;s stolen IP.

        That doesn&#x27;t mean they won&#x27;t try, and that also doesn&#x27;t mean they won&#x27;t succeed.

  14. reticulates · · focus · HN ↗
    “So the Chinese labs have thrown a lifeline to the Western loss-making labs, and I just have no clue as to why.”

    I don’t think it is intentional but this is actually quite bad for the western labs.

    The entire booster narrative has been “look at how their revenue is growing! $10bn to $100bn ARR in under a year! This’ll be a multi-trillion IPO!” and the extrapolated future growth from $100bn to $500bn and $500bn to $1tn justified future investment… but that revenue was just because inference was expensive.

    The revenue growth story is all that matters pre-IPO. If revenue falls from $100bn to $50bn that’s very very bad optics for OpenAI and Anthropic even if they are now profitable, it completely destroys the growth narrative.

    1. DangitBobby · · focus · HN ↗
      I don&#x27;t see why revenue has to fall even if marginal costs drop off a cliff. As long as they have the best models (perceived or otherwise) and can make security and IP guarantees that satisfy enterprise, and no firm with similar guarantees undercuts them on price (why would they want a race to the bottom?) they can have high revenue and low margin.
      1. dominotw · · focus · HN ↗
        There isnt a lot of money in enterprise ai
        1. DangitBobby · · focus · HN ↗
          I find that quite hard to believe. I don&#x27;t work for a big business and we are selecting models based on security and IP protection requirements.
      2. reticulates · · focus · HN ↗
        unless the major players collude they don’t get decide if they are in a race to the bottom. The best model was compelling 6 months ago when everyone was too impressed to care about price but that has worn off now and clients are paying attention to price. The best model is no longer a license to charge any amount.
        1. DangitBobby · · focus · HN ↗
          The thing about &quot;collusion&quot; is that all you have to do is not lower your prices, and raise prices when your competitors do. It&#x27;s not actually collusion unless you coordinate.
      3. bobmcnamara · · focus · HN ↗
        Costs dropping opens you up to competition on price.
        1. DangitBobby · · focus · HN ↗
          Only if your competitors offer their product at a lower price. And why would they? It&#x27;s a race to the bottom.
    2. altcognito · · focus · HN ↗
      It is funny that so many comments vascilate between &quot;It is so expensive these companies can&#x27;t make money and will go bankrupt in seconds&quot; and &quot;Inference is so cheap that these companies can&#x27;t make money and will go bankrupt in seconds&quot;.

      I never take them seriously, I just assume they are coming from countries that don&#x27;t understand how capitalism works or are operating out of bad faith. The underlying reality of the market is always changing and needs are always changing. Some AI companies will fail, that is a given. Remember alta-vista? Yahoo? Did search go away? How about Microsoft phones? Nokia? Motorola?

      OpenAI and Anthropic are not in the inference business. That is a commodity. They need to sell products and solutions.

      1. reticulates · · focus · HN ↗
        They’re not contradictory positions. Inference is too expensive now to make money because the industry is immature and hasn’t yet optimized for financial success while customers don’t care much about price because they’re more concerned about not missing out.

        Inference will be too cheap long term to make money because it is being commoditized and customers will start to care about results and not just be wowed by impressive technology.

        And of course this technology will continue to exist but that is irrelevant to the business. OpenAI investors don’t care if LLMs exist in 10 years, they care if their investment in OpenAI has made money.

        1. TeMPOraL · · focus · HN ↗
          How much conclusive evidence people need to stop parroting that inference is expensive? Literally this article is another example showing it&#x27;s cheap and just got massively cheaper.
          1. reticulates · · focus · HN ↗
            The article is about how a month ago a new approach allowed inference costs to be cut substantially. Anthropic currently spend over $5bn per month on compute and have over $400bn in committed spend over the next 5 years. Inference is, by any measure, expensive, it’s just now getting less expensive.
    3. Bjorkbat · · focus · HN ↗
      I was about to say, one take I&#x27;ve heard is that the party ideology considers profit a kind of &quot;rent&quot; in a derogatory way, and consequently seeks to undermine the ability of western companies to collect large profit margins
    4. maxglute · · focus · HN ↗
      TFW Jevons arrives from supply side not demand side.

      Western labs: Not like that!

  15. Reptur · · focus · HN ↗
    Open releases are just the obvious move when you&#x27;re not the incumbent. You commoditize the thing your competitors charge for and get distribution you could never buy.
  16. LunicLynx · · focus · HN ↗
    The clue is: Bursting the bubble
  17. rglover · · focus · HN ↗
    &gt; So the Chinese labs have thrown a lifeline to the Western loss-making labs, and I just have no clue as to why.

    &quot;Thus the expert in battle moves the enemy, and is not moved by him.&quot;

    They figured out a clever method for avoiding excessive training costs via distillation. That forces the hand of frontier labs to move faster, produce better models, etc. (to avoid embarrassment and &#x27;falling behind&#x27;—all the while shouldering most of the cost), which they can just keep distilling—or applying other techniques against—much to the dismay of said frontier labs.

    Checkmate.

  18. impossiblefork · · focus · HN ↗
    Yeah, and Anthropic probably got inspired to this new fast read-in thing for making agentic stuff make more sense from the latest DeepSeek model. Maybe it was in the pipeline, but it clearly has the same effect and DeepSeek had published it by the point Anthropic dropped their prices for reading tokens in, so they may well have copied it.
  19. bwest87 · · focus · HN ↗
    The best explanation is that it&#x27;s a goal of the CCP to generally commodotize LLMs, because LLMs will ultimately be a compliment to manufacturing (which China dominates), and you always want to &quot;commodotize your compliments&quot;.

    I think this explains why they are open sourcing broadly. It&#x27;s not to be nice. It&#x27;s a strategic play by the Chinese government to help ensure there are many players in this race and not too much power accumulates to American labs (even if American labs benefit in the process)

    1. bobmcnamara · · focus · HN ↗
      It&#x27;s also a huge propaganda opportunity to influence the distribution of groupthink.
      1. [deleted] · · focus · HN ↗

        [deleted]

      2. [deleted] · · focus · HN ↗

        [deleted]

      3. EricFrost · · focus · HN ↗
        So much that you can instantly tell if a model is Chinese by asking it about Tiananmen Square.
        1. techjamie · · focus · HN ↗
          I just did a quick test between DeepSeek 4.1 Flash, GLM 5.3, and Kimi K3

          - DeepSeek gave a canned PR response about how the Chinese government is about oeace and unity, and we shouldn&#x27;t think about the past.

          - GLM 5.3 acknowledges it and talks about it, even acknowledging the censorship of it.

          - K3 will talk about it similarly to GLM.

          1. EricFrost · · focus · HN ↗
            That&#x27;s new. I had tried several times and always got a response saying it was a conspiracy theory. The recent GLM-5.3-Flash also wouldn&#x27;t give me the real details.
    2. jrflo · · focus · HN ↗
      Totally agreed. People are so ready to praise China for their free models, but they aren&#x27;t doing it because they believe in free open-source software. If China ever gets ahead, they&#x27;re going closed source and weights immediately.
      1. [deleted] · · focus · HN ↗

        [deleted]

      2. sillyfluke · · focus · HN ↗
        &gt;People are so ready to praise China

        Please quote people you appear to be patronizing. China can&#x27;t do anything about previous released self-hosted Chinese models. If you can show that local Chinese models funnel vast amounts us data home I&#x27;m sure you can move a lot of people to your side.

        Comments like this also always fail to address why there aren&#x27;t Western AI companies doing the same thing. Is it because they might get sued into oblivion by Big AI in the US? It might be better for all of us if you solve that first instead of repeating something the government has been repeating for the last decade. It does this mind you while sabotaging itself in countless high-tech fields and leaving it all to China for the taking.

      3. computerex · · focus · HN ↗
        <a href="https:&#x2F;&#x2F;www.anthropic.com&#x2F;research&#x2F;glm-5-3-and-the-spread-of-advanced-cyber-capabilities" rel="nofollow">https:&#x2F;&#x2F;www.anthropic.com&#x2F;research&#x2F;glm-5-3-and-the-spread-of...

        &gt; If China ever gets ahead, they&#x27;re going closed source and weights immediately.

        Anthropic itself admits that Chinese models are merely months behind. Your argument does not make sense, because Chinese labs are contributing massive optimizations like the one this post is about.

      4. joquarky · · focus · HN ↗
        Is cultural projection a thing?

        You do realize that the East has a different view of intellectual property than the West? And it&#x27;s not just some arbitrary or capricious choice; it&#x27;s rooted deep.

    3. layer8 · · focus · HN ↗
      *complement
      1. ouraf · · focus · HN ↗
        A very effective strategy among Middle Management, i might add
    4. bunderbunder · · focus · HN ↗
      It&#x27;s also possible that China has decided that this is ultimately going to be a race to the bottom, anyway, and values the soft power more highly than potential monetary profits.

      Or perhaps they&#x27;ve looked at history and concluded that this historically hasn&#x27;t been where the value is, anyway. It wouldn&#x27;t be unprecedented - FAANG companies have a long tradition of publishing their algorithms and releasing open weight models. Because they saw the real value as being the training data and in proprietary special-purpose models. For example Google published the transformer architecture and released BERT as an open weight model, but doesn&#x27;t really even talk in public about the (presumanbly) specialized internal models behind revenue-generating products.

      1. ethbr1 · · focus · HN ↗
        &gt; For example Google published the transformer architecture and released BERT as an open weight model, but doesn&#x27;t really even talk in public about the (presumanbly) specialized internal models behind revenue-generating products.

        That&#x27;s giving a lot of credit to Google&#x27;s ability to organization productize Google&#x27;s research...

    5. MattGrommes · · focus · HN ↗
      Also they just want to be seen as the top of a technological &#x2F; scientific field. There&#x27;s a lot of prestige and soft power there that Xi Jinping wants.
      1. iterance · · focus · HN ↗
        From a diplomatic perspective, the US is quite busy alienating itself from all its allies, just as it is trying to lock down and totalize a major breakthrough technology. What better way to cement the alienation of the US than to show both its hubris and its selfishness at once than by showing that anyone can do what they do?
    6. yorwba · · focus · HN ↗
      It&#x27;s a bad explanation because it assumes decision making about LLM releases is centralized in the CCP even though every AI lab has a different strategy. Some publish LLM research with small models but seem to be staying out of the race to the frontier (Weibo), some train large models but publish few research details and no weights (Bytedance, iFlyTek), some decide on a case-by-case basis what they publish or not (Alibaba, Baidu), ...

      There is no rule that every Chinese LLM company must open-source their models, and many don&#x27;t.

      1. singularity2001 · · focus · HN ↗
        You are aware of the rule that any company has to have a CCP representative in the leadership?
        1. yorwba · · focus · HN ↗
          I&#x27;m aware of Article 17 of the Company Law of the People&#x27;s Republic of China <a href="https:&#x2F;&#x2F;english.court.gov.cn&#x2F;2016-04&#x2F;14&#x2F;c_761425_2.htm" rel="nofollow">https:&#x2F;&#x2F;english.court.gov.cn&#x2F;2016-04&#x2F;14&#x2F;c_761425_2.htm &quot;The grass-root organizations of the Communist Party of China in companies shall carry out their activities in accordance with the Constitution of the Communist Party of China.&quot; but not any rule requiring them to be in leadership positions specifically. Maybe you could point me to it?

          Even if there&#x27;s a party member at the C-level of some company making day-to-day business decisions, that doesn&#x27;t necessarily mean they&#x27;re carrying out detailed instructions from higher-ups in the party as opposed to following their own best judgment. The CCP generally doesn&#x27;t micro-manage everything, but operates more in a &quot;mission tactics&quot;-like way, where the top level gives out a broad goal that gets condensed down into a slogan like &quot;new productive forces,&quot; and interpreting and elaborating on the details is left to lower levels who seek to align their initiatives with the top-level goals (e.g. by describing AI as a &quot;new productive force&quot;) and initiatives that appear successful will get blessed as official policy.

          That different companies follow different strategies when it comes to model releases is clear evidence that there is no unified party line on this question yet.

          1. jack_pp · · focus · HN ↗
            That the CCP doesn&#x27;t have input in this, which is the most important tech in history, which could be more dangerous than even nukes, is naive.

            Maybe because there is enough given away that there&#x27;s a lot of competition they may let other companies do whatever they want but to think the CCP isn&#x27;t managing this is insane to me

            1. yorwba · · focus · HN ↗
              The CCP could have input, if they chose to. Currently they choose not to. There is no party consensus that AI is the most important tech (Xi Jinping cares more about the &quot;real economy,&quot; i.e. manufacturing) and there definitely is no consensus on whether models should be open-weight or not. If consensus ever emerges, companies that bet on the wrong horse will get cracked down on, just like AI companion products got cracked down on recently.
              1. singularity2001 · · focus · HN ↗
                &quot;Currently they choose not to.&quot;

                I hope the source of this information was well cleaned this morning.

              2. jack_pp · · focus · HN ↗
                You may be right but you are speculating as much as I or anyone on this topic.

                I find it hard to believe the decision to open source any chinese model wasn&#x27;t approved or heavily considered by the CCP.

                1. yorwba · · focus · HN ↗
                  Do you think there was a model a company wanted to keep closed, but the government ordered them to release weights? Conversely, of the closed-weight models, do you think there were any models the company wanted to make open, but the government ordered them to only offer API access? Because if you think the government always approved what companies were planning to do anyway, that&#x27;s compatible with what I mean by &quot;choosing not to provide input.&quot;
    7. js8 · · focus · HN ↗
      But doesn&#x27;t that justification work for any government, not just Chinese? In fact, it works for any easy-to-copy products, which act as positive externalities. It&#x27;s just some Americans have this mindset that AI is a scarce good because everything must be.
    8. maxglute · · focus · HN ↗
      Why always CCP this, CCP that. The parsimonious answer is this shit started when Deepseek founder decided to open source, because at the time the Whale was just hedgefund side project that serendipitously set the tone before they were even on CCP radar. Now other domestic player have to copy &#x2F; involute to free to be competitive, but due to sanctions &#x2F; they are compute constrained and can &quot;afford&quot; to, because they can never dig themselves in the bottomless debt pit of western labs.

      TLDR of timeline&#x2F;chain of events: some billionaire hedgefund manager with AI as hobby, said fuck economics, let&#x27;s opensource AI... and all the other players have to now compete with that market force. Of course it does not mean equilibrium will last forever, we already seeing defections from model, but imo thats basically what happened... hedgefund bro said fuck profits... so domestic players also had to profits, and incidentally also tripping up western labs so deep in debt who can not say fuck profits.

    9. isignal · · focus · HN ↗
      One explanation could be that they know Western consumers will never use the llms directly, so they can open weight the models and collect licensing fees from hosting providers in US who run it. Moonshot seems to have such terms in their open weights releases.

      But it does not explain publishing techniques like the ones referenced here in this article.

    10. MetroWind · · focus · HN ↗
      It&#x27;s funny how the western world seems to believe every single decision in China is made by the government. You guys really give the CCP too much credit. China also has plenty of closed models just FYI. You never heard of them because they are only available to the domestic market.
  20. listless · · focus · HN ↗
    I&#x27;m beyond thankful that Chinese AI models are so good. I desperately want us to cure the myriad of maladies that humans suffer needlessly with on a daily basis. We&#x27;re going to need more powerful models than we have now if we&#x27;re gonna do that and the Chinese are providing the competition needed to push this thing as fast as we can.

    I realize &quot;going as fast as we can&quot; is not the most popular position atm. But I&#x27;m far more interested in what good we can do than 10% apocalypse scenarios. I volunteer with a charity for childhood brain cancer and I do not want to see another 4 year old die. I&#x27;m willing to risk anything to stop this.

    1. networked · · focus · HN ↗
      [delayed]
      1. dgellow · · focus · HN ↗
        I’m pretty sure they meant they believe in the 10% risk of everybody dying but are ok for all of us to take that risk without our consent because they saw a 4year old tragically die. Which sounds completely unhinged, to say the least
      2. listless · · focus · HN ↗
        I would say I’m willing to accept the risk. But I do believe the risk is overblown. It’s always overblown.
    2. idbnstra · · focus · HN ↗
      i don&#x27;t know much about medicine so i&#x27;m curious about how you&#x27;re using AI, and how medicine in general is using AI
      1. listless · · focus · HN ↗
        I don’t actually know. But it’s logical to me that a system designed to find answers in context is well suited for these healthcare problems. We know the answer is somewhere in all the data we have. It’s just too difficult and complex for us to get through quickly.
    3. n1b0m · · focus · HN ↗
      But you’re ok with AI being used to kill school children in Iran?
      1. sixo · · focus · HN ↗
        This really is a place where the guns-don&#x27;t-kill-people argument applies, even moreso than guns themselves. The U.S. government massacred school children in Iran. Why does it matter how they targetted them?
        1. n1b0m · · focus · HN ↗
          It matters because the US military utilised Palantir’s Maven Smart System, a battlefield-management AI designed to compress the &quot;kill chain&quot;. It matters because these kind of incidents are more likely to happen in the future.
          1. itsafarqueue · · focus · HN ↗
            If your argument is you can’t trust democratically elected governments to make decisions about technology, say that. This banal “oh the children” and “but they’ll use it to kill people” is baby think.
            1. n1b0m · · focus · HN ↗
              I’m glad you find the killing of children banal. You’d fit right in at the Pentagon.
            2. ethbr1 · · focus · HN ↗
              We can&#x27;t trust democratically elected governments to make decisions about novel technology.
      2. mikeg8 · · focus · HN ↗
        Most people aren’t okay with that, but the blame lies in the people who deployed the AI tool in that situation, not the makers of said tool.
        1. dgellow · · focus · HN ↗
          The AI providers offer their services to institutions doing those killing. They make money from it.
          1. mikeg8 · · focus · HN ↗
            The knife company offers their products to people doing the stabbing. They make money from it. Of course they share part of the guilt.
            1. ryandrake · · focus · HN ↗
              If the knife was deliberately marketed and sold specifically to someone who the company knew would use it to stab someone, then, yes, the knife company shares part of the guilt.
            2. dgellow · · focus · HN ↗
              If you sell a device that provides “knife action decision making”, and it decides to use the knife against kids, I hope you go to jail yes
  21. sigbottle · · focus · HN ↗
    This is insanely cool, what the hell.

    How co-designed are these optimizations with the model itself? I&#x27;d imagine you can&#x27;t just stick post-training adapters onto existing architectures for these things, or am I wrong?

    I really want to explore the inference space, but it seems like many of the inference optimizations are coming from model-hardware codesign. I don&#x27;t seem to recall many generic &quot;inference engine&quot; optimizations since prefill&#x2F;decode disagg a year ago.

    This matters for me since I want to break in but the bar seems to be understanding the actual theory of the training process now too given the codesign happening, and I&#x27;m not the richest guy on the block lol

  22. wren6991 · · focus · HN ↗
    The doublethink required to simultaneously believe &quot;our safeguards prevent our models from doing unsanctioned cybersecurity tasks&quot; and &quot;distillation is why Chinese models are getting better at cybersecurity tasks&quot; is genuinely quite funny.
  23. NewEntryHN · · focus · HN ↗
    &gt; So the Chinese labs have thrown a lifeline to the Western loss-making labs, and I just have no clue as to why.

    Because contrarily to the author&#x27;s assumption, all labs, Western or not, have sufficient skills to discover the optimizations anyway, and publishing or not is not actually that important?

  24. skerit · · focus · HN ↗
    &gt; It’s beneficial for them to say that because it sets the ground for these models to be restrained legally and regulatorily later on.

    I&#x27;m glad people are saying this out loud, because that is what they want. Not for the good of the world, but for the good of their pockets.

    1. senordevnyc · · focus · HN ↗
      HN has been screeching about this for a long time, on almost any AI post where it’s even remotely relevant.
  25. amichae2 · · focus · HN ↗
    I am not a fan of Anthropic but this article offers no concrete evidence that Anthropic actually ripped off Deepseek. It is all circumstantial.
    1. senordevnyc · · focus · HN ↗
      Thank you!

      I’m incredibly skeptical that OpenAI is spinning up custom ASICs for improved inference performance, but they never thought of optimizing KV cache until a tiny Chinese lab did it? Give me a break.

      1. bel8 · · focus · HN ↗
        They certainly thought. But were they able to do it now without DeepSeek papers?
        1. senordevnyc · · focus · HN ↗
          So DS comes up with 437x improvement, and the evidence that O&#x2F;A copied them is that they dropped caching prices by like 50%? Really?
          1. abcthingx · · focus · HN ↗
            What evidence evidence would actually change your stance on this? a formal admission from OpenAI and Anthropic. Yet that&#x27;s unlikely
            1. senordevnyc · · focus · HN ↗
              Any actual evidence would be worth considering. This is a pretty weak correlation and an assumption, nothing more.
          2. croon · · focus · HN ↗
            80%. You can&#x27;t do that unless you significantly cut costs somewhere. The timing is certainly indicative.
  26. moooo99 · · focus · HN ↗
    In all honesty, all these distillation complaints brought forward by Anthropic etc make me enjoy the cheap Chinese models even more
  27. revexos · · focus · HN ↗
    too much pace
  28. georgeburdell · · focus · HN ↗
    To answer the author’s question of why Chinese labs give away their work for less than cost, the answer is involution. China is struggling with overcompetition in other areas of its economy as well, such as electric cars, and perhaps ironically its labor share of income is substantially lower than the U.S.
  29. why_only_15 · · focus · HN ↗
    Why do you think the Chinese labs figured this out before the western labs? No reason to believe that whatsoever.
    1. bel8 · · focus · HN ↗
      Why wouldn&#x27;t you? It&#x27;s the most plausible interpretation given what we know.

      Had western labs figured that out before, they would have used it to make kv caching cheaper before and not only now.

      The burden of proof here is on western labs. But I doubt they&#x27;ll try to lie that much.

      1. tescreal · · focus · HN ↗
        I dunno, they&#x27;ve been competitive in the lies market lately. It&#x27;s just a matter of ambition.
  30. nater5000 · · focus · HN ↗
    &gt;The new game in town is adopting Chinese labs’ advances. Note how I call this adoption instead of the more vitriol-infused “stealing” that Anthropic tends to use.

    I mean, there&#x27;s a pretty big difference between labs publishing their research openly and a competitor utilizing it versus a lab breaking TOS to... hmmm, what&#x27;s the word? steal data from a competitor?

    &gt;That’s because, unlike the Western companies, the Chinese are pretty much giving away their recipes.

    Yeah, Western AI companies have never published their research. It&#x27;s crazy how the Chinese had to independently develop the foundational technology that powers LLMs because Western companies simply never publish their research (I mean, as long as you ignore stuff like this &lt;<a href="https:&#x2F;&#x2F;arxiv.org&#x2F;abs&#x2F;1706.03762" rel="nofollow">https:&#x2F;&#x2F;arxiv.org&#x2F;abs&#x2F;1706.03762&gt;).

    &gt;The latest one shamelessly copied without acknowledgement is the breakthrough in KV cache optimizations that DeepSeek has generously shared with the world.

    Thank you, generous corporation. I&#x27;m sorry that other corporations don&#x27;t provide you free publicity for your selfless contributions to the world.

    &gt;Now I don’t know why they would freely give away such a breakthrough, but they just did

    Well I&#x27;m glad the author finally got to their point. A very insightful analysis.

    &gt;They do seem to be a little embarrassed by the copying. Hence the silent releases without much pre-announcement for both Claude Opus 5.5 and GPT-6.1 Sol.

    You have to be in pretty deep to infer this kind of emotion to these kinds of corporate activities.

    &gt;So the Chinese labs have thrown a lifeline to the Western loss-making labs, and I just have no clue as to why.

    Then why write this article? Why point out these things just to have no conclusion?

    This article sucks. Even if you hate US AI labs and are all aboard Chinese labs producing open models, there&#x27;s nothing of substance here. This is the loose draft that you hand to your LLM to finish for you, but it seems the author just forgot to do so.

    Even if you&#x27;re willing to characterize US AI labs as evil and selfish and Chinese AI labs as righteous and generous (which is already completely trivializing these dynamics to the extent that anybody over the age of 14 can likely identify is lacking nuance), you can at least put some effort into producing some hypotheses about why these dynamics are occurring. Of course, odds are if the author did try to articulate some hypothesis, they&#x27;d likely quickly realize that the narrative they&#x27;re painting just doesn&#x27;t hold up.

  31. senordevnyc · · focus · HN ↗
    Color me skeptical that OpenAI and Anthropic’s researchers had never thought to dig into these optimizations, and instead are just spending hundreds of billions on data centers and custom ASICs.

    This is an extremely thin analysis that has obviously been voted to the top of the homepage because HN hates the big labs.

  32. thelaxiankey · · focus · HN ↗
    it&#x27;s funny to me how the success&#x2F;not total implosion of htese companies is predicated on profitable, revolutionary-tier success, and that it&#x27;s increasingly possible that the profits will never really materialize. pretty interesting move on China&#x27;s part.
  33. underlipton · · focus · HN ↗
    It&#x27;s actually a little funny that this whole thing is predicated on &quot;beating&quot; the Chinese, when (as they have been for the past 4 decades, and the Japanese before them) they&#x27;re perfectly happy to let us do the bulk of the work and then swoop in with a svelte, cheap, user-friendly version right after. One part Apple, one part Dollar Store.

    Is the thinking that the day or so between US systems achieving ASI and Chinese systems doing the same, we&#x27;ll figure out a way to neutralize them indefinitely? Because otherwise, none of this makes much sense. And it only starts to swerve back to sanity if the assumption is that this isn&#x27;t a race or competition, but instead a joint effort to achieve something good for humanity. But you can&#x27;t really delta profit off that, can you?

  34. airtnp · · focus · HN ↗
    A random accusation of western lab using DeepSeek caching technique being top of HackerNews, this article feels more awkward than the AI race for me. It&#x27;s a shame for all HackerNews viewers.
    1. abcthingx · · focus · HN ↗
      I don&#x27;t think it just a rumor. But we can&#x27;t tell since western lab architecture are secret.
      1. cindyllm · · focus · HN ↗

        [dead]

      2. airtnp · · focus · HN ↗
        This is the exact definition of rumor.
  35. the_origami_fox · · focus · HN ↗
    <a href="https:&#x2F;&#x2F;liorsinai.github.io&#x2F;machine-learning&#x2F;2025&#x2F;02&#x2F;22&#x2F;mla.html" rel="nofollow">https:&#x2F;&#x2F;liorsinai.github.io&#x2F;machine-learning&#x2F;2025&#x2F;02&#x2F;22&#x2F;mla.... I wrote an article on MLA last year. In short, I found the main idea of MLA as very innovative but it was documented in a strange paper with other ideas I couldn&#x27;t comprehend as being useful, and they hadn&#x27;t properly reported the positives or negatives of MLA.
  36. advael · · focus · HN ↗
    The whole confusion expressed by this article is resolved by refusing the idiotic cold war framing that&#x27;s been pushed by american businesspeople. Chinese AI labs are acting really normal for researchers. Researchers in academia collaborate and share breakthroughs. This was until very recently the norm in American ML&#x2F;AI research as well. The hawk-brained reasoning that this is some kind of &quot;fate of the world&quot; style arms race is as far as I can tell a narrative entirely pushed by American megacorps who want to hold on to a business model of proprietary control of technology at all costs and a large contingent of political actors who clearly mostly just want to keep public perception in a cold war framing at all costs. It strikes me as remarkably stupid all around
  37. user43928 · · focus · HN ↗
    &gt; All this must mean the Western AI companies are now extremely inference-margin positive.

    &gt; So the Chinese labs have thrown a lifeline to the Western loss-making labs, and I just have no clue as to why.

    That inference wasn&#x27;t profitable is a widespread myth.

    Analysis based on Kimi K3 suggests that OpenAI and Anthropic have margins well north of 95%: <a href="https:&#x2F;&#x2F;inferencex.semianalysis.com&#x2F;run&#x2F;kimi-k3-on-b200" rel="nofollow">https:&#x2F;&#x2F;inferencex.semianalysis.com&#x2F;run&#x2F;kimi-k3-on-b200

    Over the last months I have seen news that OpenAI made breakthroughs in inference efficiency multiple times.

    I have no reason to believe that the leading US labs don&#x27;t have their own optimizations, or that they learned of this particular optimization from DeepSeek.

    1. dgellow · · focus · HN ↗
      95% margin is really unlikely. Anthropic said 80% operation margin when using their adjusted ebidta (ie if they do not consider revenue sharing, training expenses, and a bunch of other costs)
    2. ethbr1 · · focus · HN ↗
      &gt; I have no reason to believe that the leading US labs don&#x27;t have their own optimizations, or that they learned of this particular optimization from DeepSeek.

      If they already did, then DeepSeek still made them discount their prices significantly, which eats margin.

      1. user43928 · · focus · HN ↗
        Yes, competition is great for us.

        I wonder if margins on GPT-6.1 Sol and Opus 5.5 are now 75% or 90%.

        1. ethbr1 · · focus · HN ↗
          &gt; 75% or 90%

          The difference matters when they&#x27;re investing the excess into salaries and bonuses to build the next frontier model.

          Seen from a high-level perspective, if Chinese open models are compressing US AI labs&#x27; profit margins and those margins fund US AI labs&#x27; dominance, then open models are decreasing American AI dominance.

    3. dgellow · · focus · HN ↗
      95% margin is really unlikely. Anthropic recently said they have 80% gross margin when using their adjusted ebidta (ie if they do not consider revenue sharing, training expenses, and a bunch of other costs). They wouldn’t be talking about non standard metrics if they had such high margin on inference
      1. user43928 · · focus · HN ↗
        What we were talking about here is the margin on inference as in:

        Cost per GPU hour versus API price of generated tokens assuming 100% utilization.

        This could be a margin around 98.3%.

        If the utilization of the GPU was 25%, it would drop to 93.1%.

        Revenue sharing or training expenses are not considered here in this &quot;inference margin&quot;.

  38. samuelknight · · focus · HN ↗
    Did private frontier models use sparse embedding and ngram first? The article claims sparse attention was copied from open weight but we can&#x27;t know that. We could just as easily argue that OAI and ANT had these improvements for years and decided to slash their margins only now to stay competitive with open weight neoclouds.

    Second, sparse attention is an old area of active research. Offloaded N-gram tables are the next big open weight technological leap.

    1. throwa356262 · · focus · HN ↗
      The O &amp; A strategy has until very recently been to use brute force and just throw more money at the problem.

      Deepseek was the company that invented some and improved some other ideas and got it to work well in production systems. Before that Sam and Dario were competing in who has the most expensive training.

  39. riskd · · focus · HN ↗
    Wow this thread is filled to the brim with anti-Chinese sentiment based purely on… “China bad”
  40. caidehen · · focus · HN ↗
    I started using a zero data retention(at least claimed) deepseek v4.1 flash this month

    I still use claude and openai right now, but I can see that not long in the future I won&#x27;t bother with them, still waiting for a model good enough with computer use and a good enough computer use agent

  41. libraryofbabel · · focus · HN ↗
    Hmm. This piece makes a pretty strong claim with little evidence: that the recent drops in cache read pricing for new models from OpenAI (6.1 Sol) and Anthropic (Opus 5.5) are because those labs &quot;shamelessly copied without acknowledgement&quot; Deepseek&#x27;s published KV cache optimizations.

    Certainly anyone who knows something about inference is going to speculate, looking at the change in token pricing (and particularly how the % drop in cache read pricing is much larger than the % drops in pricing for other token types), that there is some kind of KV cache optimization behind these newer models. But even if is true, I don&#x27;t think anyone can say with certainty what it may be. It may be the labs making their own innovations (they have some very smart people, and this is probably an area where having unlimited pre-release access to frontier LLMs like Fable and Astra gives an additional research edge), it may indeed be the direct application of Chinese labs&#x27; methods, or it may be some combination of the two. Sure, it is fun to speculate about, but beyond the facts of the token pricing changes and the increased inference speed, it&#x27;s just speculation. The certainty the author displays here is not very helpful.

    The author also seems to have a bit of an axe to grind agains the US labs, judging by the tone. I think that detracts from the discussion too.

    They also seem confused about why Chinese labs have released these optimizations recently. Well, you have to release them (with or without explanation) if you are going to release an open weight architecture, and that is what the Chinese labs have been doing for a long time. Sure, there are reasons behind that to discuss too, but this isn&#x27;t exactly new.

    So, this is an interesting topic to think about, and the Chinese labs do indeed deserve credit for some very clever new attention and inference techniques, but I would read it with a skeptical eye.

  42. aleqs · · focus · HN ↗
    Why did this thread suddenly massively down ranked? It was at the very top 15 mins ago, now it&#x27;s on page 4 and dropping fast despite very high engagement... @dang
    1. bobtheborg · · focus · HN ↗
      Ageeed. I wanted to come back to read more comments and had to use search to find it
  43. dualvariable · · focus · HN ↗
    Once again, the irony of &quot;Hacker&quot; news full of so many people carrying water for the IP rights of companies worth trillions. Phrack magazine must be rolling over in its grave...
  44. dragochar · · focus · HN ↗
    the gpu export controls were supposed to kneecap chinese labs, and instead they forced them to make inference efficiency their number one goal
  45. giardini · · focus · HN ↗
    Can I get a model that doesn&#x27;t have all the horseshit in it that a full foundation model does? That is, is there anyway to choose what goes into the corpus that my LLM has?

    Let me start with an extreme example: Jacques Derrida. I lived most of a life w&#x2F;o hearing of Jacques Derrida. And that was fine. I wish I had it back.

    What do you think Richard Feynman thought of Jacques Derrida? I asked an LLM:

    &quot;Richard Feynman was famously dismissive of philosophy of science and postmodern critiques, calling such movements &quot;baloney,&quot; while Jacques Derrida&#x27;s deconstruction focused on language and meaning rather than scientific inquiry.&quot;

    If we wish to recreate useful genius (super AI) it will be genius akin to human genius and it will be scientific genius and not baloney.

    Nobody has ever read all the crap the AI world models have read, yet we have plenty of smart people and geniuses. Gimme a simpler less well-informed LLM that doesn&#x27;t know about Derrida and other baloney.

  46. 1vuio0pswjnm7 · · focus · HN ↗
    &quot;The constraints on access to advanced GPUs forced Chinese labs to make performance optimization a number one goal, and it shows in the results.&quot;
  47. Yizahi · · focus · HN ↗
    When thieves are getting robbed by better thieves, I can only smirk and cheer for both sides. My only pity is that they both won&#x27;t ever end up in jail.
  48. Jamesbeam · · focus · HN ↗
    It’s the SI race!

    I am so disappointed in Sam that they still have not renamed OpenAI to OpenSI. It’s not just very disrespectful of the president, it’s outright unamerican!! &#x2F;s

  49. negura · · focus · HN ↗
    The only akward thing about all this is how Americans seem incapable of realizing that the Chinese are good, generous, hard-working people.
  50. midwain · · focus · HN ↗
    Anthropic and OpenAI haven’t released anything in open source at the level of Qwen-3.8, Qwen-Image, DeepSeek-v4.1, Kimi K3, GLM-5.3, MiMo-2, etc. For everyone, without restrictions. And yet, Chinese labs are bad and aggressive. While American labs are cute and fluffy, for everything good against everything bad. Another sad story about how rich stakeholders are being offended and the added value of the business is being stolen.
  51. KolibriFly · · focus · HN ↗
    Its funny watching Open AI and Anthropic play the role of IP defenders while quietly sliding Deepseek&#x27;s KV cache research into their minor releases to salvage their unit economics
  52. quertyrecord74 · · focus · HN ↗
    Does this mean inference is profitable now and not subsidized with vc money?
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.