‹ BackHN Continuity

Thread

Special Projects (2016)

70 points · 42 comments · vinhnx

  1. samayashar · · focus · HN ↗
    The secret to success for OpenAI, Anthropic and labs is the vision that they saw 10 years back and kept working on it. We're in awe of how models like GPT-6 Astra and Claude Opus 5.5 are performing today, but it's important to understand that they've been working on this before we knew about AI.

    The next big thing is Robots and some stealth company building today is going to be a trillion-dollar giant in few years time.

    1. owebmaster · · focus · HN ↗
      > they've been working on this before we knew about AI

      Don't confuse AI with LLMs. "We" know about AI for a long time. We even have a term for when AI fails expectations, AI winters.

      1. bonoboTP · · focus · HN ↗
        Right. If you're new to the area, it may seem like AI came out of nowhere in 2022. If you dig a little, you'll be amazed that it somehow came from nowhere in ~2012. If you dig even more, you realize there was a wave in the late 90s, early 2000s about "machine learning" (e.g. SVMs) and before it there was an 80s wave of both neural nets, agent models, and logic-based AI, probabilistic graphical models. Then you dig more and you realize AI originated from that Dartmouth workshop by Minksy and others in the 50s. Then you dig more and realize McCulloch and Pitts already modeled neural nets as little logic circuits in the 1940s. Then you realize the role of Shannon, Turing etc. Then you realize that computers actually arose in a milieu with a much more AI-shaped vision, cybernetics etc. than what we today think of as computing (PCs etc). And the precursors in the thought-formalization and mechanization trend in math and philosophy at the start of the 20th century. And even more back Leibniz's calculus ratiocinator and "calculemus!" slogan to settle debates by reducing argumentation to computation.

        The point is, typically when something seems like it came out of nowhere, it just means you didn't dig deep enough. Ideas don't come at an instant, fully formed like Athene from Zeus' forehead. It's brick by brick, one twist on an existing idea and zeitgeist at a time.

        1. mapBasketWand · · focus · HN ↗
          I was digging into this recently with ChatGPT. I’ve loosely followed the progression of ML and NN over the past 20 years, but struggled to put it into context of where an LLM lives. The big inflection point was the 2017 Attention Is All You Need paper [1].

            Artificial Intelligence
            |
            +-- Symbolic / rule-based AI
            |   +-- expert systems
            |   +-- search / planning
            |   +-- logic / knowledge representation
            |
            +-- Machine Learning
                |
                +-- classical statistical ML
                |   +-- regression
                |   +-- decision trees
                |   +-- SVMs
                |   +-- Bayesian methods
                |
                +-- Neural Networks / Deep Learning
                    |
                    +-- computer vision
                    +-- speech
                    +-- Natural Language Processing
                        |
                        +-- Transformers
                            |
                            +-- Large Language Models
                                |
                                +-- chat systems
                                +-- multimodal models
                                +-- tool-using systems
                                +-- agents
          
          [1] <a href="https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Attention_Is_All_You_Need" rel="nofollow">https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Attention_Is_All_You_Need
          1. goldenbrillianc · · focus · HN ↗
            My interpretation is that before the transformer, most everything under the domain of &#x27;AI&#x27; was either an academic curiosity or only applicable in very narrow fields. GPT-3 was when the &#x27;magic&#x27; that people had always dreamed of with AI began to emerge, and it&#x27;s only really this year that we are starting to be seriously confronted with the possibility of a general intelligence emerging from LLMs (albeit, not quite the same thing as &#x27;true&#x27; AI which would necessarily be more of a biological exercise).
            1. bonoboTP · · focus · HN ↗
              This sounds right to me as well but the point of this thread here is about how amazed one should be that this page was written in 2016. I&#x27;d say moderately, because maybe you woke up to AI being a thing this year or the last few, but the ideas obviously go back far and in surprisingly prescient ways. But this is not something unique to these AI labs. Obviously most laypeople will see the press releases and the media announcements as the big milestones and flagposts, but the academic research was inseparable from it. The AI labs obviously took on PhDs who trained in academia and had deep roots from those ideas. The story that some genius Sama invented this back when nobody thought about AI is the same exaggeration as Bill Gates inventing personal computing in his garage out of nowhere.
              1. goldenbrillianc · · focus · HN ↗
                Honestly, I was always one of the skeptics who thought AI was a probably-not-in-our-lifetimes moonshot (imagining that it couldn&#x27;t be done on existing computer architectures) and that stuff like this article was just a bunch of puffery from people who read too much sci-fi.

                But in the end, a lot of it was proven right, even if it wasn&#x27;t quite in the way we imagined.. A lot of pre-LLM interpretations of AI imagine it as some sudden 0 to 100 breakthrough, like one day someone writes an AGI program in their basement and takes over the world with it. There&#x27;s shades of that in here too, talking about worries of organizations secretly developing AI capabilities and needing to track public data for patterns to discover it. In the end there was a &#x27;magic program&#x27; in the transformer, but it doesn&#x27;t seem like they really foresaw how the program would be useless on its own, and the &#x27;AI&#x27; would come from ingesting as much data as possible, a process that has built incrementally over years and been very much exposed to the public.

                1. bonoboTP · · focus · HN ↗
                  Exactly. What wasn&#x27;t really foreseen is that it would come not through some big modeling insight but mainly through massive data scale. People much more imagined some clever general, compact learning algorithm, not quite just gradient descent, but something more intellectually satisfying, and something where you&#x27;d feel like &quot;you cracked the mechanism&quot; and you&#x27;d see clearly that some critical missing piece had to be invented that unlocked the &quot;understanding&quot; in the model, maybe some kind of fancy Hofstadterian self-referential loop or something. But it just turned out to be data and compute and of course engineering the algorithms to be efficient (which I don&#x27;t want to discount of course).

                  That&#x27;s basically the bitter lesson. Academics only reluctantly swallowed that pill and still aren&#x27;t satisfied with this answer. It&#x27;s ugly and feels like it shouldn&#x27;t work because intuition would say there are too many combinations, curse of dimensionality, etc. But it turns out it&#x27;s just line go up, extrapolate Moore&#x27;s law and don&#x27;t worry too much about philosophical-level breakthroughs just count the flops and bits. Ray Kurzweil&#x27;s scifi extrapolations turned out closer to the truth, whether deservedly or by luck.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.