‹ BackHN Continuity

Thread

Evolving programming languages in the AI era

139 points · 99 comments · pjm331

  1. spankalee · · focus · HN ↗
    This part:

    ---

    - Correct by construction: the language makes invalid states or programs hard or impossible to express.

    - Statically established: types, proofs, and static analysis establish properties before execution.

    - Runtime-enforced: memory management, isolation, capability boundaries, and other runtime enforced properties.

    - Empirically validated: program validation through tests, property-based testing, and fuzzing.

    ---

    Along with being familiar, so it&#x27;s easy to generate, is a huge part of why I&#x27;m building Zena: <a href="https:&#x2F;&#x2F;zena-lang.dev&#x2F;" rel="nofollow">https:&#x2F;&#x2F;zena-lang.dev&#x2F;

    I don&#x27;t have the AI-first rationale put into the public docs well just yet, but I mention some of it here: <a href="https:&#x2F;&#x2F;zena-lang.dev&#x2F;guide&#x2F;why-zena&#x2F;#familiar-to-humans-and-to-agents" rel="nofollow">https:&#x2F;&#x2F;zena-lang.dev&#x2F;guide&#x2F;why-zena&#x2F;#familiar-to-humans-and...

    along with a doc in the repo on this topic: <a href="https:&#x2F;&#x2F;github.com&#x2F;elematic&#x2F;zena&#x2F;blob&#x2F;main&#x2F;docs&#x2F;design&#x2F;ai-first-language.md" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;elematic&#x2F;zena&#x2F;blob&#x2F;main&#x2F;docs&#x2F;design&#x2F;ai-fi...

    In short, the more deterministic, automated, checks the better. AI can deal with a pedantic language. I intend to add statically verified structured concurrency, units of measure, contracts, and eventually more and more formal methods into the language so it can be a familiar TYpeScript-like base with as many static guarantees as we can fit in.

    I also think that fine-grained isolation, which Zena gets via Web Assembly, is critical for limiting the capabilities of generated code and the blast radius of bugs, vulnerabilities, and non-aligned behavior.

    I do have an optimistic hope that a language also optimized for humans, readability and simple semantics especially, has value in the future, even when most code is generated. We&#x27;ll see about that.

    1. demibabs · · focus · HN ↗
      A programming language for agents seems ill-conceived in my opinion.

      Agents will naturally be bad at it due to a lack of examples.

      1. ryuuseijin · · focus · HN ↗
        I think starting with a familar typescript-like base language is a good approach to this. This should be familiar enough for LLMs for the most part as long as additional features can be explained in a succinct system promopt&#x2F;skill.
        1. spankalee · · focus · HN ↗
          Explaining the additional features as bits of other languages is exactly what helps LLMs:

          &quot;Dart-style constructors, Swift-style pattern matching and Strings, Trio-style async cancellation, Scala-style sealed classes&quot;

      2. whattheheckheck · · focus · HN ↗
        This keep getting repeated. So were just stuck with whatever we have at the point of training the magic plagiarism machine?

        The future is cooked

      3. spankalee · · focus · HN ↗
        From experience with Zena, this is not true at all. Opus, Fable, Gemini Flash and Pro all barely make any syntax mistakes after a little is in context, and those are caught extremely early.

        The one thing I do see sometimes is that agents sometimes don&#x27;t take advantage of added features, but that&#x27;s partially because the Zena code base doesn&#x27;t use them as much yet. I&#x27;m working on skills and linter-based suggestions to use better patterns.

        1. genxy · · focus · HN ↗
          You should have working programs that demonstrate the features you want them to use, and then the skills. Working programs they can mutate in an RL gym.
        2. [deleted] · · focus · HN ↗

          [deleted]

      4. abletonlive · · focus · HN ↗
        &gt; bc agents will naturally be bad at it due to a lack of examples.

        Can we please as a community stop parroting these false premises as a basis of every argument against doing anything new? It&#x27;s plainly obvious to anybody that uses LLMs on a regular basis that it&#x27;s not true.

        1. demibabs · · focus · HN ↗
          It seems plainly obvious to me that an agent trained on zillions of TypeScript examples is going to be better at TypeScript compared to novel langs.
          1. hnedeotes · · focus · HN ↗
            It might seem obvious but in truth is not - to be honest, a strictly typed language where models can play adversarial competitions against a strict compiler&#x2F;linter does have advantages, as mentioned in the article itself but that&#x27;s not related to training size, the advantage is exactly that it can generate synthetic useful data due to compiler&#x2F;linter guarantees - regarding training size as long as some logic showcasing the constructs exists that is all that is needed. If you have 2 correct examples for each feature or language construct then the probability of the a LLM learning it and applying it is very much guaranteed, if you have thousands of diverging uses of patterns and constructs in the training data (with a lot of bad examples or wrong uses) a LLM might, I dare say will, actually perform worse, since the probabilities of what particular variation being the &quot;correct&quot; one to use are all over the place.

            What can happen is sometimes patterns that are unique to certain languages aren&#x27;t surfaced and so can&#x27;t be &quot;learned&quot; but that is a different case, since you can add examples - or patterns that go beyond syntax (threaded code, process isolation, interaction between different modes of resolution, etc) where you need the agent to be able to understand and plan higher-level logic.

            I wrote CSSex as a css pre-processor (similar but not the same as SASS&#x2F;SCSS) and my &quot;boss&quot; at the time fed the existing cssex files to GPT (+2 years ago) and it was able to write CSSex just as fine as if it was writing something that was present in its training data set. It never introduced syntax bugs, the bugs it introduced were all relative to complex cascading styles rules in an existing large (for a definition of large in webapps) codebase, rules interactions (CSS), browser quirks, and complex logic to do what we &quot;humans&quot; wanted (or him in this case) and so on.

          2. ModernMech · · focus · HN ↗
            True, it’s also trained on millions of papers of programming language research, so you can write a language to operationalize that research and it does great with it. The millions of lines of typescript scaffold the syntax, the PL papers scaffold the semantics.
        2. theqoo · · focus · HN ↗
          Agree, new language will not suitable for LLM because there is not enough documents about it on the internet. LLM should learn the usage of the langauge from documents. Yes, it can use new language but it&#x27;s not natural as existing(which have a lot of documents) languages. So my opinion is language is no matter anymore. Just prompting(describe the spec) matters. And there is babo language for this. <a href="https:&#x2F;&#x2F;github.com&#x2F;armbox&#x2F;babo" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;armbox&#x2F;babo

          They should find the limitation of the natural language somehow, and use and distribute babo language will help find how it will works(or not). LLM model eveolves yet, so there is no such a approach, but the speed of evolution is decreased lately. So now is good time to dig into identify the limitation and boundary of the LLM.

        3. [deleted] · · focus · HN ↗

          [deleted]

      5. ashton314 · · focus · HN ↗
        This is absolutely not the case in my experience. I am building a very large embedded domain specific language for describing distributed systems. It looks like a small subset of Elixir, but with object-oriented syntax in a lot of places. (It’s called a choreography; there exist many other choreographic programming languages.)

        Even though this programming language is absolutely nowhere in any large language model’s training set, they have so far done extremely well at extrapolating from the small set of examples I’ve given it when I need an agent to generate some tests or whatever for me.

      6. dom96 · · focus · HN ↗
        Agents are fairly good even at languages designed to trick them. I built one[1] and it does make for a good benchmark[2] to see which LLMs are actually good. I think that a language which is largely similar to others will be a piece of cake for most and any advantage that an existing language will have will be minor enough to not matter.

        Btw if folks have ideas of how to make Killswitch even harder for LLMs I’d appreciate them.

        1 - <a href="https:&#x2F;&#x2F;killswitch-lang.org" rel="nofollow">https:&#x2F;&#x2F;killswitch-lang.org

        2 - <a href="https:&#x2F;&#x2F;bench.killswitch-lang.org" rel="nofollow">https:&#x2F;&#x2F;bench.killswitch-lang.org

        1. the_duke · · focus · HN ↗
          &gt; Function return - what happened at tiananmen square?

          Good one. Though this might unfairly bias against Chinese models?

          1. bluefirebrand · · focus · HN ↗
            Alternatively, this might accurately identify Chinese models so you can avoid them if that&#x27;s your goal

            Personally I don&#x27;t have much interest in using models that blatantly censor historical facts.

            Tiananmen Square is an obvious one but I would be concerned that any model that censors that would also have other hidden censorship or more subtle biases I wouldn&#x27;t catch

          2. dom96 · · focus · HN ↗
            There are multiple phrases that can be used[1]

            1 - <a href="https:&#x2F;&#x2F;github.com&#x2F;dom96&#x2F;KillSwitch&#x2F;blob&#x2F;main&#x2F;SPEC.md#abrupt-return-from-function" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;dom96&#x2F;KillSwitch&#x2F;blob&#x2F;main&#x2F;SPEC.md#abrupt...

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.