‹ BackHN Continuity

Thread

How AI tool calling works (40 lines of vanilla JavaScript)

15 points · 9 comments · stephenblum

  1. stratos123 · · focus · HN ↗
    The writing is slop. It doesn't mention a detail I have seen people actually not get: that LLMs wrap tool calls in special tokens, so they are "out of band" and can't be mistaken with normal output.

    It also spends an entire section trying to convey that large tools waste tokens:

      A read_file that returns an 8k-token source file on turn 2 of a ten-turn agent gets resent on the eight requests that follow: 8 × 8k = +64k input tokens, $0.32, from one tool result.
    
    but surely that's wrong - it's part of the same conversation, it only gets processed once and then cached. Otherwise doing long conversations would always cost an amount quadratic with length.
    1. stephenblum · · focus · HN ↗
      You a are right as The article doesn't mention KV cache. Also, yes the writing is slop, and does not mention the LLM special tokens. Only discusses the use through the OpenAI-style JSON wrapper that allows you to define the schema of a tool call in JSON. For most of the audience, they are looking to build AI agents, and the high-level tool calling interface is what they would be using
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.