‹ BackHN Continuity

Thread

Tokens too cheap to meter

354 points · 227 comments · teoruiz

  1. meatmanek · · focus · HN ↗
    I just want to rant about these Artificial Analysis charts that you see everywhere:

    The "most attractive quadrant" is completely meaningless. The whole point of a Pareto curve is that each point on the curve is better than everything else on at least one dimension, and that you can make these comparisons without placing a value judgement on the relative importance of the different metrics. If you make a composite score of the two metrics (any monotonically non-decreasing function, e.g. a weighted sum with non-negative weights), that score will always be maximized by one of the points on the Pareto frontier.

    So going by the numbers in the 2nd chart (1st AA chart) from TFA alone:

       - there's no reason one would choose Deepseek V4 Pro 0813 (max) even though it's in the "most attractive quadrant", because GLM-5.3-Flash is both cheaper and scores better.
       - Claude Fable 5.1 (max with fallback) on the top right* could be your most attractive option if you need the best scoring model and don't care about cost, even though it isn't in the "most attractive quadrant"
       - The un-shown model off the left side of the chart could be your most attractive option if you just need lots of cheap tokens and don't care about quality.
    
    (Obviously if you start including other factors in your score that aren't represented on the chart, then you might choose differently.)

    * I also dislike the way they place the labels, and that grey line connecting the label to the point is way too subtle.

    1. Hugsun · · focus · HN ↗
      Sound critique. I'll add that the Artificial Analysis intelligence index is not considered a good metric for intelligence anymore. Most of the benchmarks that it comprises are saturated or considered low signal today.
      1. Gander5739 · · focus · HN ↗
        What is considered a good metric?
        1. nijave · · focus · HN ↗
          Personally, a combination of low tech and semi scientific tests based on what you normally do (redo the same task you used an older model for with a newer one).

          So far, Simon Willison's pelican bike "benchmark" is the only one I've found that shows Fable 5.1 beating Opus 5.5. My personal experience has been Opus is unusable on design work it's so terrible. Evaluating whether we should consolidate AWS DMS tasks (Postgres full load and change data capture) into fewer tasks with more tables, Opus 5.5 was factually wrong and needed correction roughly every other turn.

          On a "help me find a sandbox solution for agents embedded in a web app to run untrusted code" research project it kept misrepresenting security boundaries and ended up recommending DuckDB which ironically specifically says it does not provide a strong security boundary in its own documentation. GPT 6 (can't remember if it was Sol or Astra) and Fable 5.1 both recommended FaaS like Cloudflare Workers and AWS Lambda which fit fairly well with the requirements.

          I switched from Opus to Fable in the session going badly sideways and told it to "Review the previous conversation and come up with a correct comparison table and corrected recommendations grounded in objectivity supported by citations. Do research as necessary to understand the current ecosystem" and that was a full 180 back to coherency...

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.