‹ BackHN Continuity

Thread

Jeeves. Reasoning improves Jev-like decision models

242 points · 95 comments · nicowaltz

  1. TN1ck · · focus · HN ↗
    I just did a run with a benchmark I just used to test other models against. (It's about detecting irony in german soccer tweets). On my M5 Pro with 48GB it took over 30min to decide on just 100 tweets, the thinking definitely takes long.

    It performed quite below Jev, but above other open decision models I tested (68 correct vs 79 correct for Jev - see [1]). I'm running it for the moderation benchmark as well, but that will probably take a few hours on my machine.

    [1] <a href="https:&#x2F;&#x2F;tn1ck.com&#x2F;blog&#x2F;jevdit" rel="nofollow">https:&#x2F;&#x2F;tn1ck.com&#x2F;blog&#x2F;jevdit

    1. TN1ck · · focus · HN ↗
      Update: Jeeves took about 2 hours to moderate 394 data points and performed really well. It’s not as good as Jev, but it’s super close! In general, it’s super cool that you can tune how strict you want content moderation to be with these models.
      1. nicowaltz · · focus · HN ↗
        cool to see!
    2. robbie-c · · focus · HN ↗
      One of Nico&#x27;s colleagues here, I&#x27;m working on an MPS port <a href="https:&#x2F;&#x2F;github.com&#x2F;PostHog&#x2F;jeeves&#x2F;pull&#x2F;1" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;PostHog&#x2F;jeeves&#x2F;pull&#x2F;1 which should speed up M5 Pro performance quite a bit
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.