‹ BackHN Continuity

Thread

LensVLM: Compressing long context as images, expanding only relevant pages

91 points · 10 comments · victormustar

  1. himata4113 · · focus · HN ↗
    I always found it weird that we don't have glacial type input for llms or any kind of active-working memory.

    There's no reason why we shouldn't be able to expose active relevant information that is only relevant for the next request: current agents running, time, etc.

    There's also no reason why we shouldn't have a cheaper lossy input which uses way less bytes per token - see deepseek flash 4.1.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.