‹ BackHN Continuity

Thread

LensVLM: Compressing long context as images, expanding only relevant pages

91 points · 10 comments · victormustar

  1. kazinator · · focus · HN ↗
    Say, what? When would a compressed image of text be smaller than just text?

    Maybe to save on doing the client-side layout and rendering?

    I'm reminded of the MSPaint IDE:

    <a href="https:&#x2F;&#x2F;www.youtube.com&#x2F;watch?v=eyH4aXlB1Js" rel="nofollow">https:&#x2F;&#x2F;www.youtube.com&#x2F;watch?v=eyH4aXlB1Js

    <a href="https:&#x2F;&#x2F;ms-paint-i.de&#x2F;" rel="nofollow">https:&#x2F;&#x2F;ms-paint-i.de&#x2F;

    1. MomsAVoxell · · focus · HN ↗
      &gt;Say, what? When would a compressed image of text be smaller than just text?

      Since every character in text shares compressible ligaments with every other character, maybe? Or, more finite, every pixel representing a ligature component in a character can be compressed against other pixels in the path set.

      I think, at scale, this is very important - folks training models on near-petabyte sized corpus would have cause to want to use this in pipelines, I imagine ..

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.