‹ BackHN Continuity

Thread

A single function Jev-like wrapper for LLMs, including vision models

157 points · 45 comments · allanrbo

  1. TeMPOraL · · focus · HN ↗
    Now this is how[0] we get some of the most magical Star Trek technology that eludes us to this day, such as automatic doors. Because if you notice, they work much, much better than real-life ones, because they seem to be doing something like this:

      if(within 10 meters of door then) {
        if(Jev(
           [A] Intends to go through, expects doors to open
           [B] Approaches with no intent to pass
           [C] Passing by, loiters, or otherwise
           [D] Other
        ) == most definitely A) {
          // open doors, +/- identity/security/interlocks check
        } else {
          // ignore
        }
      }
    
    Keywords: ambient awareness, understanding of intent.

    Most interactive tech on Star Trek is like this - from phasers to consoles to communicators to voice interactions with the ship's computer. The computer seems to be aware of the user and surrounding, and actively infers intent from context, to DWIM ("do what I mean") and when they mean it, instead of doing dumb things[1] on simple triggers.

    --

    [0] - The direction, not final implementation - surely we can work out how to do it more efficiently than wrapping around final stage of LLM. But the point is, multimodal.

    [1] - Obviously it's a fictional show, but in this, both Watsonian and Doylist explanations align near-perfectly: this is/portrays advanced technology, that Just Works and doesn't do stupid shit. Same intent recognition algorithm is there - fictionally in the computer, in reality in the minds of on-set technicians.

    1. moregrist · · focus · HN ↗
      > Most interactive tech on Star Trek is like this - from phasers to consoles to communicators to voice interactions with the ship's computer.

      Almost like the Star Trek mechanisms can infer perfect intent.

      Like there’s a hidden script or something.

      More seriously, I think there’s real value in an automatic door that behaves consistently rather than one that tries to infer messy human intent. Real life isn’t a TV show and there’s both ambiguity in how people behave and how they even intend to behave. It’s mostly not hard to understand how a proximity sensor door will function. Using a black-box classifier to improve that won’t necessarily make people like it more. And calling up to the cloud for every sensor event, ignoring privacy issues, adds weird latency and a huge failure mode during data center outages.

      1. philbo · · focus · HN ↗
        Spoken like a person who's never had to queue in a shop with an automatic door and then the queue reaches too close to the door and then you're the unfortunate person who keeps on accidentally opening the door while standing at the back of the queue and then everyone else in the queue glares at you.
        1. qlte · · focus · HN ↗
          So now you turn around to check if anyone else is coming in behind you and the door triggers. Or alternatively, you finally get sick of waiting and turn around but the door takes 5 seconds to realize you're not just doing aimless human fidgeting and actually want to leave.

          Also .15% of the time the door just randomly opens or stays shut with no rhyme or reason. Bug previously filed with cloud door vendor but told it's within door SLA and sent previously signed copy of agreement that "all neural nets are a black box and inherently probabilistic so errors may occur at any time and not a product defect"

          1. TeMPOraL · · focus · HN ↗
            > So now you turn around to check if anyone else is coming in behind you and the door triggers. Or alternatively, you finally get sick of waiting and turn around but the door takes 5 seconds to realize you're not just doing aimless human fidgeting and actually want to leave.

            You're literally describing how current automated doors work. That's because they can't infer what you actually want to.

            Intuition pump: imagine the doors are actually manually operated by a set technician (like they are on Star Trek sets). Such a person can trivially tell, in split second, whether you're turning around to leave or just to check if someone is behind you, and open the door at the right time. They are inferring intent. This is what I'm saying transformer models are finally allowing us to implement in silico.

            > Bug previously filed with cloud door vendor

            Who is even talking about cloud here? Giving operational control over hardware to a third party located far away from deployment site is just monumentally stupid in nearly every case.

            This is one of the top use cases for local models, because you need video feed processing with sub-second reaction time and real-time guarantees.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.