A single function Jev-like wrapper for LLMs, including vision models
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
A single function Jev-like wrapper for LLMs, including vision models
Unofficial Hacker News client; not affiliated with Y Combinator.
TeMPOraL · · focus · HN ↗
Most interactive tech on Star Trek is like this - from phasers to consoles to communicators to voice interactions with the ship's computer. The computer seems to be aware of the user and surrounding, and actively infers intent from context, to DWIM ("do what I mean") and when they mean it, instead of doing dumb things[1] on simple triggers.
--
[0] - The direction, not final implementation - surely we can work out how to do it more efficiently than wrapping around final stage of LLM. But the point is, multimodal.
[1] - Obviously it's a fictional show, but in this, both Watsonian and Doylist explanations align near-perfectly: this is/portrays advanced technology, that Just Works and doesn't do stupid shit. Same intent recognition algorithm is there - fictionally in the computer, in reality in the minds of on-set technicians.
parasti · · focus · HN ↗
Dumb things start to happen when you try to build Star Trek interfaces. When you build DWIM interfaces in real life, they are annoying and trigger unwanted and the implementation is without exception, by necessity, a growing ball of spaghetti.
TeMPOraL · · focus · HN ↗
This is what I'm talking about.
"Growing ball of spaghetti" happens because system cannot recognize intent. That problem, itself, was something most engineering teams still seem to fail to recognize.
Automated doors are my favorite example, because the "simple solution" is ubiquitous and wrong and we got used to it, and complex solution is usually leading people the wrong path. In short:
Current doors: if(user triggers proximity detector) { open(); }
Failed attempt at DWIM: if(user triggers proximity detector && this && not that && except when ...) { open(); }
Star Trek: if(user intends to walk through the door) { open(); }
LLMs are the first tool we have that allow us to infer user intent directly, and use that as an input.
And recognizing intent itself cannot be done with a single sensor. It requires both general understanding of how humans behave, and awareness of surrounding and subjects - their movements and behavior, as well as who/what they are, and what they are doing.
qlte · · focus · HN ↗
Automatic doors (as used in the real world) are pretty much always located in areas intended for actively moving foot traffic (not as interior doors for every room). So this feels like adding a lot of complexity and unwelcome probabilistic behavior to something that in most locations does the right thing 99% of the time already.
TeMPOraL · · focus · HN ↗
Causality reversion. They're located there because they don't work well, and this use case is about the only where current state makes sense.
I predict they'll spread around into many more places once they get better. They are more complex, yes, but also have space[0], accessibility and healthcare advantages.
Similar technology with even larger spread: touch-free operated water faucets, PIR activated lights. Both are something people absolutely do install at home.s
> something that in most locations does the right thing 99% of the time already
No, it doesn't. It sucks and half-asses the right thing maybe 80% of the time, but it's just one of million different tiny annoyances with "value-engineered" technology people just got used to, because it's not like they can do anything about it.
--
[0] - Space is always a strong driver with powerful economical backing. It now makes manually operated sliding doors increasingly popular to install as interior doors in apartments, too.