As a proof of concept I've been working on a Jev style vision model that can rapidly answer boolean and choice questions based on images (See [0] but not the point of this comment). One of the reasons I wanted to make this concept is because I was also thinking about exactly what the parent comment suggests. I just haven't gotten around to getting (3d printing) a robot arm yet :)
Nice work! Vision on the qwen models works very good. Jev is a really neat idea. I tried replicating jev style with qwen models using max 1 token output, which works great. How does this compare to that in terms of speed?
The problem is that with max 1 token it is not 100% guaranteed that it will be a yes/no or a score. This model (and engine) solves that by taking the probability of the yes/no token. For 1 question it is faster but not a whole lot, the power comes from being able to ask more then 1 questions for merely a few ms extra per question.
E.g. 1 question would be 300ms but 30 questions would be 360ms total. Most of the time comes from the image decoding.
I've seen on X that someone made a self driving car driven by Jev. I didn't check in details what it was about. I suppose you could give it some finite set of answers. Just wondering if someone tried it.
With a car I can imagine asking questions like "am I on the left side of the lane", "am I closer then 10 meter to the next car" which feel like questions the current models could answer. I see models understanding the calibration of the robot. Maybe you could have a jev style model make decisions like "move left", "activate gripper" and do the inverse kinematics in a different system.
I know someone that's mucking about with it in a hierarchical control setup. I think the awkward thing is that you need to have prespecified options for it to choose from, rather than have a model that can specify arbitrary joint positions or target poses.
Mystery-Machine · · focus · HN ↗
thomasikzelf · · focus · HN ↗
WildGreenLeave · · focus · HN ↗
[0]: <a href="https://huggingface.co/MeerDevelopment/Qevi-2B" rel="nofollow">https://huggingface.co/MeerDevelopment/Qevi-2B
thomasikzelf · · focus · HN ↗
WildGreenLeave · · focus · HN ↗
E.g. 1 question would be 300ms but 30 questions would be 360ms total. Most of the time comes from the image decoding.
systemerror · · focus · HN ↗
Mystery-Machine · · focus · HN ↗
thomasikzelf · · focus · HN ↗
ainch · · focus · HN ↗