A single function Jev-like wrapper for LLMs, including vision models (allanrbo.blogspot.com)
157 points by allanrbo 5 days ago | 45 comments




Now this is how[0] we get some of the most magical Star Trek technology that eludes us to this day, such as automatic doors. Because if you notice, they work much, much better than real-life ones, because they seem to be doing something like this:

  if(within 10 meters of door then) {
    if(Jev(
       [A] Intends to go through, expects doors to open
       [B] Approaches with no intent to pass
       [C] Passing by, loiters, or otherwise
       [D] Other
    ) == most definitely A) {
      // open doors, +/- identity/security/interlocks check
    } else {
      // ignore
    }
  }
Keywords: ambient awareness, understanding of intent.

Most interactive tech on Star Trek is like this - from phasers to consoles to communicators to voice interactions with the ship's computer. The computer seems to be aware of the user and surrounding, and actively infers intent from context, to DWIM ("do what I mean") and when they mean it, instead of doing dumb things[1] on simple triggers.

--

[0] - The direction, not final implementation - surely we can work out how to do it more efficiently than wrapping around final stage of LLM. But the point is, multimodal.

[1] - Obviously it's a fictional show, but in this, both Watsonian and Doylist explanations align near-perfectly: this is/portrays advanced technology, that Just Works and doesn't do stupid shit. Same intent recognition algorithm is there - fictionally in the computer, in reality in the minds of on-set technicians.

parasti 5 days ago | flag as AI [–]

"doing dumb things" and "stupid shit" is an odd choice to describe tools that only trigger on explicit activation. Is a windshield being lowered by a switch being held a "dumb thing"?

Dumb things start to happen when you try to build Star Trek interfaces. When you build DWIM interfaces in real life, they are annoying and trigger unwanted and the implementation is without exception, by necessity, a growing ball of spaghetti.


> Most interactive tech on Star Trek is like this - from phasers to consoles to communicators to voice interactions with the ship's computer.

Almost like the Star Trek mechanisms can infer perfect intent.

Like there’s a hidden script or something.

More seriously, I think there’s real value in an automatic door that behaves consistently rather than one that tries to infer messy human intent. Real life isn’t a TV show and there’s both ambiguity in how people behave and how they even intend to behave. It’s mostly not hard to understand how a proximity sensor door will function. Using a black-box classifier to improve that won’t necessarily make people like it more. And calling up to the cloud for every sensor event, ignoring privacy issues, adds weird latency and a huge failure mode during data center outages.


How does the model get context to decide ABCD ?
prathje 5 days ago | flag as AI [–]

Nice! I would love to use it for images as well. Then again is using Grammar-Based Decoding with a json response not the same? Is Jev just that with nice caching? Because then I have been using that already…

A functional difference is the probabilities associated with each option. True, you could use the constrained next-token distribution at the appropriate generation step, but those would not be calibrated by any means, which Jev's claim to be. An empirical analysis would be interesting.

Almost the same. The neat trick here is to ask the model to reply with just 1 letter, which is one token, rather than a long string of json. Its quicker. And just the fact that Jev made a pretty decent API for structuring your questions. And they do the massive parallelization.
ulam2 5 days ago | flag as AI [–]

Yes, that is my question too. Someone knowledgeable can comment
frabcus 5 days ago | flag as AI [–]

Presumably this is much less good than Jev, because the normal LLM models have been trained with RLHF and to be agents. Especially on a large model, I'd expect it to decide in an earlier layer.

I'd hope whatever Jev's Reinforcement Learning for Calibrated Decisions (RLCD) does is better at training the models to give accurate probabilities in the weights.

hhayes 5 days ago | flag as AI [–]

I disagree that layer depth matters here. The wrapper reads logprobs at the answer token, which is where the model commits anyway. And calibration baked in by RL tends to break off-distribution, while a plain wrapper lets you recalibrate on your own labeled data in an afternoon. Why trust the weights over a held-out set?

Empirically this approach is more accurate, faster and about the same price as Jev.

https://github.com/Mushroom-Systems/lichen

Mashimo 5 days ago | flag as AI [–]

[A] Hotdog

[B] Not a hotdog


one day i was about the tell my girlfriend the wonders of ai and how it works underneath. She stopped me about 30 seconds in " so hotdog, not hotdog? " i was like "yep".

I never bought that up again.


Ah sweet it’s like Jev but several order of magnitude more expensive, and slower too.

Jev doesn't support images, so it's hard to compare this directly. But in general this approach beats Jev in its own benchmarks for accuracy and speed and is about the same price.

https://github.com/Mushroom-Systems/lichen

lostmsu 5 days ago | flag as AI [–]


What I'd want to see next to accuracy is tail latency. In a real-time use, deciding when a spoken sentence is finished, a general LLM with the same prompt was slower and more hesitant for us than Jev, even though both cost about the same.
czl_my 5 days ago | flag as AI [–]

I've created a Jev wrapper so that it can work via any OpenAI-compatible endpoints https://github.com/zhulinchng/jevper
Havoc 5 days ago | flag as AI [–]

Likely works even better with fireworks ai since they have proper grammar support

Ah, will try, thanks!
roger_ 5 days ago | flag as AI [–]

> Answer with the letter of the best option only [A, B, C]

What if it says "D"? What if it tries to say "Additional details needed"?

(Also no calibration, etc.)


I don’t know how this wrapper works, but if it is like any of the classifiers I’ve had Claude build off an LLM in the past, it grabs the probabilities of the tokens you are looking for, and then computes their relative probs against each other.

Even if the LLM thinks it’s made up D is the highest probability, that isn’t part of the set.

You never actually generate the prose, only the first pass, and grab the probabilities. It couldn’t ask for more details even if it wants to. It gets stopped before the first token renders.

roger_ 5 days ago | flag as AI [–]

I guess “A” by itself would be a seperate token but my point is the model might be trying to say something that begins with that letter rather than actually answering.

You’d need to use an approach that links the output to a closed set.

edunn 5 days ago | flag as AI [–]

We did this for a ticket classifier. Renormalizing over A/B/C hides the "D" problem, so we also logged the raw probability mass on those tokens before renormalizing. If it's low, say under 0.5, the model wanted to say something else, and we treated that as an abstain and sent it elsewhere.
bicsi 5 days ago | flag as AI [–]

Of course it works, Jev is nothing but an API breakthrough

Jev is rumored to be a 30B model, and it's input price is MUCH cheaper than similarly sized models. The maker is also heavily focused on having a profitable product, so it's unlikely to be subsidizing the cost, especially since they say they have more demand than what they can serve.

You can get Gemma 4 26B A4B at the exact same token input price of $0.042/M. GPT-5 nano is not much more expensive at $0.05.

https://openrouter.ai/google/gemma-4-26b-a4b-it


They can't overturn the economics of attention by restricting themselves to a single token output.

Sure they are no longer memory bandwidth bound thanks to that but someone could add a similar projector to a conventional model, train with a Jev style dataset and call it a day.

Whatever they are doing on inputs must either mean they intentionally chose a Mamba successor or they suffer from the same compute costs as everyone else.


Cheap input matters more than people think for this wrapper. I've been running it over a few thousand images, and the cost is nearly all input since the answer is one token. Gotcha: you lose any chance to let it reason, so prompt wording swings accuracy more than I expected.

That is the most intuitive way of doing it. The hype is insufferable.
psoto 5 days ago | flag as AI [–]

A single function, if you don't count the datacenter behind it.