Turning GLM-5.3-Flash into a Jev-like decision model (privatemode.ai)
138 points by flxflx 5 days ago | 59 comments




Everyone is doing this to emulate Jev, but...

I took a random book excerpt with 23,000 words (±30k input tokens) and used it as context. Jev still responds in 800ms, sometimes 500ms. That's in the neighbourhood of 20-50,000 tok/s prefill, which is obviously not possible with normal LLMs, not even Cerebras is this fast.


This has “/dev/null as a service” vibes…

Do we have a reliable way to count tokens for Jev yet btw?

Also, it processes all questions you ask it in parallel, which is also not possible with normal LLMs.

Did you try changing the very first few tokens between runs? We got burned by this once. Prefix caching (vLLM's APC, SGLang's RadixAttention) makes a repeated 30k prompt nearly free, so 800ms might be mostly cache hits. Prepend a random string and re-measure, you'll see the real prefill number.

20k tok/sec prefill on B200/B300 isn't particularly noteworthy for medium-sized models like GLM-5.3-Flash, vLLM and SGLang achieve it on a reasonable number of models, especially at NVFP4.

50k tok/sec is pretty impressive though.

But... when you were doing your measurements, were you using the same random book excerpt? If you were potentially getting even partial cache hits for your 50k tok/sec measurement, it would taint the benchmark: pretty much any inference provider running any LLM would be able to hit those numbers.


I have run tests with qwen 3.8 and gemma 4 in a way similar to this post (based on an open source project that also does this with gemma4).

Getting competitive accuracy with Jev is fairly easy, if by accuracy you mean that the highest weighted answer is the right one. GLM 5.3 is complete overkill, much smaller llms will do

What Jev brings to the table, beyond speed, is that the reported probabilities match actual likelyhoods. If you present three options, with A and B equally likely and C impossible, jev will approximately answer with 0.5, 0.5, 0. Stock LLMs don't

m4y0u 5 days ago | flag as AI [–]

My question is why not use Jev instead? It's faster and cheaper.

> My question is why not use Jev instead? It's faster and cheaper.

Because it's proprietary? By using an open weight model you're guaranteed that you can access it forever; if one provider bans you then you can go to another one (or you can self-host). With a proprietary, single-provider model locked behind an API if your access is revoked you're screwed.


There's some speculation that Jev is an open weight model with novel post-training (RLCD). So, if these folks have competitive accuracy with just the base model, it may raise some questions about the necessity of Jev's architecture. You generally don't want to find yourself competing only on price.

Fyi, I haven't tested this yet.


I mean just from what's known of the funding and timeline it pretty much has to be based on open weights.

But it is likely more than just a fine tune + novel training. At the very least the LM head is swapped out for a classifier one and then or also idk, bidirectional attention for the encoding pass I'm out of my depth at this point and will stop guessing. The training is probably where they have the biggest moat though, not that it's necessarily huge.

I have a project that fits jev as advertised almost comically well and I've been playing with it, and the various hacks and open versions. Jev doesn't necessarily perform better overall but it is quite different. It's sensitive to prompt phrasing in ways the others aren't, it's easy to generate questions where all the other models cluster in confidence but jev is an outlier. Not necessarily more correct, but it does feel like it's getting its answers in a different way.

I'm guessing just as much as anyone else but I've been spending a ton of time on this the last couple weeks, it landed right when I was most ready to dig into it.


These questions are answered by the OP (Same speed, image support) - additionally, GLM is open weight.

Compliance. You can’t put Jev in a HIPAA compliant service, for example.

You can turn any sufficiently smart LLM into yes/no decision model or equivalent. I already have an existing workflow with a two paragraph detailed prompt, that sends pages of stuff to an LLM and asks it to return only 7 JSON objects. Several of those objects are binary "yes or no" choices of like, whether the content contains certain things.

You can even do it with small not particularly hard to host local LLMs like a variant of Qwen 3.6 35B A3B or 3.8 27B.


Format is the easy half, constrained decoding handles that. Has anyone measured the other half? Run the same page through twice, does the "contains X" flag flip? On pages-long input I'd guess a 27B model drifts on borderline cases, and nobody notices because the JSON is valid.

But does it always 100% of the time sticks to the schema? We have a prompt that is explicitly instructed to return a single html tag with the response inside it and it sometimes hallucinates

yes, it does, with appropriate tuning/testing of the llm's parameters (temperature, top p, top k, using the right model, and the right prompt). You have to give it a rigid and very specific prompt to only answer in the JSON form. Also test it with LLMs that will handle being given a very low temperature to be very 'literal'. You're not asking for creative writing.

I should also add that the source comes from one of about 400 possible places and in a variety of messed up formats, it's the raw feed from a news scraper...


Yes, you can use constrained decoding
cedar 4 days ago | flag as AI [–]

Syntactic validity is guaranteed, as far as I know, since invalid tokens get masked at decode time. What isn't guaranteed is that the content inside is right. A schema-valid wrong answer is still wrong, and tight constraints can nudge quality down a bit on some models.
gf000 4 days ago | flag as AI [–]

I'm fairly sure it is 100%, and has been available for "ages". That's how every tool call and whatnot works:

https://developers.openai.com/api/docs/guides/structured-out...


Qwen3-Next-80B-A3B can already run on a 16GB M1 MacBook at around 3–5 tok/s using aggressive memory management. Could a Jev-style controller push that to 100 tok/s on an M1?
prjkt 4 days ago | flag as AI [–]

how is Jev cheaper if I can run locally. 0.5% prefill, 0.1% decode, 99.4% cached, latency is <20ms
ttoinou 4 days ago | flag as AI [–]

Isnt this obvious ? I would have thought people would try such things before deciding they need something like Jev

It is obvious. I was going to build my own little toy doing just that ten days ago, and then found four pre-existing projects, so wrote those up instead: https://sgnt.ai/p/jev/
e12e 4 days ago | flag as AI [–]

Great article, thanks for posting.

I think I'm gradually getting a picture of what jev does, and why it might be good:

1) one shot classifiers are great, and are well-known.

2) LLMs are also great classifiers, but are slow at output - and not great at guaranteed structured output.

3) everyone focused on LLMs, and forgot about classifiers.

4) any decent one shot classifier should be easy to run "fast"/"in realtime".

5) combining 4) with LLMs and a harness (a loop) open up some clever use-cases that LLMs alone are too slow for.

As I understand it, jev isn't long for this world (we'll have foss one shot classifiers we can run on modest hardware) - but it seems likely the legacy will be that classifiers will re-gain some prominence alongside generative models.

prism94 4 days ago | flag as AI [–]

Same shape as Viola-Jones cascades in 2001: cheap classifier up front, escalate only the uncertain cases. Gets renamed every few years. Four existing projects doesn't surprise me. The hard part was never the idea, it's calibrating the confidence threshold.

It is.

What isn't obvious is why people keep shouting "Jev Jev Jev" all the time.

Astroturf.


I am just this far from adding Jev posts to my blocklist.

I cannot fathom how people are not seeing the multiple daily posts as anything but the spam they are.


don’t worry once they get acquired for 30 billion usd then the spam will stop.

Not for people with llm brain antrophy

If you are using an autoregressive decoder (which glm is) it is not “jev-like”. You lose all of the speed advantages that Jev has.
yogthos 4 days ago | flag as AI [–]

> We measured latency in separate runs with one request at a time, because timings taken under load measure the queue rather than the model.

> As Privatemode is hosted in the EU and Jev is hosted in the US, we ran four of the datasets from Germany and from the US at the same time. From Germany, Privatemode answered in 180 ms and Jev in 264 ms. From the US, the order reverses: 164 ms for Jev against 299 ms for Privatemode.

turns out there is a trick to keeping the context filled and only evaluating a handful of choice tokens https://www.youtube.com/watch?v=bcGO7xre46o


Right, so it is double the latency and will no longer feel real-time to the end user.

I personally love Privatemode's approach. Having Jev-like speed for confidential ai use cases is a huge enabler.
k__ 4 days ago | flag as AI [–]

Is it non-autoregressive?
Jabrov 4 days ago | flag as AI [–]

Is this a joke? “Jev-like” properties? People have been using LLMs as classifiers or rankers in a similar way for ages. I feel like we’re losing our minds

Jev's particular architecture is such that context can be filled faster than with LLMs.
hbrn 4 days ago | flag as AI [–]

I have a bridge to sell you.
yogthos 4 days ago | flag as AI [–]

RIP Jev

Fine-tuned yes/no gate works until input distribution drifts. Then it fails silently, confidently wrong, nobody notices till customers do. Where's the alerting? Log disagreements against a boring rules-based fallback, or you learn this at 3am.