Kev: Tiny decision models you can train and run on your own
Kev: Tiny Jev-like family of decision models built on top of Qwen3.5
Kev is a family of small decision models (0.8B, 4B, 9B) built on Qwen3.5 that answer yes/no, multiple-choice, and rating questions in a single request. They run on CUDA and Apple Silicon, and you can train your own using the provided code and data. The API matches TypeSafe's System One, so you can point their Python SDK at your local server. A web playground lets you test inputs and check how option order affects answers.
That is the point of getting probabilities back instead of a single label.
- nico
If you only need classification, and you can provide some training data, you can ask Codex/Claude to build an embeddings + logistic classifier model for you
For emails, I get 95% accuracy with this method, with only 50-100 examples for training
Training the model takes less than 5 minutes on a CPU
The resulting model is <1MB, and inference is sub 100ms
Some other cool things about this approach:
* the model doesn’t train on some “ideal” or general classification, instead it learns your preferences
* the model runs on pretty much any mobile device and can be retrained online on the device
* privacy, the whole training and inference is 100% local, no data goes anywhere (except whatever you feed codex/claude while building the model)
Note: to do a more general test, I made a classifier for the Banking77 dataset. The model is <10MB, trains in <30s on CPU and gets 94.5% accuracy, which puts it in the top 5?models by accuracy for that set (the best one is at 94.86%, but it’s 350MB in size and takes hours to train on a GPU).
- prodigycorp
Man, I'm already burnt out on all this jev talk.
The one thing jev has going for it is a dedicated company focused entirely on making the product good and keeping it maintained. I haven't been willing to jump on board with all these jev-shaped projects because their releases feel driven mostly by opportunism. I'm fine waiting a bit for the opportunists to shake out so we can see who is genuinely committed to bringing something valuable to the open-weight community.
Jev is much better than the traditional ML crowd gives it credit for, but my enthusiasm hits a wall when it comes to their data policy. It is completely draconian. Whatever you feed into the system, they retain.
The jev team needs to release a ZDR product, or their platform is dead on arrival. An open, jev-shaped model will win out solely on that basis.
- oscarfr
Found this benchmark for Jev-class models: https://benchmarkheaven.com/jev-models
There are already many Jev-like models in there.
Edit: No affiliation. Just found it and thought others might find it interesting.
- hbarka
If Jev is fundamentally trained using RLCD while you’re building on a Qwen model that was trained using RLHF, how can the resulting model be considered Jev-like?
- monkeydust
Bit of a Jev explosion going on. Is it because it's taking us back to a simpler time we understand better? Classification models have been around for a while.
- nullbio
I think a great use case for these will be when they have large context windows and are able to enforce styling rules for frontend development, and component creation rules for react. You can then ditch the styles guides and styling skills and create a decision tree for enforcing styling, so that you can't run into drift issues or duplication issues. That's where I'm wasting most of my time right now, constantly correcting all of the UX/UI issues that are created for every single feature.
- IronWolve
I wrote a small proxy that points points to a jev api and a frontier api.
My harness connects to the proxy and only sees the frontier api, models, commands, etc. When I send prompts with tool calls, proxy routes to jev, jev narrows the tools, proxy cleans/sends to the frontier api.
So far in my tests, about 60% less tool calls. I'm also going to implement model switching, so it can use cheaper models. I think my workbench harness needs its prompts cleaned up.
- aetherspawn
I hope these get small and good enough to create “pet like” AIs for games. You know, like scream “follow me” at an NPC, STT stack translates it and feeds it to a local Jev-like model that then picks a number of things for the NPC to do.