I've spent endless hours playing this game as a child so building this was a ton of fun.
I open sourced everything in case you want to hack on it yourself here: https://github.com/christianmat/jev-pokemon
The game is being streamed live including the tokens and cost - hopefully we get all the badges and don't get stuck in a cave :)
That being said, one thing having been unrealistic 10 years ago and just about possible today doesn't mean that it's going to change the world the same way another technically related, previously-impossible thing did. The Jev hype gives me a bit of the "you're still early to crypto" vibes of some later altcoins. I really like the idea, I think it's going to open up possibilities for using classifiers where we wouldn't or couldn't have trained one before. I'm crossing my fingers for an open weights version to drop. But it's still just a classifier, people have built similar things before Jev, the one thing that really stands out about it is their ability to generate hype.
That's the point. It can't. And it's not even close.
why
A bad universal classifier does suggest a good one later. And that is exactly what I would call "theoretically doable"
That said, I don't think that Jev is a magic breakthrough or anything. I think it is just a particularly good narrative with an easy way to try it out.
But I've seen nothing to indicate that the upper bound on classification tasks of a Jev-like model can exceed a frontier LLM with reasoning tokens. That seems nearly impossible even in principle (since Jev-style models are still based on LLM pretraining).
So while they're definitely on the Pareto frontier, which is valuable, they're at the "cheap" end of the spectrum more than the "good" end and I don't expect that to change.
The magic moment for me from the Jev release was not that there was some system playing doom: Rather it was the moment, they just changed a part of the prompt to "don't shoot, just dodge" and the behavior changed immediately.
This means you can have a system with fast decision-making but still interact with it via language.
They even named the model after "Jevons Paradoxon" - They anticipate that their model lowers the cost of adopting this kind of AI significantly, unlocking a lot of use cases.
I work in manufacturing. I think this will be fantastic for stuff like SPC.
But good enough for pennies in an instant is very useful.
Ie the famous "Hotdog" clip from Silicon Valley [0]
Example (2022):
https://developers.openai.com/cookbook/examples/zero-shot_cl...
> get stuck in strange loops of going in and out of the same door to no end
Math.random is statistically unlikely to do this.
> the most interesting thing about this jev stuff
> is that people are seemingly like
> completely disinterested in how smart it actually is
> I haven't even heard it mentioned a single time how it actually compares to other LLMs coming up with their own classifications. Just: it's fast and cheap
After watching a few minutes of this it makes me think that maybe we should be a little more interested in how smart it is.
My main wonder is the difference between it and having a small llm no thinking output a single number only as a choice. Isn't that nearly the same here?
For me in my day job, having extremely fast low quality decision makers over noisy inputs is very valuable. I work in security and having something that can help triage alerts, classify items and group things together is extremely valuable. It doesn't need to be perfect. Just being able to take a set of inputs from deterministic tooling and to be make general priority classifications goes a long way on helping humans look at the most important items first.
Why would you be confident in feeding garbage to a “cheap and fast” classifier with unknown domain-specific performance?
I know we kind assume omniscience for frontier models, but at this point the evidence is kind of out there.
Its very cool, but the harness is doing a _ton_ of heavy lifting.
I think if it was combined with a regular vLLM it could be really interesting, especially watching the reasoning logs.
Bonus points if it was one of the latest open models that somehow had all prior training knowledge of Pokemon abliterated so it was reasoning as an intelligent persona that had no knowledge of even the concept of Pokemon.
Props to OP for getting a working version, but it does not seem that this model is capable enough to play Pokémon at this point in time.
I predict it will keep trying and failing without changing strategy until the team is so overleveled that brute force works.
Unfortunately, the only usable move is flamethrower with 15 PPs that will be exhausted.
Remember, if it starts using items, it can run out of money. And what then?
But yes, brute force might prevail in the end. Let's say it started elite four at 13k jev calls or 1.3 USD. I'd call it a "loss" if it crosses the 2 USD threshold. It's still quite cheap though.
- establish a composition
- train the selected composition to an appropriate level
- ration PPs through 5 fights
- ration items through 5 fights + globally (you can run out of money)
- maybe even switch the order of pokemon on the list? Poor Graveler
I'm not sure it can follow-through on this. Would be funny if it had to re-do the rock boulder puzzle. Guess I'll keep watching!
You could front this with an image -> text model but that would be much lower quality vs latency, and the whole point of doing it with a decision model is remove the latency.
Games are a really interesting testing ground for robotics; if we can solve game playing (incl 3d) we could embody "system one" intelligence into robots that have something emulating general reflexes without needing to fine tune.
Take a look at a snake demo with streaming image inputs: https://x.com/spillai/status/2103630735425089957
Did I miss something? I thought one of the demo videos was it doing pretty decent at the first level of Doom?
Teaching a World Model to Play Pokemon - https://news.ycombinator.com/item?id=49849907
Very cool proof of concept!
The interesting thing here imo is the cost and latency. So far we're at 4 badges for less than $0.5
However, https://x.com/TynanSylvester/status/2096965749369720970 Astra was able to beat RimWorld. So LLMs are definitely able to drive these sorts of games to completion with their current abilities.
Has anyone tried a combined LLM + Jev? So the LLM directs the high level goal ("reach the door of the pokemon center while avoiding NPCs") then allow it to instruct Jev to do the actual movements? That seems like a good balance between the high-level, slow planning of the LLM to the fast but limited Jev. That kind of mirrors how humans work too. The harness could even allow Jev to delegate back to the LLM when it doesn't have a confident answer.
Its almost like when humans drive, we kind of let our subconscious take over once we know where to go. But if we see something unexpected in the road, we can go back to a conscious planning mode to decide what to do.
The "Jev calls" counter only seems to increment at junction points like battles, conversation prompts, menus etc.
Is something else moving the character around?
All in the OSS repo if you wanna play around with it: https://github.com/christianmat/jev-pokemon
Super cool to see it do the whole game. I spent my fable budget building something similar this week but only drove it to Brock. I like the "current focus" framing too.
The harness reads the ROM/RAM, builds a collision grid, computes reachable paths with A-star, constructs a region graph, calculates distance to the current objective, determines whether an exit leads toward or away from the objective, identifies Pokémon Centers and Marts, detects NPCs and items, and then turns all of that into natural-language choices for Jev.
Jev is given a list of "choices", already annotated information that tells Jev whether the choice should be taken or not.
For example, Jev might receive options like:
* Enter ROCK_TUNNEL_1F — Leads toward the objective (2 areas away)
* Go west to ROUTE_9 — Leads away from the objective (5 areas away)
* Talk to NPC X — Mentioned in the current objective
* Enter POKECENTER — toward nearest Pokémon Center
Those "toward the objective" judgments are not Jev figuring out the map. The harness computes shortest path distances through its region graph and literally annotates the choices with that answer.Once Jev says "Enter X" the harness already has the exact A* path and walks it automatically, including replanning around moving NPCs and retrying when blocked.
Puzzle solving is also done by the harness - it gives Jev a set of choices, annotated with sentences like "brings the boulder closer to the floor switch", "after this push no sequence of pushes can bring this boulder onto a floor switch anymore".
Story progression contains hardcoded instructions for Jev, e.g.: 'Go south from Cerulean through Route 5, the Underground Path and Route 6 to Vermilion City; board the S.S. Anne at the dock and get HM01 (Cut) from the captain.'
And even with all of those hardcoded decisions, it still plays terrible. I just saw it get stuck in a loop - but of course harness annotates repeated actions, so eventually Jev is steered to escape the loop.
You should be ashamed of yourself.
This seems like a technology heading in the right direction but not quiet there yet. Excited for what they are cooking up but probably won't start building around it yet.