Show HN: Jevman – AI decision models play Pac-Man

opper.ai

76 points by felix089 2 days ago

Openai just launched their decisions endpoint, cloudflare launched clef the other week, and many more jev alternatives are out there.

We wanted to put the popular ones to the test and thought Pac-Man is a good benchmark for simple and fast decision making.

So we let jev 1.13, kev, clef, clef flash, GPT-6 Luna and Laya play Pac-Man against bot ghosts.

The low latency of these models allows for real time play. We had each model play 100 games, published a leader board and open-sourced the repo so anyone can run their own model and join the ranking. Link to repo: https://github.com/opper-ai/jevman-benchmark/blob/main/CONTR...

You can also join the game and play as Pac-Man yourself, and the ghosts are the models, either a mix of models or all jev, kev, clef etc. A game costs about 2 cent, all models are running via my startup opper, and we added free credits for everyone to try.

It's pretty fun to play and surprisingly difficult to beat jev's highscore. Any feedback is more than welcome!

hung 1 day ago

Shouldn't it be running at every tick of the game rather than just the junctions? Pac-Man should be able to change direction at any point.

  • hyperhello 1 day ago

    It may not matter if the ghosts' next move is predictable over that maximum distance.

    • theowaway213456 1 day ago

      That requires much more thinking (because of calculation) though doesn't it? Which decision models are not optimized for.

      • hyperhello 1 day ago

        Speculating, but it seems like needing to be able to change direction at any moment would be more calculating, not less.

  • felix089 1 day ago

    Ideally yes but that would not work with the model latency. But it is already partly covered when a ghost enters its corridor it is asked again. Here's how the decisions work in detail, also below a copy https://github.com/opper-ai/jevman-benchmark/blob/main/READM...

    How decisions work

    When a character commits to a corridor its next junction is known, so the game asks jev about it straight away (one System One request per frame, one choice question per character, options = legal directions described with computed facts: distances to pellets, power pellets, fruit and ghosts, whether the nearest ghost is coming closer, and whether a ghost can reach the end of the corridor before Pac-Man). Up to three requests are in flight at once. If the character reaches the junction before the answer, it waits there (its panel card says "thinking…"). After 2 s, or on an error, it uses a greedy rule.

    Pac-Man can also get a second question mid-corridor: when a dangerous ghost is in the corridor ahead, or can reach the junction at its end before he does, jev is asked whether to keep going or turn back right now (pacman_escape). Pac-Man keeps moving while it is open, and each situation is asked once.

NichoPaolucci 1 day ago

This is neat. At least I'm still better than the models at pac-man.

Also love the idea of a shared pool for users to try things out. I was considering more of a crowdfunded approach for one of my toy projects, something like... Giving it $10 in credits to begin and somehow allowing users to feed a buck in if they wanted.

felix089 1 day ago

Quick update, we doubled the models tested based on community contributions, and a new model took the #1 spot. also well over 1000 games have been played against decision model ghosts. https://opper.ai/jevman-benchmark/#leaderboard

  • gsandahl 1 day ago

    i guess a model trained on this game will end up taking the #1 at some point

nico 1 day ago

Very cool, are you also testing local, cpu-runnable/trainable models/classifiers?

I trained some to do some interesting things, including playing doom: https://github.com/nicobrenner/jeffy

I’ll try training one for this benchmark, seems like fun

ozozozd 1 day ago

Super cool!

I wish the controls were a little easier on mobile.

  • felix089 1 day ago

    Thank you and yea agreed fixing it now, try again in an hour from now :)

bayarearefugee 1 day ago

Paying 2 cents per game to avoid playing it.

What even is this reality.

  • gsandahl 1 day ago

    haha thats one way to look at it :)

ismailmaj 1 day ago

it would be nice to have a regular LLM for reference

  • felix089 23 hours ago

    Someone actually ran the benchmark on Qwen 3.8 Flash Next NVFP4, which is a chat model and submitted it via PR. We merged and it's live on the leader board in the top 3.