An experiment exploring whether local AI decision models running through Ollama can control the Chrome Dino game in real time.
Watch the AI controlling Chrome Dino:
dino.game.with.jev.style.model.mp4
The project combines:
- 🖥️ Chrome Dino
- 📸 PyAutoGUI for screen capture
- 👁️ OpenCV for obstacle detection
- 🧠 Ollama decision models
- 🤖 Nimble / Tev1
- ⌨️ PyAutoGUI for keyboard automation
- 🐍 Python
The main goal of this project is not simply to make the Dino jump.
The goal is to investigate an important question:
Can a local AI decision model make decisions quickly enough to control a real-time environment?
The initial idea was simple:
Chrome Dino
↓
Screenshot
↓
OpenCV
↓
Extract game state
↓
AI Decision Model
↓
JUMP / DUCK / WAIT
↓
Keyboard
The AI receives a structured representation of the game state instead of the raw screenshot.
For example:
{
"game": "Chrome Dino",
"dino": {
"on_ground": true,
"jumping": false
},
"obstacle": {
"type": "cactus",
"distance": 288,
"height": 58,
"width": 28
}
}The model then selects an action:
JUMP
DUCK
WAIT
Chrome Dino
│
▼
PyAutoGUI Screenshot
│
▼
OpenCV
│
▼
Obstacle Detection
│
▼
Game State
│
▼
Ollama Decision Model
│
┌──────┼──────┐
▼ ▼ ▼
JUMP DUCK WAIT
│
▼
PyAutoGUI
│
▼
Chrome Dino
The first model tested was Nimble, a local decision model available through Ollama.
The reasoning behind choosing Nimble was straightforward:
If the task is decision making rather than text generation, a specialized decision model should theoretically be a good fit.
The model was given a structured game state and asked to choose between:
jump
duck
wait
The measured request was:
HTTP status: 200
Response time: 17636 ms
Approximately:
17.6 seconds
Nimble correctly returned:
{
"choice": "jump"
}However, the response arrived far too late for the game.
Chrome Dino is a real-time environment.
Suppose the obstacle is detected at:
distance = 288 px
The controller sends the state to Nimble:
288 px → Nimble
But the model takes approximately:
17.6 seconds
to respond.
By that time:
Dino → Cactus → 💀
The model may still return:
JUMP
but the decision is no longer useful.
This demonstrated an important principle:
A correct decision that arrives too late is effectively an incorrect decision in a real-time system.
The next question was:
Is there a smaller/faster decision model that can perform the same task?
This led to testing Tev1.
The goal was not simply to find a model with better reasoning.
The goal was to find a model with a better latency-to-decision tradeoff.
The first Tev1 model tested was:
tev1:0.8b
The model was running entirely on the RTX 5050 GPU:
PROCESSOR: 100% GPU
Measured response:
2525 ms
After simplifying the input:
2347 ms
So the approximate latency was:
~2.35 seconds
Compared with Nimble:
Nimble: ~17.6 seconds
Tev1 0.8B: ~2.35 seconds
That is roughly a 7.5× reduction in measured response time in this particular setup.
However, 2.35 seconds was still too slow for a real-time reflex controller.
Latency wasn't the only problem.
We also tested whether Tev1 changed its decision according to obstacle distance.
Test states:
distance = 400
distance = 300
distance = 250
distance = 200
distance = 150
distance = 100
The results were:
400 → JUMP
300 → JUMP
250 → JUMP
200 → JUMP
150 → JUMP
100 → JUMP
The model's confidence remained low:
~0.12–0.14
This suggested that simply providing distance wasn't enough for the model to learn the game's timing relationship.
We also tested:
tev1:4b-q4_K_M
The measured response time was:
73224 ms
Approximately:
73.2 seconds
The model returned:
JUMP
but the response time made it unusable for this application.
This was an important reminder:
A larger model is not automatically a better model for a specific real-time task.
These are measurements from this project, not universal benchmarks.
Hardware and software configuration can significantly affect inference performance.
| Model | Approx. Runtime Size | Measured Latency | Suitable for Real-Time Reflex? |
|---|---|---|---|
| Nimble 9B | ~10 GB runtime | 17.64 s | ❌ |
| Tev1 0.8B | ~893 MB | 2.35 s | ❌ |
| Tev1 4B Q4_K_M | ~2.7 GB | 73.22 s | ❌ |
CPU: Intel Core 7 240H
GPU: NVIDIA GeForce RTX 5050
VRAM: 8 GB
RAM: 24 GB DDR5
Runtime: Ollama
OS: Windows
The decision to switch from Nimble to Tev1 was based on application requirements, not simply model popularity or parameter count.
The following conditions were considered.
For a real-time game:
Decision latency << reaction window
If:
model latency > available reaction time
the model should not be placed directly in the control loop.
For example:
Obstacle detected
↓
Available reaction time: ~500 ms
↓
Model latency: 2.3 seconds
↓
Model is too slow
A model should not only respond quickly.
It should also produce useful decisions.
For example:
distance = 400 → WAIT
distance = 200 → JUMP
is more useful than:
distance = 400 → JUMP
distance = 300 → JUMP
distance = 200 → JUMP
distance = 100 → JUMP
A model that produces the same action regardless of meaningful state changes may not be suitable for the task.
Local inference depends heavily on:
- GPU VRAM
- GPU compute performance
- CPU performance
- RAM
- model size
- quantization
- context length
- runtime configuration
For example, this machine has:
RTX 5050
8 GB VRAM
Tev1 0.8B was able to run:
100% GPU
while the larger Nimble model was split between CPU and GPU.
Therefore, model size and hardware compatibility are important considerations when selecting a local model.
A language/decision model is useful when the task actually requires reasoning.
Chrome Dino is different.
The decision space is extremely small:
JUMP
DUCK
WAIT
The environment can be represented using only a few numerical values:
distance
height
width
speed
dino position
This means using a multi-billion-parameter language model may be unnecessary for the final controller.
The biggest lesson from this experiment was:
AI capability and AI usability are different things.
A model can make the correct decision but still fail at the application if the decision arrives too late.
For real-time agents, the important metric isn't just:
"What answer did the model give?"
It is:
"How quickly did the model give a useful answer?"
The project currently has a working fast controller based on OpenCV.
The simplified logic is:
if distance <= JUMP_DISTANCE:
jump()This reacts much faster than the language-model-based controller.
The experiment demonstrated that deterministic logic is currently much better suited to the critical reflex loop.
Instead of:
OpenCV
↓
LLM
↓
Keyboard
the next architecture will be:
Chrome Dino
↓
OpenCV
↓
Game State
↓
Fast Controller
↓
JUMP
with AI operating at a higher level:
Tev1
↓
Strategy / Analysis
↓
Fast Controller
↓
JUMP
This creates a hybrid architecture:
┌──────────────────────────────┐
│ AI / Strategy Layer │
│ │
│ Tev1 / LLM │
│ │
└──────────────┬───────────────┘
│
│ high-level decisions
▼
┌──────────────────────────────┐
│ Real-Time Layer │
│ │
│ OpenCV + Fast Controller │
│ │
└──────────────┬───────────────┘
│
▼
PyAutoGUI
│
▼
Chrome Dino
The next goal is to build a tiny specialized AI controller rather than using a general-purpose language model.
Possible input:
distance
obstacle width
obstacle height
obstacle velocity
Dino Y position
Dino jump state
Output:
WAIT
JUMP
DUCK
A tiny neural network could potentially perform this classification much faster than a 0.8B language model.
The experiment will then compare:
Rule-based Controller
vs
Tiny Neural Network
vs
Tev1 0.8B
vs
Nimble 9B
using metrics such as:
- Decision latency
- Survival time
- Number of obstacles cleared
- Decision accuracy
- CPU/GPU utilization
This project started with a simple question:
Can a local AI decision model play Chrome Dino?
The answer is:
Yes, technically — but model latency becomes the limiting factor.
Nimble produced valid decisions but was too slow.
Tev1 0.8B was substantially faster but still too slow for the critical reflex loop.
Tev1 4B Q4_K_M was significantly slower in this particular setup.
The experiment therefore shifted from:
"Which LLM should control the Dino?"
to:
"What is the smallest and fastest intelligent controller that can make useful real-time decisions?"
And that's where the next phase of the project begins. 🚀
dino-game-with-nimble/
│
├── detect.py
│
├── dino_controller.py
│
├── nimble_agent.py
│
├── nimble_connection_test.py
│
├── tev1_connection_test.py
│
├── tev1_fast_test.py
│
├── tev1_distance_test.py
│
├── tev1_4b_test.py
│
├── requirements.txt
│
└── README.md
- Python
- OpenCV
- PyAutoGUI
- Requests
- Ollama
- Nimble
- Tev1
- Chrome Dino
- NVIDIA CUDA/GPU acceleration
All latency numbers in this README were measured during this experiment on the listed hardware and software environment.
They should not be interpreted as official benchmarks for Nimble or Tev1.
Actual performance can vary depending on:
- Hardware
- GPU VRAM
- CPU
- Ollama version
- Model quantization
- Context size
- Background processes
- Drivers
- Operating system
- Model/runtime configuration
The purpose of the measurements is to compare the models within this particular experiment and environment.