Local inference · real measurements
Which GPU for your local LLMs ?
Tokens/s, actual VRAM usage, real power draw: every card is measured with the same llama.cpp protocol. Numbers you can finally compare, before you buy.
3
GPUs measured
21
Verified runs
6
Models tested
Our reference model: Meta Llama 3.1 8B Instruct Q4_K_M — tested by every capable card, it's what makes the leaderboard comparable.
Generation leaderboard
Average tokens/s · select one or more models
| GPU | VRAM | Generation | Prompt | Avg power | VRAM used | Runs | Price |
|---|---|---|---|---|---|---|---|
| 1RTX 5070 Ti | 16 Go | 151 tok/s | 6 818 tok/s | 161 W | 5,4 Go | 1 | 1 400 € (24/09/26) |
| 2RTX 4090demo | 24 Go | 132 tok/s | 5 900 tok/s | 285 W | 6,2 Go | 1 | |
| 3RTX 5080demo | 16 Go | 124 tok/s | 5 100 tok/s | 250 W | 6,0 Go | 1 | 1 850 € (24/09/26) |
| 4RTX 4080demo | 16 Go | 103 tok/s | 4 700 tok/s | 230 W | 6,1 Go | 1 | 780 € (04/08/26) |
| 5RTX 4070 Ti | 12 Go | 89,6 tok/s | 5 960 tok/s | 158 W | 5,6 Go | 2 | 927 € (04/08/26) |
| 6RTX 4070 | 12 Go | 82,3 tok/s | 4 336 tok/s | 138 W | 6,8 Go | 1 | 595 € (04/08/26) |
| 7RTX 3060demo | 12 Go | 48,0 tok/s | 1 500 tok/s | 155 W | 5,9 Go | 1 | 260 € (04/08/26) |
Generation = how fast the answer is written (the metric you actually feel). A "—" = card not tested on this model (often: it doesn't fit in its VRAM). “Buy” = links to retailers (LDLC, Materiel.net, TopAchat, Cybertek, Grosbill, Rue du Commerce, Amazon).
90 tok/s — what does that actually feel like? Put cards side by side and watch the same answer being written at their real speed.One protocol, three steps
No manual input: the CLI runs the benchmark and collects telemetry itself. Everything in the database was measured, never declared.
Download
llama.cpp (official binary) and the reference GGUF model.
Run the CLI
One command: it measures tokens/s, VRAM, temperatures and power draw.
Publish
The verified run joins the database and appears in the leaderboard.