Platformer Slice Benchmark
Complete benchmark
A browser platformer and tool-use benchmark for comparing how local AI models build, revise, and safely execute bounded tasks.
Current buildDetails
- Screenshots
- 7 images
- Source access
- Private
Stack
Source
Benchmark
Benchmark results
Scores and generation speed for each run.
Test hardware
All listed runs were completed on the same machine.
- Machine
- MacBook Pro
- Chip
- Apple M5 Max
- Memory
- 128 GB unified memory
| Model | Final | Shot 1 | Shot 2 | Shot 3 | Tokens/sec | Wall time |
|---|---|---|---|---|---|---|
| ds4-100k-nothink | 79/100 | 72 | 72 | 79 | 18.6 | 638.2s |
| qwen27-mtp-fast | 75/100 | 73 | 73 | 75 | 18.3 | 347.7s |
| qwen122-q4xl-vision-64k-think | 75/100 | 67 | 68 | 75 | 26.7 | 271.4s |
| step37-unsloth-iq4xs-text-mtp2-r2048 | 75/100 | 66 | 74 | 75 | 28.1 | 524.3s |
| qwen35-a3b-no-think | 74/100 | 47 | 67 | 74 | 51.6 | 115.6s |
| qwen122-q4xl-vision-64k | 74/100 | 71 | 71 | 74 | 26.3 | 296.5s |
| nex-n2-mini-q8-vision-64k | 74/100 | 64 | 63 | 74 | 55.3 | 222.7s |
| step37-unsloth-iq4xs-vision-r2048 | 71/100 | 62 | 68 | 71 | 23.4 | 522.6s |
| laguna-s21-q4km-64k-think | 39/100 | 32 | 34 | 39 | 24.0 | 2833.6s |
Playable Builds
Try a benchmark build
Playable final builds from each benchmark run.
Screenshots

Final outputs from local-model runs, including Laguna's blank canvas after three capped attempts.

Laguna scored 88 across the full 69-scenario Tool-Eval run: ahead of Qwen27, Step, and DS4, but behind Qwen35.

Plain 32K decoding averaged 37.2 tokens per second; DFlash was prompt-dependent and slower on the representative 64K workflow set.

Laguna's long-form artifact failed despite three high-cap attempts, finishing at 39/100.

Laguna's final browser capture: the page loaded, but repeated empty tile data left the game canvas blank and non-interactive.

Playtest capture from the top-scoring DS4 run after automated movement input.

Playtest capture from the Nex run after automated movement input.
Overview
Each model received the same platformer brief and three chances to build, test, and revise a playable browser level. The July Laguna run also used the full 69-scenario Tool-Eval suite and representative Unity safety artifacts.
Now
The July 2026 Laguna S 2.1 run is complete. It was strong at structured text and tools, but failed the long-form game artifact and did not replace Qwen35 or Qwen27 in the local workflow.
Why I'm making it
A playable test exposes differences that a chat transcript cannot: movement, layout, bugs, and polish all have to work on screen.