Chingie, the Studio Chingie mascot

Studio Chingie

Platformer Slice Benchmark

Complete benchmark

A browser platformer and tool-use benchmark for comparing how local AI models build, revise, and safely execute bounded tasks.

AI benchmarkBrowser prototypePlatformer sliceLocal model test
Contact sheet comparing platformer-slice benchmark outputs from several AI modelsCurrent build
Final outputs from the benchmark runs, including Laguna S 2.1's failed three-shot artifact.

Details

Screenshots
7 images
Source access
Private

Stack

HTML CanvasPlaywrightTool-EvalLocal modelsBenchmark reports

Source

platformer-slice-world1 benchmark reportsPrivate | Markdown / HTML | Updated 2026-07-21Local benchmark reports and generated browser platformer slices.

Benchmark

Benchmark results

Scores and generation speed for each run.

Test hardware

All listed runs were completed on the same machine.

Machine
MacBook Pro
Chip
Apple M5 Max
Memory
128 GB unified memory
ModelFinalShot 1Shot 2Shot 3Tokens/secWall time
ds4-100k-nothink79/10072727918.6638.2s
qwen27-mtp-fast75/10073737518.3347.7s
qwen122-q4xl-vision-64k-think75/10067687526.7271.4s
step37-unsloth-iq4xs-text-mtp2-r204875/10066747528.1524.3s
qwen35-a3b-no-think74/10047677451.6115.6s
qwen122-q4xl-vision-64k74/10071717426.3296.5s
nex-n2-mini-q8-vision-64k74/10064637455.3222.7s
step37-unsloth-iq4xs-vision-r204871/10062687123.4522.6s
laguna-s21-q4km-64k-think39/10032343924.02833.6s

Playable Builds

Try a benchmark build

Playable final builds from each benchmark run.

Loading playable build...

Screenshots

Contact sheet comparing several AI-generated browser platformer benchmark outputs
Current build

Final outputs from local-model runs, including Laguna's blank canvas after three capped attempts.

Bar chart comparing full Tool-Eval scores for Qwen35, Laguna S 2.1, Qwen27, Step 3.7, and DS4
Benchmark chart

Laguna scored 88 across the full 69-scenario Tool-Eval run: ahead of Qwen27, Step, and DS4, but behind Qwen35.

Bar chart comparing Laguna plain decoding and DFlash decoding speed
Benchmark chart

Plain 32K decoding averaged 37.2 tokens per second; DFlash was prompt-dependent and slower on the representative 64K workflow set.

Bar chart comparing three-shot platformer benchmark scores with Laguna S 2.1
Benchmark chart

Laguna's long-form artifact failed despite three high-cap attempts, finishing at 39/100.

Blank Laguna S 2.1 platformer canvas with only a small HUD visible
Failed artifact

Laguna's final browser capture: the page loaded, but repeated empty tile data left the game canvas blank and non-interactive.

DS4 model browser platformer benchmark capture after playtest input
Current build

Playtest capture from the top-scoring DS4 run after automated movement input.

Nex model browser platformer benchmark capture after playtest input
Current build

Playtest capture from the Nex run after automated movement input.

Overview

Each model received the same platformer brief and three chances to build, test, and revise a playable browser level. The July Laguna run also used the full 69-scenario Tool-Eval suite and representative Unity safety artifacts.

Now

The July 2026 Laguna S 2.1 run is complete. It was strong at structured text and tools, but failed the long-form game artifact and did not replace Qwen35 or Qwen27 in the local workflow.

Why I'm making it

A playable test exposes differences that a chat transcript cannot: movement, layout, bugs, and polish all have to work on screen.