← All work
In progress · not ready yetLabNode + Playwright2026

LLM Visual Bench

A small, honest benchmark for one-shot visual code from local open-weight models.

In progress This project is not ready yet. Details below describe where it is heading and may change.
Overview

Built for 4-bit quantisations of 27B to 35B models on a laptop. It grades self-contained HTML files (animations, playable games, animated marketing pages) automatically in a real browser and produces a scorecard, a percentage per category, a methodology write-up per run, and a leaderboard. Three prompt tiers, the smallest short enough that a 30B Q4 model can finish it, because a benchmark where everything scores 2/10 ranks nothing.

Like this?

Let's build yours.

Message me on X ↗