LLM Visual Bench
A small, honest benchmark for one-shot visual code from local open-weight models.
In progress This project is not ready yet. Details below describe where it is heading and may change.
Overview
Built for 4-bit quantisations of 27B to 35B models on a laptop. It grades self-contained HTML files (animations, playable games, animated marketing pages) automatically in a real browser and produces a scorecard, a percentage per category, a methodology write-up per run, and a leaderboard. Three prompt tiers, the smallest short enough that a 30B Q4 model can finish it, because a benchmark where everything scores 2/10 ranks nothing.
Like this?
Message me on X ↗