Open benchmark · local models + subscription · no paid API

Which AI draws the best capybara on a bicycle?

We asked 19 models to draw the same thing in SVG code: 10 running on a computer and 9 through Claude and ChatGPT subscriptions. A judge compared them head to head without knowing who made what. The kit is open: run it with your own models.

Capybara Arena: AIs draw a capybara on a bicycle
What it is

A silly test that reveals a lot

Simon Willison evaluates each new model by asking it to create “an SVG of a pelican riding a bicycle.” He doesn’t trust leaderboards; he prefers his own test, one he understands himself. Google already showed his pelican in a keynote, so we changed the animal.

Same prompt, blind judge, both orders, ranking, local models, build your own

🦫 Difficult by design

A text model doesn’t draw: it writes code. Bicycles are hard to recall, capybaras are hard to draw, and capybaras don’t pedal. You can immediately see who understands shape and composition.

⚖️ Blind judge, both orders

The judge sees only “A” and “B,” with no names. Each pair appears twice, with the sides swapped, to cancel out the tendency to prefer the left. A second judge from another company rechecks a sample.

💻 Runs at home

Local models run through Ollama; Claude and GPT use the claude and codex CLIs with the subscription you already pay for. No API key required.

Round 1 · 05/10/2026

The podium

Claude Fable 5.1 won 97% of matchups. The surprise: Qwen 3.8 27B, running on a computer, came in 5th out of 18, ahead of GPT-5.5 and Claude Sonnet 5.5. Command R 35B did not produce valid SVG. The second judge (GPT-6 Luna) agreed with the first in 92% of the 64 matchups it rechecked.

Claude Fable 5.1
1st Claude Fable 5.1 · 97% wins
GPT-6.1 Sol
2nd GPT-6.1 Sol · 94% wins
GPT-6 Astra
3rd GPT-6 Astra · 91% wins
First-place model versus last-place model
First versus last. The judge: “A shows a recognizable capybara pedaling a well-drawn bicycle in a complete scene, while B is just a collection of abstract shapes with overlaid text, with no identifiable capybara or bicycle.”
All the drawings
All valid drawings, from 1st to last. Full ranking, times, and code for each SVG →
How it works

Five scripts, one per step

Each step saves its output to resultados/ and skips work that is already done. You can stop and continue later.

prompt.txt→ gerar.py→ renderizar.py→ julgar.py→ ranking.py→ galeria.py

Generate

Sends the same prompt to each model in modelos.json, saves the raw response, and extracts the <svg>. Local models run one at a time to fit in memory.

Render and judge

Converts SVG to PNG without fetching anything from the internet, arranges images as A|B, and asks the judge for JSON with the winner and the reason.

Ranking and gallery

Win rate (the most honest measure), average ELO from 200 randomized orderings, side bias, and agreement between judges. Generates the public page in PT, EN, and ES.

Prerequisites

What needs to be installed

Only the first item is required. Use the engines you have; remove the others from modelos.json.

Python 3.10+

With cairosvg (SVG to PNG) and Pillow.

pip install cairosvg pillow

Ollama (local)

Any model you’ve already downloaded. An 8B runs on a laptop; a 70B needs around 48 GB of memory.

ollama pull llama3.1:8b
ollama list

Claude Code and/or Codex

Optional. Log in with your subscription; the kit calls claude -p and codex exec.

claude --version
codex --version
User guide · step by step

Your arena in six commands

Start small: two or three models. Then add more.

1

Download the kit

The results from our round are included in resultados/. Delete the folder if you want to start from scratch.

git clone https://github.com/inematds/arena-capivara
cd arena-capivara
rm -rf resultados && mkdir -p resultados/{svgs,respostas,png,pares}  # optional: clean arena
2

Choose your competitors

Edit arena/modelos.json. Each line has a motor (ollama, claude, or codex) and the modelo name, just as it appears in ollama list or the CLI.

{"id": "llama3-1-8b", "motor": "ollama", "modelo": "llama3.1:8b",
 "rotulo": "Llama 3.1 8B", "familia": "Local"}
3

Generate the drawings

Claude and Codex run in parallel; local models run one at a time. If a model does not return SVG, it is recorded as “no output,” which is also a result.

python3 arena/gerar.py  # or --so llama3-1-8b claude-haiku-4-5
4

Render and set up the matchups

python3 arena/renderizar.py  # "N valid images, M matchups created"
5

Call the judge (and a second judge)

The default is Claude Sonnet 5.5. The second judge rechecks a sample so you can see how much to trust the first.

python3 arena/julgar.py --paralelo 4
python3 arena/julgar.py --juiz codex --modelo gpt-6-luna --amostra 60
6

View the ranking and publish

python3 arena/ranking.py --tabela  # table in terminal + resultados/ranking.json
python3 arena/galeria.py          # resultados/index.html (+ en/ es/)
python3 -m http.server 8000      # open http://localhost:8000/resultados/
Build your own benchmark

The capybara is just an example. The method works for your own work.

The lesson from the talk isn’t “model X draws better.” It’s this: don’t outsource your model choice to a leaderboard. Build a small test you understand and run it whenever a new model comes out.

1
Use a task from your own workSummarize a contract, reply to a customer, write SQL. Or pick something silly but difficult, like the capybara, that models were never trained to get right.
2
Use the same prompt for everyone and save it to a fileIf the prompt changes between models, you’re measuring the prompt, not the model. Here, it’s arena/prompt.txt.
3
Keep the raw responseWhen the result surprises you, you’ll want to see what the model actually wrote (resultados/respostas/).
4
Compare pairs instead of assigning scores“A or B?” is much more consistent than “give it a score from 0 to 10,” for people and models alike.
5
Use a blind judge in both orders and measure biasHide the names, swap the sides, and report how often the left side won. Use a second judge and report their agreement.
6
One round is a snapshotRun it again and the drawings will change. To make a serious decision, run each model 3 to 5 times.
7
Change the subject when it gets famousA test that makes the news ends up in the training data. The pelican became a capybara; the capybara will have to change too.

Adapt it to another task

Change the text in arena/prompt.txt and the judge’s question (PERGUNTA in arena/julgar.py). If the output is text instead of a drawing, skip renderizar.py and have the judge read both responses. To use Claude Code’s own history as a test, see personal-benchmark.

Upcoming rounds

What comes next

The arena is open. Submit your round as a pull request.

Done
Round 1 · 05/10/202619 models, one attempt each, Claude Sonnet 5.5 judge + sample reviewed with GPT-6 Luna.
Next
Three attempts per modelTo separate good models from lucky drawings.
Idea
Community roundsEach person runs the benchmark on their own machine with the models they have; the ranking combines the results.
Sibling
Lethal TrifectaFrom the same talk: when an AI agent can leak your data. View the checklist