An SVG repair benchmark

Vector-Bench

Can models surgically edit SVG code?

Natural-language repairs, hidden structural targets, and a binary specification reward that rejects collateral changes.

Results

Leaderboard

A repair can be approximate. Its side effects cannot. 34 model endpoints receive one scored outcome per task with no fallback routing. A full pass requires every requested repair within tolerance, a valid SVG, and semantic preservation outside the requested fields.

Updated 2026-07-20Evaluator semantic-perceptual-binary-2026-07-21Corpus 2a62410b5de7
Submit your run
top ten / specification gateshigher is better
full table34 entries
#SolverProviderFull passNearProgressCleanSourceValid UCRValidTargetTruncatedErrorsCostTasksDate
#1
Claude Sonnet 5
Anthropic15.0%2.5%43.7%42.5%42.5%0.8%62.5%0.0%40.0%0.0%$2.032402026-07-20
2
KAT Coder Air V2.5
KwaiPilot12.5%0.0%27.3%25.0%25.0%0.3%42.5%0.0%57.5%0.0%$0.128402026-07-20
3
MiniMax M3
MiniMax10.0%0.0%33.0%25.0%25.0%1.0%45.0%0.0%55.0%0.0%$0.240402026-07-20
4
HY 3
Tencent7.5%2.5%56.8%37.5%37.5%1.8%100.0%0.0%0.0%0.0%$0.092402026-07-20
5
Gemma 4 31B IT
Google5.0%5.0%54.8%47.5%47.5%0.8%87.5%2.5%5.0%0.0%$0.084402026-07-20
6
Qwen3.5 397B A17B
Qwen5.0%0.0%39.8%32.5%32.5%1.4%65.0%0.0%32.5%0.0%$0.618402026-07-20
7
DeepSeek V4 Pro
DeepSeek5.0%2.5%30.6%20.0%20.0%1.2%47.5%0.0%52.5%0.0%$0.508402026-07-20
8
Kimi K2.6
Moonshot AI5.0%5.0%16.0%15.0%15.0%0.3%25.0%0.0%72.5%2.5%$1.077402026-07-20
9
Qwen3 Coder
Qwen2.5%0.0%54.7%45.0%45.0%1.4%100.0%0.0%0.0%0.0%$0.295402026-07-20
10
Gemini 3.1 Flash Lite
Google2.5%0.0%54.5%25.0%25.0%2.5%100.0%0.0%0.0%0.0%$0.269402026-07-20
11
Devstral 2512
Mistral2.5%5.0%46.0%57.5%57.5%0.6%97.5%0.0%0.0%0.0%$0.358402026-07-20
12
Inkling
Thinking Machines2.5%0.0%26.6%27.5%27.5%0.8%45.0%0.0%55.0%0.0%$0.806402026-07-20
13
GPT-5
OpenAI2.5%0.0%9.1%7.5%7.5%0.3%15.0%0.0%85.0%0.0%$1.945402026-07-20
14
Nemotron 3 Ultra
NVIDIA2.5%0.0%4.9%5.0%5.0%0.5%7.5%0.0%92.5%0.0%$0.622402026-07-20
15
Claude Haiku 4.5
Anthropic0.0%0.0%54.0%27.5%27.5%2.0%100.0%0.0%0.0%0.0%$0.789402026-07-20
16
Seed 2.0 Mini
ByteDance0.0%2.5%51.2%35.0%35.0%2.0%92.5%0.0%2.5%0.0%$0.174402026-07-20
17
GPT-OSS 120B
OpenAI0.0%5.0%49.9%42.5%42.5%1.1%95.0%0.0%5.0%0.0%$0.057402026-07-20
18
Llama 4 Maverick
Meta0.0%0.0%46.9%20.0%20.0%2.5%92.5%0.0%0.0%0.0%$0.143402026-07-20
19
Mistral Small 2603
Mistral0.0%0.0%39.0%5.0%5.0%4.7%95.0%0.0%0.0%0.0%$0.119402026-07-20
20
GPT-5.4 Nano
OpenAI0.0%0.0%38.0%22.5%22.5%2.2%97.5%0.0%0.0%0.0%$0.177402026-07-20
21
GPT-OSS 20B
OpenAI0.0%2.5%36.5%27.5%27.5%1.6%72.5%0.0%27.5%0.0%$0.030402026-07-20
22
Phi-4
Microsoft0.0%0.0%30.7%12.5%12.5%3.8%92.5%0.0%0.0%0.0%$0.026402026-07-20
23
Laguna M.1
Poolside0.0%2.5%23.3%25.0%22.5%1.4%42.5%0.0%55.0%0.0%$0.101402026-07-20
24
Ling 2.6 Flash
InclusionAI0.0%0.0%19.4%15.0%15.0%4.8%87.5%0.0%5.0%0.0%$0.006402026-07-20
25
Granite 4.1 8B
IBM0.0%0.0%18.5%10.0%10.0%14.3%77.5%0.0%2.5%0.0%$0.019402026-07-20
26
GLM 4.7
Z.ai0.0%0.0%15.1%12.5%12.5%0.9%30.0%0.0%70.0%0.0%$0.549402026-07-20
27
MiMo V2.5 Pro
Xiaomi0.0%0.0%6.4%5.0%5.0%0.8%10.0%0.0%87.5%0.0%$0.384402026-07-20
28
Qwen3.6 35B A3B
Qwen0.0%0.0%4.9%2.5%2.5%1.4%7.5%0.0%92.5%0.0%$0.236402026-07-20
29
Trinity Large Thinking
Arcee AI0.0%0.0%1.8%0.0%0.0%0.0%2.5%0.0%97.5%0.0%$0.176402026-07-20
30
Gemini 3.1 Pro Preview
Google0.0%0.0%0.0%0.0%0.0%n/a0.0%0.0%100.0%0.0%$2.494402026-07-20
31
Nemotron 3 Nano
NVIDIA0.0%0.0%0.0%0.0%0.0%n/a0.0%0.0%80.0%0.0%$0.051402026-07-20
32
OLMo 3 32B Think
AI20.0%0.0%0.0%0.0%0.0%n/a0.0%0.0%0.0%100.0%$0.000402026-07-20
33
Reka Flash 3
Reka0.0%0.0%0.0%0.0%0.0%n/a0.0%0.0%12.5%0.0%$0.040402026-07-20
34
Step 3.7 Flash
StepFun0.0%0.0%0.0%0.0%0.0%n/a0.0%0.0%100.0%0.0%$0.234402026-07-20

paper analysis

The binary score separates four distinct outcomes.

Full paper
Mutually exclusive decomposition of benchmark outcomes by evaluation gate
Each output is assigned once: full pass, completed repair with side effects, valid incomplete repair, or invalid artifact.
Scatter plot of repair progress against valid-output unintended change rate
Repair progress and collateral change on valid outputs are measured independently.
Scatter plot of model repair progress against benchmark run cost
Provider cost does not fully predict repair quality.

Team

Shared research, shared material.

Vector-Bench and its paper were created by Yug Gupta and Prannay Hebbar.