vtikz-eval
LLM Code Customization with Visual Results: A Benchmark on TikZ — Reux et al. (2025) (arXiv:2505.04670, 2025)
What this evaluates
Evaluates Large Language Models' ability to customize and edit TikZ code based on a specified visual intent. The benchmark probes three core capabilities: locating relevant code features, synthesizing correct code variants, and validating that the modified code produces the intended visual output.
Datasets
- vTikZ — total 100; splits: test (100); repo https://github.com/IV2C/VTikZ
Metrics
visual result validation(primary) — range: [0, 1]- The proportion of tasks where the model's generated TikZ code, when rendered, correctly matches the target visual intent. Ground truths are parameterized to accept multiple syntactically valid code solutions rather than requiring exact string matches.
Input / output format
Input: Original TikZ code snippet paired with a description of the desired visual modification or target output.
Output: Modified TikZ code.
Scoring recipe
def compute_metric(predictions, golds):
correct = 0
for pred, gold in zip(predictions, golds):
pred_img = render_tikz(pred)
gold_img = render_tikz(gold)
if is_visual_match(pred_img, gold_img):
correct += 1
return correct / len(predictions)
Common pitfalls
- Exact string matching fails because the benchmark uses parameterized ground truths that account for multiple valid code solutions.
- Visual validation depends on the rendering environment; differences in TikZ compilers or backend renderers may cause false negatives if not standardized.
Evidence (verbatim from paper)
It addresses three core challenges—feature location, code variant synthesis, and visual result validation—through 100 manually curated tasks with parameterized ground truths that account for multiple valid solutions.
Citation
@misc{reux2025vtikz,
title={LLM Code Customization with Visual Results: A Benchmark on TikZ},
author={Reux et al. (2025)},
year={2025},
note={arXiv:2505.04670}
}
- arXiv: 2505.04670