geneval-eval
Nucleus-Image: Sparse MoE for Image Generation — Akiti et al. (2026) (arXiv:2604.12163, 2026)
What this evaluates
Evaluates text-to-image alignment by testing whether models correctly render specified objects, attributes, and spatial relations across six categories: single object, two objects, counting, colors, position, and attribute binding.
Datasets
- GenEval — total ?; splits: test (-1)
Metrics
Overall(primary) — range: [0, 1]- Mean accuracy across six categories: single object, two objects, counting, colors, position, and attribute binding.
Input / output format
Input: Text prompt specifying objects, attributes, and spatial relations to be rendered.
Output: Generated image at 1024x1024 resolution, 50 inference steps, CFG scale 8.0.
Scoring recipe
category_scores = {cat: [] for cat in ['single', 'two', 'count', 'colors', 'position', 'attr']}
for prompt in prompts:
img = model.generate(prompt, steps=50, cfg=8.0, res=1024)
cat = get_category(prompt)
correct = vqa_model.verify(img, prompt)
category_scores[cat].append(correct)
overall = mean(mean(v) for v in category_scores.values())
return overall
Common pitfalls
- Models frequently struggle with spatial position and counting categories compared to single-object or color rendering.
- Automated evaluation relies on external VQA models that may introduce measurement noise or bias.
Evidence (verbatim from paper)
Nucleus-Image achieves an overall score of 0.865 (reported as 0.87), matching Qwen-Image and surpassing all other reported models including GPT Image 1 High (0.84) and Seedream 3.0 (0.84).
Citation
@misc{akiti2026nucleusimage,
title={Nucleus-Image: Sparse MoE for Image Generation},
author={Akiti et al. (2026)},
year={2026},
note={arXiv:2604.12163}
}
- arXiv: 2604.12163