oneig-bench-eval
Nucleus-Image: Sparse MoE for Image Generation — Akiti et al. (2026) (arXiv:2604.12163, 2026)
What this evaluates
Measures broader image generation capabilities across five axes: alignment, text rendering, reasoning, style, and diversity.
Datasets
- OneIG-Bench — total ?; splits: test (-1)
Metrics
Overall(primary) — range: [0, 1]- Mean score across five axes: alignment, text rendering, reasoning, style, and diversity.
Input / output format
Input: Text prompt designed to test alignment, text rendering, reasoning, style, and diversity.
Output: Generated image at 1024x1024 resolution, 50 inference steps, CFG scale 8.0.
Scoring recipe
dims = ['alignment', 'text', 'reasoning', 'style', 'diversity']
scores = {d: [] for d in dims}
for prompt in prompts:
img = model.generate(prompt, steps=50, cfg=8.0, res=1024)
for d in dims:
scores[d].append(evaluator.score(img, prompt, dimension=d))
overall = mean(mean(v) for v in scores.values())
return overall
Common pitfalls
- Diversity scores tend to be lower across most models, indicating a known limitation in generating varied outputs from similar prompts.
- Text rendering and style evaluation may rely on specialized models or human-like scoring that can vary in strictness.
Evidence (verbatim from paper)
Nucleus-Image achieves an overall score of 0.522, placing it among the top tier of open and proprietary models and ahead of Imagen4 (0.515) and Recraft V3 (0.502).
Citation
@misc{akiti2026nucleusimage,
title={Nucleus-Image: Sparse MoE for Image Generation},
author={Akiti et al. (2026)},
year={2026},
note={arXiv:2604.12163}
}
- arXiv: 2604.12163