evalalign-eval
EVALALIGN: Supervised Fine-Tuning Multimodal LLMs with Human-Aligned Data for Evaluating Text-to-Image Models — Zhiyu Tan et al. (arXiv:2406.16562, 2024)
What this evaluates
Evaluates text-to-image generation models on two core capabilities: image faithfulness (consistency with real-world commonsense) and text-image alignment (adherence to the conditioning prompt). It probes whether generated images accurately reflect both visual realism and prompt instructions using a fine-grained, human-aligned framework.
Datasets
- EvalAlign — total ?; splits: validation (-1), test (-1)
Metrics
EvalAlign_f(primary) — range: other- Average score over the image faithfulness question set S_f. Each question's autoregressive response is parsed via regex to extract a chosen option, which maps to a predefined score. Scores are averaged across all questions in S_f.
EvalAlign_a(primary) — range: other- Average score over the text-image alignment question set S_a. Computed identically to EvalAlign_f but using the alignment-specific question set S_a.
Input / output format
Input: A fine-grained question/instruction (Q) and a multimodal input (M) consisting of the generated image and its conditioning text prompt.
Output: Autoregressive text response (R) from the MLLM, followed by rule-based filtering and regular expressions to extract the selected option.
Scoring recipe
def compute_evalalign(responses, question_set):
scores = []
for R in responses:
option = extract_option_via_regex(R)
score = map_option_to_score(option)
scores.append(score)
return sum(scores) / len(scores)
# EvalAlign_f = compute_evalalign(test_responses, S_f)
# EvalAlign_a = compute_evalalign(test_responses, S_a)
Common pitfalls
- Zero-shot MLLMs perform poorly on generated images due to data bias; supervised fine-tuning on human-aligned data is required.
- Prompts describing impossible scenarios (e.g., 'a man on Mars without a spacesuit') should not penalize image faithfulness scores, as the model is correctly following the prompt.
- Applying LoRA to ViT and projection modules during fine-tuning causes overfitting; only Q and K attention weights should be adapted.
Evidence (verbatim from paper)
We devise two holistic and detailed question sets, $S_{f}$ and $S_{a}$, that encompass every aspect of image faithfulness and text-image alignment, respectively. Consequently, our metric, EvalAlign, can be defined by averaging the scores of the questions in the two sets: ... where $ ext{EvalAlign}{ ext{f}}$ and $ ext{EvalAlign}{ ext{a}}$ indicate the image faithfulness score and the text-image alignment score evaluated by our method, respectively.
Citation
@misc{tan2024evalalign,
title={EVALALIGN: Supervised Fine-Tuning Multimodal LLMs with Human-Aligned Data for Evaluating Text-to-Image Models},
author={Zhiyu Tan et al.},
year={2024},
note={arXiv:2406.16562}
}
- arXiv: 2406.16562