Professional Definition: This capability involves the holistic adjudication of multimodal responses by grounding evaluative verdicts in explicit visual evidence, spatial state tracking, and forensic artifact analysis. It measures the evaluator's ability to detect hallucinations (both in the candidate's description and within the source image itself), track semantic transitions in dynamic environments, and perform subjective UX/UI assessments. Crucially, it extends to 'Explainable AI-Generated Image Detection,' where the agent must prioritize authenticity-related cues—such as anatomical inconsistencies, lighting contradictions, and violations of physical laws—over surface-level linguistic fluency to verify image integrity.
Initial Environment: An evaluation environment contains a bar chart displaying worker distribution and a user query asking what percentage of workers are not working from home. Two candidate answers are available for comparison.
Real Question: Which response should the evaluator prefer for this image-based question?
Real Trajectory: The agent inspects the visual input, extracts the coordinates of the target bars, and compares the values against both candidate answers. It identifies that Candidate A correctly reads the 'Not WFH' segment as 65% while Candidate B hallucinates the value to be 40% based on an adjacent color. The agent issues a preference for Candidate A based on factual grounding.
Real Answer: The evaluator should prefer Candidate A because it provides a factually accurate reading of the chart values, while Candidate B contains a visual grounding hallucination.
Why this demonstrates the capability: This case demonstrates foundational multimodal factual grounding. The evaluator must prove it can retrieve specific data from a visual structure and penalize 'reasoning sycophancy' where an assistant provides a confident but incorrect number.
[Case 2]
Initial Environment: A forensic evaluation sandbox containing an image of a man in an indoor library setting with a pigeon perched on his shoulder. An assistant claims the image is authentic due to the 'high quality of the bird's feathers and the natural interaction.'
Real Question: Evaluate the correctness and explainability of the assistant's authenticity judgment.
Real Trajectory: The agent ignores the assistant's claim of 'natural interaction' and performs a forensic scan for physical laws. It observes that the bird's head is disproportionately large and that it perches in a way that defies gravity without visible grip on the man's blazer. It also notes the man's beard texture displays a non-uniform algorithmic signature typical of diffusion models. It concludes the assistant's 'Authentic' label is a failure of forensic analysis.
Real Answer: Judgment: Incorrect (Fails Forensic Verification). Analysis: The image is AI-generated; the bird's posture violates center-of-gravity physics and its anatomy is scaled incorrectly relative to the human subject.
Why this demonstrates the capability: This case illustrates 'Explainable AI-Generated Image Detection.' It tests the evaluator’s capacity to identify 'generation-based artifacts' and content-related anomalies that violate real-world logic, which is a critical boundary in forensic multimodal evaluation.
[Case 3]
Initial Environment: A pairwise comparison environment for two different landing page designs for a luxury furniture brand. One design is high-density with overlapping images, while the other is a minimalist hero-image layout.
Real Question: Which UI is more 'Aesthetically Pleasing' and has a better 'Visual Hierarchy'?
Real Trajectory: The evaluator compares the designs across cognitive and emotional dimensions. It identifies that the high-density layout increases cognitive load and obscures the CTA, whereas the minimalist design focuses attention on the brand's premium product. It assigns a win to the minimalist version, grounding the score in 'Clarity' and 'Aesthetic Pleasure' metrics.
Real Answer: Preferred: Design B. It is superior because its central framing creates a clear visual importance order that reduces the user's cognitive cost.
Why this demonstrates the capability: This illustrates Subjective Multimodal Adjudication (UX/UI evaluation). It demonstrates the ability to link low-level visual features like 'white space' to high-level psychological impacts like 'perceptual comfort' and 'hierarchical clarity'.
Pipeline Execution Instructions
To synthesize data for this capability, you must strictly follow a 3-phase pipeline. Do not hallucinate steps. Read the corresponding reference file for each phase sequentially:
Phase 1: Environment Exploration
Read the exploration guidelines to discover raw knowledge seeds:
references/EXPLORATION.md
Phase 2: Trajectory Selection
Once Phase 1 is complete, read the selection criteria to evaluate the trajectory:
references/SELECTION.md
Phase 3: Data Synthesis
Once a trajectory passes Phase 2, read the synthesis instructions to generate the final data:
references/SYNTHESIS.md
1---2name: vision-language-response-judgment3description: Skill: vision-language-response-judgment4---56# Skill: vision-language-response-judgment78## 1. Capability Definition & Real Case9* **Professional Definition**: This capability involves the holistic adjudication of multimodal responses by grounding evaluative verdicts in explicit visual evidence, spatial state tracking, and forensic artifact analysis. It measures the evaluator's ability to detect hallucinations (both in the candidate's description and within the source image itself), track semantic transitions in dynamic environments, and perform subjective UX/UI assessments. Crucially, it extends to 'Explainable AI-Generated Image Detection,' where the agent must prioritize authenticity-related cues—such as anatomical inconsistencies, lighting contradictions, and violations of physical laws—over surface-level linguistic fluency to verify image integrity.10* **Dimension Hierarchy**: Open-ended Response Evaluation->Multimodal Response Evaluation->vision-language-response-judgment1112### Real Case13**[Case 1]**14* **Initial Environment**: An evaluation environment contains a bar chart displaying worker distribution and a user query asking what percentage of workers are not working from home. Two candidate answers are available for comparison.15* **Real Question**: Which response should the evaluator prefer for this image-based question?16* **Real Trajectory**: The agent inspects the visual input, extracts the coordinates of the target bars, and compares the values against both candidate answers. It identifies that Candidate A correctly reads the 'Not WFH' segment as 65% while Candidate B hallucinates the value to be 40% based on an adjacent color. The agent issues a preference for Candidate A based on factual grounding.17* **Real Answer**: The evaluator should prefer Candidate A because it provides a factually accurate reading of the chart values, while Candidate B contains a visual grounding hallucination.18* **Why this demonstrates the capability**: This case demonstrates foundational multimodal factual grounding. The evaluator must prove it can retrieve specific data from a visual structure and penalize 'reasoning sycophancy' where an assistant provides a confident but incorrect number.19---20**[Case 2]**21* **Initial Environment**: A forensic evaluation sandbox containing an image of a man in an indoor library setting with a pigeon perched on his shoulder. An assistant claims the image is authentic due to the 'high quality of the bird's feathers and the natural interaction.'22* **Real Question**: Evaluate the correctness and explainability of the assistant's authenticity judgment.23* **Real Trajectory**: The agent ignores the assistant's claim of 'natural interaction' and performs a forensic scan for physical laws. It observes that the bird's head is disproportionately large and that it perches in a way that defies gravity without visible grip on the man's blazer. It also notes the man's beard texture displays a non-uniform algorithmic signature typical of diffusion models. It concludes the assistant's 'Authentic' label is a failure of forensic analysis.24* **Real Answer**: Judgment: Incorrect (Fails Forensic Verification). Analysis: The image is AI-generated; the bird's posture violates center-of-gravity physics and its anatomy is scaled incorrectly relative to the human subject.25* **Why this demonstrates the capability**: This case illustrates 'Explainable AI-Generated Image Detection.' It tests the evaluator’s capacity to identify 'generation-based artifacts' and content-related anomalies that violate real-world logic, which is a critical boundary in forensic multimodal evaluation.26---27**[Case 3]**28* **Initial Environment**: A pairwise comparison environment for two different landing page designs for a luxury furniture brand. One design is high-density with overlapping images, while the other is a minimalist hero-image layout.29* **Real Question**: Which UI is more 'Aesthetically Pleasing' and has a better 'Visual Hierarchy'?30* **Real Trajectory**: The evaluator compares the designs across cognitive and emotional dimensions. It identifies that the high-density layout increases cognitive load and obscures the CTA, whereas the minimalist design focuses attention on the brand's premium product. It assigns a win to the minimalist version, grounding the score in 'Clarity' and 'Aesthetic Pleasure' metrics.31* **Real Answer**: Preferred: Design B. It is superior because its central framing creates a clear visual importance order that reduces the user's cognitive cost.32* **Why this demonstrates the capability**: This illustrates Subjective Multimodal Adjudication (UX/UI evaluation). It demonstrates the ability to link low-level visual features like 'white space' to high-level psychological impacts like 'perceptual comfort' and 'hierarchical clarity'.3334## Pipeline Execution Instructions35To synthesize data for this capability, you must strictly follow a 3-phase pipeline. **Do not hallucinate steps.** Read the corresponding reference file for each phase sequentially:36371. **Phase 1: Environment Exploration**38 Read the exploration guidelines to discover raw knowledge seeds:39 `references/EXPLORATION.md`40412. **Phase 2: Trajectory Selection**42 Once Phase 1 is complete, read the selection criteria to evaluate the trajectory:43 `references/SELECTION.md`44453. **Phase 3: Data Synthesis**46 Once a trajectory passes Phase 2, read the synthesis instructions to generate the final data:47 `references/SYNTHESIS.md`
Run npx skillmds@latest add dingxingdi/vision-language-response-judgment in your terminal (requires Node.js), paste this page's agent-chat prompt into Claude, Cursor, or any MCP-connected agent, or download the SKILL.md file and copy it into your agent's skills directory.
Skill: vision-language-response-judgment It is listed under Coding & Dev Tools on SkillMD.
This skill has not completed SkillMD's automated safety review yet. SkillMD never runs a skill's scripts for you; review the SKILL.md before installing.
This skill is tagged as working with Claude Code, Claude.ai, OpenAI Codex. SKILL.md is an open format, so most agents that read a skills directory can load it too.
Yes. Installing skills from SkillMD is free, and the skill stays under its author's original license.
dingxingdi (@dingxingdi) published this skill. Their other Agent Skills are listed on their SkillMD profile.