Remoteshield Eval

Probes the robustness and cross-condition consistency of multimodal large language models on Earth observation tasks under realistic visual and textual perturbations. Evaluates performance degradation and behavioral stability across clean and perturbed inputs for scene classification, VQA, and visual grounding. Use when the user wants to benchmark on RemoteShield clean-perturbed benchmarks, or asks about evaluating this task. Reports RPD, CCA.

qhjqhj00 4127cf1 3.9 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/remoteshield-eval commit 4127cf1d3e

Frequently asked questions

npx skillmds add qhjqhj00/remoteshield-eval