Maps Multilingual Agent Eval

Evaluates the performance and security robustness of agentic AI systems when operating in multilingual settings. It measures how task completion accuracy and vulnerability to adversarial prompts degrade or shift when instructions are translated from English into 11 typologically diverse languages. Use when the user wants to benchmark on GAIA, SWE-bench, MATH, ASB, or asks about evaluating this task. Reports accuracy.

qhjqhj00 9c185f7 3.9 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/maps-multilingual-agent-eval commit 9c185f7e1d

Frequently asked questions

npx skillmds add qhjqhj00/maps-multilingual-agent-eval