Multiagentfraudbench Eval

Evaluates the ability of LLM-powered multi-agent systems to collude and execute financial fraud in simulated social platform environments. It probes how agents coordinate through public and private channels, adapt to content warnings, and amplify fraud risks based on interaction depth and activity levels. Use when the user wants to benchmark on MultiAgentFraudBench, or asks about evaluating this task. Reports fraud_success.

qhjqhj00 cc9ecec 3.3 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/multiagentfraudbench-eval commit cc9ececf86

Frequently asked questions

npx skillmds add qhjqhj00/multiagentfraudbench-eval