Malware Family Classification Eval

Evaluates the ability of LLMs and their ensembles to correctly classify malware samples into one of ten canonical families based on their behavior or code semantics. It probes robustness to class imbalance and the effectiveness of hierarchical decision-making under obfuscation. Use when the user wants to benchmark on Gold-standard malware family dataset, or asks about evaluating this task. Reports Macro F1-score.

qhjqhj00 0f8a7d0 3.6 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/malware-family-classification-eval commit 0f8a7d006d

Frequently asked questions

npx skillmds add qhjqhj00/malware-family-classification-eval