Detecting Data And Model Poisoning

mukul975/detecting-data-and-model-poisoning · Agent Skill (multi-file)

by mukul975 · bundle

Published · Last updated


Detect poisoned training data and backdoored models across the ML pipeline using statistical analysis, activation clustering, and spectral signatures.

SKILL.md

Files

This skill is a package of 5 files. Install with the command above, or download the folder.

  • 📄SKILL.md entry
  • 📁references
  • 📄api-reference.md 2.2 KB
  • 📄standards.md 1.6 KB
  • 📁scripts
  • ⚙️agent.py 4.9 KB
  • 📄LICENSE 11.0 KB

Related

  1. AI Security · alirezarezvani bundle
    Assess AI/ML systems for prompt injection, jailbreak vulnerabilities, model inversion risk, data poisoning exposure, and agent tool abuse, with MITRE ATLAS mapping and guardrail recommendations.
    20.4k
    repo stars
  2. Defending Llms With Guardrails · mukul975 bundle
    Deploy Llama Guard, NeMo Guardrails, and LLM Guard as runtime input/output scanners to block jailbreaks, prompt injection, and toxic content in production LLM applications.
    24.6k
    repo stars
  3. Detecting Model Extraction Attacks · mukul975 bundle
    Detect model stealing, model inversion, and membership inference performed through inference-API abuse by monitoring query patterns, applying output perturbation, and red-teaming your own model's extractability.
    24.6k
    repo stars
  4. Testing For System Prompt Leakage · mukul975 bundle
    Test LLM applications for system prompt leakage using manual payloads, garak, and Promptfoo to extract embedded secrets and routing logic.
    24.6k
    repo stars
  5. Detecting Indirect Prompt Injection · mukul975 bundle
    Detect and defend against prompt injection hidden in documents, web pages, and images consumed by an agent.
    24.6k
    repo stars
  6. Testing Prompt Injection In RAG Pipelines · mukul975 bundle
    Probe RAG applications for prompt injection via poisoned retrieved context and embedding manipulation.
    24.6k
    repo stars

Frequently asked questions

How do I install the Detecting Data And Model Poisoning skill?

Run npx skillmds add mukul975/detecting-data-and-model-poisoning in your terminal (requires Node.js), paste this page's agent-chat prompt into Claude, Cursor, or any MCP-connected agent, or download the SKILL.md file and copy it into your agent's skills directory.

What does the Detecting Data And Model Poisoning skill do?

Detect poisoned training data and backdoored models across the ML pipeline using statistical analysis, activation clustering, and spectral signatures. It is listed under Security, AI & ML, DevOps & Infra, Model Training & Fine-tuning, Vulnerability Scanning on SkillMD.

Is Detecting Data And Model Poisoning safe to use?

SkillMD's automated safety review verdict for this skill is CAUTION. Independent scanners report: SkillSpector: CAUTION, Skill Scanner: PASS. Capability flags: executes scripts, makes network calls, reads secrets. SkillMD never runs a skill's scripts for you; review the SKILL.md before installing.

Which AI agents work with Detecting Data And Model Poisoning?

This skill is tagged as working with Claude Code, Claude.ai, OpenAI Codex. SKILL.md is an open format, so most agents that read a skills directory can load it too.

Is Detecting Data And Model Poisoning free to use?

Yes. Installing skills from SkillMD is free. This skill is licensed under Apache-2.

Who published Detecting Data And Model Poisoning?

mukul975 (@mukul975) published this skill. Their other Agent Skills are listed on their SkillMD profile.