Content Moderation Gate

Check user-generated text against a written policy before it is published. One multi-label call scores every policy category independently, the band sets severity, only the middle band reaches a human queue, and hard-block categories are held whatever the score says. Use when a comment box, feed or chat product needs "moderation", "policy enforcement", "flag bad posts", or an answer to "what do we do with the borderline ones".

mrmps Updated

File contents

mrmps/classifier-dev/tree/main/skills/content-moderation-gate commit b42841125b

Frequently asked questions

npx skillmds@latest add mrmps/content-moderation-gate