Language Id Score Filter

语言识别分数过滤器。过滤器以保留置信度得分大于特定最小值的特定语言的样本。 本SKILL使用依赖data_juicer,请在调用前安装好python环境并安装data_juicer,你可用以下指令进行安装: pip install py-data-juicer 当用户提到语言识别过滤、语言检测、文本语言筛选、按语言过滤等需求时使用此skill。

cas-bigdatalab e0d941d 5 files · 41.5 KB Updated

File contents

cas-bigdatalab/piflow/tree/main/workspace/skills/language_id_score_filter commit e0d941de73

Frequently asked questions

npx skillmds@latest add cas-bigdatalab/language-id-score-filter