Ray Document Deduplicator Ray分布式文档去重 Skill
功能概述
本skill基于Ray分布式框架,使用MD5哈希算法进行文档级别的精确匹配去重。 适用于大规模数据集的并行处理场景。
触发条件
当用户请求以下任务时,应使用此skill:
- 文档去重
- 删除重复文档
- 样本去重
- 分布式去重
- 大规模数据去重
- Ray并行处理
核心参数说明
必需参数
| 参数 | 说明 |
|---|---|
--input |
输入JSON文件路径 |
--output |
输出JSON文件路径 |
可选参数
| 参数 | 说明 | 默认值 |
|---|---|---|
--lowercase |
是否将文本转为小写进行比对 | False |
--ignore_non_character |
是否忽略非字母字符(空格、数字、标点) | False |
--backend |
分布式后端类型 | ray_actor |
--redis_address |
Redis服务器地址 | redis://localhost:6379 |
--text_key |
文本字段的键名 | text |
输入文件格式
{
"data": [
{"text": "文档内容1"},
{"text": "文档内容2"},
{"text": "文档内容1"}
]
}
或直接使用文档列表:
[
{"text": "文档内容1"},
{"text": "文档内容2"},
{"text": "文档内容1"}
]
使用方法
基本用法
python scripts/run_ray_document_deduplicator.py \
--input ./input.json \
--output ./output.json
使用Redis后端
python scripts/run_ray_document_deduplicator.py \
--input ./input.json \
--output ./output.json \
--backend redis \
--redis_address redis://localhost:6379
忽略大小写去重
python scripts/run_ray_document_deduplicator.py \
--input ./input.json \
--output ./output.json \
--lowercase True
输出示例
命令行输出:
[OK] Ray document deduplication completed!
Backend: ray_actor
Original documents: 5
Deduplicated documents: 4
Removed duplicates: 1
Input file: ./input.json
Output file: ./output.json
输出JSON格式:
[
{"text": "文档内容1"},
{"text": "文档内容2"},
{"text": "文档内容3"},
{"text": "文档内容4"}
]
环境要求
安装依赖: 本SKILL使用依赖 data_juicer 和 Ray,请在调用前安装好python环境并安装依赖:
pip install py-data-juicer ray redis
启动Ray:
ray start --head
或启动Redis:
redis-server
与普通DocumentDeduplicator的区别
| 特性 | DocumentDeduplicator | RayDocumentDeduplicator |
|---|---|---|
| 处理方式 | 单机处理 | 分布式并行处理 |
| 适用场景 | 小规模数据 | 大规模数据 |
| 依赖 | 仅data_juicer | data_juicer + Ray/Redis |
| 性能 | 一般 | 高 |
注意事项
- 分布式环境:使用前需要启动 Ray 或 Redis
- 后端选择:小规模数据用
ray_actor,超大规模用redis - 输入格式:输入JSON需要包含
data字段或直接是文档列表 - 处理逻辑:保留首次出现的文档,删除后续重复项