> For the complete documentation index, see [llms.txt](https://doc.duaer.com/zh/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://doc.duaer.com/zh/snippets/remove-duplicates.md).

# 在 Duaer 里去重：本次重复和以前处理过的

Duaer 的 Remove Duplicates 能去掉这一批里重复的数据，也能记住以前处理过的，定时任务只处理新的。
## 去掉这一批里的重复

1. 接 [Remove Duplicates](/zh/build/remove-duplicates.md)，Operation 选 Remove Items Repeated Within Current Input。
2. Compare 选 Selected Fields，Fields To Compare 填判断重复的字段，例如 email 或 order_id。整条完全一样才算重复时选 All Fields。
3. 点执行此步骤。重复的只留第一条。

## 只处理以前没处理过的

1. 定时拉数据（每 5 分钟查一次新订单、新邮件）时，Operation 选 Remove Items Processed in Previous Executions。
2. Keep Items Where 选 Value Is New，Value to Dedupe On 填 {{ $json.id }}。
3. 自增编号或时间可以选 Value Is Higher than Any Previous Value 或 Value Is a Date Later than Any Previous Date，只放行更大的。
4. Options 里的 Scope 决定记录归这个节点还是整条数字组织；History Size 是最多记多少条。

要重新处理一遍时，临时把 Operation 改成 Clear Deduplication History 执行一次，再改回来。

## Duaer 在线工具

- [Excel / CSV 表格去重（在线，按指定列）](https://doc.duaer.com/zh/tools/remove-duplicates.md)
## Questions

### Duaer 定时任务怎么避免每次都重复处理同一批数据？

在 Duaer 里加 Remove Duplicates，选 Remove Items Processed in Previous Executions，按 id 去重，只放行新的。

### Duaer 的去重记录会一直增长吗？

不会超过 Duaer Remove Duplicates 选项里的 History Size，超过后最早的记录被丢掉。

## 相关

- [在 Duaer 里用 Remove Duplicates 去重](https://doc.duaer.com/zh/build/remove-duplicates.md)
- [在 Duaer 里用 Schedule Trigger 定时运行](https://doc.duaer.com/zh/build/schedule-trigger.md)
- [在 Duaer 里按关键字段合并两张表（类似 VLOOKUP）](https://doc.duaer.com/zh/snippets/merge-lists-by-key.md)

