Duaer

Snippets

Remove duplicates in Duaer, within a run and across runs

Remove Duplicates in Duaer drops repeats within one batch, and can remember what earlier runs handled so a schedule processes only new items.

Drop repeats within one batch

  1. Add Remove Duplicates with Operation Remove Items Repeated Within Current Input.
  2. Set Compare to Selected Fields and Fields To Compare to the key, such as email or order_id. Use All Fields when only identical items count.
  3. Select Execute step. The first of each repeat stays.

Process only what earlier runs did not

  1. When a schedule pulls data (new orders or emails every 5 minutes), set Operation to Remove Items Processed in Previous Executions.
  2. Set Keep Items Where to Value Is New and Value to Dedupe On to {{ $json.id }}.
  3. For increasing IDs or times, use Value Is Higher than Any Previous Value or Value Is a Date Later than Any Previous Date.
  4. Scope under Options keeps the history per node or per digital organization; History Size caps how many values it keeps.

To reprocess everything, run once with Operation Clear Deduplication History, then switch back.

Questions

How does a Duaer schedule avoid reprocessing the same data?

Add Remove Duplicates in Duaer with Remove Items Processed in Previous Executions on the id, so only new items pass.

Does the Duaer dedupe history grow forever?

No. It stays within History Size in Duaer Remove Duplicates options; the oldest values drop off.

In this section