First checks: duplicate records

Illustrative project · synthetic data

Goal

Check a small list for repeated IDs and preserve the first occurrence of each value. This is a teaching example, not evidence of a real project or business result.

Setup

Python 3; no additional packages are needed to run the cells. The input records are included below.

Steps

1. Inspect the input

These IDs are deliberately repeated. In a real dataset, repeated IDs may represent legitimate events; establish the intended row grain before removing anything.

In [1]:
records = ["A", "B", "A", "C", "B"]
print("Input:", records)
print("Record count:", len(records))
Input: ['A', 'B', 'A', 'C', 'B']
Record count: 5

2. Preserve the first occurrence

The dictionary keeps insertion order, so the resulting list follows the order of the original IDs.

In [2]:
def unique_values(values):
    return list(dict.fromkeys(values))

clean = unique_values(records)
print("Unique IDs:", clean)
print("Repeated entries:", len(records) - len(clean))
Unique IDs: ['A', 'B', 'C']
Repeated entries: 2

Checks

Test the expected order, an empty input, and idempotence: cleaning an already-clean list should leave it unchanged.

In [3]:
assert clean == ["A", "B", "C"]
assert unique_values([]) == []
assert unique_values(clean) == clean
print("All three checks passed.")
All three checks passed.

Next steps

The five input entries contain three unique IDs and two repeated entries. Before applying this pattern to a real dataset, define which columns identify a record and whether repeated records should be retained, merged, or removed.