I will match, merge and deduplicate your messy datasets

Sommige informatie wordt in het Engels weergegeven.

Mexico

Ik spreek Spaans, Engels, Frans

PhD Candidate in Astrophysics, Statistical Consulting

PhD candidate in astrophysics. I spend my days doing statistics on messy, incomplete, correlated data — galaxy catalogs where nothing is clean and every error bar has to be defended. I bring that sta...
Over deze dienst

Two lists of the same people, products or places, no shared ID: names spelled differently, typos, missing fields. Excel's VLOOKUP gives up immediately.


I do probabilistic record linkage: instead of "match / no match", every candidate pair gets a score for how likely it is to be the same entity, based on how much agreement on each field is actually worth. Agreeing on a rare surname is strong evidence; agreeing on a common one is weak. The method knows the difference.


Why this matters to you: you get to choose the threshold. High confidence for automatic merging, and a separate list of ambiguous pairs for a human to look at, instead of silently merging two different customers, or keeping the same one twice.


I use this on astronomical catalogs, where merging the wrong two objects invalidates the science. Your data deserves the same care.


What I deliver: your merged dataset with a confidence score on every match; a separate "needs review" list for ambiguous cases; a quality report on how many matched, how many didn't, and why; and duplicates found within each file, not just across them.


Send a small sample first if you'd like to see it work before ordering.

Technologie:

Python

Expertise:

Gegevensmanipulatie

Data validatie

etl

Normalisatie