VSThiran

Find near-duplicate records even when the names differ slightly

Quick answer

Upload your spreadsheet, choose the column that should be unique, and set how close counts as a match. Records like "Jon Smith" and "John Smith" - the same person, entered slightly differently - are grouped as likely duplicates for you to review.

  1. 1Upload your spreadsheet.
  2. 2Choose the column to check for near-duplicates.
  3. 3Set how close counts as a match.
  4. 4Review and resolve each group.

Free, no sign-up, and your file is read on your own device rather than uploaded.

A dataset built up over time from more than one source - a CRM import, a form, a manual entry - collects near-duplicates that an exact match will never catch: "Jon Smith" and "John Smith", a phone number with and without formatting, a company name with and without "Ltd". They are the same underlying record, entered slightly differently.

This compares every record against every other one for similarity, not just exact equality, and groups the ones that are probably the same thing so you can decide - keep one, merge them, or confirm they really are different after all.

Step by step

  1. Upload your spreadsheet

    Excel or CSV, read locally in your browser.

    Find Similar Records
  2. Choose the column to check

    A name, company or reference field - whatever should be unique but might have been entered inconsistently.

    Find Similar Records
  3. Set how close counts as a match

    Controls how similar two values need to be before they are grouped as likely duplicates.

    Find Similar Records
  4. Review and resolve each group

    Every group of likely duplicates is shown together - confirm, merge or dismiss each one before downloading.

    Find Similar Records

Tips

  • Run the plain Duplicate Finder first if you haven't - it is faster and catches every genuinely identical row, leaving fuzzy matching to focus on the harder, inexact cases.
  • A tighter threshold produces fewer, more confident groups; a looser one catches more but needs more review. Start default and adjust based on what the review queue looks like.
  • This is the same underlying comparison Fuzzy Lookup uses between two files - here it is applied within one file, row against row.

Common problems

Two records you know are duplicates were not grouped together.

Lower the threshold slightly and re-run. If they still are not grouped, double check you chose the column where the actual difference is - similarity is only measured in the column you select.

The review queue contains records that are actually different people or companies.

That is expected at a looser threshold - short or common values (like short names) are more likely to score as similar by coincidence. Reject those specific groups rather than loosening the threshold further.

Frequently asked questions

What's the difference between this and the plain Duplicate Finder?
Duplicate Finder catches rows that are exactly identical. This catches rows that are probably the same thing but entered slightly differently - different spelling, formatting, or punctuation.
Can I control how strict the matching is?
Yes - a single setting controls how close two values need to be before they count as likely duplicates.
Will it automatically merge or delete rows?
No - every group goes to a review queue first. Nothing is merged or removed until you confirm it.
Is my spreadsheet uploaded anywhere?
No. Everything runs in your browser, on your own device.

Related data tools

Related articles

Ready to do it?

Free, no sign-up, and nothing is uploaded to a server.