Homehome

Windows · Duplicates

Three files, one name.
Only two are the same.

invoice.pdf, invoice (1).pdf and invoice (2).pdf are the most Windows thing there is, and the names tell you nothing about which are actually identical. A duplicate finder that matches on names and sizes will guess, and eventually it will guess wrong on something you needed.

Download for WindowsBetaWindows

VaultSort for Windows is in beta. Free to try, no account required.

Why duplicate finders lose files

Name and size matching is a guess

Two photos from the same camera can share a byte count exactly. Two drafts of a document can share a name. Matching on metadata is fast because it never opens the file, and wrong for exactly the same reason.

Something has to choose which copy survives

With four hundred groups you are not deciding each one by hand. If the tool will not tell you its rule, its rule is whatever it happened to scan first — which is a coin toss you are not present for.

Some of your duplicates are load-bearing

Package caches, node_modules and Windows’ own component store are full of byte-identical files that are supposed to be identical. Deleting from WinSxS breaks Windows Update. A scan of your whole user folder will surface thousands of these next to the ones you actually want gone.

The real test is what happens after a bad run

Every one of these tools works when it is right. The question is what you do when it removed four hundred files it should not have, and "restore them from the bin one at a time, working out where each belonged" is not an answer.

How VaultSort finds them

Verified byte for byte

Files are grouped by size, then by a hash of their first 64 KB, and only survivors of both get a full-content hash. Nothing is ever called a duplicate on the strength of its name — the fast tiers exist to avoid reading files that cannot possibly match, not to avoid checking the ones that might.

Nothing is deleted by default

The default action moves duplicates into a folder you nominate, so the result is a pile you can look through, restore from, or bin yourself. Secure deletion is available as a deliberate second choice.

A keep rule you can predict

The copy that survives is the oldest by modification time, with ties broken by the shortest path and then alphabetically. Deterministic, so two runs over the same folder never disagree about which file was the original.

Hardlinks are collapsed, not counted

Two paths sharing the same storage are recognised as one file by volume identity, so you are never offered space that deleting could not free.

Similar images, kept as a separate question

An optional perceptual mode finds the same photo saved at a different size, format or quality — a 64-bit difference hash with its own threshold. It is a different claim from "identical" and is presented as one, rather than blended into a single number.

Organize first, deduplicate second

Sorting a folder by type before hunting duplicates makes the results far easier to judge — forty PDFs rather than four hundred mixed files — and organizing is fully reversible, so the sequence stays safe.

What it does not do yet

  • The scanner skips filesystem bookkeeping — $RECYCLE.BIN, System Volume Information and their Mac equivalents — but it does not carry an exclusion list for AppData, package caches or WinSxS. Point it at Downloads or a project folder rather than at your whole user directory.
  • OneDrive Files On-Demand placeholders are read rather than skipped, so scanning a folder full of them pulls their contents down. Make them available offline first, or scan somewhere else. (The equivalent iCloud placeholders are skipped on macOS; the Windows case is not handled yet.)
  • Network drives are scanned but the results are less reliable than on local storage, and slow enough that a large share is best left to run rather than watched.
  • Symbolic links are never followed, so a duplicate reachable only through a link outside the folder you scanned will not appear.

Questions people ask

Does it match renamed copies?

Yes. Matching is on content, so a file renamed, moved, or saved with a different extension is still recognised. The name is never part of the comparison.

Can it find photos that are the same but not identical?

Yes, through the optional similar-image mode — the same picture at a different resolution or compression level. It is deliberately opt-in and kept separate from exact matching, because "looks the same" and "is the same" are different claims and should not share a checkbox.

How fast is it on a large folder?

Most files are never opened. Anything with a unique size is eliminated without being read, and everything else is settled by its first 64 KB unless that also collides. The full hash — the expensive part — runs on genuine candidates only.

What if I change my mind after moving duplicates?

They are in the folder you chose, with their names intact. Nothing was deleted, so putting them back is a file operation rather than a recovery job.

Try it on your own files

One purchase, no subscription. $24.99 once, and every feature is included.

Stay Updated with VaultSort

Updates, security tips, and feature announcements. No noise.

No spam. Unsubscribe at any time.