The Shopify media ID that hid 500 rows
Keith Pillay · 29 September 2026 · 2 min read
The most useful bug I hit while building the Product Health Scanner was one where nothing looked wrong.
What happened
I ran my first large scan on a development store with 1,014 products. The summary was exactly right: 1,014 products, 506 images too small, one of unknown size. Every number matched what I'd seeded.
Then I looked at the results table. It held 7 rows instead of 507.
The totals were computed as the scan stream passed through, so they were right. The rows were written to the database keyed by the image's media ID, and most of them were silently overwriting each other.
Why
Shopify uses one media (file) ID for every product that references the same file. If you create a thousand products from the same image URL, you get a thousand image entries pointing at a single media ID.
My test data had 1,015 image entries and only 15 unique media IDs. I had assumed that a media ID identified one image in one product. It identifies one file, and a file can appear anywhere.
In a real store this happens whenever products share an image: bundles, variants of the same item sold as separate products, a placeholder used across a catalogue, or imports from a supplier feed that reuses assets.
The fix
The result key became the combination of shop, product ID and media ID. A given file in a given product is one row, and the same file in another product is another row.
That's a one-line change. The more valuable part was what I added around it.
Three rules I took from it
1. Assert that stored rows equal the scan totals. The summary and the table were computed on different paths, so they could disagree, and they did. I now have an invariant check: the number of stored rows must equal the totals reported by the scan. Had it existed, the first large run would have failed immediately.
2. Test with duplicated IDs, not just clean ones. My unit fixtures all had unique IDs, because that's how you naturally write fixtures. The bug only appears with realistic duplication. I added tests where several products share a media ID, and a test that the totals and the stored rows agree.
3. Be suspicious of "it all added up". Matching totals felt like proof. They were evidence that one code path was right, and said nothing about the other.
The general lesson
When you key data by an identifier from an external system, find out what that identifier is unique across, from the documentation or from an experiment, rather than from how the objects feel. Ask: can the same ID legitimately appear twice? If you can't answer from the docs, test it with real data at realistic scale.
A scale test is also a correctness test. I wouldn't have found this on a 20-product store. It took a thousand products that reused a handful of files to expose it.
This is one of the lessons from Part 2 of the build story.

