← All writing
Shopify Apps

Testing a Shopify app: unit, integration, scale, and breaking your own tests

Keith Pillay · 3 October 2026 · 3 min read

By the time I submitted the Product Health Scanner, it had 182 unit tests and 37 integration and scale tests. The number isn't the point. What matters is what each layer caught, and the habit that made them believable.

The layers

Unit tests cover the pure logic: classifying an image as too small or unknown, parsing the bulk-operation JSONL file, aggregating a product's issues, building the CSV exports, computing the health score. These are fast and numerous. They also make refactoring safe.

Integration tests run the scan lifecycle against a real Postgres database with fake shops and a fake Shopify: start, import, store, replace on rescan, delete on uninstall, process queued webhooks. They exercise the code that talks to the database, and the transactions, which unit tests can't.

Scale tests push a synthetic 100,000-product dataset through the real parser, collector and database, and check time and memory against a budget (here 120 seconds and 400 MB). They exist to catch the problems that only appear at size.

The real store. I also ran everything against a real Shopify development store: a real bulk operation on a thousand products, and later six thousand. Tests can all pass while the product is broken.

What each layer caught

  • Unit tests caught edge cases in the price and compare-at-price truth table, and made the duplicate-SKU logic safe to change.
  • Integration tests caught a database performance problem. My first webhook-queue manager used one transaction per product, and a 200-item batch took over a minute.
  • The scale test caught the shared media ID bug. Small fixtures never would have.
  • The real store caught a page that never loaded, because a component imported a server-only module. Typecheck, lint and unit tests were all green. More in the embedded UI traps.

Each layer found something the others couldn't. That's the argument for having all of them.

Break your own tests

A test that passes on its first run proves very little. It might be testing nothing.

So after writing tests for a piece of logic, I mutate the code on purpose and check that the right test fails:

  • Change a < to <= in the size classifier. Two tests should fail.
  • Remove a guard in the import step. Exactly the "replaces previous results" test should fail.
  • Bypass encryption in the session store. Exactly the "ciphertext, not plaintext" tests should fail.

Then I restore the code. If nothing fails, the test isn't doing its job.

Make fixtures realistic

Clean fixtures hide bugs. Datasets with duplicated IDs, products missing every optional field, draft products priced at zero (which should not be flagged), and one product that deliberately triggers each check in isolation all earned their keep.

When I added more checks, my "isolated" datasets stopped being isolated, because the new checks fire on absent data too. Every "good" product needed a complete baseline. A reminder that tests are code and need maintenance.

Know what you couldn't test

I kept an honest list of what wasn't verified: the pager click in the embedded frame, a few automation limits around iframes, an empty-store screen I'd written but not seen. Listing them is better than implying total coverage. For things I couldn't click, I verified the data layer directly, and reused patterns already proven elsewhere.

A routine

  1. After every route change, run a build and load the page.
  2. Run the unit suite on every change.
  3. Run integration and scale tests before merging.
  4. Mutate new logic to check the tests bite.
  5. Do a real-store pass before submission.

None of it is glamorous. All of it is why I'm comfortable with what I submitted.

Part of the build story.

Hiring a senior Shopify developer?

I'm open to remote roles worldwide. Send a message and I'll reply within a day.