← All writing
Platform & APIs

Shopify bulk operations: how to scan a whole catalogue without fighting rate limits

Keith Pillay · 30 September 2026 · 3 min read

When I started building a scanner that reads every product image in a Shopify store, the first decision was how to read a whole catalogue. There are two options, and they behave very differently as stores get larger.

The paginated approach, and where it hurts

The obvious method is a normal GraphQL query: ask for products, 25 at a time, follow the cursor, repeat.

It works fine for small stores. The pressure comes from Shopify's query cost system:

  • Every field has a cost, and connections are priced by the page size you request.
  • A single query is capped at a maximum cost of 1,000 points. Ask for too many products with too many nested images and the whole query is rejected.
  • Your app draws from a rate-limit bucket that refills at a fixed rate that depends on the store's plan. Standard plans refill at 100 points per second.

Do the arithmetic for a store with 10,000 products and several images each, and you are looking at hundreds of requests and, in the worst case, more than an hour of throttled paging. For 100,000 products it stops being practical.

Bulk operations

Shopify provides a different mechanism for exactly this job: bulk operations. Instead of asking for pages, you hand Shopify one query and it runs the export asynchronously on its side. When it's done, you download the results as a file.

The shape of the flow:

  1. Start the job with the bulkOperationRunQuery mutation.
  2. Wait for it to finish. Shopify can notify you with the bulk_operations/finish webhook.
  3. Download the result from the operation's url.

Two properties make this attractive. The mutation that starts the job is the only part counted as a normal request, so the export itself doesn't eat your rate limit. And there is no cursor bookkeeping on your side.

What the result looks like

The output is JSONL: one JSON object per line. Parent and child records come as separate lines, and each child carries a __parentId pointing at its parent. So a product and its images arrive as a product line followed by image lines that reference it.

There is an option, groupObjects, that nests children under parents for you. The documentation warns it raises the chance of timeouts, so I leave it off and do the joining myself while reading the file.

The practical consequence: stream the file line by line. For a large catalogue the file can be tens of megabytes. Loading it into memory in one go is the kind of shortcut that works in testing and falls over in production.

The rules of the query

A bulk query is not unconstrained GraphQL. At the time of writing, per Shopify's documentation:

  • It must contain at least one connection, and no more than five in total.
  • Nesting is limited to two levels of connections.
  • Pagination arguments like first are ignored. The whole connection is exported.
  • Top-level node and nodes queries can't be used.

Design your query to fit those rules up front. For a product-and-media scan, it fits comfortably.

Don't trust a single notification

Webhooks are the natural way to learn that a job finished, but Shopify's own documentation points out that delivery isn't guaranteed. So the reliable pattern is webhook plus polling fallback: subscribe to the finish event, and also check the operation's status on a timer. Whichever arrives first wins; the other is a no-op.

Also remember that the result URL is signed and expires, roughly a week after the job completes. Download and process it promptly rather than storing the link.

One code path

The trade-off with bulk is that results are asynchronous, so your interface needs a proper "scan running" state. In return you get one scan implementation that works at every catalogue size, instead of a paginated path for small stores and a bulk path for large ones.

I kept paginated queries only for small, targeted lookups, such as re-checking a single product when a webhook says it changed.

That is how Storko Product Health Scanner reads a 1,000-product test store in around ten seconds and keeps working on far larger catalogues.

API details change between Shopify API versions. Check the current documentation before relying on any specific limit.

Hiring a senior Shopify developer?

I'm open to remote roles worldwide. Send a message and I'll reply within a day.