Shopify Integrations

Designing a highly available Shopify integration: what keeps working when something upstream does not

Queues instead of direct calls, idempotency as a requirement rather than a nicety, graceful degradation, and knowing which failures you can absorb and which you cannot.

A long aisle of white server cabinets in a data centre
Photo: PiDatacenters, CC BY-SA 4.0, via Wikimedia Commons

A highly available Shopify integration is not one that never fails. It is one where a failure in any single part does not stop orders being taken, does not lose data, and does not require somebody to reconstruct what happened from memory afterwards.

That distinction matters because a highly available Shopify integration is rarely what gets built first: most are written for the happy path and then hardened reactively, usually after an incident that cost real money. The patterns below are the ones worth having from the start, in rough order of how much availability they buy per hour spent.

A highly available Shopify integration puts a queue between event and work

This is the highest value decision in a highly available Shopify integration, and everything else on this page assumes it.

The naive arrangement is direct: a webhook arrives, your code receives it and immediately writes to an ERP, a warehouse system and an accounting package. If any of those three is slow or down, the webhook handler fails. The order is now somewhere between systems and nobody knows where.

The durable arrangement separates receipt from processing. The handler does one thing: write the payload to durable storage and acknowledge. A separate worker reads from that store and does the real work, retrying as needed.

What this buys you. An ERP maintenance window stops being an outage and becomes a backlog that drains afterwards. A deploy of your own code stops being a risk to order capture. A slow third party stops being your problem for the duration.

What it costs. Eventual consistency. The order exists in your queue before it exists in the ERP, so anything reading the ERP will briefly not see it. That is almost always an acceptable trade, and it has to be an explicit one because somebody will ask why the warehouse cannot see an order that the customer already has a confirmation for.

The queue has to be durable, not in memory. A process that holds pending work in memory loses it on restart, and restarts happen for ordinary reasons. Durable means it survives the process dying.

Idempotency in a highly available Shopify integration

Once a highly available Shopify integration retries, it has duplicate delivery. This is not a possibility to guard against occasionally, it is a certainty to design for.

Webhooks can be delivered more than once. A worker can crash after doing its work but before recording that it did. A replay after an incident reprocesses a window deliberately. All three produce the same message twice.

The mechanism is a uniqueness key the receiving system enforces. Carry the Shopify order identifier onto the downstream record and check for it before creating anything. Not a lookup followed by an insert, which has a race between the two, but a constraint the database itself enforces so that the second attempt fails cleanly and the handler treats that failure as success.

Partial work is the harder case. A job that creates an order, then a shipment, then a ledger entry can fail after the first two. Retrying must not duplicate those. Either make each step individually idempotent with its own key, or record progress so a retry resumes rather than restarts. The first is usually simpler to reason about.

Test it deliberately. Send the same webhook twice on purpose, in a test environment, and confirm exactly one order exists. A highly available Shopify integration that has never had this tested is one duplicate delivery away from double shipping.

Decide what degrades and what stops

Not every dependency in a highly available Shopify integration deserves equal protection. The useful exercise is to list each one and decide, in writing, what happens when it is unavailable.

Must not block order capture: the ERP, the accounting system, the loyalty platform, the email tool, the analytics. If any of these being down prevents a customer from completing a purchase, the design is wrong. All of them can be caught up later.

Can degrade visibly: live stock figures. If the inventory source is unreachable, you can serve a cached figure, or fall back to a conservative one, or hide the quantity and keep selling. All three are better than an error page, and which you choose is a commercial decision about whether you would rather risk an oversell or lose a sale.

Genuinely must work: payment authorisation. There is no useful degradation for this and nothing to design around it.

Writing this list takes an hour and it changes the architecture, because it tells you which calls belong on the critical path and which belong behind the queue. Most teams discover that almost nothing needs to be synchronous.

Timeouts, and why missing ones cause outages

An unbounded call is the commonest cause of cascading failure in a highly available Shopify integration, and it is usually a default nobody changed.

When a downstream system slows rather than fails, requests to it pile up. Each one holds a connection, a thread or a worker slot. The slot pool exhausts, and now your service is unavailable to everything, not just to the slow dependency. The dependency never returned an error, so nothing in your error handling fired.

Every outbound call needs an explicit timeout. Set it from what the call actually takes at the high end, not from what feels generous. A call that normally completes in 300 milliseconds does not benefit from a 60 second timeout; it just holds a slot for a minute while failing.

Retries need backoff and jitter. Immediate retries turn a brief problem into a sustained one, and synchronised retries across many workers arrive as a thundering herd exactly when the dependency is weakest. The standard approach is exponential backoff with randomness added so attempts spread out.

Retries need a budget. Infinite retrying of a message that will never succeed consumes capacity forever. After a bounded number of attempts the message belongs in a dead letter queue where a person can see it, with the reason attached.

And retries need a breaker above them. Backoff limits how fast you retry; it does not stop you retrying something that is comprehensively down. That is the job of a separate pattern, covered in circuit breaker patterns for Shopify APIs.

Failure domains: do not let one tenant or one job sink everything

A highly available Shopify integration should not run every kind of job through one worker pool, because that pool shares its fate across all of them. One poisonous message that takes a long time to fail, or one unusually large bulk job, consumes the pool and everything else waits.

Separate queues by work type. Order capture, stock updates, catalogue publishing and reporting exports have different urgency and different failure profiles. Order capture should never queue behind a catalogue export.

Separate by tenant if you serve several stores. One store’s traffic spike should not starve another’s. This is the difference between an integration platform that scales and one that gets a support ticket every time somebody runs a promotion.

Cap concurrency per dependency. If an ERP tolerates four concurrent writes, allow four, and let the rest wait in the queue. Pushing forty at it produces rejections, retries and a self inflicted outage. A highly available Shopify integration respects the limits of the things it talks to rather than discovering them under load.

Health checks for a highly available Shopify integration

The failure that costs a highly available Shopify integration most is the one where nothing errors because nothing ran. Credentials expired, a scheduled job was not scheduled, a worker died quietly, a webhook subscription was removed. Error alerting is silent, because there are no errors.

So measure arrival against expectation. You know roughly how many orders arrive in an hour at this time on this day. A check that compares actual against expected and alarms on a floor catches a stopped pipeline within the hour.

Reconcile counts end to end, daily. Orders in Shopify against orders in the downstream system for the same period. This is the single most reliable detector of a quiet gap, and it is a scheduled report rather than anything elaborate.

Track queue age, not just queue depth. Depth tells you how much is waiting. The age of the oldest item tells you whether anything is moving at all, which is the question that matters when a worker has stopped.

Alert on the dead letter queue being non empty. Anything that landed there needs a human. A dead letter queue nobody looks at is a list of things you have silently decided not to do.

Define what good looks like numerically. A stated target such as orders reaching the ERP within fifteen minutes gives you something to measure and something to alarm on. Objectives of this kind are what turn availability from an opinion into a number.

Deploys, credentials and the boring operational causes

In practice, most availability incidents on a highly available Shopify integration are not exotic. They are these.

Credentials expiring. Tokens have lifetimes and nothing announces their end. Track the expiry date somewhere a person reads, and rehearse rotation once while nothing is urgent.

API version deprecation. Shopify retires versions on a published schedule. An integration pinned to a version and never reviewed will break on a date that was knowable months in advance.

Deploys during trading hours. With a queue in front, a deploy drops throughput briefly rather than losing orders. Without one, it drops orders.

Configuration drift. A webhook subscription removed during unrelated app housekeeping, a changed field validation downstream, a renamed account. All invisible until volume reveals them.

No named owner. The most common root cause behind a long outage is that nobody was watching and nobody knew who should. One accountable person and a documented stand in is not engineering, and it prevents more downtime than most engineering does.

Building a highly available Shopify integration from scratch

Build a highly available Shopify integration in this order, because the first three are a few days of work and buy most of the benefit.

First, a durable queue between webhook receipt and downstream processing. Second, idempotency keys enforced by a constraint on every downstream write. Third, explicit timeouts on every outbound call, with bounded retries and a dead letter queue. Fourth, a daily count reconciliation and an arrival check that alarms on silence. Fifth, separate queues per work type. Everything beyond that is refinement.

Related reading: rate limit recovery strategies covers what to do when Shopify itself pushes back, performance bottlenecks covers where time actually goes, and real time inventory synchronisation covers the hardest consistency problem in this space.

If you are carrying an integration that falls over and you want a view on where to spend first, describe the setup and we will say which of the five above would help you most.

Straight answers

Frequently asked questions

What is the single most valuable change?

Putting a durable queue between the event and the work. It converts an outage in any downstream system from lost orders into delayed orders, which is a completely different class of problem and the one change that most improves availability.

Why does idempotency matter so much?

Because any system with retries will eventually deliver the same message twice. Without a uniqueness check, a replay creates a duplicate order, a duplicate shipment or a double stock movement. Idempotency is what makes retrying safe, so it is a precondition for everything else.

Should the storefront ever call an ERP directly?

No. A synchronous call from a checkout to a backend system makes the checkout as unavailable as the slowest thing in the chain. Write the fact of the order somewhere durable, confirm to the customer, and process onwards asynchronously.

How do we know the integration is healthy?

Measure arrival, not errors. The costly failure mode is silence, where nothing errors because nothing ran. A check that expects a volume in a window and alarms when it does not arrive catches a stopped integration; error alerting never will.

Is high availability worth the cost for a small store?

The low cost parts are, and they are most of the benefit. A queue, idempotency keys, sane timeouts and an arrival check are modest work. Multi region redundancy and hot standbys are a different budget and most stores do not need them.

Sources

  1. Shopify.dev: API reference accessed 7 October 2026
  2. Shopify.dev: Webhooks accessed 7 October 2026
  3. Shopify.dev: API rate limits accessed 7 October 2026
  4. Google SRE Book: Service Level Objectives accessed 7 October 2026
  5. AWS: Timeouts, retries and backoff with jitter accessed 7 October 2026

Fixed price, in writing

Send your brief. Get a scope and a price within 45 minutes.

  • One fixed number, agreed in writing before work starts
  • No obligation, and no pressure to sign
  • English and Arabic work, with proper right to left layout
  • One team for design, marketing, web, media and copy

Get your fixed price quote

Written scope and price within 45 minutes in business hours. No obligation.

By sending this you agree to be contacted about your enquiry. Privacy policy

Keep reading

More articles for UAE businesses

All articles
Call WhatsApp Get a quote