Real time inventory synchronisation: the hardest consistency problem in ecommerce integration
Source of truth, out of order updates, available versus on hand, reservations, and the reconciliation loop that catches what events miss.
Read the articleShopify Integrations
Why retrying a dead dependency amplifies the failure, the three states, choosing thresholds you can defend, fallbacks worth having, and testing the breaker before you need it.

A circuit breaker for the Shopify API exists to answer one question that retry logic cannot: should anybody be calling this at all right now. Retries decide how patiently a single caller waits. A breaker decides whether the call is worth making, and that distinction is what prevents a dependency failure from becoming an outage of your own.
This is the pattern, the decisions it requires, and the mistakes that make it fire at the wrong times.
Consider a system with no circuit breaker for the Shopify API. A dependency starts failing. Every request to it times out after thirty seconds, then retries three times with backoff.
Each unit of work now occupies a worker for two minutes instead of a second. The worker pool fills with jobs waiting on something that will not answer. Work for healthy dependencies queues behind them. Your service is now unavailable for everything, caused by one failing component.
Meanwhile the failing dependency is receiving more traffic than usual. Everybody is retrying. If it was struggling under load, the retries hold it down, and a system that might have recovered in thirty seconds is held under for as long as the retries continue. This is the textbook cascade, and it is common precisely because each individual retry looks reasonable.
A circuit breaker for the Shopify API cuts both problems at once. It stops your workers waiting on something that will not answer, and it stops you contributing traffic to a system that needs quiet to recover.
A circuit breaker for the Shopify API has three states. Closed. Normal operation. Calls pass through and outcomes are recorded. If failures cross the configured threshold, the breaker opens.
Open. Calls are rejected immediately without being attempted. No waiting, no timeout, no load on the dependency. The caller gets an instant failure and takes whatever fallback path exists. After a cooling period the breaker moves to the third state.
Half open. A small number of trial calls are allowed through. If they succeed, the breaker closes and normal service resumes. If they fail, it opens again and the cooling period restarts.
The half open state is what makes this self healing, and it is the part implementations get wrong. Letting all queued traffic through the moment the cooling period ends slams the recovering dependency and reopens the breaker immediately, which produces an oscillation that looks worse than a steady outage. Allow a trickle, and only widen when it succeeds.
What trips a circuit breaker for the Shopify API determines whether it helps or interferes.
Trip on timeouts. A dependency not answering within its budget is the clearest signal, and it is the failure mode that consumes your resources.
Trip on server errors. Repeated errors from the far end mean something is wrong there, and hammering it will not help.
Trip on connection failures. Refused or unreachable means there is nothing to talk to.
Do not trip on throttling. Being told to slow down is the platform working correctly and telling you something useful. That is a pacing problem handled by rate limit recovery, and a breaker that opens on throttled responses will fire during every busy period and block work that would have succeeded moments later.
Do not trip on client errors. A malformed request or a missing record is your bug or your data, and it will fail identically on every attempt. These belong in a dead letter queue, not in the breaker’s failure count, where they would open it against a perfectly healthy dependency.
That distinction, between the far end being unwell and your request being wrong, is the single most important thing to get right in this pattern.
Copying numbers from an example is how a circuit breaker for the Shopify API ends up either useless or hostile.
Use a proportion over a window, with a minimum volume. A rate such as half of the last twenty calls failing is meaningful. A rule based on consecutive failures alone trips on a quiet dependency that received three requests, two of which happened to fail.
Set the timeout from measurement. Look at what the call actually takes at the high end in normal conditions and set the budget a little above it. A timeout far longer than the realistic worst case is the same as having none, because the worker is held for the duration either way.
Make the cooling period long enough to be useful. A few seconds does not give anything time to recover. Tens of seconds is a reasonable starting point, and it should be long enough that the dependency gets genuine quiet.
Review the numbers against real incidents. After the first time it fires, ask whether it fired early, late or correctly, and adjust. A circuit breaker for the Shopify API configured once and never revisited is usually wrong in one direction or the other.
An open circuit breaker for the Shopify API fails fast, which already beats hanging. Failing fast with something useful behind it is better.
Queue it, if it can wait. Most write operations can. An order that cannot reach an ERP right now can sit in a queue and arrive in ten minutes. This is the best fallback and it is available for the majority of integration work.
Serve a cached value, if staleness is acceptable. A stock figure from five minutes ago is usually better than an error, and the business can tell you what the acceptable age is.
Degrade the feature, not the page. If a recommendation panel depends on a failing service, hide the panel. Do not fail the page that contains it.
Say something honest, if nothing else works. A clear message that this part is temporarily unavailable is better than a generic error, and far better than a spinner that never resolves.
Decide these per dependency and write them down. The decision is commercial rather than technical: how wrong may this be, and for how long, is a question for the business.
A circuit breaker for the Shopify API stops bad calls. A bulkhead stops one dependency’s problems consuming resources that others need, and the two patterns work together.
Separate breakers per dependency. The ERP, the courier, the tax service and the platform itself each get their own. One shared breaker means a failing courier blocks order capture, which is the opposite of what you want.
Separate resource pools too. If every outbound call draws from one connection pool or one worker pool, a slow dependency starves the others regardless of the breakers. Capping concurrency per dependency is what makes the isolation real.
Name the breakers in your monitoring. A dashboard that shows which breaker is open turns an incident call from guesswork into a sentence. This is a meaningful operational benefit of the pattern beyond the resilience itself.
An untested circuit breaker for the Shopify API is a configuration file, not a safety mechanism. Three tests are worth running.
Force it open. Point a dependency at something that refuses connections, confirm the breaker opens after the expected number of failures, and confirm the fallback behaves as designed. Most teams discover at this point that the fallback was never implemented.
Confirm it closes again. Restore the dependency and watch the half open trial succeed and the breaker close. Verify it does not flood the recovering dependency with the whole backlog at once.
Introduce latency rather than failure. This is the case that catches people. A dependency that is slow rather than dead produces no errors, so the breaker never trips unless the timeout is doing its job. If a slow dependency can still exhaust your workers, the timeout is too long.
Run these in a test environment, write down what happened, and repeat after any significant change. The point of the exercise is that the first time the breaker fires should not be the first time anyone has seen it fire.
A circuit breaker for the Shopify API is one layer among several, and they solve different problems.
Timeouts bound how long one call may take. Retries with backoff handle transient failures. A circuit breaker for the Shopify API handles sustained failures by stopping the calls. Bulkheads stop one failure consuming shared resources. Queues mean that work deferred during all of this is not lost.
Implementing the breaker without the queue means you fail fast and drop the work, which is faster but no more correct. Implementing the queue without the breaker means you fill the queue with work for a system that cannot take it. They belong together, and the architecture that holds both is described in designing highly available Shopify integrations.
Related reading: performance bottlenecks and real time inventory synchronisation. If an integration of yours falls over when something else does, describe what happens and we will say which of these layers is missing.
Straight answers
Backoff controls how fast one caller retries. A breaker decides whether anyone should be calling at all. When a dependency is comprehensively down, every caller backing off politely still produces a stream of doomed requests; the breaker stops them and lets the dependency recover.
No. Being told to slow down means the platform is healthy and you are asking too fast, which is a pacing problem rather than a failure. Trip on timeouts and server errors. Treating throttling as a failure will open the breaker during normal busy periods.
Fail fast and do something useful. Queue the work for later if it can wait, serve a cached value if a stale answer is acceptable, or return a clear message if neither applies. Failing fast with no fallback is still better than hanging.
From your own measurements rather than from a blog post. Look at the normal error rate and normal latency, then set the trip condition above the noise and the timeout just above the realistic high end. Thresholds copied from elsewhere either never fire or fire constantly.
One per dependency, at least. A single breaker covering every outbound call means one failing system stops calls to healthy ones. Separate breakers also make the dashboard tell you which dependency is the problem.
Sources
Fixed price, in writing
Got it. Your quote is being written now.
In business hours you will have it within 45 minutes. Check your inbox for the confirmation.
Keep reading

Source of truth, out of order updates, available versus on hand, reservations, and the reconciliation loop that catches what events miss.
Read the article
What each stage produces, which one overruns, and the handovers to insist on before paying the next invoice.
Read the article
Queues over direct calls, idempotency as a requirement, graceful degradation, and health checks that notice silence rather than errors.
Read the article