July 24, 2026 · 7 min read

An API integration doesn't end when it returns 200 OK

How to design reliable API integrations for retries, idempotency, webhooks, rate limits, monitoring, and recovery beyond the happy path.

← Back to blog

Connecting an API can give a misleading sense of progress: we make a request, the data arrives, and the server returns 200 OK. All this proves is that two systems can talk under normal conditions. A working API isn't the same as a working integration.

If you need to apply these decisions to real systems, we also provide API integration services that connect existing tools without replacing them by default.

The difference shows up when we stop testing the happy path.

200 OK only says that one request went well

A successful response confirms something pretty limited. The server received a request and decided to respond correctly. It doesn't necessarily mean that:

It can even happen that an API responds correctly and the real process keeps going asynchronously behind the scenes. From our side, everything looks finished. On the other side, it just started.

The easy errors are the least interesting ones

A 500 is relatively simple: something failed, so we log it, retry it, or show an error. The trickier problems are the ones that look like successes.

Say we send an order to an external system. The API responds 200. But the associated product no longer exists. The system accepts the order, leaves it in an intermediate state, and an hour later an internal process fails.

Our integration already marked the operation as completed. Now we have two systems with different versions of reality. That kind of failure is a lot harder to catch than a straightforwardly wrong response.

You need to define what success means for each operation

Before integrating, it's worth defining what result you actually need to confirm. It could be: "the server received the request." Or:

"the record was created." Or: "the payment was processed." Or: "the order reached the fulfillment system and can now be prepared." Those aren't the same thing. An HTTP response only covers part of the flow. The integration needs to know when an operation can actually be considered done.

APIs have limits

Many services limit how many requests we can make in a given period. During development, with ten products, twenty users, and five orders, that almost never matters.

Five orders. In production, a hundred thousand records show up and the same strategy stops working. 429 Too Many Requests starts appearing. That's when we need to think about:

An integration designed around test data can behave completely differently once real volume shows up.

Retrying doesn't mean blindly repeating

If a request fails due to a timeout, the natural reaction is to send it again. But there's an important question:

do we know the first operation didn't happen?

This could have happened:

  1. We send a request
  2. The external server processes the operation correctly
  3. The response gets lost due to a network issue
  4. Our system interprets it as a failure
  5. We send the exact same thing again

Now we have two operations. If it was a read, probably nothing happens. If it was creating a payment, an order, or a subscription, we have a different problem.

Idempotency

Idempotency lets you repeat an operation without the result getting duplicated. Some APIs offer idempotency keys. We can send something like:

order-5821 and even if the request arrives twice, the service recognizes that both represent the same operation. When the API does not offer that mechanism, you may need to store external identifiers and log processed operations on your side.

Check states before creating new records. The exact mechanism changes. The idea doesn't. An integration should be able to survive the same operation happening more than once.

Webhooks aren't perfect messages either

Webhooks are useful for reacting quickly when something changes. If the transport is still an open decision, this comparison of webhooks and polling covers their costs and failure modes. But we should not assume they will arrive exactly once, in the right order, immediately, or at all. A provider might resend a webhook because our server took too long to respond. Two events might arrive reversed.

A temporary outage can make a notification show up several minutes later. That's why a webhook should communicate that something changed, but the real state should be verifiable against the original source when needed.

Order can break things

Imagine two events:

  1. Order created
  2. Order canceled

For some reason they arrive at the receiving system in reverse order. We process the cancellation first. Then the creation arrives. If we apply the events without validating anything, we end up with an active order that's canceled at the source.

Integrations that manage state need to consider that a network doesn't necessarily guarantee information arrives in the same order it happened. Sometimes timestamps are enough. Other times we need versions, sequences, or to re-query the object before updating it.

Syncing doesn't always mean copying everything

A tempting strategy is to request all the records every so often and compare. That can work with a hundred items. With hundreds of thousands, it starts getting expensive. A mature integration usually needs some way of knowing what changed.

Changes can be detected by modification date, cursor, events, incremental pages, or processed identifiers. The goal is to avoid repeatedly checking information we already know has not changed.

Logs should tell a story

Storing:

API error doesn't help much. A good log should let you reconstruct what happened. Which system started the operation. Which entity was involved.

What we tried to do. What local identifier it had. What external identifier we received. When it happened. What the service responded. How many times it was retried. That doesn't mean storing sensitive information without control. It means recording enough context to understand the process.

When an integration fails three months after being built, that context is worth a lot more than a generic error description.

We also need to know when something stopped working

Some failures produce errors. Others simply stop producing activity. A process that normally syncs a thousand records a day can start syncing zero. There's no 500.

There's no exception. It just stopped doing what it was supposed to do. That's why some integrations need monitoring based on behavior, not just on errors. When was the last sync?

How many records did we process? How many failed? Is a queue growing? Has it been too long since the last webhook?

A silently stalled integration can be worse than one that fails visibly.

You have to think about how it gets fixed

Detecting the problem isn't enough. You also need to be able to recover the system. Reprocess an order. Re-run a date range.

Recovery might involve resending an operation, draining a queue, re-syncing an entity, or fixing a relationship before continuing.

If the only way to fix an integration is manually editing the database, you don't have much recovery capacity. The internal tools for operating an integration rarely show up in the demo. They're the ones that end up mattering once the system has been running for months.

APIs change

APIs release new versions, deprecate fields, retire authentication methods, and remove endpoints.

Changes to usage limits. An external integration always depends on decisions we don't control. That's why it's worth knowing which version we're using, following the provider's announcements, and avoiding reliance on undocumented, accidental behavior. If an API is critical to the business, maintaining it is also part of the product.

The best integration isn't the one that never fails

That integration does not exist: external services go down, networks fail, data arrives malformed, credentials expire, and people change settings. What matters is what happens next. Do we know it failed?

Can we understand why? Can the operation be retried? Is there a risk of duplicating it? Can we recover the correct state?

That's where the real quality of an integration starts to show. Getting the first 200 OK can take an afternoon. Designing how to detect, understand, and recover from failures is what turns that connection into an integration the team can operate.