Here's the failure: a checkout.session.completed event reaches your endpoint. The handler marks the order paid, calls the fulfillment service, renders and sends a confirmation email, and then runs past the provider's response window before it returns 200. Stripe logs that delivery as failed and schedules a retry. In live mode it will keep retrying with exponential backoff for up to three days. The second delivery goes through the same handler and succeeds, so the customer gets two emails, fulfillment ships twice, and the order has two 'paid' transitions in its audit log. Nothing in this sequence is exotic. It's the default outcome of payment integration in website projects that treat the gateway as a form and the webhook as a notification.
My position, after eleven years of shipping web platforms at EltexSoft, much of it on Laravel, React and PostgreSQL with Stripe and Stripe Connect behind it: payment integration is a distributed state machine that happens to have a checkout page attached. You take the payment-provider event stream as the source of truth. You make every write on both sides idempotent. You record money in a ledger rather than in a status column. You reconcile every night. What follows goes through the usual consensus one claim at a time and shows where each claim fails.
Claim 1: Hosted Checkout Means Payments Are Solved
The true part: a hosted page like Stripe Checkout, or an embedded field like Payment Element, keeps card data off your servers. That typically puts you in the lightest PCI self-assessment tier (SAQ A for a full redirect). It also gives you 3D Secure, wallets and localised payment methods without writing any of it yourself. Use it. There is almost never a good reason to handle raw PANs in a website integration.
Where it fails: hosted checkout solves card capture and nothing else. Your application still has to answer questions the hosted page can't. Is order 48213 paid? How much of it was refunded? Is there an open dispute? Did the SEPA debit settle, or did it bounce four days later? Those answers live in your database, and they only get there through the integration you write. Across the payment work we've shipped, the checkout UI has been the smallest line item. Most of the engineering time goes into webhook handling, refunds and disputes flowing back into product state, and reconciliation.
Claim 2: The Success Redirect Means the Customer Paid
The redirect to success_url is a browser navigation, and the payment state lives elsewhere. The customer can close the tab after 3D Secure finishes but before the redirect lands. Mobile browsers kill background tabs. A flaky connection drops the GET. In any of those cases the payment succeeded and your success handler never ran. If fulfillment hangs off that handler, you now have a paid order your system doesn't know about.
The reverse failure is worse. Asynchronous payment methods such as SEPA Direct Debit, ACH and some bank redirects complete the Checkout Session with payment_status set to unpaid. Settlement arrives later as checkout.session.async_payment_succeeded, or as async_payment_failed. A success page that says 'thanks, you're all set' and triggers provisioning at that point is provisioning against money that may never arrive.
Our rule is that the success page is read-only. It fetches the session, shows the current state (paid, processing, or failed) and polls or subscribes for changes. Fulfillment is triggered only from the webhook path. That puts all state-changing logic on a single code path, which is also the path you can retry and test.
Claim 3: Webhooks Are Notifications You Handle Inline
Stripe documents that delivery is at-least-once and unordered. The same event can arrive more than once. payment_intent.succeeded can land before checkout.session.completed. A charge.refunded can reach you while an earlier event for that charge is still in your retry queue. A handler that assumes exactly-once, in-order delivery works in development, because development traffic is one event at a time. It fails in production under concurrency.
The handler we build has four steps, in this order. First, verify the signature against the raw request body. Second, insert the event ID into a processed_events table with a unique constraint, using something like INSERT ... ON CONFLICT (event_id) DO NOTHING in PostgreSQL. If no row was inserted, return 200 and stop. Third, enqueue a job carrying the event ID. Fourth, return 2xx. All the real work happens in the queued job: status transitions, email, fulfillment. A slow SMTP server can't push the endpoint past the response window, and a job that fails is retried by your queue, not re-sent by the provider.
Signature verification is where the first production bug usually appears, and the error message is precise: 'No signatures found matching the expected signature for payload.' Nine times out of ten, a body-parsing middleware ran before the verifier. In Express, a global express.json() re-serialises the body, and a single whitespace difference breaks the HMAC. The fix is express.raw({ type: 'application/json' }) scoped to the webhook route. In Laravel, read $request->getContent() rather than the parsed input, and exclude the route from CSRF middleware. Keep the default 300-second timestamp tolerance. It's there to block replay of captured payloads.
Out-of-order delivery is handled in the job, not the endpoint. Treat the event as a signal that something changed, then re-fetch the object from the API (the PaymentIntent, the Session, the Charge) and apply its current state. Guard every write with an allowed-transitions check: succeeded can move to refunded or disputed, but nothing moves from succeeded back to processing. Re-fetching costs one API read per event. Stripe's default live-mode limit is on the order of 100 operations per second, while test mode is considerably lower. That's plenty for almost any website's event volume, and it removes a whole class of ordering bugs.
Claim 4: Idempotency Keys Are the Gateway's Problem
Deduplicating inbound webhooks covers half the surface. The other half is your own outbound calls. Consider this sequence. Your server calls create PaymentIntent (or create Refund, or create Transfer on Connect), and the request reaches Stripe and is executed. The response is lost to a network timeout. Your HTTP client, or your queue worker, retries. Without an idempotency key, the retry creates a second object. For a refund or a Connect transfer, that is real money moving twice.
Every mutating call we make to a payment API carries an Idempotency-Key header derived from our own domain identifiers, for example refund:{order_id}:{refund_request_id}, and never a random UUID generated inside the retry loop. A UUID generated inside the loop is new on every attempt, so it deduplicates nothing. Two properties of the key contract matter in practice. Stripe keeps keys for at least 24 hours, so a retry scheduled days later isn't protected by the key and needs its own check in your database first. And reusing a key with different parameters returns an error rather than silently replaying the call, which surfaces bugs where the amount was recalculated between attempts. Every PR at EltexSoft is reviewed by at least one other senior engineer before merge. On payment code, 'where does the idempotency key come from' is one of the first questions a reviewer asks.
Claim 5: A Status Column on the Order Is Enough
orders.status = 'paid' holds up until the first partial refund. After that you need to know that the order was paid 120.00, refunded 30.00, has an open dispute on the remainder, and that the Connect application fee was partially reversed. A single enum can't represent that. Teams usually respond by adding columns: refunded_amount, disputed, then refund_count. Each new column is another value that can drift from what the provider actually holds.
The model we default to is an append-only payment_transactions table. Each row is one money movement: charge, refund, dispute, dispute reversal, fee, transfer. It carries the provider object ID (with a unique index, which gives you idempotency a second time at the storage layer), a signed amount in integer minor units, a currency, and the provider timestamp. Order state is derived from it: the net captured amount, whether a dispute is open, whether the order is fully refunded. Cache the derived value on the order row if list views need it. The ledger remains the source.
Integer minor units aren't optional. In JavaScript, 19.99 * 100 evaluates to 1998.9999999999998, and a Math.floor on that charges one cent less than intended. Zero-decimal currencies like JPY take the amount as-is, so a currency-aware conversion function, called in exactly one place, prevents a 100x overcharge on your first Japanese customer. The ledger costs a day or two up front. Retrofitting it after real refunds and disputes have started accumulating in ad-hoc columns takes weeks.
Claim 6: If It Works in Test Mode, It Works
Test mode exercises your happy path against a provider that behaves unrealistically well. It doesn't reproduce issuer-specific 3D Secure behaviour, realistic webhook latency under load, live-mode rate-limit headroom, dispute timelines measured in weeks, or the restricted-key permissions you should be using in production. A suite that only ever sees one clean event per checkout says little about how the handler behaves under concurrency.
What we test deliberately:
Duplicate delivery. Resend the same event with the Stripe CLI (stripe events resend) and assert that nothing changes the second time.
Reversed order. Feed the job payment_intent.succeeded before checkout.session.completed, and charge.refunded before the charge's success event.
Asynchronous failure. Use the test payment methods that settle to failure and assert that provisioning is rolled back or never happens.
Outbound timeouts. Inject a timeout after the provider call, retry, and assert that exactly one object exists.
stripe listen --forward-to and stripe trigger make these tests cheap to wire into CI. Most of them can also run as integration tests against recorded fixtures without a network. Our QA practice keeps these as permanent regression tests. The payment code path doesn't change often, but when it does, the change tends to be large.
Claim 7: Reconciliation Is Accounting's Job
Even with everything above in place, events get missed. Someone deploys a broken handler on a Friday. A DNS change sends the endpoint to the wrong host. A firewall rule drops the provider's IP range. The provider will retry, and if failures persist for days it will warn you and can disable the endpoint. The retry window is finite, and after it closes, the missed events aren't delivered again. Without a check against the provider, those orders sit in the wrong state until a customer complains.
A nightly job closes the gap. Page through the provider's balance transactions (or PaymentIntents and refunds) for the previous 48 hours, the overlap guarding against boundary timing. Diff them against the ledger by provider object ID. Missing rows are pushed through the same queued job the webhook path uses, not written directly, so there is still only one code path that changes state. Mismatched amounts raise an alert instead of being auto-corrected, because an amount mismatch means a bug, and hiding it makes the next one harder to find. At volume, this job is what keeps the integration correct. On a marketplace like the one we took over for Nautical Commerce (a one-year-old codebase, delivered against a 90-day go-live, and since then processing 200K+ transactions a month), a 0.1% miss rate would mean around 200 wrong orders every month. At that level the reconciliation job belongs to the product's correctness, not to finance.
What It Costs to Build Correctly
For a single-merchant website on hosted checkout, the full design (a read-only success page, a deduplicating webhook endpoint with queued jobs, idempotency keys on outbound writes, a transactions ledger, and nightly reconciliation) comes to about two to four engineer-weeks with a senior developer who has done it before. The range is driven mainly by two things. One is how many asynchronous payment methods you accept, since each adds pending and failure states to the state machine. The other is whether refunds and disputes have to change product state, such as revoking access or restocking inventory, or only need to be recorded. Stripe Connect or any split-payment marketplace roughly doubles that estimate. Transfers, application fees, connected-account onboarding events and payout failures each add their own event streams and their own reconciliation.
The version that skips these pieces ships in a week and looks fine in the demo. The missing work shows up later as support tickets, double shipments and manual refunds, and each one has to be traced by hand through logs that were never designed to answer the question.
This is the kind of integration our engineers build and review day to day, usually in Laravel or Node with PostgreSQL behind it and Stripe or Stripe Connect in front. If you have a payment flow in production and aren't sure how it behaves under duplicate or out-of-order events, a discovery week is a reasonable way to find out. We run the resend and reorder tests against a staging copy and share what breaks.
Last updated September 30, 2026