An order that can be run again
- Provisioning
- Distributed systems
The product is deliberately simple to buy. One page, one button, card details, done — and a working phone number on the other side as quickly as we can manage it. There is no configuration step, no options, nothing to choose.
That simplicity is entirely on the customer's side. Behind the button is a sequence of things that have to happen in order, against several systems, some of which are not ours. A number has to be allocated from a range. An account has to exist to attach it to. The telephony platform has to be told about both. Mailboxes and defaults have to be set up so the thing actually answers when somebody rings it.
Every one of those can fail, and some of them fail in ways that have nothing to do with us. So the design question is not how to make the sequence reliable. It is what the system does when part of it does not work.
Steps, not a function
The thing that makes this tractable is refusing to treat provisioning as one operation. It is a list of named steps, each recorded as it completes.
That sounds like bookkeeping overhead until the first failure. If provisioning is a single function that either worked or did not, a failure three-quarters of the way through leaves you with no idea which quarter is outstanding. You cannot re-run it, because the parts that succeeded will be attempted again, and allocating a second number to somebody who already has one is worse than the original failure. So you end up unpicking it by hand, from logs.
With the steps recorded, re-running is the obvious move: start at the top, skip what is already done, carry on from the first thing that is not. The order matters and it has to be deterministic, because a re-run that executes the steps in a different sequence is a different operation.
Splitting them finely is worth it too. A step that does two things is a step you cannot half-complete, which means any failure inside it re-does both.
Nothing tells you it stalled
The failure I care most about is not the one that throws. It is the one where an order simply stops.
A step calls something external, the call does not come back, nothing raises, and the order sits at step four indefinitely. Nobody finds out. The customer has paid, the subscription is active, and from the billing side everything looks correct — because from the billing side it is. There is no error anywhere, and no alert, because nothing failed. It just did not finish.
The answer is not cleverness, it is asking the question. There is a query for orders that started and have not finished within a sensible window, it is exposed so support can see it, and the same orders can be pushed back onto the queue from the same place. Most of the value of that is not the requeue button. It is that somebody finds out on the day rather than when the customer rings up.
The two halves have to agree
The related problem is slower and quieter. Payment and telephony are two systems with two opinions, and they can drift.
A subscription is active but the number was never provisioned. A subscription has been cancelled but the service is still live. Neither system is wrong about its own half. There is no error on either side, and nothing brings the two views together unless something is built specifically to do it.
So something is: a report of accounts whose billing state and telephony state disagree. It is not sophisticated — it is a comparison — but it catches the class of problem that produces no signal at all. A customer paying for something they do not have, or holding something they stopped paying for, are both bad and neither announces itself.
Fast, and honest when it is not
The other half of a one-button purchase is that the customer is watching. They have just paid and they are waiting for a number, so the sequence should finish in seconds rather than minutes, and normally does.
When it does not, the temptation is to hide it — keep a spinner going, show a number that is not quite live yet, tell them it is ready and hope the last step lands. That converts a delay into a support ticket, because the customer will try to use it.
An order that is still in progress says so. An order that stalled is visible to the people who can do something about it before the customer has to explain it to them. Being quick in the normal case and truthful in the abnormal one is a better trade than appearing quick in both.
What I took from it
Record the steps, not the outcome. A process that can only report success or failure cannot be resumed, and anything that cannot be resumed gets finished by a person reading logs.
Design the re-run before you need it. Skipping completed work and running in a fixed order are what make a retry safe. Retrofitting that onto a process that already has live orders in it is considerably less pleasant.
The stall is worse than the error. Failures announce themselves. Something that stops halfway produces no signal at all, so you have to go and ask — a query for things that started and never finished is a small amount of code for a whole category of problem.
Two authoritative systems need a third thing comparing them. When each is correct about its own half, disagreement between them is invisible from inside either one.