What you cut when the date is fixed

  • Product
  • Releases

I was moved onto a product launch about two weeks before the date. The date was not moving — it had been committed to outside engineering, which is the usual reason a date does not move.

What existed at that point was less than the schedule implied. There was no working payment integration. There was no onboarding flow around it, no customer dashboard, no handling for any path where the customer does something other than the ideal thing, and nothing at all for the support team who would have to answer for it.

So the interesting work was not writing code. It was deciding, quickly and repeatedly, what was going to exist on launch day and what was not.

Cutting the wrong things is easy

Re-scoping had already started when I arrived, and some of what was being dropped should not have been.

The clearest example: the ability to pick a checkout session back up after abandoning it. On paper that is a nice-to-have — the customer can always start again. In practice it is what stops a dropped connection halfway through payment leaving somebody with an account that exists, cannot be paid for, and cannot be re-created because the email address is already taken. That is not a missing feature, it is a support ticket per occurrence, and it costs about a day.

It has since grown the half that makes it pay for itself: an email that takes somebody back to the checkout they left. Recovering the session was the cheap, defensive version of that idea, and cutting it would have made the useful version impossible to add later without doing the same work anyway.

This is the failure mode of scope-cutting under pressure. The things that get cut are the ones that are easiest to describe as optional, and "can the customer recover from their own mistake" always sounds optional right up until the first customer makes one. Cheap resilience work is disproportionately likely to be cut, because its value is counterfactual and its cost is visible.

Some of that came down to confidence. Enough of the plan had slipped by then that there was not much appetite for anything that sounded like an addition, whatever it cost. A fair amount of what I put back went in on my own judgement, on the basis that it was a day's work and I would rather explain it afterwards than argue about it beforehand. I would do that again. It is only defensible when you are close enough to the work to know the real cost, and you have to be right.

And some cuts were correct

Plenty of the original scope genuinely could not be built in the time, and a couple of pieces should not have been built as designed at all.

Some of it needed changes far below the surface — behaviour the underlying platform does not currently have, where the honest estimate is not "two weeks" but "a quarter, and it touches other products". Those are easy calls once someone says them out loud.

One was a security problem rather than a scheduling one. The design had a one-time-code sender reachable from the public site with nothing in front of it — no authentication, no rate limit, no cooldown. Left as drawn, that is a way for anyone to send messages to arbitrary numbers at our expense: a fraud surface and a per-message bill at the same time.

The reflex is to harden it. Put a cooldown on resends, cap how many codes one address can ask for, rate limit the endpoint. The better question was what the verification was buying us, and the answer was not much — so it came out altogether rather than getting a set of controls built around it.

Deleting a feature is a harder argument to make than securing one, because hardening looks like diligence and removal looks like retreat. It is usually the cheaper answer. Something that is gone needs no rate limit, no monitoring, no alert threshold and no incident response, and it cannot be got wrong later by someone who does not know why the controls are there.

The distinction worth holding onto is that these are different decisions wearing the same clothes. "We cannot build this in time" and "this should not be built like this" both come out as a line through a requirement, and it matters a great deal which one you are making.

Not getting attached to anything

Requirements kept changing, right up to the end. That is what happens when a launch gets close and people who have been thinking about it abstractly start seeing it work.

There was no time to defend against that with process. A scope freeze needs a meeting to declare it, another to handle the first exception, and by then you have spent more of the fortnight on the freeze than on the work. So the approach was the opposite: assume the shape will change, and keep everything cheap to change.

In practice that meant not getting sentimental about a particular schema or a particular sequence of steps. Build the working version, let it be reshaped, and resist the urge to generalise early — a clever abstraction over a process that is about to change is worth less than an obvious one you can rewrite in an hour.

Where the work happens

One correction came up repeatedly, in different disguises: something that belonged on the backend had been implemented in front of it.

Order confirmation emails were the clearest case. They were being triggered from the web layer, on the basis that the web layer is what the customer just interacted with. But the customer interacting with a page is not the event you care about. The event is the payment provider confirming a payment, which arrives as a webhook, to the service that owns the order data. Trigger from the page and you send nothing when the customer closes the tab at the wrong moment, and you send twice when they refresh.

Moving that kind of thing back is not architectural purity. It is that the front end cannot tell you whether something actually happened — only that somebody was looking at a page when it might have.

Build it so anyone could have

One constraint I kept deliberately, even though it cost time: wherever the new product needed something the wider platform already does, it had to go through the same public interfaces everybody else uses, rather than reaching inside.

Under deadline that is the expensive choice every single time. The shortcut — special-case it, write directly to the table, add a flag that only this product reads — is always faster on the day.

It is also how you end up with a product only its author can maintain. A launch is not the end of the work, and the team that picks it up afterwards should not have to learn a private set of rules that exist only here. Going through the front door meant a handful of shortcuts were unavailable to me, and meant the result behaves like the rest of the estate.

The part that was not engineering

A meaningful share of the fortnight was not spent writing anything.

There was a standup with the technical contributors every morning to work out what was actually blocked as opposed to merely unfinished, a written update to stakeholders every day, and a more or less continuous conversation with individuals throughout. The point of the daily written update was less to inform than to force the question: what changed since yesterday, and does the date still hold.

The other thing worth the time was making sure QA always had something to test. A build that is three days stale is not a test environment, it is an argument generator — every bug report has to be checked against whether it still exists. Keeping the turnaround short meant the testing feedback was about the product rather than about the build.

What I took from it

Scope-cutting has a bias, and you have to correct for it. Cheap resilience work is cut first because its benefit is invisible and its cost is not. Check what has been dropped for things that are a day's work and save a class of support ticket.

Separate "cannot build in time" from "should not build". They look identical on a plan and they are completely different decisions. Conflating them means you either ship something you should not have, or permanently lose something you could have had later.

When change is certain, optimise for being cheap to change. Process that tries to stop late change costs more than absorbing it, at this timescale. Do not generalise a process that is still moving.

The deadline is the worst possible reason to skip the front door. Everything you special-case to save a day becomes something the next person cannot predict from how the rest of the system behaves.

Judgement calls need you close to the work. Putting things back without asking is defensible when you know the real cost and you carry the consequence. It is not a general licence, and it stops being defensible the moment you are guessing.