Support needs it on day one
- Internal tools
- Product
Nothing in the scope for our new product said what the support team would need. Not as an oversight anybody argued for — it simply was not there. The requirements covered what the customer would see and what the business wanted to sell, and stopped.
That is an easy thing to leave out, because the argument for leaving it out is genuinely reasonable. A new product launches with no customers. Volume builds slowly. Whatever support needs can be added once there is enough demand to know what they actually need, and in the meantime the engineers can answer the occasional question directly.
I did not agree with it, and the reason is that low traffic is not the same as low stakes.
One is enough to be a problem
On day one you might have a handful of customers. If one of them has a problem and the person answering the phone cannot see their order, cannot tell whether they were charged, and cannot see whether their number was ever set up, that is not a minor gap. It is an escalation, immediately, on the day the product launched, about the only customers you have.
The engineers-will-answer-it plan also does not survive contact with reality. It means every question goes to whoever built it, the answer takes as long as it takes them to notice, and it is delivered by reading a database. That is fine for one question. It is unworkable by the third, and the third arrives sooner than the volume projections suggest because early customers are disproportionately likely to hit edges.
The asymmetry is what settles it. Building the tooling up front costs a few days. Not building it costs nothing at all, right up until it costs an escalation about a customer nobody can help, in the launch week, which is the worst possible week for it.
Specifying it myself
Since nobody had written down what support needed, I worked from having been the person on the other end of it. I spent two years in technical support on this platform before moving to engineering, and I maintain the internal portal the support team uses now.
That turns out to be the whole trick. The question is not "what data could we expose" — it is "what is the first thing somebody types when a customer rings up, and does it work". The answer is almost always an identifier the customer can read out. So the first thing built was order search, because the order number is the thing on the confirmation email in front of the customer.
From there it is the questions that actually get asked. Did they pay, and for what — a read-only view of the subscription, enough to answer without exposing anything the agent has no business changing. Did their number get set up, and if it did not, can we make it happen now — the provisioning state, with the ability to re-run an order that stalled part-way. Are they being billed for something they have not got — the report comparing billing state against telephony state.
And the mundane one that gets forgotten: the ticketing system has to know the product exists. A customer of a new brand raising a ticket needs it routed to the right place, with the right branding on the reply, or the answer they get is addressed from the wrong company.
Read-only is a feature
Most of what I built is read-only, which was deliberate rather than a shortcut.
At launch the only two actions were re-running a stalled provisioning order and raising a ticket. Everything else answered a question. That is not because support cannot be trusted; it is that a launch-week tool built quickly is exactly the wrong place to put a write path into billing. Being able to see the state resolves most of the calls. Being able to change it adds a way to make things worse at the moment there is least time to notice.
The re-run is the exception because it is safe by construction — the order is a sequence of recorded steps, so running it again skips what already happened. An action that is idempotent is a very different proposition from one that is not.
Seeing what the customer sees
The thing that closed the last gap arrived around launch: letting an agent open the customer's own account and look at the screens the customer is looking at.
Descriptions over the phone only go so far. Someone saying their calls are not forwarding, and an agent reading a row in a database, are describing the same account in two different languages. The gap between those two descriptions is where the call time goes.
It is more awkward than it sounds, for reasons that have nothing to do with the feature. The internal portal and the customer's dashboard are different origins. A server-to-server call from one to the other cannot sign the agent in: the session cookie would be set on the portal's own response rather than in the agent's browser, and a cookie scoped to one host cannot be written for another anyway.
So it goes through the browser instead, in two hops. The portal asks the customer application, server to server, for a one-time ticket standing in for a session. It gets back an identifier and nothing else — no cookie on that response, deliberately. It then redirects the agent to a URL on the customer host, which spends the ticket, sets the cookie for the host the browser actually visited, and lands them on the dashboard.
Two things made that cheap. The redeeming half already existed for a different cross-domain sign-in, so only the minting of the ticket is new. And the ticket is single-use with a life measured in seconds, which is most of what stands between "support can help" and "an identifier in a log is an account".
The part worth saying plainly: an endpoint whose job is to create a session has no session of its own to authenticate against. It cannot inherit the protection everything around it gets for free, so its constraints have to be deliberate and written down. That is a different security posture from the rest of the tooling, and it earned more scrutiny than an internal convenience would normally get.
The part that is not software
The tooling is half of it. The other half is that nobody knows the product exists.
A new brand means the support team has never seen it, does not know what the customer bought, does not know what the dashboard looks like from the customer's side, and does not know which of their existing habits do not apply. Handing them a set of screens without that is handing them a puzzle.
So I wrote the guide as well: what the product is, where it appears in the portal they already use, what each screen is telling them, and what to do about the cases that have an answer. Writing it also found gaps — explaining how an agent would handle a scenario is a quick way to discover there is no way to handle it.
What I took from it
Low volume is not low stakes. The argument for deferring internal tooling is always about traffic, and traffic is the wrong measure. The cost of one unanswerable customer does not scale down with how few of them there are.
Start from the identifier the customer can read out. Tooling designed from the data model gives you lookups by the wrong key. The first thing built should match the first thing said on the phone.
Read-only answers most of it. Visibility resolves the majority of calls without adding a way to make things worse. Add write paths deliberately, starting with the ones that are safe to repeat.
A session-minting endpoint is its own security boundary. It cannot inherit authentication from the session it exists to create, so whatever constrains it has to be explicit. Single use and a short life do more work than they look like they do.
Write the guide, and let it find your gaps. Describing how somebody would handle a case is the cheapest way to notice that they could not.