Nobody Priced the Environment You Already Have
Three things get said in three different meetings at the same firm, usually within a fortnight of each other.
The investment team says onboarding a new manager takes a quarter, and asks why. The controller says the break list never gets to zero, and asks for another pair of hands. The COO says operations headcount has grown every year the book has grown, and asks whether that is really necessary.
Nobody connects them, because they arrive separately, from different people, in different forums, and each one has a plausible local answer. Hire someone. Chase the custodian. Push the manager for better files.
They are one problem. The firm has reached the ceiling of what its investment operations environment can carry, and each of those three complaints is the ceiling pressing on a different part of the organization.
The reason nobody fixes it is not that the fix is unknown. It is that the business case gets built against zero.
Legacy was not a mistake
Start here, because the conversation dies if you skip it.
The environment you are running was almost certainly a sound choice when it was chosen. It fit the book you had. Somebody did the work, ran the selection, made the case and signed for it, and they were right.
What changed is the shape of the book, not the quality of that decision. A general account that was ninety percent public fixed income in 2015 now carries private credit, structured product, commercial mortgages and a derivatives overlay. Each of those arrived one allocation at a time, each was absorbed by operations without anyone declaring a systems problem, and each added a class of data the environment was never designed to hold.
That is not a failure of foresight. It is what growth looks like from the inside. But it means the argument for replacing the environment cannot be that it is bad, because the person you are making the argument to is the person who bought it. The argument has to be that the book outgrew it, and that the outgrowing has a price.
Three diagnostics
Each of these is a question with a measurable answer. Not an opinion, not a maturity score, a number you can go and get this week.
How many elapsed weeks from decision to first clean valuation? Note that this is elapsed time, not effort. Effort is what shows up on a timesheet and it is usually modest. Elapsed time is what the capital experiences, and it is usually a multiple of the effort, because the work sits in queues waiting on someone else. Measure the calendar, not the hours.
What share of reference data breaks reach valuations, payments or client reports before anyone catches them? Most firms can tell you their break count. Very few can tell you the leakage rate, which is the number that actually matters, because a break caught in the queue costs an hour and a break caught by a client costs a great deal more than an hour.
How many people are on the exception queue, and has that queue become somebody's job title? The second half of that question is the diagnostic. A rotating duty is a process with friction. A permanent role is a process the firm has given up on and started staffing.
If you cannot answer all three, that is itself the finding. An environment nobody measures is an environment nobody is managing.
The four costs
Here is the part that gets skipped, and skipping it is why these projects lose.
A modernization business case almost always compares the cost of a new platform to nothing at all. License, implementation, integration, some contingency, set against a savings number that the finance function correctly regards as soft. The status quo enters the comparison at zero, because nobody ever priced it.
The status quo is not free. It has four costs, and they are separable.
Exception labor. The share of operations time absorbed by breaks, rekeying and chasing reference data. This is the visible one and usually the only one anybody counts.
Onboarding delay. Elapsed weeks per onboarding, times how often onboarding happens. Two components: the operations effort itself, and the spread given up while capital waits on operations to be ready.
Remediation. What bad reference data costs once it has already reached a valuation, a payment or a client report. Restatements, client credits, audit findings, and the labor to unwind them. Low frequency, high severity, and almost never in anybody's model.
The ceiling. Positions per FTE as a hard constraint. In a legacy environment, headcount grows roughly with the book. The cost is the hiring that growth forces on you.
Run those four on an illustrative mid-size insurer. Eight billion of assets, thirty five thousand positions, twelve operations FTE, ten percent position growth, forty five percent of operations time on exception work, six onboardings a year at fourteen weeks elapsed:
| Cost line | Annual |
|---|---|
| Exception labor | $783,000 |
| Onboarding delay | $679,038 |
| Remediation | $209,150 |
| The ceiling | $169,167 |
| Cost of current state | $1,840,355 |
Two point three basis points of assets. Fifty two dollars and change per position, per year.
The ceiling line deserves a word, because it is smaller than people expect. It charges the average staffing the year actually demands, not the staffing the book demands on the thirty first of December. Growth that forces a fourteenth hire in October is not a fourteenth salary for twelve months. The year-end run rate on that same line is $290,000, which is the right number for a headcount plan and the wrong number for this year's cost.
That is the number a platform has to beat. Not zero.
One discipline matters here, and getting it wrong is how these arguments get dismissed in the first meeting. This is not your total operations cost. A team you would employ regardless is not a cost of the current state. Only the portion attributable to the environment belongs in the comparison, and if you overstate it, someone will find the overstatement and stop listening to the rest.
I know that because I overstated it myself, in the first version of the worksheet that accompanies this article.
What survived scrutiny
The worksheet went through three rounds of it, twice from me and once from somebody reading it as a chief financial officer trying to disprove the case rather than approve it. Every round made the number smaller. That is the direction corrections run when you let somebody attack them.
The first version said the five year NPV was $2.79 million. It now says $906,998. Two thirds of the original case was never real, and every piece of it was removed by an argument I would rather hear from a colleague than from a committee.
Round one. A soft assumption doing too much work. The original onboarding line valued deferred capital at the full net investment spread. Seventy five million waiting fourteen weeks at a hundred and twenty basis points, six times a year. That produced a cost of current state of $2.87 million, of which $1.45 million, slightly over half, came from one input. A number with fifty one percent of its weight on a single soft assumption is not an argument. It is a target.
Capital waiting on operations is rarely sitting in cash. It is parked in something liquid and lower yielding, or the mandate simply starts later. The honest input is the difference between what it earns while it waits and what it would earn once deployed, not the whole spread.
Round two. Benefits that arrived before the thing was built. The model charged the entire implementation cost in year one and handed the firm a full year of modernized operating economics in that same year, while the input sheet plainly said the implementation took fourteen months. Both cannot be true. For fourteen months you are still running the old environment, still carrying all four of its costs, and paying the platform fee on top.
Benefits now phase in rather than switch on. Breakeven moved from year three to year four.
This is the most common way a modernization case gets built wrong, and it is rarely deliberate. Spend is easy to schedule because somebody invoices you for it. Benefit is not, so it quietly gets booked from the date of signature rather than the date of go live.
Round three. Two places the model was inflating quietly. Both came from the outside reading, and both are worth knowing because they are easy to commit and hard to spot.
The ceiling line took the year-end position count, worked out the headcount it required, and charged that headcount for the whole year. You do not need the fourteenth person on the first of January. You grow into them. Charging mean staffing rather than the year-end requirement took roughly $483,000 out of the status quo across five years, all of which had been flattering the case.
The second is subtler and I think more instructive. Remediation grew with the book, which is right for the routine unwind labor because break volume scales with positions. But the severity events were inside the same line. So a user who entered two audit or restatement events a year found the model quietly using more than two by year five, purely because the book had grown sixty one percent. If somebody types a frequency into a box, the model has no business overriding it. Routine remediation now scales and severity does not.
One thing that is not a correction but a design choice. Sixty nine percent of the first sensitivity grid returns a negative answer. Enter a modest book, a small exception queue and a large platform cost and the worksheet tells you plainly that your current environment is cheaper.
A worksheet that can only say yes is a sales sheet. The value of pricing the status quo is that sometimes the status quo wins, and you want to find that out in a spreadsheet rather than eighteen months into an implementation.
The second grid attacks the onboarding assumption directly, across the spread given up and the elapsed weeks you expect to achieve afterwards. It runs from negative $195,825 to positive $3.33 million, and one column in it is worth the whole exercise. If modernization does not shorten onboarding at all, if you still take fourteen weeks afterwards, the case comes out at negative $195,825 and it does so at every single spread assumption. The opportunity cost of waiting capital contributes nothing whatever unless the waiting gets shorter.
That is the sentence to take into the vendor meeting. Not how much the spread is worth. Whether the elapsed time actually falls.
Why it arrives sooner for some asset classes
The ceiling is not a function of assets under management. It is a function of how much reference data each position demands and how little of it arrives in a usable form.
Private credit. No clean security master, no standard identifier, and cashflows that arrive as a PDF from an agent. Every position is a small ongoing data maintenance obligation rather than a one time setup.
OTC derivatives. The documentation and the lifecycle events are the position. Resets, amendments, novations, collateral. A system that models the trade but not its life leaves the difference to a person.
Structured product. Factors and paydowns have to be right before anything downstream can be. When they are late, the error does not sit still. It propagates into valuation, into effective duration, into the hedge ratio.
A firm that added any of these to a book the environment predates will hit the ceiling earlier than its asset total suggests, and will be puzzled about why, because the peer group it benchmarks against is holding something simpler.
Where the ceiling actually binds
Roughly twenty thousand to three hundred thousand positions.
Below twenty thousand, manual process absorbs the work. It is inelegant and it is fine. People who tell you otherwise are selling something.
Above three hundred thousand, the firm was forced to solve this years ago, because at that size the environment does not bend, it breaks, visibly, in front of a regulator or a client.
Inside that band is where it hurts, and the reason is that the break sits upstream of execution. Trading works. Settlement works. The systems everybody watches are all fine. The failure is in reference data and onboarding, which nobody has a dashboard for, and which surface only as the three complaints in three meetings that this article opened with.
Treat those two numbers as a rule of thumb rather than a threshold. They come from where I have seen the problem bite, not from a study, and the worksheet labels them the same way. The shape of the argument does not depend on them. If your book carries heavy per position data obligations, the band starts lower for you.
Sequencing
If you do act, the order is not negotiable, and most programs get it backwards.
Golden source first. One authoritative record for each instrument, with a named owner and a defined path for corrections. Then workflow, which is only trustworthy once the data underneath it is. Then reporting.
Programs run it in reverse because reporting is the visible part and the part the board asked about. Reporting built on unreconciled reference data is a faster way to distribute the same wrong number, and the first time that happens the program loses the credibility it needed to finish.
The sequencing question and the selection question are the same question asked twice. Buy the Exit, Not the Demo is the other half of this: that piece interrogates the replacement, this one prices what you already have. Neither argument works without the other.
Before the forcing event
The projected saving in the illustrative case is a five year NPV of $906,998, breaking even in year four, with operations headcount at twenty under the current environment against twelve under a modernized one by year five. Eight people of avoided hiring, which at the loaded cost in that scenario is $1.16 million a year of run rate that never gets added.
Those figures are illustrative and yours will differ. The structural point does not.
A merger, a regulatory examination, a new asset class arriving with a mandate and a deadline, an acquirer running diligence. Any of those turns this from a planned program into an emergency, and the same work costs materially more under duress, because you lose the ability to sequence it and you lose the ability to walk away.
The cheapest version of this project is the one you start while you still have the option not to.
The test
Three questions. Ask them in your own shop this week.
Can anybody tell you the elapsed time from allocation decision to first clean valuation? Not the effort. The calendar.
Can anybody tell you what share of reference data breaks reach a valuation, a payment or a client report before being caught? Not the break count. The leakage rate.
Is the exception queue somebody's job? Not a rotation. A role.
If the answer to all three is no, you do not have a technology problem yet. You have a measurement problem, and it is the cheaper of the two to fix.
The calculator
The worksheet prices the four costs on your own environment and sets them against a modernization case you enter yourself. It names no vendor and assumes no vendor pricing, because platform and implementation cost are the two things only you know. Six tabs, every figure a live formula, no macros and nothing hidden, so any number can be traced back to the input it came from.
It reports the cost of current state annually, per position and in basis points of assets, then projects five years against your growth rate and returns a breakeven year and an NPV. The implementation elapsed time you enter drives the projection rather than sitting beside it, so benefits phase in over the build while you carry the old environment and the platform fee together. Two sensitivity grids follow, both in NPV so they answer the same question the headline does: one across platform cost and residual rate, one across the onboarding assumption.
The conventions that matter are on the face of it rather than buried. The ceiling charges mean staffing across each year, not the year-end requirement. Severity events hold at the frequency you enter instead of quietly scaling with the book. Cash flows are discounted year-end, which is conventional and modestly favorable to modernization, and it says so beside the NPV.
It assumes no headcount reduction, only avoided hiring, which understates the case for any firm that does intend to reduce. It does not model the extra people a parallel run usually needs, nor migration risk, nor an implementation that fails. Those are real and they are not in there.
It is on its own page: Cost of Current State Worksheet
The figures in this article are illustrative and do not describe any client or engagement. They are produced by the worksheet from the inputs stated above. Nothing here is investment advice.