Draw the boundaries before you parallelize
Decomposing a monolith into services and parallelizing its computation look like one project. Doing them in the wrong order makes the second one much harder.
When Yield & Spread moved from analyzing single securities to analyzing whole portfolios, the obvious-sounding plan is: break the monolith into services, then parallelize the computation across them, and the latency wins should follow from both. In practice, the order those two things happen in matters more than it looks like it should.
Parallelizing computation inside a system whose internal boundaries are still fuzzy just distributes the fuzziness. If two components are informally sharing assumptions about data shape, ordering, or state — the way pieces of a monolith often do without anyone deciding to design it that way — splitting them across services and running them concurrently doesn't remove those assumptions. It just makes them harder to find, because now they're implicit contracts between processes instead of implicit contracts between function calls in the same stack trace.
So the decomposition work that actually mattered wasn't the parallel computation itself — it was defining the service boundaries and the data contracts between them before writing the code that would run them in parallel. A precise contract is what lets a service be optimized, scaled, or reasoned about independently, because you've made explicit what used to be an assumption. Skip that step, and "distributed" mostly just means "the same coupling, now with network calls."
The pattern generalizes past this one system: parallelism is a multiplier. It amplifies whatever correctness properties your boundaries already have — including their absence. Get the contracts right first, and the performance work becomes comparatively mechanical. Get them wrong, and every latency win comes with a matching increase in how hard the system is to debug.