Every scaling problem I’ve been called in to fix has eventually turned out to be a boundary problem. Not enough servers, not the wrong database — a boundary drawn in the wrong place, usually years earlier, by someone solving a smaller problem than the one the system now has.
Boundaries are cheap to draw and expensive to move. The team that draws them well isn’t the one with the most experience — it’s the one willing to ask “what will this decision cost us to reverse” before asking “does this work today.”
This is the first in a short series on what actually breaks as systems grow, and where the fixes really live.