Common failures
Where this usually goes wrong
The three we are called in to fix most often. If any of these describe your current setup, that is the first thing we would change.
Measurement
How you will know it is working
Agreed before we start, so progress is visible in the months before it shows up in revenue.
- p95 response time
- The average hides the requests that lose customers. The 95th percentile is what a bad experience actually looks like.
- From launch
- Memory over 24 hours
- A flat line means no leak. A staircase means a restart is masking one.
- Week 1 post-launch
- Known vulnerabilities
- High and critical advisories across the dependency tree, checked in CI so the number cannot drift upward unnoticed.
- Every build