Use Cases
In a distributed system, the service that got slow is rarely the one that broke. The map shows which edge degraded first, and the trace shows how far the wait spread.
signal api-gateway p95 1.9 s · was 0.9 s, no deploy
The path
payment-svc@1.4.2 shipped four hours before latency went up. The map, the trace and the logs below are all filtered to it.
api-gateway is the service being paged about, but its own spans are fine. The degraded edge is two hops in: payment-svc to user-db, p95 up from 12 ms to 210 ms.
Open a slow request and the db span is almost all of it. The query itself takes 4 ms, but the span takes 210 ms. The other 206 ms was spent waiting before the query ran.
Lines from that span read pool wait 206 ms, acquired after 3 attempts. Filter by service.version and they start at 1.4.2, which raised concurrency without raising the pool.
outcome
Pool size raised to match the new concurrency. The p95 edge latency returned to 14 ms without a rollback.
What it does
Keep going
One endpoint, one key. Traces, logs, metrics and sessions, linked from the first request.