Use Cases

Checkout is failing. Which orders, and why.

During a sale you need to know which customers lost their basket and why. Filter by one order id and the alert, the trace and the logs all show the same failure.

signal POST /checkout error rate 8.4% · threshold 1%

The path

One order, from the alert to the timeout

order_8421 is one failed order. Each screen below is filtered to it, so you never line up timestamps across two tools.

  1. alert T+0:00 · order order_8421

    The alert fires on the checkout route only

    Error rate on POST /checkout crosses its threshold while the rest of the site is healthy. The sale traffic is fine. The payment path is not.

    An alert rule with its threshold, evaluation window and firing history
  2. trace T+0:25 · order order_8421

    The trace ends at Stripe

    Open one failed order. Three retries sit at the bottom of the waterfall in red, each timing out at exactly 1.75 s, the client's own deadline.

    An 18-span API trace waterfall with nested application and SQL operations
  3. logs T+2:10 · order order_8421

    The logs say the pool was empty

    The log lines from those spans share the same trace id. They show the retry budget running out and the connection pool at capacity. The requests timed out waiting for a connection.

    The StripeTimeout issue with its affected-order count and the sessions that produced them

outcome

Pool raised and the retry budget cut to two. Checkout error rate back to baseline within one evaluation window.

2 min
alert to root cause
3
retries per failed order
1.75 s
timeout, every time

What it does

Why one order id was enough

order.id
Business ids are just attributes
Set order.id on the span once and it is filterable everywhere spans are: the trace list, the error issue, the log search. No separate index to maintain.
SpanKind.CLIENT
Outbound calls are spans too
The Stripe call is a client span with its own duration and status. You see the third-party timeout directly instead of guessing at it from your own error rate.
trace_id
Logs are joined to the span
You open a span's logs from the span. There is no text query to narrow by hand to the same minute, which is usually the slowest part of an investigation.

Keep going

The same data, other jobs

Point OTLP at Maple.

One endpoint, one key. Traces, logs, metrics and sessions, linked from the first request.