t3-code-android-nightly/docs/operations/relay-observability.md
Julius Marminge f33b060caf
feat(relay,server,web,mobile): opt-in to hold webhooks while offline (#15487)
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-05 15:14:53 -07:00

5.7 KiB

Relay observability

For maintainers. Using T3 Code? See docs/user.

The relay Alchemy stack owns a shared Axiom trace setup:

  • t3-code-relay-traces-prod, the OpenTelemetry trace dataset shared by the Worker, mobile app, and first-party relay clients
  • t3-code-relay-otel-ingest-prod, the dataset-scoped Worker ingest token
  • t3-code-mobile-otel-ingest-prod, the dataset-scoped mobile ingest token
  • t3-code-relay-client-otel-ingest-prod, the dataset-scoped first-party relay-client ingest token
  • t3-code-relay-recent-spans-prod, a view of recent request and endpoint spans

Alchemy stages append their sanitized stage name to isolate resources, for example t3-code-relay-traces-dev-julius for a personal stage.

Deploy from infra/relay with the normal Alchemy workflow:

vp run deploy

Alchemy resolves account-level Axiom deployment credentials through its provider. At runtime, the Worker receives only its scoped ingest token. Mobile and relay clients use their own separately provisioned scoped ingest tokens.

The Worker emits Effect's built-in HTTP server spans plus endpoint and database child spans. Effect's OpenTelemetry exporter stores semantic HTTP attributes below the attributes. prefix. For example:

['t3-code-relay-traces-prod']
| where name startswith 'http.server'
| extend endpoint = column_ifexists('attributes.http.route', ''),
    customAttributes = column_ifexists('attributes.custom', dynamic({}))
| project _time, name, trace_id, duration,
    ['attributes.http.request.method'],
    ['attributes.url.path'],
    ['attributes.http.response.status_code'],
    endpoint,
    relayOperation = customAttributes['relay']['operation']
| order by _time desc
| limit 200

The provisioned view also reads the endpoint from attributes.http.route. Relay-specific span annotations are stored under attributes.custom; relay.operation is one of the emitted custom attributes.

Agents should prefer the provisioned view or APL queries for completed incidents instead of tailing the Cloudflare Worker. The stack does not provision a separate query token. Responders who need scripted query access use the authorized account-level AXIOM_TOKEN together with AXIOM_ORG_ID; scoped ingest tokens remain write-only credentials for their producers.

DPoP proof failures include the stable relay.dpop.failure_code span attribute. A time_window failure means that a signed proof was too old or too far in the future for the relay's allowed window. It can point to a date or time problem on either device, but it can also result from a delayed request. The client uses this category, and the absence of a category from an older relay, to decide whether clock skew is confirmed or only one possible cause.

Webhooks

A public webhook request is one relay.hooks.forward span. Its relay.hook.outcome says what happened: forwarded, held, rate_limited, inbox_full, not_found, payload_too_large, environment_unavailable, or environment_timeout. relay.hook.endpoint_key identifies the managed endpoint, and with it the environment. On a forward, relay.hook.upstream_status or relay.hook.upstream_error records the environment's answer, and relay.hook.upstream_outcome what it did with the request (accepted, duplicate, prompt_too_long, queue_full, rejected_signature, expired, disabled, ...), from its x-t3-hook-outcome response header. A held request's delivery records the same on its relay.inbox.deliver span. The relay sends its own traceparent with each forward and drops any a sender supplied, and signs each forward (x-t3-relay-delivery) so the environment trusts the relay's delivery id, receive time and trace context only from the relay, so on environments that export to T3 Connect, the environment's ScheduledTaskService.triggerWebhook span lands in the same trace. relay.hook.rate_limit says which budget ran out: endpoint or hook. relay.hook.rate_limiter_failed_open is set when the Cloudflare rate limiter was unavailable and the request went through unlimited.

Held requests are handled in the endpoint's HookInboxObject, and each call into it is its own trace, not a child of the forward span: relay.inbox.hold, relay.inbox.wake, and relay.inbox.deliver for each alarm run. Each carries relay.hook.endpoint_key, so they join to forwards. The details are on the child spans: HookInboxStore.hold carries relay.inbox.refused when the inbox refused a request (max_per_hook, max_requests, max_bytes, or already_held). HookInboxStore.deliverDue carries relay.inbox.run_result (drained, more_pending, busy, or unreachable), how many requests it sent and delivered, the consecutive failures behind the retry delay, the longest time a delivered request waited, and the backlog left. A run that errored outright has relay.inbox.run_result = failed on the root relay.inbox.deliver span instead.

Questions these answer:

// Webhook outcomes per hour
['t3-code-relay-traces-prod']
| where name == 'relay.hooks.forward'
| summarize count() by bin(_time, 1h), outcome = tostring(['attributes.custom']['relay.hook.outcome'])

// Inboxes stuck behind an unreachable environment, by endpoint
['t3-code-relay-traces-prod']
| where name == 'HookInboxStore.deliverDue'
| extend c = ['attributes.custom']
| where tostring(c['relay.inbox.run_result']) == 'unreachable'
| project trace_id, failures = toint(c['relay.inbox.consecutive_failures']),
    backlog = toint(c['relay.inbox.held_count'])
| join kind=inner (
    ['t3-code-relay-traces-prod']
    | where name == 'relay.inbox.deliver'
    | project trace_id, endpoint = tostring(['attributes.custom']['relay.hook.endpoint_key'])
  ) on trace_id
| summarize runs = count(), failures = max(failures), backlog = max(backlog) by endpoint