23 KiB
Observability
For maintainers. Using T3 Code? See docs/user.
T3 Code has one server-side observability model:
- pretty logs go to stdout for humans
- completed spans go to a local NDJSON trace file
- traces, metrics, and logs can also be exported over OTLP to a real backend like Grafana LGTM
The local trace file is the persisted source of truth for normal local launches. Those launches do not
write a separate server log file, but SSH-managed launches also persist the remote process's
stdout/stderr at ~/.t3/ssh-launch/<state>/server.log.
Where To Find Things
Logs
Logs are human-facing:
- destination: stdout
- format:
Logger.consolePretty() - normal local persistence: none
- SSH-managed launch persistence:
~/.t3/ssh-launch/<state>/server.log - remote export: OTLP only, when configured
If you want a log message to show up in the trace file, emit it inside an active span with Effect.log.... Logger.tracerLogger will attach it as a span event.
Configuring a logs endpoint takes over that job. The server then exports log records, which cover
every message instead of only the ones inside an active span and carry the trace and span ids so
they still line up with the trace. Logger.tracerLogger is dropped in that mode, so the same
message is not exported twice and the trace file stops carrying log messages. stdout output and
SSH-managed launch persistence stay unchanged either way.
Traces
Completed spans are written as NDJSON records to serverTracePath. The default depends on how the
server starts: production and explicitly configured homes use
<home>/userdata/logs/server.trace.ndjson (so ~/.t3/userdata/... by default, or
/custom/path/userdata/... with --home-dir /custom/path), a linked worktree dev run uses
<worktree>/.t3/userdata/logs/server.trace.ndjson, and an implicit dev run outside a linked
worktree uses ~/.t3/dev/logs/server.trace.ndjson.
Important fields common to both record types:
type:effect-spanorotlp-spanname: span nametraceId,spanId,parentSpanId: correlationdurationMs: elapsed timeattributes: structured contextevents: embedded logs and custom events
effect-span records also contain exit with Success, Failure, or Interrupted. otlp-span
records instead carry OTLP resource, scope, and optional status fields.
The TraceRecord, EffectTraceRecord, and OtlpTraceRecord schemas live in
packages/shared/src/observability.ts.
DPoP proof failures include the safe environment.dpop.failure_code span
attribute. A time_window failure means that a signed proof was too old or too
far in the future for the environment server's allowed window. It can point to
a date or time problem on either device, but it can also result from a delayed
request.
Summarize the trace file
t3 trace summary reads the trace file and its rotated backups directly, so it works while the
server is stalled or stopped. It prints counts, rates, and latency percentiles per span name. Use
it to measure background work or to compare two builds.
t3 trace summary --since 30m --limit 40
It reads T3CODE_TRACE_FILE if set, else <home>/userdata/logs/server.trace.ndjson for
--base-dir or T3CODE_HOME, plus the T3CODE_TRACE_MAX_FILES rotated backups. For a dev run or
a copied file, set T3CODE_TRACE_FILE. --since 30m keeps spans that ended in the last 30
minutes. The rate is per minute between the first and last span end.
Metrics
Metrics are not written to a local file.
- local persistence: none
- remote export: OTLP only, when configured
- current definitions:
apps/server/src/observability/Metrics.ts
If OTLP is not configured, metrics still exist in-process, but you will not have a local artifact to inspect.
Event Loop Stalls
apps/server/src/observability/EventLoopMonitor.ts samples the server's event loop every 30 s. When
the loop stalled for more than 2 s since the previous sample, it records a root
server.eventLoop.stall span with a warning. The span has trace level Warn, so it stays when
T3CODE_TRACE_MIN_LEVEL is Warn. The warning shows in Settings > Diagnostics unless OTLP logs are
on. The span time is when the sample ran, not when the stall happened.
Some delay is not recorded:
delayMaxMsis the longest stall, and can undercount it by up to 1 s. The 2 s threshold applies to this value, so a stall over 3 s is normally recorded, and a shorter one can be missed. A stall that ends just as a sample runs can be missed too.- Time the computer spends asleep reads as delay on macOS and Windows. So a sample only counts when
the loop was busy, not waiting for events, for at least
delayMaxMs. Busy time covers the whole window, so a short sleep in an otherwise busy window can still record a false stall. The span then shows CPU time far belowdelayMaxMs. - The first sample after launch is skipped. Startup work such as migrations and projection bootstrap can block the loop for seconds on a large database.
CPU times and page faults cover the whole process over the whole window since the previous sample.
The window is nominally 30 s, but a long stall delays the sample and makes the window longer. Other
work in the window can hide a wait, so only CPU time far below delayMaxMs proves the thread was
waiting. Read CPU together with page faults:
- High
cpuSystemMswith many page faults means memory pressure. Major faults are reads from disk or swap. On macOS, reads from compressed memory are minor faults plus system CPU. - High
cpuUserMswith few page faults means JavaScript work or garbage collection. - Low CPU with few major page faults points at synchronous disk I/O, such as SQLite reads or trace file writes.
- Many
involuntaryContextSwitchesmean other processes were competing for the CPU.
Related Artifacts
Provider event NDJSON files still exist for provider runtime streams. Those are separate from the main server trace file.
Run The Server In Instrumented Mode
There are two useful modes:
- local-only: stdout + local
server.trace.ndjson - full local observability: stdout + local trace file + OTLP export to Grafana/Tempo/Prometheus
The local trace file is always on. OTLP export is opt-in.
Option 1: Local Traces Only
You do not need any extra env vars. Just run the app normally and inspect server.trace.ndjson.
Examples:
npx t3
node --run dev
node --run dev:desktop
Option 2: Run With A Local LGTM Stack
1. Start Grafana LGTM
docker run --name lgtm \
-p 3000:3000 \
-p 4317:4317 \
-p 4318:4318 \
--rm -ti \
grafana/otel-lgtm
Then open http://localhost:3000.
Default Grafana login:
- username:
admin - password:
admin
2. Export OTLP env vars
export T3CODE_OTLP_TRACES_URL=http://localhost:4318/v1/traces
export T3CODE_OTLP_METRICS_URL=http://localhost:4318/v1/metrics
export T3CODE_OTLP_LOGS_URL=http://localhost:4318/v1/logs
export OTEL_RESOURCE_ATTRIBUTES=deployment.environment.name=development
Optional:
export T3CODE_TRACE_MIN_LEVEL=Info
export T3CODE_TRACE_TIMING_ENABLED=true
3. Launch the app from that same shell
CLI:
npx t3
Monorepo web/server dev:
node --run dev
Monorepo desktop dev:
node --run dev:desktop
Packaged desktop app:
Launch the actual app executable from the same shell so the desktop app and embedded backend inherit T3CODE_OTLP_*.
macOS app bundle example:
T3CODE_OTLP_TRACES_URL=http://localhost:4318/v1/traces \
T3CODE_OTLP_METRICS_URL=http://localhost:4318/v1/metrics \
T3CODE_OTLP_LOGS_URL=http://localhost:4318/v1/logs \
"/Applications/T3 Code.app/Contents/MacOS/T3 Code"
Direct binary example:
T3CODE_OTLP_TRACES_URL=http://localhost:4318/v1/traces \
T3CODE_OTLP_METRICS_URL=http://localhost:4318/v1/metrics \
T3CODE_OTLP_LOGS_URL=http://localhost:4318/v1/logs \
./path/to/your/desktop-app-binary
Do not rely on launching from Finder, Spotlight, the dock, or the Start menu after setting shell env vars. Those launches usually will not pick them up.
4. Fully restart after changing env
The backend reads observability config at process start. If you change OTLP env vars, stop the app completely and start it again.
How To Use Traces And Metrics To Debug The Server
Start With The Local Trace File
The trace file is the fastest way to inspect raw span data.
Resolve the path for the launch mode once. Production and explicitly configured homes store runtime
state under the base directory's userdata folder:
TRACE_FILE="${T3CODE_HOME:-$HOME/.t3}/userdata/logs/server.trace.ndjson"
A dev server started from a linked worktree defaults to that worktree's local home:
TRACE_FILE="$WORKTREE/.t3/userdata/logs/server.trace.ndjson"
Only an implicit dev run outside a linked worktree uses the shared dev directory:
TRACE_FILE="$HOME/.t3/dev/logs/server.trace.ndjson"
Tail the selected file:
tail -f "$TRACE_FILE"
Show failed spans:
jq -c 'select(.type == "effect-span" and .exit._tag != "Success") | {
name,
durationMs,
exit,
attributes
}' "$TRACE_FILE"
Show slow spans:
jq -c 'select(.durationMs > 1000) | {
name,
durationMs,
traceId,
spanId
}' "$TRACE_FILE"
Inspect embedded log events:
jq -c 'select(any(.events[]?; .attributes["effect.logLevel"] != null)) | {
name,
durationMs,
events: [
.events[]
| select(.attributes["effect.logLevel"] != null)
| {
message: .name,
level: .attributes["effect.logLevel"]
}
]
}' "$TRACE_FILE"
Follow one trace:
jq -r 'select(.traceId == "TRACE_ID_HERE") | [
.name,
.spanId,
(.parentSpanId // "-"),
.durationMs
] | @tsv' "$TRACE_FILE"
Filter orchestration commands:
jq -c 'select(.attributes["orchestration_v2.command_type"] != null) | {
name,
durationMs,
commandType: .attributes["orchestration_v2.command_type"],
threadId: .attributes["orchestration_v2.thread_id"]
}' "$TRACE_FILE"
Filter git activity:
jq -c 'select(.attributes["git.operation"] != null) | {
name,
durationMs,
operation: .attributes["git.operation"],
cwd: .attributes["git.cwd"],
hookEvents: [
.events[]
| select(.name == "git.hook.started" or .name == "git.hook.finished")
]
}' "$TRACE_FILE"
Use Tempo When You Need A Real Trace Viewer
Tempo is better than raw NDJSON when you want to:
- search across many traces
- inspect parent/child relationships visually
- compare many slow traces
- drill into one failing request without hand-joining by
traceId
Recommended flow in Grafana:
- Open
Explore. - Pick the
Tempodata source. - Set the time range to something recent like
Last 15 minutes. - Start broad. Do not begin with a very narrow query.
- Look for spans from the
t3code-serverort3code-desktopservice, then narrow by span name or attributes.
Good first searches:
- service name
t3code-serverort3code-desktop, plus a resource attribute such asdeployment.environment.name - span names like
sendTurnor a Git operation such asGitVcsDriver.statusDetails.status - Git spans whose
git.operationattribute identifies the operation - orchestration spans with attributes like
orchestration_v2.command_type
Once you know traces are arriving, narrower TraceQL queries for names such as sendTurn or Git
operation names become useful.
Use Metrics To See Systemic Problems
Traces are best for one request. Metrics are best for trends.
Good metric families to watch:
t3_rpc_request_durationt3_provider_turn_duration(how long the provider adapter takes to start a turn, not the turn's run time)t3_git_command_duration
Counters tell you volume and failure rate:
t3_rpc_requests_totalt3_provider_turns_totalt3_git_commands_total
Webhooks have their own families:
-
t3_webhook_deliveries_totalbyoutcomeandsource(relayordirect). Beyond what the sender sees,queue_fullmeans a task already had its limit of deliveries waiting,prompt_too_longmeans the filled-in prompt passed the provider limit, andduplicatemeans the relay delivered a request this environment had already run. -
t3_webhook_runs_totalbyoutcome(started,skipped,failed) for the runs those deliveries start, which happen after the sender has its answer. -
t3_webhook_held_delayfor how long requests the relay held waited before arriving. -
t3_secret_requests_totalbystatus(saved,declined,cancelled,timed_out) for secrets agents asked users for, andt3_secret_refs_consumed_totalbyresult(used,rejected) for tools redeeming them. Neither ever carries a value.
ScheduledTaskService.triggerWebhook spans carry the same outcome per request, and each run
started from a delivery is its own ScheduledTaskService.runWebhookDelivery trace. For a request
the relay forwarded, the span also goes to the T3 Connect trace export as a child of the relay's
span; requests that reach the environment directly never join a sender's trace.
Use metrics when the question is:
- "is this always slow?"
- "did this get worse after a change?"
- "which command type is failing most often?"
Use traces when the question is:
- "what happened in this specific request?"
- "which child span caused this one slow interaction?"
- "what logs were emitted inside the failing flow?"
Common Workflows
"Why did this request fail?"
- Start with the local NDJSON file.
- Find
effect-spanrecords whereexit._tag != "Success". - Group by
traceId. - Inspect sibling spans and span events.
- If needed, move to Tempo for the full trace tree.
"Why is the UI feeling slow?"
- Search for slow top-level spans in the trace file or Tempo.
- Check child spans for sqlite, git, provider, or terminal work.
- Look at the matching duration metrics to see whether the slowness is systemic.
"Are git hooks causing latency?"
- Filter
git.operationspans. - Inspect
git.hook.startedandgit.hook.finishedevents. - Compare hook timing to the enclosing git span duration.
"Why do I have spans locally but nothing in Grafana?"
Usually one of these is true:
T3CODE_OTLP_TRACES_URLwas not set- the app was launched from a different environment than the one where you exported the vars
- the app was not fully restarted after changing env
- Grafana is looking at the wrong time range or service name
If the local NDJSON file is updating, local tracing is working. The problem is almost always OTLP export configuration or process startup.
How To Think About Adding Tracing To Future Code
Prefer Boundaries Over Tiny Helpers
Good span boundaries:
- RPC methods
- orchestration command handling
- provider adapter calls
- external process calls
- persistence writes
- queue handoffs
Avoid tracing every tiny helper. Most helpers should inherit the active span rather than create a new one.
Reuse Effect.fn(...) Where It Already Exists
The codebase already uses Effect.fn("name") heavily. That should usually be your first tracing boundary.
For ad hoc work:
import { Effect } from "effect";
const runThing = Effect.gen(function* () {
yield* Effect.annotateCurrentSpan({
"thing.id": "abc123",
"thing.kind": "example",
});
yield* Effect.logInfo("starting thing");
return yield* doWork();
}).pipe(Effect.withSpan("thing.run"));
Put High-Cardinality Detail On Spans
Use span annotations for IDs, paths, and other detailed context:
yield *
Effect.annotateCurrentSpan({
"provider.thread_id": input.threadId,
"provider.request_id": input.requestId,
"git.cwd": input.cwd,
});
Keep Metric Labels Low Cardinality
Good metric labels:
- operation kind
- method name
- provider kind
- aggregate kind
- outcome
Bad metric labels:
- raw thread IDs
- command IDs
- file paths
- cwd
- full prompts
- full model strings when a normalized family label would do
Detailed context belongs on spans, not metrics.
Use Logs As Span Events
Logs inside a span become part of the trace story:
yield * Effect.logInfo("starting provider turn");
yield * Effect.logDebug("waiting for approval response");
Those messages show up as span events because Logger.tracerLogger is installed.
Use The Pipeable Metrics API
withMetrics(...) is the default way to attach a counter and timer to an effect:
import { someCounter, someDuration, withMetrics } from "../observability/Metrics.ts";
const program = doWork().pipe(
withMetrics({
counter: someCounter,
timer: someDuration,
attributes: {
operation: "work",
},
}),
);
Detailed API Reference
Runtime Wiring
The server observability layer is assembled in apps/server/src/observability/Observability.ts.
It provides:
- pretty stdout logger
Logger.tracerLogger- local NDJSON tracer
- optional OTLP trace exporter
- optional OTLP metrics exporter
- optional OTLP log exporter
- Effect trace-level and timing refs
The desktop main process is a second producer, assembled in
apps/desktop/src/app/DesktopObservability.ts. It reads the same T3CODE_OTLP_* names and the same
Settings entries as the backend it supervises, and covers work the backend cannot see: app startup,
window and menu handling, backend supervision, and updates. It reports as service
t3code-desktop, so a collector shows it alongside the backend rather than mixed into it. It
exports traces and logs only; the main process records no metrics, so the metrics endpoint applies
to the backend alone.
Env Vars
Local trace file:
T3CODE_TRACE_FILE: override trace file pathT3CODE_TRACE_MAX_BYTES: per-file rotation size, default10485760T3CODE_TRACE_MAX_FILES: rotated file count, default10T3CODE_TRACE_BATCH_WINDOW_MS: flush window, default1000T3CODE_TRACE_MIN_LEVEL: minimum trace level, defaultInfoT3CODE_TRACE_TIMING_ENABLED: enable timing metadata, defaulttrue
OTLP export:
T3CODE_OTLP_TRACES_URL: OTLP trace endpointT3CODE_OTLP_METRICS_URL: OTLP metric endpointT3CODE_OTLP_LOGS_URL: OTLP log endpointT3CODE_OTLP_EXPORT_INTERVAL_MS: export interval, default10000T3CODE_OTLP_HEADERS: extra headers for all three exporters, same format asOTEL_EXPORTER_OTLP_HEADERS: comma-separatedkey=valuepairs with percent-encoded values.T3CODE_OTLP_PROTOCOL:http/json(default) orhttp/protobuf
The server and the desktop app also read the standard
OTEL_EXPORTER_OTLP_{TRACES,METRICS,LOGS}_ENDPOINT and generic OTEL_EXPORTER_OTLP_ENDPOINT (with
/v1/traces, /v1/metrics, or /v1/logs appended), for a collector expecting those instead. A
non-blank T3CODE_OTLP_*_URL wins over either, and a per-signal endpoint wins over the generic one
for its signal. A blank value counts as unset. A signal with an OTEL endpoint takes its headers from
OTEL_EXPORTER_OTLP_HEADERS and its protocol from OTEL_EXPORTER_OTLP_PROTOCOL (default
http/protobuf, read case-insensitively), and a per-signal
OTEL_EXPORTER_OTLP_{TRACES,METRICS,LOGS}_HEADERS or _PROTOCOL wins over the generic one for its
signal. T3CODE_OTLP_HEADERS and T3CODE_OTLP_PROTOCOL never apply to it. An endpoint that is not
an http or https URL, a protocol other than http/protobuf or http/json such as grpc, or
headers that are not key=value pairs with percent-encoded values turn that signal's export off
with a startup warning, rather than sending it to the Settings endpoint.
Service names are fixed: t3code-server for the backend and t3code-desktop for the desktop main
process, both in service.namespace t3code. OTEL_SERVICE_NAME and a service.name or
service.namespace in OTEL_RESOURCE_ATTRIBUTES are ignored. Tell installations apart with other
resource attributes, such as OTEL_RESOURCE_ATTRIBUTES=deployment.environment.name=development.
If the OTLP URLs are unset, local tracing still works, metrics stay in-process only, and logs stay on stdout only.
The Kill Switch
T3CODE_OTEL_SDK_DISABLED and OTEL_SDK_DISABLED turn off every OTLP export in both the server and
the desktop main process, overriding any endpoint from the environment or Settings. Local trace
files and stdout logs are unaffected.
T3CODE_OTEL_SDK_DISABLED wins when set, so T3CODE_OTEL_SDK_DISABLED=false re-enables export on a
machine that sets OTEL_SDK_DISABLED for everything else. It accepts the usual boolean spellings
(true/false, yes/no, on/off, 1/0, y/n). OTEL_SDK_DISABLED follows the
OpenTelemetry specification and only true disables export, so OTEL_SDK_DISABLED=1 does not.
Values are case-insensitive and trimmed. An unrecognized value is ignored with a startup warning.
OTEL_TRACES_EXPORTER, OTEL_METRICS_EXPORTER, or OTEL_LOGS_EXPORTER set to none turns off
just that signal, overriding an OTEL endpoint and the Settings endpoint. A T3CODE_OTLP_*_URL still
wins for its signal. otlp is the default, and any other exporter name, such as console or
prometheus, is ignored with a startup warning.
What Is Instrumented Today
Current high-value span and metric boundaries include:
- WebSocket RPC request spans (
ws.rpc.<method>) and metrics inapps/server/src/observability/RpcInstrumentation.ts - startup phases
- orchestration command processing
- provider session and turn operations
- git command execution and git hook events
- terminal session lifecycle
- sqlite query execution
- event loop stalls (
server.eventLoop.stall)
Current Constraints
- logs outside spans are not persisted in the trace file; SSH-managed launch stdout/stderr is still captured in its launcher log
- metrics are not snapshotted locally
Heap Snapshots
To see what a long-running server holds in memory, send it SIGUSR2. The server writes a V8 heap
snapshot to its logs dir and logs the path. This works for desktop, npx t3, and service installs
on macOS and Linux. Windows has no SIGUSR2.
Send the signal to the server pid in server-runtime.json, which sits in the server's state dir
next to the logs dir. For a dev server or a --home-dir launch, use that server's state dir from
Traces. Do not send it to the desktop app or the service launcher: a process without the
handler exits on SIGUSR2. After a crash the file can keep a stale pid that now belongs to a
different process, so check the pid first.
pid="$(jq .pid "${T3CODE_HOME:-$HOME/.t3}/userdata/server-runtime.json")"
ps -p "$pid" -o command=
If ps shows the T3 Code server, send the signal:
kill -USR2 "$pid"
The file is <logsDir>/server-<pid>-<timestamp>.heapsnapshot, next to server.trace.ndjson. To
open it, use the Memory tab in Chrome DevTools and select Load.
Before you take one:
- The server stops while it writes the file. For a large heap this can take a minute or more. Connected clients can reconnect during the pause, and an event loop monitor, if the server has one, records the pause as a stall. Send the signal once. A second signal sent during a write takes another snapshot after the first one finishes.
- The write needs about as much free memory as the heap uses. On a machine that is already swapping, it can make the problem worse or crash the server.
- The file contains everything in server memory, including tokens, secrets, and thread content. Do not share it publicly. Delete it when you are done, because storage cleanup does not remove it.