docs/operations/observability.md

Observability

For maintainers. Using T3 Code? See docs/user.

T3 Code has one server-side observability model:

  • pretty logs go to stdout for humans
  • completed spans go to a local NDJSON trace file
  • traces, metrics, and logs can also be exported over OTLP to a real backend like Grafana LGTM

The local trace file is the persisted source of truth for normal local launches. Those launches do not write a separate server log file, but SSH-managed launches also persist the remote process's stdout/stderr at ~/.t3/ssh-launch/<state>/server.log.

Where To Find Things

Logs

Logs are human-facing:

  • destination: stdout
  • format: Logger.consolePretty()
  • normal local persistence: none
  • SSH-managed launch persistence: ~/.t3/ssh-launch/<state>/server.log
  • remote export: OTLP only, when configured

If you want a log message to show up in the trace file, emit it inside an active span with Effect.log.... Logger.tracerLogger will attach it as a span event.

Configuring a logs endpoint takes over that job. The server then exports log records, which cover every message instead of only the ones inside an active span and carry the trace and span ids so they still line up with the trace. Logger.tracerLogger is dropped in that mode, so the same message is not exported twice and the trace file stops carrying log messages. stdout output and SSH-managed launch persistence stay unchanged either way.

Traces

Completed spans are written as NDJSON records to serverTracePath. The default depends on how the server starts: production and explicitly configured homes use <home>/userdata/logs/server.trace.ndjson (so ~/.t3/userdata/... by default, or /custom/path/userdata/... with --home-dir /custom/path), a linked worktree dev run uses <worktree>/.t3/userdata/logs/server.trace.ndjson, and an implicit dev run outside a linked worktree uses ~/.t3/dev/logs/server.trace.ndjson.

Important fields common to both record types:

  • type: effect-span or otlp-span
  • name: span name
  • traceId, spanId, parentSpanId: correlation
  • durationMs: elapsed time
  • attributes: structured context
  • events: embedded logs and custom events

effect-span records also contain exit with Success, Failure, or Interrupted. otlp-span records instead carry OTLP resource, scope, and optional status fields.

The TraceRecord, EffectTraceRecord, and OtlpTraceRecord schemas live in packages/shared/src/observability.ts.

DPoP proof failures include the safe environment.dpop.failure_code span attribute. A time_window failure means that a signed proof was too old or too far in the future for the environment server's allowed window. It can point to a date or time problem on either device, but it can also result from a delayed request.

Summarize the trace file

t3 trace summary reads the trace file and its rotated backups directly, so it works while the server is stalled or stopped. It prints counts, rates, and latency percentiles per span name. Use it to measure background work or to compare two builds.

bash
1t3 trace summary --since 30m --limit 40

It reads T3CODE_TRACE_FILE if set, else <home>/userdata/logs/server.trace.ndjson for --base-dir or T3CODE_HOME, plus the T3CODE_TRACE_MAX_FILES rotated backups. For a dev run or a copied file, set T3CODE_TRACE_FILE. --since 30m keeps spans that ended in the last 30 minutes. The rate is per minute between the first and last span end.

Metrics

Metrics are not written to a local file.

  • local persistence: none
  • remote export: OTLP only, when configured
  • current definitions: apps/server/src/observability/Metrics.ts

If OTLP is not configured, metrics still exist in-process, but you will not have a local artifact to inspect.

Related Artifacts

Provider event NDJSON files still exist for provider runtime streams. Those are separate from the main server trace file.

Run The Server In Instrumented Mode

There are two useful modes:

  • local-only: stdout + local server.trace.ndjson
  • full local observability: stdout + local trace file + OTLP export to Grafana/Tempo/Prometheus

The local trace file is always on. OTLP export is opt-in.

Option 1: Local Traces Only

You do not need any extra env vars. Just run the app normally and inspect server.trace.ndjson.

Examples:

bash
1npx t3
bash
1node --run dev
bash
1node --run dev:desktop

Option 2: Run With A Local LGTM Stack

1. Start Grafana LGTM

bash
1docker run --name lgtm \
2 -p 3000:3000 \
3 -p 4317:4317 \
4 -p 4318:4318 \
5 --rm -ti \
6 grafana/otel-lgtm

Then open http://localhost:3000.

Default Grafana login:

  • username: admin
  • password: admin

2. Export OTLP env vars

bash
1export T3CODE_OTLP_TRACES_URL=http://localhost:4318/v1/traces
2export T3CODE_OTLP_METRICS_URL=http://localhost:4318/v1/metrics
3export T3CODE_OTLP_LOGS_URL=http://localhost:4318/v1/logs
4export OTEL_RESOURCE_ATTRIBUTES=deployment.environment.name=development

Optional:

bash
1export T3CODE_TRACE_MIN_LEVEL=Info
2export T3CODE_TRACE_TIMING_ENABLED=true

3. Launch the app from that same shell

CLI:

bash
1npx t3

Monorepo web/server dev:

bash
1node --run dev

Monorepo desktop dev:

bash
1node --run dev:desktop

Packaged desktop app:

Launch the actual app executable from the same shell so the desktop app and embedded backend inherit T3CODE_OTLP_*.

macOS app bundle example:

bash
1T3CODE_OTLP_TRACES_URL=http://localhost:4318/v1/traces \
2T3CODE_OTLP_METRICS_URL=http://localhost:4318/v1/metrics \
3T3CODE_OTLP_LOGS_URL=http://localhost:4318/v1/logs \
4"/Applications/T3 Code.app/Contents/MacOS/T3 Code"

Direct binary example:

bash
1T3CODE_OTLP_TRACES_URL=http://localhost:4318/v1/traces \
2T3CODE_OTLP_METRICS_URL=http://localhost:4318/v1/metrics \
3T3CODE_OTLP_LOGS_URL=http://localhost:4318/v1/logs \
4./path/to/your/desktop-app-binary

Do not rely on launching from Finder, Spotlight, the dock, or the Start menu after setting shell env vars. Those launches usually will not pick them up.

4. Fully restart after changing env

The backend reads observability config at process start. If you change OTLP env vars, stop the app completely and start it again.

How To Use Traces And Metrics To Debug The Server

Start With The Local Trace File

The trace file is the fastest way to inspect raw span data.

Resolve the path for the launch mode once. Production and explicitly configured homes store runtime state under the base directory's userdata folder:

bash
1TRACE_FILE="${T3CODE_HOME:-$HOME/.t3}/userdata/logs/server.trace.ndjson"

A dev server started from a linked worktree defaults to that worktree's local home:

bash
1TRACE_FILE="$WORKTREE/.t3/userdata/logs/server.trace.ndjson"

Only an implicit dev run outside a linked worktree uses the shared dev directory:

bash
1TRACE_FILE="$HOME/.t3/dev/logs/server.trace.ndjson"

Tail the selected file:

bash
1tail -f "$TRACE_FILE"

Show failed spans:

bash
1jq -c 'select(.type == "effect-span" and .exit._tag != "Success") | {
2 name,
3 durationMs,
4 exit,
5 attributes
6}' "$TRACE_FILE"

Show slow spans:

bash
1jq -c 'select(.durationMs > 1000) | {
2 name,
3 durationMs,
4 traceId,
5 spanId
6}' "$TRACE_FILE"

Inspect embedded log events:

bash
1jq -c 'select(any(.events[]?; .attributes["effect.logLevel"] != null)) | {
2 name,
3 durationMs,
4 events: [
5 .events[]
6 | select(.attributes["effect.logLevel"] != null)
7 | {
8 message: .name,
9 level: .attributes["effect.logLevel"]
10 }
11 ]
12}' "$TRACE_FILE"

Follow one trace:

bash
1jq -r 'select(.traceId == "TRACE_ID_HERE") | [
2 .name,
3 .spanId,
4 (.parentSpanId // "-"),
5 .durationMs
6] | @tsv' "$TRACE_FILE"

Filter orchestration commands:

bash
1jq -c 'select(.attributes["orchestration.command_type"] != null) | {
2 name,
3 durationMs,
4 commandType: .attributes["orchestration.command_type"],
5 aggregateKind: .attributes["orchestration.aggregate_kind"]
6}' "$TRACE_FILE"

Filter git activity:

bash
1jq -c 'select(.attributes["git.operation"] != null) | {
2 name,
3 durationMs,
4 operation: .attributes["git.operation"],
5 cwd: .attributes["git.cwd"],
6 hookEvents: [
7 .events[]
8 | select(.name == "git.hook.started" or .name == "git.hook.finished")
9 ]
10}' "$TRACE_FILE"

Use Tempo When You Need A Real Trace Viewer

Tempo is better than raw NDJSON when you want to:

  • search across many traces
  • inspect parent/child relationships visually
  • compare many slow traces
  • drill into one failing request without hand-joining by traceId

Recommended flow in Grafana:

  1. Open Explore.
  2. Pick the Tempo data source.
  3. Set the time range to something recent like Last 15 minutes.
  4. Start broad. Do not begin with a very narrow query.
  5. Look for spans from the t3code-server or t3code-desktop service, then narrow by span name or attributes.

Good first searches:

  • service name t3code-server or t3code-desktop, plus a resource attribute such as deployment.environment.name
  • span names like sendTurn or a Git operation such as GitVcsDriver.statusDetails.status
  • Git spans whose git.operation attribute identifies the operation
  • orchestration spans with attributes like orchestration.command_type

Once you know traces are arriving, narrower TraceQL queries for names such as sendTurn or Git operation names become useful.

Use Metrics To See Systemic Problems

Traces are best for one request. Metrics are best for trends.

Good metric families to watch:

  • t3_rpc_request_duration
  • t3_orchestration_command_duration
  • t3_orchestration_command_ack_duration
  • t3_provider_turn_duration
  • t3_git_command_duration

Counters tell you volume and failure rate:

  • t3_rpc_requests_total
  • t3_orchestration_commands_total
  • t3_provider_turns_total
  • t3_git_commands_total

Use metrics when the question is:

  • "is this always slow?"
  • "did this get worse after a change?"
  • "which command type is failing most often?"

Use traces when the question is:

  • "what happened in this specific request?"
  • "which child span caused this one slow interaction?"
  • "what logs were emitted inside the failing flow?"

What The New Ack Metric Means

t3_orchestration_command_ack_duration measures:

  • start: command dispatch enters the orchestration engine
  • end: the first committed domain event for that command is published by the server

That is a server-side acknowledgment metric. It does not measure:

  • websocket transit to the browser
  • client receipt
  • React render time

If you need those later, add client-side instrumentation or a dedicated server fanout metric.

Common Workflows

"Why did this request fail?"

  1. Start with the local NDJSON file.
  2. Find effect-span records where exit._tag != "Success".
  3. Group by traceId.
  4. Inspect sibling spans and span events.
  5. If needed, move to Tempo for the full trace tree.

"Why is the UI feeling slow?"

  1. Search for slow top-level spans in the trace file or Tempo.
  2. Check child spans for sqlite, git, provider, or terminal work.
  3. Look at the matching duration metrics to see whether the slowness is systemic.

"Did this command take too long to acknowledge?"

  1. Check t3_orchestration_command_ack_duration by commandType.
  2. If it is high, inspect the corresponding orchestration trace.
  3. Look at child spans for projection, sqlite, provider, or git work.

"Are git hooks causing latency?"

  1. Filter git.operation spans.
  2. Inspect git.hook.started and git.hook.finished events.
  3. Compare hook timing to the enclosing git span duration.

"Why do I have spans locally but nothing in Grafana?"

Usually one of these is true:

  • T3CODE_OTLP_TRACES_URL was not set
  • the app was launched from a different environment than the one where you exported the vars
  • the app was not fully restarted after changing env
  • Grafana is looking at the wrong time range or service name

If the local NDJSON file is updating, local tracing is working. The problem is almost always OTLP export configuration or process startup.

How To Think About Adding Tracing To Future Code

Prefer Boundaries Over Tiny Helpers

Good span boundaries:

  • RPC methods
  • orchestration command handling
  • provider adapter calls
  • external process calls
  • persistence writes
  • queue handoffs

Avoid tracing every tiny helper. Most helpers should inherit the active span rather than create a new one.

Reuse Effect.fn(...) Where It Already Exists

The codebase already uses Effect.fn("name") heavily. That should usually be your first tracing boundary.

For ad hoc work:

ts
1import { Effect } from "effect";
2
3const runThing = Effect.gen(function* () {
4 yield* Effect.annotateCurrentSpan({
5 "thing.id": "abc123",
6 "thing.kind": "example",
7 });
8
9 yield* Effect.logInfo("starting thing");
10 return yield* doWork();
11}).pipe(Effect.withSpan("thing.run"));

Put High-Cardinality Detail On Spans

Use span annotations for IDs, paths, and other detailed context:

ts
1yield *
2 Effect.annotateCurrentSpan({
3 "provider.thread_id": input.threadId,
4 "provider.request_id": input.requestId,
5 "git.cwd": input.cwd,
6 });

Keep Metric Labels Low Cardinality

Good metric labels:

  • operation kind
  • method name
  • provider kind
  • aggregate kind
  • outcome

Bad metric labels:

  • raw thread IDs
  • command IDs
  • file paths
  • cwd
  • full prompts
  • full model strings when a normalized family label would do

Detailed context belongs on spans, not metrics.

Use Logs As Span Events

Logs inside a span become part of the trace story:

ts
1yield * Effect.logInfo("starting provider turn");
2yield * Effect.logDebug("waiting for approval response");

Those messages show up as span events because Logger.tracerLogger is installed.

Use The Pipeable Metrics API

withMetrics(...) is the default way to attach a counter and timer to an effect:

ts
1import { someCounter, someDuration, withMetrics } from "../observability/Metrics.ts";
2
3const program = doWork().pipe(
4 withMetrics({
5 counter: someCounter,
6 timer: someDuration,
7 attributes: {
8 operation: "work",
9 },
10 }),
11);

Detailed API Reference

Runtime Wiring

The server observability layer is assembled in apps/server/src/observability/Layers/Observability.ts.

It provides:

  • pretty stdout logger
  • Logger.tracerLogger
  • local NDJSON tracer
  • optional OTLP trace exporter
  • optional OTLP metrics exporter
  • optional OTLP log exporter
  • Effect trace-level and timing refs

The desktop main process is a second producer, assembled in apps/desktop/src/app/DesktopObservability.ts. It reads the same T3CODE_OTLP_* names and the same Settings entries as the backend it supervises, and covers work the backend cannot see: app startup, window and menu handling, backend supervision, and updates. It reports as service t3code-desktop, so a collector shows it alongside the backend rather than mixed into it. It exports traces and logs only; the main process records no metrics, so the metrics endpoint applies to the backend alone.

Env Vars

Local trace file:

  • T3CODE_TRACE_FILE: override trace file path
  • T3CODE_TRACE_MAX_BYTES: per-file rotation size, default 10485760
  • T3CODE_TRACE_MAX_FILES: rotated file count, default 10
  • T3CODE_TRACE_BATCH_WINDOW_MS: flush window, default 200
  • T3CODE_TRACE_MIN_LEVEL: minimum trace level, default Info
  • T3CODE_TRACE_TIMING_ENABLED: enable timing metadata, default true

OTLP export:

  • T3CODE_OTLP_TRACES_URL: OTLP trace endpoint
  • T3CODE_OTLP_METRICS_URL: OTLP metric endpoint
  • T3CODE_OTLP_LOGS_URL: OTLP log endpoint
  • T3CODE_OTLP_EXPORT_INTERVAL_MS: export interval, default 10000
  • T3CODE_OTLP_HEADERS: extra headers for all three exporters, same format as OTEL_EXPORTER_OTLP_HEADERS: comma-separated key=value pairs with percent-encoded values.
  • T3CODE_OTLP_PROTOCOL: http/json (default) or http/protobuf

The server and the desktop app also read the standard OTEL_EXPORTER_OTLP_{TRACES,METRICS,LOGS}_ENDPOINT and generic OTEL_EXPORTER_OTLP_ENDPOINT (with /v1/traces, /v1/metrics, or /v1/logs appended), for a collector expecting those instead. A non-blank T3CODE_OTLP_*_URL wins over either, and a per-signal endpoint wins over the generic one for its signal. A blank value counts as unset. A signal with an OTEL endpoint takes its headers from OTEL_EXPORTER_OTLP_HEADERS and its protocol from OTEL_EXPORTER_OTLP_PROTOCOL (default http/protobuf, read case-insensitively), and a per-signal OTEL_EXPORTER_OTLP_{TRACES,METRICS,LOGS}_HEADERS or _PROTOCOL wins over the generic one for its signal. T3CODE_OTLP_HEADERS and T3CODE_OTLP_PROTOCOL never apply to it. An endpoint that is not an http or https URL, a protocol other than http/protobuf or http/json such as grpc, or headers that are not key=value pairs with percent-encoded values turn that signal's export off with a startup warning, rather than sending it to the Settings endpoint.

Service names are fixed: t3code-server for the backend and t3code-desktop for the desktop main process, both in service.namespace t3code. OTEL_SERVICE_NAME and a service.name or service.namespace in OTEL_RESOURCE_ATTRIBUTES are ignored. Tell installations apart with other resource attributes, such as OTEL_RESOURCE_ATTRIBUTES=deployment.environment.name=development.

If the OTLP URLs are unset, local tracing still works, metrics stay in-process only, and logs stay on stdout only.

The Kill Switch

T3CODE_OTEL_SDK_DISABLED and OTEL_SDK_DISABLED turn off every OTLP export in both the server and the desktop main process, overriding any endpoint from the environment or Settings. Local trace files and stdout logs are unaffected.

T3CODE_OTEL_SDK_DISABLED wins when set, so T3CODE_OTEL_SDK_DISABLED=false re-enables export on a machine that sets OTEL_SDK_DISABLED for everything else. It accepts the usual boolean spellings (true/false, yes/no, on/off, 1/0, y/n). OTEL_SDK_DISABLED follows the OpenTelemetry specification and only true disables export, so OTEL_SDK_DISABLED=1 does not. Values are case-insensitive and trimmed. An unrecognized value is ignored with a startup warning.

What Is Instrumented Today

Current high-value span and metric boundaries include:

  • Effect RPC websocket request spans from effect/rpc
  • RPC request metrics in apps/server/src/observability/RpcInstrumentation.ts
  • startup phases
  • orchestration command processing
  • orchestration command acknowledgment latency
  • provider session and turn operations
  • git command execution and git hook events
  • terminal session lifecycle
  • sqlite query execution

Current Constraints

  • logs outside spans are not persisted in the trace file; SSH-managed launch stdout/stderr is still captured in its launcher log
  • metrics are not snapshotted locally

Heap Snapshots

To see what a long-running server holds in memory, send it SIGUSR2. The server writes a V8 heap snapshot to its logs dir and logs the path. This works for desktop, npx t3, and service installs on macOS and Linux. Windows has no SIGUSR2.

Send the signal to the server pid in server-runtime.json, which sits in the server's state dir next to the logs dir. For a dev server or a --home-dir launch, use that server's state dir from Traces. Do not send it to the desktop app or the service launcher: a process without the handler exits on SIGUSR2. After a crash the file can keep a stale pid that now belongs to a different process, so check the pid first.

bash
1pid="$(jq .pid "${T3CODE_HOME:-$HOME/.t3}/userdata/server-runtime.json")"
2ps -p "$pid" -o command=

If ps shows the T3 Code server, send the signal:

bash
1kill -USR2 "$pid"

The file is <logsDir>/server-<pid>-<timestamp>.heapsnapshot, next to server.trace.ndjson. To open it, use the Memory tab in Chrome DevTools and select Load.

Before you take one:

  • The server stops while it writes the file. For a large heap this can take a minute or more. Connected clients can reconnect during the pause, and an event loop monitor, if the server has one, records the pause as a stall. Send the signal once. A second signal sent during a write takes another snapshot after the first one finishes.
  • The write needs about as much free memory as the heap uses. On a machine that is already swapping, it can make the problem worse or crash the server.
  • The file contains everything in server memory, including tokens, secrets, and thread content. Do not share it publicly. Delete it when you are done, because storage cleanup does not remove it.