For maintainers. Using T3 Code? See docs/user.
T3 Code has one server-side observability model:
The local trace file is the persisted source of truth for normal local launches. Those launches do not
write a separate server log file, but SSH-managed launches also persist the remote process's
stdout/stderr at ~/.t3/ssh-launch/<state>/server.log.
Logs are human-facing:
Logger.consolePretty()~/.t3/ssh-launch/<state>/server.logIf you want a log message to show up in the trace file, emit it inside an active span with Effect.log.... Logger.tracerLogger will attach it as a span event.
Configuring a logs endpoint takes over that job. The server then exports log records, which cover
every message instead of only the ones inside an active span and carry the trace and span ids so
they still line up with the trace. Logger.tracerLogger is dropped in that mode, so the same
message is not exported twice and the trace file stops carrying log messages. stdout output and
SSH-managed launch persistence stay unchanged either way.
Completed spans are written as NDJSON records to serverTracePath. The default depends on how the
server starts: production and explicitly configured homes use
<home>/userdata/logs/server.trace.ndjson (so ~/.t3/userdata/... by default, or
/custom/path/userdata/... with --home-dir /custom/path), a linked worktree dev run uses
<worktree>/.t3/userdata/logs/server.trace.ndjson, and an implicit dev run outside a linked
worktree uses ~/.t3/dev/logs/server.trace.ndjson.
Important fields common to both record types:
type: effect-span or otlp-spanname: span nametraceId, spanId, parentSpanId: correlationdurationMs: elapsed timeattributes: structured contextevents: embedded logs and custom eventseffect-span records also contain exit with Success, Failure, or Interrupted. otlp-span
records instead carry OTLP resource, scope, and optional status fields.
The TraceRecord, EffectTraceRecord, and OtlpTraceRecord schemas live in
packages/shared/src/observability.ts.
DPoP proof failures include the safe environment.dpop.failure_code span
attribute. A time_window failure means that a signed proof was too old or too
far in the future for the environment server's allowed window. It can point to
a date or time problem on either device, but it can also result from a delayed
request.
t3 trace summary reads the trace file and its rotated backups directly, so it works while the
server is stalled or stopped. It prints counts, rates, and latency percentiles per span name. Use
it to measure background work or to compare two builds.
| 1 | t3 trace summary --since 30m --limit 40 |
It reads T3CODE_TRACE_FILE if set, else <home>/userdata/logs/server.trace.ndjson for
--base-dir or T3CODE_HOME, plus the T3CODE_TRACE_MAX_FILES rotated backups. For a dev run or
a copied file, set T3CODE_TRACE_FILE. --since 30m keeps spans that ended in the last 30
minutes. The rate is per minute between the first and last span end.
Metrics are not written to a local file.
apps/server/src/observability/Metrics.tsIf OTLP is not configured, metrics still exist in-process, but you will not have a local artifact to inspect.
apps/server/src/observability/EventLoopMonitor.ts samples the server's event loop every 30 s. When
the loop stalled for more than 2 s since the previous sample, it records a root
server.eventLoop.stall span with a warning. The span has trace level Warn, so it stays when
T3CODE_TRACE_MIN_LEVEL is Warn. The warning shows in Settings > Diagnostics unless OTLP logs are
on. The span time is when the sample ran, not when the stall happened.
Some delay is not recorded:
delayMaxMs is the longest stall, and can undercount it by up to 1 s. The 2 s threshold applies to
this value, so a stall over 3 s is normally recorded, and a shorter one can be missed. A stall
that ends just as a sample runs can be missed too.delayMaxMs. Busy time covers the whole
window, so a short sleep in an otherwise busy window can still record a false stall. The span then
shows CPU time far below delayMaxMs.CPU times and page faults cover the whole process over the whole window since the previous sample.
The window is nominally 30 s, but a long stall delays the sample and makes the window longer. Other
work in the window can hide a wait, so only CPU time far below delayMaxMs proves the thread was
waiting. Read CPU together with page faults:
cpuSystemMs with many page faults means memory pressure. Major faults are reads from disk or swap.
On macOS, reads from compressed memory are minor faults plus system CPU.cpuUserMs with few page faults means JavaScript work or garbage collection.involuntaryContextSwitches mean other processes were competing for the CPU.Provider event NDJSON files still exist for provider runtime streams. Those are separate from the main server trace file.
There are two useful modes:
server.trace.ndjsonThe local trace file is always on. OTLP export is opt-in.
You do not need any extra env vars. Just run the app normally and inspect server.trace.ndjson.
Examples:
| 1 | npx t3 |
| 1 | node --run dev |
| 1 | node --run dev:desktop |
| 1 | docker run --name lgtm \ |
| 2 | -p 3000:3000 \ |
| 3 | -p 4317:4317 \ |
| 4 | -p 4318:4318 \ |
| 5 | --rm -ti \ |
| 6 | grafana/otel-lgtm |
Then open http://localhost:3000.
Default Grafana login:
adminadmin| 1 | export T3CODE_OTLP_TRACES_URL=http://localhost:4318/v1/traces |
| 2 | export T3CODE_OTLP_METRICS_URL=http://localhost:4318/v1/metrics |
| 3 | export T3CODE_OTLP_LOGS_URL=http://localhost:4318/v1/logs |
| 4 | export OTEL_RESOURCE_ATTRIBUTES=deployment.environment.name=development |
Optional:
| 1 | export T3CODE_TRACE_MIN_LEVEL=Info |
| 2 | export T3CODE_TRACE_TIMING_ENABLED=true |
CLI:
| 1 | npx t3 |
Monorepo web/server dev:
| 1 | node --run dev |
Monorepo desktop dev:
| 1 | node --run dev:desktop |
Packaged desktop app:
Launch the actual app executable from the same shell so the desktop app and embedded backend inherit T3CODE_OTLP_*.
macOS app bundle example:
| 1 | T3CODE_OTLP_TRACES_URL=http://localhost:4318/v1/traces \ |
| 2 | T3CODE_OTLP_METRICS_URL=http://localhost:4318/v1/metrics \ |
| 3 | T3CODE_OTLP_LOGS_URL=http://localhost:4318/v1/logs \ |
| 4 | "/Applications/T3 Code.app/Contents/MacOS/T3 Code" |
Direct binary example:
| 1 | T3CODE_OTLP_TRACES_URL=http://localhost:4318/v1/traces \ |
| 2 | T3CODE_OTLP_METRICS_URL=http://localhost:4318/v1/metrics \ |
| 3 | T3CODE_OTLP_LOGS_URL=http://localhost:4318/v1/logs \ |
| 4 | ./path/to/your/desktop-app-binary |
Do not rely on launching from Finder, Spotlight, the dock, or the Start menu after setting shell env vars. Those launches usually will not pick them up.
The backend reads observability config at process start. If you change OTLP env vars, stop the app completely and start it again.
The trace file is the fastest way to inspect raw span data.
Resolve the path for the launch mode once. Production and explicitly configured homes store runtime
state under the base directory's userdata folder:
| 1 | TRACE_FILE="${T3CODE_HOME:-$HOME/.t3}/userdata/logs/server.trace.ndjson" |
A dev server started from a linked worktree defaults to that worktree's local home:
| 1 | TRACE_FILE="$WORKTREE/.t3/userdata/logs/server.trace.ndjson" |
Only an implicit dev run outside a linked worktree uses the shared dev directory:
| 1 | TRACE_FILE="$HOME/.t3/dev/logs/server.trace.ndjson" |
Tail the selected file:
| 1 | tail -f "$TRACE_FILE" |
Show failed spans:
| 1 | jq -c 'select(.type == "effect-span" and .exit._tag != "Success") | { |
| 2 | name, |
| 3 | durationMs, |
| 4 | exit, |
| 5 | attributes |
| 6 | }' "$TRACE_FILE" |
Show slow spans:
| 1 | jq -c 'select(.durationMs > 1000) | { |
| 2 | name, |
| 3 | durationMs, |
| 4 | traceId, |
| 5 | spanId |
| 6 | }' "$TRACE_FILE" |
Inspect embedded log events:
| 1 | jq -c 'select(any(.events[]?; .attributes["effect.logLevel"] != null)) | { |
| 2 | name, |
| 3 | durationMs, |
| 4 | events: [ |
| 5 | .events[] |
| 6 | | select(.attributes["effect.logLevel"] != null) |
| 7 | | { |
| 8 | message: .name, |
| 9 | level: .attributes["effect.logLevel"] |
| 10 | } |
| 11 | ] |
| 12 | }' "$TRACE_FILE" |
Follow one trace:
| 1 | jq -r 'select(.traceId == "TRACE_ID_HERE") | [ |
| 2 | .name, |
| 3 | .spanId, |
| 4 | (.parentSpanId // "-"), |
| 5 | .durationMs |
| 6 | ] | @tsv' "$TRACE_FILE" |
Filter orchestration commands:
| 1 | jq -c 'select(.attributes["orchestration.command_type"] != null) | { |
| 2 | name, |
| 3 | durationMs, |
| 4 | commandType: .attributes["orchestration.command_type"], |
| 5 | aggregateKind: .attributes["orchestration.aggregate_kind"] |
| 6 | }' "$TRACE_FILE" |
Filter git activity:
| 1 | jq -c 'select(.attributes["git.operation"] != null) | { |
| 2 | name, |
| 3 | durationMs, |
| 4 | operation: .attributes["git.operation"], |
| 5 | cwd: .attributes["git.cwd"], |
| 6 | hookEvents: [ |
| 7 | .events[] |
| 8 | | select(.name == "git.hook.started" or .name == "git.hook.finished") |
| 9 | ] |
| 10 | }' "$TRACE_FILE" |
Tempo is better than raw NDJSON when you want to:
traceIdRecommended flow in Grafana:
Explore.Tempo data source.Last 15 minutes.t3code-server or t3code-desktop service, then narrow by span name or
attributes.Good first searches:
t3code-server or t3code-desktop, plus a resource attribute such as
deployment.environment.namesendTurn or a Git operation such as GitVcsDriver.statusDetails.statusgit.operation attribute identifies the operationorchestration.command_typeOnce you know traces are arriving, narrower TraceQL queries for names such as sendTurn or Git
operation names become useful.
Traces are best for one request. Metrics are best for trends.
Good metric families to watch:
t3_rpc_request_durationt3_orchestration_command_durationt3_orchestration_command_ack_durationt3_provider_turn_durationt3_git_command_durationCounters tell you volume and failure rate:
t3_rpc_requests_totalt3_orchestration_commands_totalt3_provider_turns_totalt3_git_commands_totalUse metrics when the question is:
Use traces when the question is:
t3_orchestration_command_ack_duration measures:
That is a server-side acknowledgment metric. It does not measure:
If you need those later, add client-side instrumentation or a dedicated server fanout metric.
effect-span records where exit._tag != "Success".traceId.t3_orchestration_command_ack_duration by commandType.git.operation spans.git.hook.started and git.hook.finished events.Usually one of these is true:
T3CODE_OTLP_TRACES_URL was not setIf the local NDJSON file is updating, local tracing is working. The problem is almost always OTLP export configuration or process startup.
Good span boundaries:
Avoid tracing every tiny helper. Most helpers should inherit the active span rather than create a new one.
Effect.fn(...) Where It Already ExistsThe codebase already uses Effect.fn("name") heavily. That should usually be your first tracing boundary.
For ad hoc work:
| 1 | import { Effect } from "effect"; |
| 2 | |
| 3 | const runThing = Effect.gen(function* () { |
| 4 | yield* Effect.annotateCurrentSpan({ |
| 5 | "thing.id": "abc123", |
| 6 | "thing.kind": "example", |
| 7 | }); |
| 8 | |
| 9 | yield* Effect.logInfo("starting thing"); |
| 10 | return yield* doWork(); |
| 11 | }).pipe(Effect.withSpan("thing.run")); |
Use span annotations for IDs, paths, and other detailed context:
| 1 | yield * |
| 2 | Effect.annotateCurrentSpan({ |
| 3 | "provider.thread_id": input.threadId, |
| 4 | "provider.request_id": input.requestId, |
| 5 | "git.cwd": input.cwd, |
| 6 | }); |
Good metric labels:
Bad metric labels:
Detailed context belongs on spans, not metrics.
Logs inside a span become part of the trace story:
| 1 | yield * Effect.logInfo("starting provider turn"); |
| 2 | yield * Effect.logDebug("waiting for approval response"); |
Those messages show up as span events because Logger.tracerLogger is installed.
withMetrics(...) is the default way to attach a counter and timer to an effect:
| 1 | import { someCounter, someDuration, withMetrics } from "../observability/Metrics.ts"; |
| 2 | |
| 3 | const program = doWork().pipe( |
| 4 | withMetrics({ |
| 5 | counter: someCounter, |
| 6 | timer: someDuration, |
| 7 | attributes: { |
| 8 | operation: "work", |
| 9 | }, |
| 10 | }), |
| 11 | ); |
The server observability layer is assembled in apps/server/src/observability/Layers/Observability.ts.
It provides:
Logger.tracerLoggerThe desktop main process is a second producer, assembled in
apps/desktop/src/app/DesktopObservability.ts. It reads the same T3CODE_OTLP_* names and the same
Settings entries as the backend it supervises, and covers work the backend cannot see: app startup,
window and menu handling, backend supervision, and updates. It reports as service
t3code-desktop, so a collector shows it alongside the backend rather than mixed into it. It
exports traces and logs only; the main process records no metrics, so the metrics endpoint applies
to the backend alone.
Local trace file:
T3CODE_TRACE_FILE: override trace file pathT3CODE_TRACE_MAX_BYTES: per-file rotation size, default 10485760T3CODE_TRACE_MAX_FILES: rotated file count, default 10T3CODE_TRACE_BATCH_WINDOW_MS: flush window, default 200T3CODE_TRACE_MIN_LEVEL: minimum trace level, default InfoT3CODE_TRACE_TIMING_ENABLED: enable timing metadata, default trueOTLP export:
T3CODE_OTLP_TRACES_URL: OTLP trace endpointT3CODE_OTLP_METRICS_URL: OTLP metric endpointT3CODE_OTLP_LOGS_URL: OTLP log endpointT3CODE_OTLP_EXPORT_INTERVAL_MS: export interval, default 10000T3CODE_OTLP_HEADERS: extra headers for all three exporters, same format as
OTEL_EXPORTER_OTLP_HEADERS: comma-separated key=value pairs with percent-encoded values.T3CODE_OTLP_PROTOCOL: http/json (default) or http/protobufThe server and the desktop app also read the standard
OTEL_EXPORTER_OTLP_{TRACES,METRICS,LOGS}_ENDPOINT and generic OTEL_EXPORTER_OTLP_ENDPOINT (with
/v1/traces, /v1/metrics, or /v1/logs appended), for a collector expecting those instead. A
non-blank T3CODE_OTLP_*_URL wins over either, and a per-signal endpoint wins over the generic one
for its signal. A blank value counts as unset. A signal with an OTEL endpoint takes its headers from
OTEL_EXPORTER_OTLP_HEADERS and its protocol from OTEL_EXPORTER_OTLP_PROTOCOL (default
http/protobuf, read case-insensitively), and a per-signal
OTEL_EXPORTER_OTLP_{TRACES,METRICS,LOGS}_HEADERS or _PROTOCOL wins over the generic one for its
signal. T3CODE_OTLP_HEADERS and T3CODE_OTLP_PROTOCOL never apply to it. An endpoint that is not
an http or https URL, a protocol other than http/protobuf or http/json such as grpc, or
headers that are not key=value pairs with percent-encoded values turn that signal's export off
with a startup warning, rather than sending it to the Settings endpoint.
Service names are fixed: t3code-server for the backend and t3code-desktop for the desktop main
process, both in service.namespace t3code. OTEL_SERVICE_NAME and a service.name or
service.namespace in OTEL_RESOURCE_ATTRIBUTES are ignored. Tell installations apart with other
resource attributes, such as OTEL_RESOURCE_ATTRIBUTES=deployment.environment.name=development.
If the OTLP URLs are unset, local tracing still works, metrics stay in-process only, and logs stay on stdout only.
T3CODE_OTEL_SDK_DISABLED and OTEL_SDK_DISABLED turn off every OTLP export in both the server and
the desktop main process, overriding any endpoint from the environment or Settings. Local trace
files and stdout logs are unaffected.
T3CODE_OTEL_SDK_DISABLED wins when set, so T3CODE_OTEL_SDK_DISABLED=false re-enables export on a
machine that sets OTEL_SDK_DISABLED for everything else. It accepts the usual boolean spellings
(true/false, yes/no, on/off, 1/0, y/n). OTEL_SDK_DISABLED follows the
OpenTelemetry specification and only true disables export, so OTEL_SDK_DISABLED=1 does not.
Values are case-insensitive and trimmed. An unrecognized value is ignored with a startup warning.
Current high-value span and metric boundaries include:
effect/rpcapps/server/src/observability/RpcInstrumentation.tsserver.eventLoop.stall)To see what a long-running server holds in memory, send it SIGUSR2. The server writes a V8 heap
snapshot to its logs dir and logs the path. This works for desktop, npx t3, and service installs
on macOS and Linux. Windows has no SIGUSR2.
Send the signal to the server pid in server-runtime.json, which sits in the server's state dir
next to the logs dir. For a dev server or a --home-dir launch, use that server's state dir from
Traces. Do not send it to the desktop app or the service launcher: a process without the
handler exits on SIGUSR2. After a crash the file can keep a stale pid that now belongs to a
different process, so check the pid first.
| 1 | pid="$(jq .pid "${T3CODE_HOME:-$HOME/.t3}/userdata/server-runtime.json")" |
| 2 | ps -p "$pid" -o command= |
If ps shows the T3 Code server, send the signal:
| 1 | kill -USR2 "$pid" |
The file is <logsDir>/server-<pid>-<timestamp>.heapsnapshot, next to server.trace.ndjson. To
open it, use the Memory tab in Chrome DevTools and select Load.
Before you take one: