first-tree ships a vendor-neutral observability stack built on pino (logs) and OpenTelemetry (traces). Tracing is off by default — configure an OTLP endpoint to enable it.
No setup required. The server emits structured logs to stdout in a
human-readable format during development, and as NDJSON in production
(NODE_ENV=production).
Configure an OTLP/HTTP endpoint and matching headers — any backend that speaks OTLP works (Logfire, Honeycomb, Jaeger, Tempo, SigNoz, Axiom, …).
# server.yaml (SaaS internal — typically injected as env vars in production)
observability:
tracing:
endpoint: https://<your-otlp-endpoint>/v1/traces
headers:
Authorization: "Bearer <write-token>"
serviceName: first-tree
environment: production
sampleRate: 1.0Or via environment variables (preferred for secret tokens):
FIRST_TREE_OTEL_ENDPOINT=https://<your-otlp-endpoint>/v1/traces
FIRST_TREE_OTEL_HEADERS="Authorization=Bearer <write-token>"Restart the server. You should see tracing enabled: exporter=otlp-http ...
in the startup log.
| Path | Env var | Default | Notes |
|---|---|---|---|
observability.logging.level |
FIRST_TREE_LOG_LEVEL |
info |
trace / debug / info / warn / error / fatal |
observability.logging.format |
— | pretty (dev), json (prod) |
Pretty for humans, JSON for log collectors |
observability.logging.bridgeToSpanLevel |
— | error |
Minimum pino level whose records are attached to the active span. warn / error / off |
observability.tracing.endpoint |
FIRST_TREE_OTEL_ENDPOINT |
"" (disabled) |
OTLP/HTTP traces URL |
observability.tracing.headers |
FIRST_TREE_OTEL_HEADERS |
"" |
key1=val1,key2=val2 format |
observability.tracing.exporter |
— | otlp-http |
otlp-http or otlp-grpc |
observability.tracing.serviceName |
— | first-tree |
Shown in trace backends |
observability.tracing.environment |
— | development |
deployment.environment.name OTel attr |
observability.tracing.sampleRate |
— | 1.0 |
0.0–1.0, ratio applied at root |
First Tree uses Sentry for product-side error monitoring on the Web Console and local Client runtime. These are separate Sentry projects because the Web Console is a hosted browser surface while the Client daemon runs on operator-owned machines.
| Surface | Sentry project | Default behavior | Disable path |
|---|---|---|---|
| Web Console | first-tree-web |
Enabled when VITE_SENTRY_DSN is configured in the web build. |
Omit VITE_SENTRY_DSN or set VITE_SENTRY_ENABLED=false for local/dev builds. |
| Client daemon/runtime | first-tree-client |
Enabled when FIRST_TREE_CLIENT_SENTRY_DSN is configured. |
Set FIRST_TREE_CLIENT_SENTRY_ENABLED=false. |
Both surfaces tag events with the git SHA used to build the artifact:
- Web reads
FIRST_TREE_WEB_BUILD_IDat build time, falling back togit rev-parse HEAD. - Client reads the git SHA baked into the published CLI package, or
FIRST_TREE_GIT_SHA,FIRST_TREE_CLIENT_GIT_SHA, orGITHUB_SHAwhen an operator override is needed.
VITE_SENTRY_DSN=https://public@example.ingest.sentry.io/1
VITE_SENTRY_ENVIRONMENT=production
VITE_SENTRY_TRACES_SAMPLE_RATE=0.1
FIRST_TREE_WEB_BUILD_ID=$GITHUB_SHAThe Vite build generates hidden source maps and uploads them through
@sentry/vite-plugin when SENTRY_AUTH_TOKEN and SENTRY_ORG are present.
The upload target defaults to first-tree-web; override with
SENTRY_PROJECT_WEB only for a temporary diagnostic build. Missing upload
credentials or upload failures emit warnings and do not fail the deployable
artifact.
The Docker release workflow passes VITE_SENTRY_DSN and
VITE_SENTRY_ENVIRONMENT from GitHub repository/environment variables into
the Web build, which makes Sentry enabled by default for the managed release
path without committing the public DSN to the repo.
FIRST_TREE_CLIENT_SENTRY_DSN=https://public@example.ingest.sentry.io/2
FIRST_TREE_CLIENT_SENTRY_ENVIRONMENT=production
FIRST_TREE_CLIENT_SENTRY_ENABLED=true
FIRST_TREE_GIT_SHA=$GITHUB_SHAThe npm publish workflow bakes FIRST_TREE_CLIENT_SENTRY_DSN from GitHub
repository/environment variables into the published CLI package. Local operator
env still wins at runtime, so a machine can override the DSN or set
FIRST_TREE_CLIENT_SENTRY_ENABLED=false to disable Client events explicitly.
For background services, put FIRST_TREE_CLIENT_SENTRY_* overrides in the
user-owned daemon.env file under FIRST_TREE_HOME. The daemon loads that file
before Sentry initialization, so the operator's choice survives daemon restarts
without baking observability settings into the launchd/systemd unit.
The Client scrubber drops user identity, request bodies, cookies, query strings, breadcrumbs, bearer tokens, OAuth/token-like fields, and local home / workspace path prefixes before sending events. Runtime content fields such as provider prompts, model output, tool output, stdout, and stderr are redacted.
The Web Console loads Microsoft Clarity for production session insights. The
project id is xj2f9syfng.
Clarity is loaded from packages/web/index.html only when
window.location.hostname is cloud.first-tree.ai. Local development hosts and
staging hosts such as dev.cloud.first-tree.ai do not fetch the Clarity SDK and
do not write into the production project. This mirrors the GA4 host gate in the
same SPA shell and avoids requiring a Vite env var, Docker build argument, or
secret for this fixed production tag.
The Web Console treats Clarity as layout/session telemetry, not content
telemetry. The React root (#root) carries data-clarity-mask="true" so chat
messages, agent traces, connect commands, tokens, repo paths, and other
customer/workspace text rendered by the app are masked before upload. Do not add
data-clarity-unmask inside the app unless the surface is public/static and the
privacy boundary has been reviewed with the exact content classes it can render.
A common deployment pattern is to point all environments (dev, staging, prod, …) at the same Logfire / Honeycomb / … project, and rely on the UI's filter/group features to tell them apart. The server supports this out of the box — no need to maintain multiple projects or tokens.
Every span the server emits carries three OTel resource attributes that trace backends treat as first-class:
| Attribute | Value | Configured by |
|---|---|---|
service.name |
first-tree (customizable) |
observability.tracing.serviceName |
deployment.environment.name |
development / staging / production / … |
observability.tracing.environment or FIRST_TREE_OTEL_ENVIRONMENT |
service.instance.id |
srv_<8-char-hex> — unique per process |
auto-generated at startup |
The SaaS Docker image runs node packages/server/dist/index.mjs; pass the
OTLP env vars to the container (e.g. via the platform's secrets store):
docker run \
-e FIRST_TREE_OTEL_ENDPOINT=https://logfire-us.pydantic.dev/v1/traces \
-e FIRST_TREE_OTEL_HEADERS="Authorization=Bearer <token>" \
-e FIRST_TREE_OTEL_ENVIRONMENT=production \
-e FIRST_TREE_DATABASE_URL=... \
ghcr.io/agent-team-foundation/first-tree:latestSwap production → staging for the staging instance; same token, different
env label.
- Logfire: top filter bar → add
deployment.environment.name = production - Honeycomb: dataset filter
deployment.environment.name = production - Jaeger / Tempo: search → Tags →
deployment.environment.name=production - SigNoz / Axiom: resource attribute filter with the same key
One project-per-environment is only necessary if:
- Strict data isolation is required (e.g. prod data must never commingle with dev for compliance) — then use distinct tokens per environment and rotate them independently.
- Retention policies differ per environment — some backends apply retention at the project level.
- Cost attribution needs to be per environment — most backends bill per-project.
Otherwise, a single project with deployment.environment.name filtering is
simpler to operate and makes cross-environment comparisons trivial.
When you scale out to multiple server instances behind a load balancer, each
process gets its own service.instance.id at startup. To drill into a
specific replica in your trace backend, filter by that attribute — this is
how you answer "which replica is spiking latency" without manual tagging.
All runtime logs — server and client — go through the same pino pipeline with the same pretty/json formats. To keep grep-ability and structured- search working across a growing codebase, follow these five rules when adding new log lines.
// ✅
log.error({ err }, "failed to reload config");
log.info({ count }, "cleaned up stale sessions");
// ❌
log.error({ err }, "Failed to reload config.");
log.info("CLEANUP DONE!");One shape for error messages makes skimming faster.
// ✅
log.error({ err, entryId }, "failed to deliver inbox entry");
// ❌
log.error({ err }, "delivery failed");
log.error({ err }, "inbox entry delivery error");// ✅
log.info({ agentId }, "bound agent");
log.info({ count }, "marked agents as stale");
// ❌
log.info({ agentId }, "binding agent...");
log.info({ agentId }, "agent will be bound");// ✅ — structured, grep / Loki-filterable
log.warn({ entryId, retries }, "retry budget exhausted");
// ❌ — a needle in a haystack at scale
log.warn(`entry ${entryId} gave up after ${retries} retries`);Rule of thumb: if a downstream consumer (Loki / Datadog / a human running
grep) might want to filter by a value, it belongs in attrs.
The [Module] prefix is emitted automatically by the pretty formatter.
const log = createLogger("Inbox");
// ✅ output: INFO [Inbox] entry expired
log.info({ entryId }, "entry expired");
// ❌ output: INFO [Inbox] inbox: entry expired
log.info({ entryId }, "inbox: entry expired");- CLI status output (
first-treeconnect / client banners,status(...)helpers) is a different channel — it's interactive, user-facing, and can use formatting / localized strings. It does not go through pino. - Existing log lines that predate these rules. Follow the style when you touch a file; don't do bulk rewrites purely for style.
- Natural exceptions: acronyms stay capitalized (
"failed to parse JWT"), proper nouns stay capitalized ("connected to Logfire").
Values below go into observability.tracing.endpoint / .headers.
endpoint: https://logfire-us.pydantic.dev/v1/traces # or logfire-eu.*
headers:
Authorization: "Bearer pylf_v1_us_xxxxxxxxxxxx"
Get a write token at Logfire → Settings → Write Tokens.
endpoint: https://api.honeycomb.io:443/v1/traces
headers:
x-honeycomb-team: "<ingest-key>"
endpoint: http://<jaeger-host>:4318/v1/traces
endpoint: https://tempo-us-central1.grafana.net/tempo/v1/traces
headers:
Authorization: "Basic <base64(user:token)>"
endpoint: https://api.axiom.co/v1/traces
headers:
Authorization: "Bearer <api-token>"
x-axiom-dataset: "<dataset-name>"
Every delivery path stamps the same attribute set. Search in your trace backend for any of:
inbox.entry.id = "<id>"— single inbox entry, enqueue → deliver → ackmessage.id = "<id>"— one business message across chats / fan-outschat.id = "<id>"— every span on a given chatagent.id = "<id>"— all spans touching a specific agent
Search ws.connection spans filtered by client.id = "<id>". The span's
duration shows connection lifetime; ws.close.code attribute shows the
reason the socket closed.
Error responses return a x-trace-id header and include traceId in the
body. Copy it into your trace backend to jump straight to the full tree.
Inject the env vars when launching the SaaS server container:
docker run \
-e FIRST_TREE_LOG_LEVEL=debug \
-e FIRST_TREE_OTEL_ENDPOINT=https://... \
-e FIRST_TREE_OTEL_HEADERS="Authorization=Bearer ..." \
ghcr.io/agent-team-foundation/first-tree:latest
| Deployment | Sample rate | Rationale |
|---|---|---|
| Local dev | 1.0 |
Full trace on every request, fast feedback |
| Staging | 1.0 |
Rare enough traffic that full sample is cheap |
| Prod, low-traffic (<10 req/s) | 1.0 |
Still cheap at this volume |
| Prod, medium-traffic | 0.1–0.25 |
Representative sample without overwhelming the backend |
| Prod, high-traffic (>100 req/s) | 0.01–0.05 |
Cap backend cost; relies on tail-based or attribute-biased sampling at the backend for error-path coverage |
Sampling is ParentBased(TraceIdRatioBased(sampleRate)) — inbound
traceparent headers from upstream services are respected.
- WebSocket traces appear "Incomplete" while connected. Our
ws.connectionspan is deliberately long-running (lives for the full socket lifetime), so Logfire and other trace UIs will show it as "Incomplete" or "missing root span" until the client disconnects. This is expected — the span finalizes with attributes likews.close.codeonce the socket closes. For short-lived per-message visibility, search byws.message.typeor theclient.id/agent.idattributes directly. - Orphan spans at async boundaries. When a message is enqueued and
delivered in different async roots (later tick, different WS connection),
the deliver span is a new trace root rather than a child of the original
request. Use attribute-based search (see "Find the full story" above).
Fixing this requires persisting W3C
traceparenton inbox rows and is tracked as tech debt until multi-replica deployment makes it necessary. - No client-side tracing. Client (
@first-tree/client) emits logs only. Agent-side work is observed indirectly via server-side spans (ws.connection,ws.message, inbox attrs). - No PG span. PostgreSQL queries are not wrapped in spans — for query
performance investigation, PostgreSQL's own
log_min_duration_statement+pg_stat_statementsis typically sufficient. If cross-query correlation becomes a need, this should be reconsidered as a dedicated feature, not a flag-gated default. - Server restart loses in-flight trace context. Messages enqueued but not yet delivered keep their business state but lose the original request's trace linkage. Attribute search still works.