Skip to content

synthkit runbook — credentials → telemetry in Grafana

This is the end-to-end path from a fresh checkout to visible synthetic telemetry in Grafana Cloud: configure credentials, sanity-check offline, push live, and verify metrics / traces / logs / RUM plus the control plane, Synthetic Monitoring, and Fleet Management registrations.

Conventions: examples use the generic stack placeholders <customer-stack> (the synthetic-data destination) and <staff-stack> (the generator's own self-observability + profiling — a separate stack with its own credentials). Replace them with your gcx context names. Never commit real stack names, IDs, or tokens — secrets live only in the gitignored .env.


0. What success looks like

When the run is healthy you will see, in the customer stack: - Mimir series for every declared construct (e.g. aws_rds_cpuutilization_average, kube_node_info, pg_stat_statements_calls_total). Blueprint-scoped constructs carry a blueprint=<name> selector; substrate-scoped ones (Kubernetes, dbo11y, CSP — including kube_node_info and pg_stat_statements_calls_total here) instead disambiguate by declared identity such as cluster or account_id and carry no blueprint label. - Tempo traces with end-to-end request correlation (service.name=<workload> → child DB span). - Loki streams for the app log stream (blueprint=<name>, source=app). - (optional) Faro/RUM beacons, Synthetic Monitoring check series, and Fleet Management collectors.

And locally/on the host: the operator UI at /control/ui with a green sink-readiness strip.


1. Prerequisites

  • Go 1.27 (for local runs) or Docker Engine 24.0+ with Docker Compose 2.24.4+ (for the containerised deploy).
  • Bash 5.0+, Python 3.11+, and just 1.58.0+ for the supported Compose path. The literal clean host bootstrap is Installation.
  • gcx 1.2.0+ configured with a context for the customer stack (and optionally the staff stack) only when remote verification is requested. See the gcx:setup-gcx skill if it is not yet set up.
  • The credential set for the customer stack. synthkit reads three independent destinations, each with its own token — never share GC_TOKEN across them: | Purpose | Env vars | Destination | |---|---|---| | Synthetic data (metrics+logs+traces) | GC_TOKEN, GC_PROM_RW(+GC_PROM_USER), GC_OTLP_ENDPOINT(+GC_OTLP_USER), GC_LOKI(+GC_LOKI_USER) | customer stack | | RUM (optional) | GC_FARO_COLLECTOR, GC_FARO_APP_KEY | customer stack | | Synthetic Monitoring provisioning (optional) | GC_SM_URL, GC_SM_TOKEN | customer stack tenant | | Fleet Management registration (optional) | GC_FM_URL, GC_FM_STACK_ID, GC_FM_TOKEN | customer stack | | Self-obs + profiling (optional) | GC_SELF_OTLP_*, GC_PYROSCOPE_* | staff stack (separate) |

The exact endpoint shapes are documented inline in .env.example.


2. Configure .env

install -m 600 .env.example .env       # then fill the values

Minimum for a live synthetic push: GC_TOKEN + GC_PROM_RW/GC_PROM_USER + GC_OTLP_ENDPOINT/GC_OTLP_USER + GC_LOKI/GC_LOKI_USER. Leave the optional blocks empty to disable RUM / SM / FM / self-obs.

DRY_RUN defaults to true — a live push is always an explicit opt-in (DRY_RUN=false). Keep comments on their own line (Docker env_file does not strip inline value # comment). BLUEPRINT_NAMES defaults empty, which is setup mode and emits no synthetic telemetry. Set one or more exact runtime names; use * only when the complete catalog is deliberate. Process self-observability is independent and may still send when SELFOBS_ENABLED=true.

To enable Fleet Management collector registration, fill the GC_FM_* triplet and ensure a blueprint declares a fleet_management construct (e.g. blueprints/k8s-full-stack.yaml). With the triplet empty, those collectors still emit metrics — they just are not registered with the FM API.


3. Sanity-check offline (always do this first)

Confirm the blueprints load and the series inventory is what you expect, with no network push:

DRY_RUN=true BLUEPRINT_NAMES=otlp-native go run ./cmd/synthkit -once -dump 2>&1 | less

Expected: a selected blueprints: 1 [otlp-native] line, its loaded blueprint line, a synthkit up: 1 blueprints line, and [dry-run promrw|loki|otlp] summaries. Spot-check a few names against signals/ — synthkit never invents names, so anything surprising is a bug.


4. Push live

Pick one path.

4a. Local foreground run

DRY_RUN=false go run ./cmd/synthkit

It binds the control plane on 127.0.0.1:8088 (loopback-safe). Open http://127.0.0.1:8088/control/ui. Wait one declared emission interval plus SEND_BATCH_DEADLINE before verifying (§5). Metric lanes use a 60-second floor unless their blueprint explicitly declares high_dpm.metric_interval. That override is bounded by MAX_DPM_PER_SERIES (6 DPM by default) and cannot run below TICK_DEFAULT; a few master ticks are not landing proof.

4b. Containerised deploy (the standing host)

The committed docker-compose.yml is secret-free. The emitter reads its full configuration from the service env file: .env by default, or the path selected by SYNTHKIT_ENV_FILE. Use the same override for every Compose command in one deployment. The one-shot SM job receives only its SM endpoint/token and explicit control flags through Compose interpolation. First-time setup on a new host — create the state bind-mount directory and give it to the container's user. The image is distroless and runs as uid 65532 (nonroot); the bind mount keeps the state file directly inspectable/editable on the host, but the dir must be writable by 65532 or every save fails (silently except for the surfaced error — see below):

Use the no-sudo, pinned-helper state preparation in Installation before the first Compose command. It rejects a symlink or non-directory and leaves /data owned by uid 65532.

For a first deployment, retain the image pin copied from .env.example, validate Compose with fake inputs, and wait for delivery-aware health:

# on the host clone (e.g. /opt/synthkit):
just compose-check
docker compose up -d --wait

For an existing deployment, do not replace .env or use a plain pull-and-up sequence. Follow the reproducible upgrade and rollback procedure: verify the exact candidate, record the running index/platform/config/image/version identities, stop and snapshot /data, compare-and-swap only SYNTHKIT_IMAGE_REF, recreate with --wait, and retain the concrete rollback target.

The host .env runs live (DRY_RUN=false) and binds 0.0.0.0:8088 inside the container so Docker's port mapping can reach it; host exposure is restricted separately by SYNTHKIT_BIND in the compose port mapping. Control state persists to the mounted /data volume (CONFIG_SNAPSHOT_PATH=/data/control-state.json, set in compose); the bind mount must be a directory owned by uid 65532 (distroless nonroot) — a single-file mount breaks the atomic save.

No control-state.json yet? That's normal until the first control-plane change — the snapshot is written lazily on the first mutation, not at startup. If a change you make in the operator UI doesn't stick across a restart, check persist.last_error in /control/status (§5.1): a permission denied there means the bind-mount dir isn't owned by uid 65532 — run the chown above. State reset is destructive maintenance. Stop the container and take a retained integrity snapshot first; never delete the file or directory contents as part of an upgrade.


5. Verify in Grafana

Before signal queries, prove that the container still runs the intended immutable release:

python3 scripts/synthkit-deploy.py inspect-running \
  --container "$(docker compose ps -q synthkit)" \
  --expected-reference "ghcr.io/rknightion/synthkit@sha256:<index>" \
  --expected-version "X.Y.Z" --expected-revision "<40-hex-source-sha>"

This closed report distinguishes the registry index, platform manifest, OCI config, running Docker image ID, binary version, and source revision without exposing configuration values.

5.1 Sink readiness (fastest signal)

curl -u control prompts without echoing the password; enter CONTROL_TOKEN (or press Enter for an intentionally token-free loopback deployment).

curl -s -u control http://127.0.0.1:8088/control/status | jq

Each sink shows last_success_ms advancing. Compare last_success_ms with last_error_ms: a newer success means the sink recovered even though cumulative failures remains non-zero. Inspect .queues separately: depth is current pressure, blocked_enqueues is cumulative backpressure, and dropped_items is authoritative loss after retry exhaustion. current_loss=false with non-zero dropped_items means new delivery recovered but historical loss remains; discarded items were not replayed. dry_run: true means you are not actually pushing — re-check DRY_RUN. The sink and queue states are also rendered in the operator UI.

5.2 Metrics (Mimir)

gcx --context <customer-stack> metrics query 'count by (blueprint) ({__name__=~"aws_rds_.+"})'
gcx --context <customer-stack> metrics query 'kube_node_info'
gcx --context <customer-stack> metrics query 'pg_stat_statements_calls_total'

Expect one series group per blueprint that declares the construct. (Use the gcx:explore-datasources skill to browse what landed.)

5.3 Traces (Tempo) — end-to-end request correlation

In Explore → Tempo (customer stack), search service.name="<your-service>" (or your workload) and confirm a trace whose root request span has a child DB span to the declared database. By default the span-derived RED metrics come from Tempo's metrics-generator; if you opt a blueprint into synthkit-native span metrics (the span_metrics_blueprints control toggle, off by default), they also appear as traces_spanmetrics_*{blueprint=<name>} in Mimir.

5.4 Logs (Loki)

gcx --context <your-stack> logs query '{blueprint="<blueprint-name>", source="app"} | json'

Expect structured app log lines (route, status, latency). High-cardinality fields are JSON payload fields, never stream labels.

5.5 Synthetic Monitoring (if selected and GC_SM_* set)

SM checks are provisioned by a one-shot control-plane command separate from the emitter. Preview performs read-only inventory calls; apply uses GC_SM_URL and GC_SM_TOKEN for remote writes:

docker compose --profile sm-provision run --rm sm-provision
SM_PROVISION_APPLY=true docker compose --profile sm-provision run --rm sm-provision
docker compose restart synthkit

Require optional_lanes[synthetic_monitoring] to be enabled/verified after the restart. Then wait the declared interval plus delivery deadline and query probe_*; no real probe execution occurs.

When rotating the SM credential or endpoint, first recreate synthkit with the new value so it writes a new suppressed snapshot. Run the provisioner preview with SM_PROVISION_MIGRATE_TARGET=true, then within 15 minutes rerun with both that flag and SM_PROVISION_APPLY=true, and restart synthkit. The migration fails closed if any recorded remote ID, key, managed specification, or revision changed.

5.6 Fleet Management

With a fleet_management blueprint, the metrics-only lane is valid without GC_FM_*. With the complete triplet, the separate registration lane registers each collector with the FM connect API at startup and heartbeats it every 45s. Open the Fleet Management app on the customer stack and confirm the fake collectors (linux/windows/darwin per the blueprint's collectors_per_os) appear; their collector_id/os/cluster attributes match the alloy_* metrics the construct emits. Verify fleet_metrics and fleet_registration separately in authenticated optional_lanes; never call the FM API for metrics-only mode.

5.7 Self-observability (if SELFOBS_ENABLED=true)

The generator's own telemetry ships to the staff stack: service.name=synthkit, metrics synthkit.* (push/tick/ledger.size/volume.multiplier/blueprint.count), per-tick traces, the operational log stream, and continuous profiles (service_name=synthkit). This is a separate data path from the synthetic telemetry above and never uses GC_TOKEN.


6. Operate (control plane)

The operator UI (/control/ui) drives the live runtime without a restart: master volume multiplier, per-blueprint incident scenarios, ad-hoc failure injection, live service/node scaling, and per-construct / per-kind / per-blueprint enable toggles. When CONTROL_TOKEN is set, HTTP Basic auth (username control) protects every mutation, all Infinity data routes, and control reads other than sanitized /control/readiness. /healthz is also public. The Infinity datasource's secure Basic-auth configuration covers server-side reads only. Browser-direct action buttons use a separate browser challenge, so verify them through a browser-trusted HTTPS endpoint. State persists across restarts via the snapshot file.

Security notes for shared-use deployments. Non-loopback startup fails unless CONTROL_TOKEN is non-empty and CONTROL_EXPOSURE_ACK is exactly trusted-network or tls-proxy. Use trusted-network only for an isolated plaintext path; use tls-proxy with a trusted HTTPS proxy and never send Basic credentials over untrusted plaintext HTTP. The UI and snapshot persist only the env-var name for git blueprint source tokens (for example MY_GIT_TOKEN); the resolved PAT is read from the environment at fetch time and never serialized. The owner-only snapshot still contains operational state and source metadata, so exclude it from untrusted backups.


7. Teardown

  • Local: Ctrl-C (graceful drain, bounded).
  • Container: docker compose down (state survives in the /data bind mount).
  • To stop emitting a blueprint without redeploying: disable it from the control UI, or delete its blueprints/*.yaml and restart — removing a blueprint affects nothing else.

8. First-value smoke checklist

  • .env filled with exact BLUEPRINT_NAMES; DRY_RUN=true … -once -dump inventory matches signals/.
  • DRY_RUN=false run; /control/status shows every sink last_success advancing, failures: 0.
  • Mimir: per-blueprint series present.
  • Tempo: end-to-end correlated trace (service → DB) present.
  • Loki: app log stream present.
  • (optional) SM checks, FM collectors, RUM beacons, self-obs on the staff stack.