Skip to content

Quick Start

Five steps from a fresh clone to verified synthetic telemetry in Grafana Cloud.


Step 1: Build

git clone https://github.com/rknightion/synthkit.git
cd synthkit
go build ./cmd/synthkit

Requires Go 1.27.0 or later. See Installation for the Docker path.


Step 2: Dry run one focused workload offline

Before touching credentials, select the bundled otlp-native reference blueprint and inspect its full selected series inventory:

DRY_RUN=true BLUEPRINT_NAMES=otlp-native ./synthkit -once -dump

BLUEPRINT_NAMES is the runtime selector. An empty or absent value starts setup mode and emits no synthetic telemetry; comma-separated exact names select only those blueprints; * is the explicit opt-in for the complete available catalog. An unknown name stops startup and lists the names that are available. -once runs a single tick and exits. -dump prints the complete series/label inventory to stdout — every metric name, label set, and example value that would be pushed. No network calls are made.

Expected output includes selected blueprints: 1 [otlp-native] before any sink can push, a loaded blueprint "otlp-native" line, synthkit up: 1 blueprints [otlp-native], and [dry-run promrw|loki|otlp] summaries.

The focused runtime identities are:

  • Blueprint: otlp-native
  • Cluster: otlp-native-prod-euw1
  • Workloads and Tempo service names: otlp-api-enriched and otlp-api-naked
  • Loki stream selector for the enriched service: {source="app",service_name="otlp-api-enriched",cluster="otlp-native-prod-euw1"}

To prove the Docker path is using your authored checkout rather than a cached published image, build with the local-source Compose override and then run the same isolated dry run:

test -e .env || install -m 600 .env.example .env
docker compose -f docker-compose.yml -f docker-compose.build.yml build synthkit
docker compose -f docker-compose.yml -f docker-compose.build.yml run --rm --no-deps \
  -e DRY_RUN=true -e BLUEPRINT_NAMES=otlp-native synthkit -once -dump

The build override tags the result synthkit:local and disables registry pulling, so this command uses the current checkout's Dockerfile, binary, and blueprints/ directory.

Use -dump to verify signal contracts

Spot-check a few metric names against the signals/ catalogue. synthkit never invents names — anything unexpected is a bug, not a configuration choice.


Step 3: Configure credentials

test -e .env || install -m 600 .env.example .env

This creates .env with mode 0600 without replacing an existing file. Keep DRY_RUN=true while you add credentials. The minimum set for a live push is:

GC_TOKEN=<your-CAP-token>
GC_PROM_RW=https://prometheus-prod-XX-<region>.grafana.net/api/prom/push
GC_PROM_USER=<mimir-instance-id>
GC_OTLP_ENDPOINT=https://otlp-gateway-<region>.grafana.net/otlp
GC_OTLP_USER=<stack-id>
GC_LOKI=https://logs-prod-XXX.grafana.net/loki/api/v1/push
GC_LOKI_USER=<loki-instance-id>
BLUEPRINT_NAMES=otlp-native
DRY_RUN=true

A single Cloud Access Policy token with metrics:write, logs:write, traces:write scopes covers all three sinks. See Credentials for the full table including optional RUM, Synthetic Monitoring, Fleet Management, and self-observability destinations.

DRY_RUN defaults to true

The shipped .env.example has DRY_RUN=true. You must explicitly set DRY_RUN=false to push live. This is intentional — a dry run can never accidentally write to a production stack.


Step 4: Push live

./synthkit -preflight
./plugins/synthkit/skills/initial-setup/scripts/set-env.sh DRY_RUN false .env
./synthkit

The preflight sends empty authenticated requests and must report ready for Prometheus, Loki, and OTLP before you disable dry-run. It prints only redacted lane states and reason codes.

synthkit loads the .env file automatically on startup. It runs the emit loop continuously (default tick: every 5 seconds). Let it run for a few ticks so cumulative counter series accumulate before querying.

To run a single tick and exit:

DRY_RUN=false ./synthkit -once

Step 5: Verify the focused workload

Fastest signal — the operator UI:

Open http://localhost:8088/control/ui in your browser. The sink-readiness strip shows the last push result for each sink (promrw, loki, otlp). Green = pushing successfully. If any sink shows failures, check the error detail there before querying Grafana.

Via the JSON API:

curl -u control prompts for CONTROL_TOKEN without echoing it. Press Enter only for an intentionally token-free loopback run.

curl -fsS -u control http://localhost:8088/control/status | jq -e '.dry_run == false'

This must print true; otherwise DRY_RUN is still set incorrectly. Each sink should also show last_success_ms advancing and failures: 0 in the operator UI.

jq is optional. If it is not installed, use the dependency-free form below and confirm that the response contains "dry_run":false:

curl -fsS -u control http://localhost:8088/control/status

In Grafana:

  1. Open Explore in your Grafana Cloud stack.
  2. In Prometheus, query the enriched service's native-OTLP histogram count:
    http_server_request_duration_seconds_count{service_name="otlp-api-enriched"}
    
    Expect one or more non-empty series. To also confirm the k8s-enriched identity, query:
    target_info{service_name="otlp-api-enriched",k8s_cluster_name="otlp-native-prod-euw1"}
    
    Expect a gauge-1 series. The native active-request metric is also non-empty and is a gauge, not a _total counter:
    http_server_active_requests{service_name="otlp-api-enriched"}
    
  3. In Loki, query:
    {source="app",service_name="otlp-api-enriched",cluster="otlp-native-prod-euw1"}
    
    Expect log lines after a few ticks.
  4. In Tempo, use TraceQL:
    { resource.service.name = "otlp-api-enriched" }
    
    Expect one or more traces; open one to confirm a root HTTP request span for the selected service. Repeat the three checks with otlp-api-naked to compare its non-enriched resource shape.

Ingestion lag

Mimir and Loki typically ingest within seconds. Tempo trace search has a short ingestion lag (30s–2m) before new traces appear in search results.


Next steps

  • RUNBOOK.md — deep verification walkthrough (sink readiness, gcx queries, end-to-end trace correlation check, log correlation, SM/FM verification)
  • Deployment — standing production deploy with docker-compose, the persistent state volume, and the host bind setup
  • Blueprints overview — write your own blueprint
  • Incidents & Scenarios — declare and activate failure scenarios