Skip to content

Capture & Tooling

synthkit ships several supporting tools for capturing real environments, generating blueprint drafts, generating Grafana dashboards, and LLM-assisted workflow skills.

skcapture — environment snapshot

skcapture inspects a live Kubernetes environment via kubectl and writes a versioned, optionally age-encrypted inventory file. It has zero synthkit imports — no blueprint, construct, or runner code. This trust boundary is enforced by internal/capture.TestCaptureTrustBoundary.

Typical use: run skcapture inside the target cluster (as a kubectl-shell container or a one-shot Job), encrypt the output, and retrieve it for processing with skforge outside the cluster.

# Encrypted output (recommended for production clusters)
skcapture \
  --passphrase-file /path/to/passphrase \
  --out capture.age

# Plain JSON (for local development or non-sensitive environments)
skcapture --plain --out capture.json

# Restrict to specific namespaces
skcapture --plain \
  --namespaces my-app,my-services \
  --out capture.json

Key flags:

FlagDefaultDescription
--out <path>capture.ageOutput file path.
--passphrase-file <path>—Path to a file containing the encryption passphrase. Required unless --plain.
--plainfalseWrite unencrypted JSON. Mutually exclusive with --passphrase-file.
--namespaces <list>(all)Comma-separated namespace allow-list.
--exclude-namespaces <list>kube-system,kube-node-lease,kube-publicComma-separated namespace deny-list.
--collectors <list>k8sComma-separated list of enabled collectors (currently: k8s).
--version—Print tool version and schema version, then exit.

The capture output is a versioned JSON Inventory envelope containing resource kinds (nodes, namespaces, deployments, statefulsets, daemonsets, services, ingresses, addons). The schema version is embedded in the output.

The current v1 focus is AWS/EKS. The Inventory struct is designed to support additional collectors in future versions.

What the zero-secret default covers

A capture leaves the cluster, so the boundary of the zero-secret default is worth stating exactly.

Covered. Secret and arbitrary ConfigMap values are never captured. The former --include-secret-data and --include-configmap-data flags were removed because they never implemented data capture, so no shipped RBAC promises that access. The default RBAC in deploy/skcapture/rbac.yaml grants no access to Secrets or ConfigMaps. The separate identity grant permits reading only one named ConfigMap's cluster and self-reporting-metric.prom keys, which provide collector identity and chart version rather than captured object data. Object annotations are reduced to a fixed allowlist of the keys the tooling consumes: Helm release identity, Argo CD tracking, deployment revision, and metric scrape hints. Every other annotation is dropped, including kubectl.kubernetes.io/last-applied-configuration, which on any cluster managed with kubectl apply embeds the object's full spec and therefore every container environment value.

Not covered. A capture is still a description of your environment. It carries namespace, workload, service and ingress names; container image references including the registry host; ingress hostnames; ExternalName service targets; pod-template labels; and node instance types and pool names. Treat the output as sensitive and use the encrypted path for anything leaving the cluster. A credential placed in a workload name, a label value or an image reference is captured, because those are identity fields the tool has to read to describe the environment.

Cluster identity and where it came from

The captured cluster name is the primary join key for everything forged from a capture: get it wrong and the resulting blueprint emits telemetry that can never join to the real cluster's dashboards, while still loading and validating cleanly. skcapture therefore records which source produced the name in the cluster's name_source field, and prints it on completion.

name_sourceMeaning
collector-release-infoRead from the in-cluster metrics collector's release-info ConfigMap. This is the name the collector applies as a label to every metric, log and trace the cluster ships, so it is the only source guaranteed to join to that cluster's real telemetry.
eks-arn-contextRecovered from an EKS ARN kubeconfig context. The cluster's AWS identity, which is not necessarily the name its telemetry carries.
kubeconfig-contextA slug of whatever the current kubeconfig context happens to be called. Describes your kubeconfig, not the cluster.
defaultNothing was discoverable, and a placeholder was used.

Resolution starts with the collector release-info ConfigMap, then uses the kubeconfig fallbacks. Anything other than collector-release-info prints a warning: verify the name against a live signal before forging a blueprint from that capture.

The collector lookup is a targeted get of a named ConfigMap in the collector's own namespace — never a namespace-wide or cluster-wide list — and reads only the two known release-info keys. The cluster key supplies the authoritative name. The self-reporting-metric.prom key is parsed only for the exact grafana_kubernetes_monitoring_build_info{version="..."} line and populates monitoring.chart_version; arbitrary metric lines and ConfigMap data are ignored. The default RBAC grants no ConfigMap access, so an in-cluster run under that role falls back and says so. To get the authoritative identity and chart version from an in-cluster run, customise deploy/skcapture/rbac-collector-identity.yaml with the observed namespace and exact release-info name before applying it; it grants only get on that one ConfigMap, never list or Secret access.

Provider and platform detection

cluster.provider records the provider label family found on the nodes: eks, gke, or aks. When no supported family is present it is undetermined, so skforge can surface the missing evidence instead of treating the cluster as AWS by default. Karpenter-only EKS nodes are recognised from the karpenter.k8s.aws/ label family even when they have no eks.amazonaws.com/ labels.

Addon recognition combines the allowlisted Helm release name with known namespace and workload names. The capture currently recognises Crossplane (crossplane-system, crossplane), external-secrets (external-secrets), the GitHub Actions runner controller (arc-systems, gha-rs-controller and gha-runner-scale-set*), the GitHub-to-OTel bridge (github2otel), and OpenCost (opencost). These entries deliberately retain an empty addon kind when there is no standalone construct. Forge keeps one narrow image fallback for Crossplane provider workloads whose name and namespace are not recognised by capture. In the forge coverage report, Crossplane, external-secrets, the runner controller, and github2otel are no matching construct gaps. OpenCost is an unmapped name: its cost surface is modelled by the registered k8s_cluster construct's k8s_monitoring.opencost option. Karpenter's construct models node autoscaler telemetry, so it does not make the Actions runner controller a modeled product.

Published in-cluster Job

The shipped Job pins both containers to the stable release image ghcr.io/rknightion/synthkit-skcapture:1.4.0, built from Dockerfile.skcapture through the shared container-publish reusable. A Git release tag named vX.Y.Z publishes the image as ghcr.io/rknightion/synthkit-skcapture:<X.Y.Z> without the leading v. Release v1.4.0 verified both images' signatures and provenance on 2026-09-06. For development against main, use the edge override form ghcr.io/rknightion/synthkit-skcapture:main-<shortsha>. Follow deploy/skcapture/README.md: apply base RBAC, create the passphrase Secret from a protected file, apply the Job, wait for the capture container to finish, and copy the encrypted file from output-hold before its ten-minute window ends. Waiting for the whole Job first loses that retrieval window. Keep decrypted output outside Git, then delete the Job and its RBAC. The local k3d recipe substitutes its local image.

The live validation on 2026-09-06 ran the capture container under the base ServiceAccount: the permission audit allowed the seven declared topology resource types, denied Secret listing, and the capture log contained no Forbidden, 403, cannot list, or cannot get hits. The encrypted output decrypted successfully and its recursive secret-field scan was empty against a cluster carrying real Secrets. Check permissions from inside the pod when a cluster proxy does not honour impersonation; an operator-context --as result alone is not proof. The API warns that webhook permissions may be absent from auth can-i --list, so also check Secret denial explicitly. Admission-injected init containers can delay startup.

Base RBAC intentionally does not read collector ConfigMaps. This live capture detected EKS, node groups and workloads, but recorded name_source: default; no permission-denial log entry means neither that collector identity was recovered nor that the forged blueprint has the correct telemetry join key. Use the separately documented named-ConfigMap grant when that identity is needed. Decrypt with just forge inspect ... --key ... while redirecting stdout to a protected file; just --quiet suppresses the recipe output, including the JSON. Retrieve first, then wait for the whole Job to complete, and finally verify both namespace and cluster RBAC cleanup.

Disposable k3d proof of the Job path

Maintainers can exercise the shipped Job, its RBAC, encryption, and non-EKS forge refusal without touching a customer or lab cluster:

just skcapture-k3d

The harness builds Dockerfile.skcapture locally, imports skcapture:dev into one disposable k3d cluster, rewrites both Job container image references to that local image, then applies only deploy/skcapture/rbac.yaml and the rewritten Job manifest. It copies the encrypted output from output-hold using the same documented kubectl cp path and runs skforge inspect followed by the default skforge prompt path. Results, including the capture and the explicit forge result, are written below artifacts/skcapture-k3d/.

This is deliberately a non-EKS contrast. A k3d capture reports undetermined; the default forge command must then refuse before writing a blueprint, and the result record names that behaviour. The v1 loader requires a cloud block whenever an environment declares a cluster, so it cannot load a cloudless pure-Kubernetes skeleton today. The refusal tells the operator the captured provider and names --assume-aws, which is available only when the operator independently knows the cluster is AWS. Set SKCAPTURE_K3D_CLUSTER_NAME only to an unused lowercase name; the harness refuses to delete a pre-existing cluster.

skforge — blueprint forge

skforge takes a captured inventory and produces a synthkit blueprint draft. It uses a deterministic skeleton mapper to translate inventory resources into blueprint declarations, then emits a self-contained LLM prompt that you feed to Claude (or another LLM) to produce the final blueprint YAML.

# Inspect a capture file
skforge inspect capture.age --key /path/to/passphrase

# Generate an LLM prompt + optional coverage report for an EKS capture
skforge prompt capture.age \
  --key /path/to/passphrase \
  --report coverage.md > blueprint-prompt.txt

# A non-EKS/undetermined capture is refused by default. Use this only when the
# operator independently knows the captured cluster is AWS.
skforge prompt capture.age --key /path/to/passphrase --assume-aws > blueprint-prompt.txt

# Validate a blueprint draft
skforge validate my-blueprint.yaml

Subcommands

inspect — decrypt (or read plain) a capture file and print it as indented JSON for inspection.

prompt — the main workflow step. An EKS capture produces the deterministic AWS/EKS skeleton and a self-contained LLM prompt. A gke, aks, or undetermined capture fails before a skeleton is written, naming the observed provider and --assume-aws. The v1 schema requires cloud identity for every cluster, so a cloudless skeleton is not loadable. Pass --assume-aws only when independent evidence establishes that an otherwise undetected cluster is AWS; the resulting AWS placeholder remains called out as an explicit assumption. Pass --report <path> to also write a coverage report showing which resources mapped to which construct kinds and which were not covered.

validate — load a blueprint file through the real registry and cardinality projection. Prints OK, Name, Cardinality, Estimated, and any Diagnostics. Exits non-zero if the blueprint is invalid. Useful for confirming an LLM-generated draft is structurally correct before running synthkit.

OK:          true
Name:        my-service
Cardinality: 847
Estimated:   true

The capture → forge → blueprint workflow

  1. Capture: run skcapture in the target environment to produce an inventory file.
  2. Forge prompt: run skforge prompt to generate a self-contained LLM prompt.
  3. LLM authoring: feed the prompt to Claude (or another LLM). The prompt is self-contained — it includes the full construct catalog description so the LLM can work without any other context.
  4. Validate: run skforge validate on the resulting blueprint.yaml to confirm it loads cleanly.
  5. Deploy: place the validated blueprint in blueprints/ and run synthkit.

See custom-blueprints.md for how to add custom blueprints at runtime.

synthkit-dash — Grafana dashboard generator

synthkit-dash generates Grafana v2 dashboard JSON for a blueprint's synthetic telemetry. It resolves the blueprint, derives the signal manifest (via internal/dashgen), runs registered per-blueprint templates, and writes dashboard JSON files to an output directory.

just dashgen \
  -blueprint blueprints/my-service.yaml \
  -out dashboards/my-service/ \
  -datasource metrics=grafanacloud-prom \
  -datasource logs=grafanacloud-logs \
  -datasource traces=grafanacloud-traces
FlagRequiredDescription
-blueprint <path>yesPath to the blueprint YAML.
-out <dir>yesOutput directory for generated dashboard JSON files.
-integrations <path>noOptional integrations config YAML for thin-index deep-links.
-folder <uid>noUID for the generated Folder resource (defaults to <blueprint>-dashboards).
-datasource <group=name>yes, repeatedExplicit datasource name for every rendered query group. Generation fails if any group is unmapped.

synthkit-dash emits a Folder resource, a thin index dashboard, a metrics dashboard, and a panel inventory. Additional per-blueprint dashboards are generated when templates are registered for that blueprint. Files are written to the explicit output directory (never stdout). Every inventory row records its dashboard, panel, datasource, ref ID, and rendered query.

For blueprints that define recording/alert rules, synthkit-dash also emits a <blueprint>-rules.json file.

Push generated dashboards to Grafana with gcx:

gcx resources push -p dashboards/my-service/

After collecting one normalized observation per inventory query, classify every panel:

just dashgen -verify-inventory dashboards/my-service/panel-inventory.json \
  -observations observations.json \
  -verification-out verification.json

synthkit-control-dash — control dashboard generator

synthkit-control-dash generates the customer-facing self-serve control dashboard: an Infinity-datasource-backed Grafana v2 dashboard with volume and scenario knobs as read panels and native action buttons.

go run ./cmd/synthkit-control-dash \
  -ds-name "My Infinity DS" \
  -out dashboards/control/
FlagRequiredDescription
-ds-name <name>yesInfinity datasource name in Grafana.
-out <dir>yesOutput directory for generated JSON.
-write-base-url <url>required with infinityBase URL for action-button POSTs; it must not end in /control because the generator appends /control/.... In fetch mode the browser must reach it; in infinity mode it must be an absolute HTTP or HTTPS URL without embedded credentials, a query or a fragment that the datasource can reach. An empty value leaves relative paths in fetch mode.
-blueprints <dir>noDirectory of blueprint YAML files for enumerating scenarios (default ./blueprints).
-action-mode <mode>nofetch (default): browser-direct POST buttons. infinity: Grafana sends each POST server-side through the Infinity datasource. Requires -ds-uid, -write-base-url, and Grafana's vizActionsAuth feature toggle.
-ds-uid <uid>with infinityInfinity datasource UID used by server-side actions.

Without -write-base-url, the buttons POST to relative /control/... paths on Grafana's own origin. They work only when a reverse proxy on that origin routes /control/ to the synthkit control plane. Otherwise pass a browser-reachable -write-base-url.

In infinity mode, plain HTTP is acceptable only for the hop from Grafana through Private Data Source Connect (PDC), or from inside the cluster, to a ClusterIP control-plane Service. Never expose the control plane over plain HTTP across the public internet; use HTTPS elsewhere.

When CONTROL_TOKEN is set, the Infinity datasource stores the credential for server-side reads and infinity actions; fetch actions use the browser's separate Basic challenge. No token is embedded in the dashboard JSON. Anyone who can query the Infinity datasource can issue authenticated control POSTs and change load or scenarios, even without viewing this dashboard or using its action buttons. Restrict datasource query access and dashboard visibility to operators. Grafana's Viewer restriction on dashboard actions does not protect the datasource from direct queries. See Grafana's feature note.

LLM-assisted skills

synthkit ships agent skills for Claude Code, Codex, and OpenCode. Skills are authored once under plugins/synthkit/skills/; just skills-sync creates symlinks in .claude/skills/ and .agents/skills/ so the same skills are available in all three harnesses. just skills-check verifies the symlink farm and is safe for CI.

Available skills

SkillDescription
/initial-setupWalks through credentials, environment variables, and first-run verification. Start here when deploying synthkit for the first time.
/create-blueprintGuided blueprint authoring — asks about your infrastructure and applications and produces a blueprint YAML.
/setup-fleet-managementConfigures Grafana Fleet Management collector registration for a blueprint.
/verify-deploymentEnd-to-end deployment verification — checks credentials, runs -once -dump, and confirms telemetry is reaching Grafana Cloud.

Using skills in Claude Code

Open the synthkit repository in Claude Code and run a skill directly:

/initial-setup
/create-blueprint
/verify-deployment

Using skills from outside the repo (plugin install)

Install synthkit as a Claude Code plugin from any directory. Enter these slash commands in Claude Code chat; they are not shell commands:

/plugin marketplace add rknightion/synthkit
/plugin install synthkit@synthkit

After installation the skills are available as /synthkit:initial-setup, /synthkit:create-blueprint, /synthkit:setup-fleet-management, and /synthkit:verify-deployment.

Cross-harness compatibility

The same skills work in Codex (reads .agents/skills/) and OpenCode (reads .claude/skills/). Both directories are populated by just skills-sync. Install the tool on a new machine and run just skills-sync to get the skills available immediately.

For more on authoring custom blueprints, see custom-blueprints.md. For Fleet Management setup, see fleet-management.md. For the full CLI reference, see cli.md.

Local documentation validation

The documentation hub generates its complete Zensical configuration externally, so a fresh clone cannot run that private build locally. Validate the repository-owned navigation, every relative Markdown/HTML link, and the intentional 404.md page with:

just docs-check

The command uses only Python 3.11+ standard-library tomllib; it does not install dependencies or contact the documentation hub. Run it before opening a documentation change. The same check runs in .github/workflows/trigger-docs-sync.yml before the cross-repository sync is dispatched.