Skip to content

Configuration

Everything is supplied through the process environment. No deployment identifier is defaulted. A default org id or tenant would be one deployment's identifiers baked into everyone else's collector, and the failure is silent rather than loud: the scan authenticates, succeeds, and writes a plausible set of series into somebody else's tenant.

Required scan configuration

collector.config.load() requires the org, write stack, both endpoints and both tenants even with --dry-run. The bucket is required for S3-backed hydration, rate-card and diff reads and for publication. A dry run suppresses writes, not reads. Local publishing is refused; use deployed ECS task definitions for production or manual publishing.

VariableMeaning
GCINSIGHT_ORG_IDthe org to scan
GCINSIGHT_WRITE_STACKthe one stack results are published to
GCINSIGHT_MIMIR_URLhttps://prometheus-prod-NN-<region>.grafana.net
GCINSIGHT_MIMIR_TENANTthe write stack's hmInstancePromId
GCINSIGHT_LOKI_URLhttps://logs-prod-NNN.grafana.net
GCINSIGHT_LOKI_TENANTthe write stack's hlInstanceId
GCINSIGHT_S3_BUCKETthe deployment bucket

Credentials

VariableRead byRealm
GCINSIGHT_READ_TOKENcollectororg
GCINSIGHT_WRITE_TOKENcollectorthe write stack alone
GCINSIGHT_PROVISION_TOKENprovisioner onlyorg
GCINSIGHT_STACK_TOKEN_PREFIXcollectorSSM path holding per-stack reader tokens

GCINSIGHT_WRITE_TOKEN falls back to the read token when unset, so a single-credential interactive run works. A deployment sets both. Per-stack reader tokens are SSM SecureString values below the configured prefix. The provisioner alone reads the provision token; the collector never receives it.

In a deployed setup the Secrets Manager object must contain separate read and write token keys before scan schedules are enabled, plus the provisioner key if that opt-in task is created. Secret values are populated outside Terraform; the module manages the container or adopts one.

Runtime policy and regions

VariableGeneric default and contract
GCINSIGHT_S3_REGIONeu-west-1; bucket region for AWS operations and generated S3 URLs
GCINSIGHT_SSM_REGIONeu-west-1; region holding per-stack reader credentials
GCINSIGHT_STACK_TOKEN_PREFIX/gcinsight/stack-token; shared by scan and provisioner
GCINSIGHT_OPT_OUTempty; comma-separated stack slugs the owner asks not to provision
GCINSIGHT_COVERAGE_SCORE_WEIGHTSequal weights; partial JSON overrides for metrics, logs, traces, profiles, dashboard, alert, slo; finite non-negative numbers with positive total
GCINSIGHT_DASHBOARD_DETAIL_ENABLEDfalse; true/1 or false/0; opt-in dashboard JSON inspection for service attribution, with no retained query text
GCINSIGHT_READER_PRODUCT_READSreports opts in to T2 configured report counts (including disabled objects), not execution/delivery/scheduling activity; exactly reports:read at reports:*, only guarded GET /api/reports on the fresh inventory HTTPS origin with HTTP 200, no query/redirects; complete bare arrays only, unchanged credential guard, all report details discarded, empty measures zero and unreadable/partial is absent; no customer grant or universal visibility claim. Default empty; comma-separated slo, synthetic-monitoring, synthetic-monitoring-query, irm-integrations, faro-apps, ml-jobs, cloud-accounts, pdc-networks and/or reports; query token requires synthetic-monitoring and explicitly opts in to exact discovered SM datasource query plus probes read and count-only T2 collection; scan and provisioner must agree; legacy SM token alone makes no SM calls; missing or unreadable coverage is absent. irm-integrations enables T2 configured counts, not usage/activity; its permission also reaches secret-bearing configuration, but only the projected counters GET is permitted; no customer grant is implied. faro-apps adds only apps read (empty scope) and Kowalski plugin access, enabling T2 point-in-time web/mobile/unknown app counts; all app details are dropped at parse, unreadable stacks remain absent, and visibility is conditional on product permissions. ml-jobs adds only forecasting read (empty scope) and ML plugin access for T2 configured forecast job counts on the named jobs GET route; job-top-level grafanaApiKey is dropped before the unchanged structural credential guard, all other details are discarded, unknown coverage is absent, and transient key receipt is an accepted named-route risk, not a customer grant. cloud-accounts adds only CSP read (empty scope) and CSP plugin access for T2 configured AWS account counts on the fixed accounts GET, with validated fresh inventory origin and numeric stack ID, no query or redirects. It requires a data-only array and unchanged structural credential validation; account details are discarded. Genuine empty data measures zero; unknown/failed/partial inputs are absent. Other providers remain unknown, not zero; this is not an all-provider total. Backend write isolation of this credential remains unproven; the GET permission does not authorize live writes or customer rollout. pdc-networks adds only private-networks read (empty scope) and PDC plugin access for T2 stack-attributed policy counts on the fixed guarded accesspolicies GET. It requires complete reads of fresh inventory regions plus control realms, validating exact current stack realm and set:pdc-signing independently of server filtering. Complete empty reads measure zero; malformed/partial/unknown coverage is absent. All details are discarded; no tokens route, connections POST, policy writes or product metric. The permission also reaches tokens GET, which the collector never uses; staff conditional visibility is not universal completeness or a customer grant

Maturity ownership attribution excludes only identities containing @grafana.com (case-insensitive) in an Admin's login or email. Vendor and partner logins otherwise remain owner candidates; there is no configurable login exclusion list.

Expected retention and Fleet scrape policy are described below. The Terraform module exposes coverage_score_weights, dashboard_detail_enabled, provision_opt_out and provisioner_product_reads for the corresponding runtime policies. Do not use these tunables to store discovered inventory. A consumer must populate every projection field, even where the generic runtime has a default.

Schedules and retention

Module defaults are T1 hourly at :05, T2 daily, T3 six-hourly, T4 daily and the opt-in provisioner daily, with schedule_timezone = "UTC". tiers and provisioner_schedule_expression can override those defaults. See the operator timetable for the canonical times and enablement gates. Set schedules_enabled = false during initial setup: the module default is true. create_provisioner defaults to false; its schedule additionally depends on provisioner_enabled. Schedule changes require corresponding deadline, staleness-alert and carry-forward review, not just a cron edit.

scan_retention_days defaults to 90 positive whole days. It governs current scans/ objects and the reserved full-key prefix views/risk_label_hygiene.json, not all views. Age starts at last publication; hydration republishes and resets it. Noncurrent versions expire after seven days and AWS lifecycle processing is asynchronous. Adopted buckets need equivalent targeted retention in their existing lifecycle policy before raw-match publication. See Security.

Dashboard build

Build-time Grafana credentials are separate from runtime credentials and should be short-lived. The build token is not a runtime secret.

VariableMeaning
GCINSIGHT_VIEWS_DIRread views from a local directory instead of S3
GCINSIGHT_WRITE_STACK_URLhttps://<slug>.grafana.net
GCINSIGHT_WRITE_STACK_IDnumeric stack id
GCINSIGHT_GRAFANA_TOKENshort-lived build token

The builder resolves the insights folder by title. An offline build supplies --ds-uid, local views and a synthetic bucket name for generated URLs; that placeholder JSON must not be published.

Immutable consumer identity

collector.identity.PROJECTION_ENVS is the exact non-secret projection contract. A deployment owns these values in its manifest; use the consumer tools rather than editing a task's environment by hand. GCINSIGHT_RUNTIME_CONFIG_DIGEST verifies the resolved projection; GCINSIGHT_REQUIRE_EXPLICIT_CONFIG=1 requires it and all non-optional projection values. See Consumer upgrades and rollback.

VariablesMeaning
GCINSIGHT_METRIC_PREFIX, GCINSIGHT_LOKI_JOB, GCINSIGHT_USER_AGENTemitted metric namespace, Loki job and publisher user agent
GCINSIGHT_ROLE_NAME, GCINSIGHT_ROLE_DISPLAY, GCINSIGHT_ROLE_GROUPper-stack custom-role identity
GCINSIGHT_READER_SA_NAME, GCINSIGHT_ADMIN_SA_NAME, GCINSIGHT_TOKEN_NAME_PREFIXpersistent reader, transient provisioning Admin and minted-token names
GCINSIGHT_DASHBOARD_UID_PREFIX, GCINSIGHT_DASHBOARD_TITLE_PREFIX, GCINSIGHT_DASHBOARD_TAGgenerated dashboard identity
GCINSIGHT_DASHBOARD_DS_NAME, GCINSIGHT_DASHBOARD_FOLDER_TITLEInfinity datasource and insights-folder lookup names
GCINSIGHT_INSIGHTS_FOLDER_UIDexplicit target folder for alert publication; dashboard publication resolves the folder by title
GCINSIGHT_PROM_DS_UIDalert Prometheus datasource uid; generic default grafanacloud-prom
GCINSIGHT_ALERT_RULE_GROUP, GCINSIGHT_ALERT_RULE_UIDS_JSONalert group and JSON mapping from rule keys to stable uids
GCINSIGHT_ALERT_TITLE_PREFIX, GCINSIGHT_ALERT_TITLE_SEPARATOR, GCINSIGHT_ALERT_SERVICE_LABELalert title and service-label identity
GCINSIGHT_GCX_CONTEXToptional context default for bin/trace.py --live; the usage probe requires explicit --context and --out and ignores this variable; not a collector credential

Generic identity defaults live in collector.identity, collector.provision, bin/dashboards.py, bin/alerts.py and the emitters. Changing provisioner names affects live reconciliation; changing output identity affects dashboard queries and alerts. Treat either as a reviewed consumer change, not cosmetic renaming.

The rate card

Optional, and read from config/ratecard.csv in the deployment bucket by the task role. Absence means volume-only panels, which is a supported state rather than a degraded one.

ratecard.example.csv in the repository is the format reference.

Ten dimensions can be priced. price() returns None, never 0.0, when a dimension is not priced - an unpriced dimension must read as unknown, not free. A deployment may price only some dimensions, but the UI discloses which components are omitted and must not present the subtotal as a complete estate total.

Configuration errors, all rejected rather than coerced:

  • mixed currencies;
  • duplicate dimensions;
  • unsupported units;
  • non-positive prices;
  • unsupported dimensions or billing bases, wrong fixed divisors, invalid included quantities, or a period other than month.

Currency and billing period come from the card. Metrics-series pricing is per 1,000 series where declared, and metrics support two explicit bases:

  • base_rate_only excludes DPM;
  • dpm_aware applies max(active_series, total_dpm / included_dpm) per stack, using live usage inputs and a dedicated dashboard calculation. It never falls back to the two-input base-series saving.

Fleet Management default scrape interval

GCINSIGHT_FLEET_DEFAULT_SCRAPE_INTERVAL is the organisation's expected scrape cadence, as a Go duration such as 60s, 1m or 1m30s. It defaults to 60s. Terraform exposes it as fleet_default_scrape_interval.

The hourly tier already lists every Fleet Management pipeline; it parses each pipeline's declared scrape_interval values, plus any explicit collection_interval on other wired OTel pull receivers, in memory and compares them with this default, so the check adds no Fleet Management calls. An enabled pipeline that reaches active collectors and declares a shorter interval raises DPM and is counted per stack, listed in risk_fleet_scrape_intervals and raises the paused fleet_fast_scrape alert rule. An omitted scrape_interval is the component default of 60s; an omitted collection_interval is not read, because its default differs per receiver. An interval that is not a literal (a module argument or an environment lookup) is counted as unparsed, never as the default. Local Alloy and collector configs outside Fleet Management are not visible.

Expected Loki retention policy

GCINSIGHT_EXPECTED_RETENTION_POLICY is an optional JSON list of selector and minimum_period objects. Terraform exposes the same value as expected_retention_policy and passes it into every scan task. It is genuine deployment policy, not an estate inventory:

[{"selector":"{service=\"example\"}","minimum_period":"14d"}]

The default is an empty list. In that state the policy metric is absent and the policy-gap view is empty. A configured selector is evaluated only against readable effective retention_stream data; an unreadable or undisclosed response is not called compliant and is not called a breach. Selectors can contain customer label names and values, so the redacted runtime configuration logs only the number of configured expectations.

Optional Firehose logs

The collector writes its own structured Loki records, but it cannot report an image-pull failure, bootstrap error, early traceback or OOM kill - by the time any of those happen there is no collector to do the writing. The optional Firehose path forwards ECS CloudWatch logs to Loki and is off by default.

Enable it in three stages, in this order:

  1. set firehose_logs_enabled=true with a dedicated adopted secret containing {"api_key":"<loki-tenant>:<logs-write-token>"};
  2. send a deliberate test record and verify it lands in Loki, and that the failed-record S3 path works;
  3. only then set firehose_log_subscription_enabled=true.

The subscription switch cannot stand alone. Failed deliveries have their own encrypted, lifecycle-bound bucket.

Optional panel plugins

The shipped dashboards need no third-party panel plugins. If you adopt panels that use them, these are the minimum versions verified as Grafana 13.3 compatible from their installed manifests:

Plugin IDMinimum verified version
volkovlabs-echarts-panel7.2.5
volkovlabs-table-panel3.6.5
volkovlabs-variable-panel5.2.0
marcusolsson-treemap-panel2.1.1

They remain optional unless an adopted panel requires one of them.