Screenshots¶
What the demo looks like once it has been running with traffic on for a few hours. Every screenshot comes from a live deployment of this repository: Grafana dark theme, the last three to twenty-four hours, and only the fictional Touchline Times data the demo generates (its readers, its five developers at touchline.example, its teams and its agents). The dashboards are the four this module provisions; the other pages are Grafana Cloud apps reading the same telemetry.
For what to click and what to say in front of an audience, see the Demo runbook. For which source feeds each tab, see Data source tiers.
Agent Observability¶
The Grafana Cloud Agent Observability app, reading the generations, evaluations and guard results that the in-app agents and the Claude Code plugin send it.
Overview, agents and tools¶
The analytics landing page: requests, error rate, evaluation pass rate, latency, cost and tokens at a glance, with what needs attention.
The agents list: the five in-app Touchline agents and the Claude Code agent, with generations, errors, latency, tokens and estimated cost for each.
The orchestrator agent: three prompt versions, with pass rate by evaluator and tokens, latency and error rate per version.
The orchestrator's Prompts tab: the system prompt Agent Observability recorded for the current version.
The orchestrator's Quality tab: evaluator pass rate per prompt version, with the latest version marked as improved.
The Claude Code agent, fed by the Agent Observability plugin on the five developer containers.
The Claude Code agent's tool inventory: built-in tools plus the MCP servers pushed by the gateway and the developers' own.
Tool performance across every agent: executions, error rate and p95 latency per tool.
Conversations¶
Every reader question to the match desk as a conversation, with the agents and models involved.
One conversation opened: the reader's question, the news agent's get_news call and replies, and the flow of generations by agent with tokens, cost and duration.
Evaluations and guards¶
The three online evaluations: answer quality and PII for the agents, content safety and PII for Claude Code, and responsible-gambling language for the compliance agent.
The guards: redaction and evaluator guards with their priorities, with the Claude Code PII gate denying requests and the injected tool result guard warning.
Experiments¶
The scheduled experiments: completed runs of the match desk comparison suite, with trials, cost and tokens.
Agentic app dashboard¶
The in-app agents dashboard, built from OpenTelemetry GenAI metrics in Mimir, Tempo spans and Agent Observability counters.
Overview: site API requests, agent generations, input and output tokens, and where the work happens.
Request path: from the browser and the site API through the agent chain, with links to recent traces.
Cost and tokens: estimated model cost split by agent and model, the token mix, and orchestrator cost per prompt version.
Latency and tools: generation and tool latency by agent, and which tools each agent calls.
Evaluations and guards: online quality scores by agent and outcome, and guard actions on tool results.
Traces: recent site and agent traces, each row linking to Tempo and to its Agent Observability conversation.
Service graph edges: calls per hour from the site to the orchestrator and on to each specialist agent.
Claude Code dashboard¶
The Claude Code dashboard, built from Claude Code's own OpenTelemetry metrics, logs and traces, which the gateway's managed settings switch on.
OpenTelemetry overview: spend, sessions, tokens and active time per developer and team.
Cost and tokens: the token mix, how much the cache saves, and spend by model and developer.
Prompts and sessions: one developer session read back as a transcript, with the prompts and responses around it.
MCP servers and tools: tool calls per MCP server, and which servers the gateway manages versus the developers' own.
MCP attribution: spend, tokens and sessions for each MCP server and tool.
Hooks, permissions and plugins: hook runs and latency, permission decisions and the plugins loaded.
Agent Observability plugin: generation cost, Cloud guard outcomes and online evaluation scores for Claude Code.
Cloud guard outcomes per rule, with the PII gate denies alongside the passes.
Traces: each Claude Code turn as a trace, with turn latency, time to first token and the developer who asked.
Bedrock invocation log, gateway caller: the same traffic from Bedrock's side, where every developer is the gateway's one IAM identity.
Claude apps gateway dashboard¶
The gateway dashboard, built from the gateway's audit log in Loki, its Postgres spend store over Private Data Source Connect, and recording rules over both.
Audit log: sign-ins, inference requests and managed settings served, per developer and model.
Spend store: the gateway's own spend ledger per developer and team, and each spend cap against what has been spent.
Derived metrics: touchline_gateway_* recording rules, so a gateway with no metrics endpoint can still be graphed and alerted on.
Operational logs: the gateway's own warn and info lines and the other agent host containers.
Alerts: the demo alert rules and their current state.
What the gateway emits: every signal the gateway produces on its own, and what it cannot see.
Bedrock dashboard¶
The Bedrock dashboard, built from the AWS/Bedrock CloudWatch metric stream and the optional model invocation log.
CloudWatch metrics: invocations, tokens, latency and quota usage per inference profile, AWS's own counters.
What Bedrock saw: the invocation log, split by team and model through the per-team application inference profiles.
App OpenTelemetry versus Bedrock log: the application's count of generations next to Bedrock's, two views of the same calls.
Application Observability and Knowledge Graph¶
The same agent spans seen as services, with no extra instrumentation.
Application Observability services: every Touchline service with its duration, errors and rate.
The orchestrator as a service: RED metrics and its inbound calls from the site API.
Knowledge Graph: the Touchline services and the call edges between them.
Entity catalog: the same services with latency, error ratio and request rate.










































