Understand the agent metrics dashboard, drill into per-step bottlenecks, configure threshold alerts, and export telemetry to Datadog, Grafana, or any OTLP backend.
The agent metrics dashboard
Each agent has a dedicated Metrics dashboard accessible from its overview page. The top-level tiles show: total runs in the selected time range, success rate (percentage of runs that completed without error), P50 and P95 run duration, total tokens consumed, and estimated cost. You can filter by time range (last hour, last 24 hours, last 7 days, last 30 days, or a custom range) and compare two time ranges side by side to spot trends. The dashboard updates in near-real-time with a 30-second data lag.
Step-level performance breakdown
Below the top-level tiles, the Step Performance table shows the average duration, error rate, and token usage for every node in the agent across all runs in the selected range. This view is invaluable for identifying bottlenecks: if one HTTP Request action accounts for 80% of the agent's total runtime, it is a candidate for parallelization or caching. The table also surfaces the most common error types per step, helping you prioritize which errors to address first. Click any row to drill into a filtered list of runs where that specific step failed.
Setting up metric alerts
Navigate to Settings > Alerts to configure threshold-based alerts for your agent. Available alert conditions include: error rate above X%, average run duration above Y seconds, run volume below Z per hour (useful for detecting a failed trigger), and token spend above $N per day. Alerts can be delivered via email, Slack, PagerDuty, or any webhook endpoint you configure. Set your alert thresholds based on observed baseline metrics during the first week after deployment, when you have established a normal operating range. Avoid setting thresholds too tight during the initial learning period, as you will receive false positives.
Integrating with external monitoring
If your organization uses an external observability platform such as Datadog, Grafana, or New Relic, Cotonity can push agent metrics to it via OpenTelemetry (OTLP). Enable the OTLP exporter in Settings > Observability, provide your collector endpoint and authentication headers, and select which metric types to export. Cotonity emits traces with span-level detail for each run, counters for run outcomes, and histograms for run duration and token usage. This integration lets you correlate agent performance with the health of the broader systems your agent interacts with on a single monitoring dashboard.