Step by Step

From alert to root cause in under 90 seconds.

No agents to install. No new data pipelines. Connect your existing telemetry stack via read-only API keys, and Devtract handles the correlation and RCA the moment an alert fires. Your observability tools stay exactly as they are.

Connect your telemetry stack in under 10 minutes.

Devtract reads from your existing telemetry systems via read-only API keys. No agents to install on your servers, no new data pipelines to build or maintain. Provide the API keys for your observability tools and Devtract starts listening immediately.

  • Observability: Datadog, Prometheus, Grafana, CloudWatch
  • Alerting: PagerDuty, OpsGenie
  • Deployment: GitHub Actions, Terraform Cloud, Argo CD
  • Read-only access, zero write permissions requested
Anomaly Detection / auth-service

Devtract clusters anomalies across signal types, not just one feed.

When an alert fires, the correlation engine opens a time window around the alert timestamp and scans for overlapping anomalies: log error-rate spikes, metric deviations from baseline, OOMKilled or CrashLoopBackOff events in Kubernetes, and deploy events from GitHub Actions, Terraform, or Argo CD. Statistical overlap is the signal, not keyword matching.

  • Correlation window opens automatically within 4 seconds of alert trigger
  • Overlapping anomaly clusters identified across all connected signal types
  • Deploy timestamps from GitHub, Terraform, and Argo CD mapped to anomaly onset
  • Full cross-signal scan completes in under 30 seconds

Ranked RCA card delivered to Slack and PagerDuty.

Within 90 seconds of the alert, your on-call engineer receives a structured card in Slack or PagerDuty. The card shows the top three root-cause hypotheses with confidence scores, the evidence chain for each, and a recommended action. No runbook lookup required.

  • Delivered to your existing incident channel or pager
  • Top 3 hypotheses with confidence percentages
  • Evidence chain: which signals drove each score
  • Actionable recommendation (rollback, scale, investigate)
RCA Delivery / Slack #incidents
#1
auth-service v2.4.1 OOM on startup
91%
#2
postgres connection pool exhaustion
34%
#3
upstream rate-limit from payments-api
12%

Ask follow-up questions in natural language.

Once the RCA card is delivered, the incident is not over. The on-call copilot stays active for the duration of the incident, ready to answer follow-up questions about your stack. All answers are grounded in the actual telemetry from the active incident window.

  • Ask about any service, metric, or event in the incident window
  • Historical pattern queries: has this happened before?
  • Rollback impact questions: what would reverting v2.4.1 affect?
  • Available in Slack, PagerDuty, and the Devtract web UI
Devtract Copilot / incident-4871
How many services depend on auth-service?
7 services have direct upstream dependency on auth-service: checkout, order-svc, user-api, admin-api, reporting, webhooks, cron-jobs.
Has v2.4.x caused OOM before?
No prior OOM events in v2.4.0 or v2.3.x. Pattern is new in v2.4.1. Rollback to v2.3.9 is the cleanest path.
Postmortem Draft / incident-4871

Postmortem draft generated automatically.

Once the incident is resolved, Devtract generates a draft postmortem from the actual incident timeline, the RCA findings, and the copilot conversation history. The timeline section is built from real event data, not reconstructed from memory. Export to Markdown or structured JSON for Confluence, Notion, or your incident tracker. The draft is a starting point, not a final document. Action items and blameless culture belong to your team.

  • Incident timeline reconstructed from actual event data
  • Root cause section auto-filled from RCA findings
  • Fully editable, exported as Markdown or structured JSON
Try the Full Workflow

Connect in 10 minutes. See results in the first incident.

14-day free trial. No credit card. Works with your existing Datadog or Prometheus setup.