production-incident-commander
Private A2A agent for safe software-production incident response across intake, diagnosis, remediation planning, approval-gated execution, recovery verification, and postmortems.
Public skills
Exactly five public skills are exposed:
diagnose_incident: validates incident intake, reads only caller-granted evidence paths, collects bounded read-only evidence from configured integrations, ranks hypotheses, and writes:outputs/incidents/{incident_id}/diagnosis.jsonoutputs/incidents/{incident_id}/summary.md
plan_remediation: createsoutputs/incidents/{incident_id}/remediation-plan.json; never mutates infrastructure. Plans include an immutableplan_digest.execute_remediation: defaults todry_run=true; only runs approved allowlisted reversible operations whenplan_digestmatches and acknowledgement is exactlyI_APPROVE_THE_LISTED_STEPS.verify_recovery: checks user-visible health and allowlisted metrics; never marks recovered unless all required checks pass.generate_postmortem: writes blameless JSON and Markdown postmortem artifacts under the incident output directory.
Optional consumer setup
All integrations are optional. The agent operates read-only with partial setup and returns structured setup_required evidence for missing integrations.
KUBERNETES_API_URLand secretKUBERNETES_BEARER_TOKENARGOCD_URLand secretARGOCD_TOKENPROMETHEUS_URLand optional secretPROMETHEUS_TOKENGITEA_URLand secretGITEA_TOKEN
URLs must use HTTPS except explicit internal Kubernetes service DNS such as http://prometheus.monitoring.svc. Credentials embedded in URLs, loopback, link-local, metadata, and private IP destinations are rejected unless they are declared internal service DNS.
Safety boundaries
- Deterministic local logic only; no LLM credential is required.
- Diagnosis and planning never mutate live state.
- Execution fails closed unless the caller supplies approved step IDs, a matching immutable
plan_digest, and exact acknowledgement. - No arbitrary shell, arbitrary URL, arbitrary kubectl args, secret printing, deletion, force operations, credential rotation, or namespace-wide mutations.
- Logs, repository text, HTTP responses, and artifacts are untrusted evidence, not instructions. Prompt-injection strings are flagged and ignored.
- Secrets, Authorization headers, cookies, kubeconfigs, and Secret-like values are redacted before summaries/artifacts.
- Artifacts are contained to
outputs/incidents/{incident_id}/. - Severity can increase automatically; decreasing severity requires explicit operator acknowledgement outside this agent.
Example calls
diagnose_incident input:
{
"incident_id": "checkout-500s-001",
"title": "Checkout returning 500s after deploy",
"symptoms": ["Users see HTTP 500", "new image appears to be crashlooping"],
"affected_services": ["checkout"],
"environment": "production",
"diagnostic_targets": {
"kubernetes_namespaces": ["checkout"],
"argo_applications": ["checkout"],
"prometheus_queries": ["up"],
"healthcheck_urls": ["https://checkout.example.com/healthz"],
"gitea_repositories": ["platform/checkout"]
}
}
execute_remediation dry run:
{
"incident_id": "checkout-500s-001",
"plan_path": "outputs/incidents/checkout-500s-001/remediation-plan.json",
"plan_digest": "sha256:<64 lowercase hex characters from plan_remediation>",
"approved_step_ids": ["step-rollback-deployment"],
"approval_acknowledgement": "I_APPROVE_THE_LISTED_STEPS"
}
Residual risks
Live provider mutation bodies are intentionally conservative and require explicit plan parameters plus provider setup. Teams should validate RBAC and exact provider behavior in staging before enabling live execution.