Defines SLIs/SLOs and error budgets, builds golden-signal monitoring and multiwindow burn-rate alerts in Prometheus, automates toil, and runs chaos experiments to validate resilience.
---
name: sre-engineer
description: Use when defining SLIs/SLOs, managing error budgets, building reliable systems at scale, incident management, chaos engineering, toil reduction, or capacity planning. Produces SLO definitions, monitoring/alerting config, and automation.
---
# SRE Engineer
Define service-level objectives, error-budget policies, monitoring, and reliability automation for production systems.
## Workflow
1. Assess reliability - review architecture, current SLOs, incidents, and toil levels.
2. Define SLOs - pick meaningful SLIs and set targets justified by user impact; calculate error budgets from them.
3. Implement monitoring - golden-signal dashboards (latency, traffic, errors, saturation) with PromQL.
4. Build alerts - multiwindow burn-rate rules in Prometheus (fast and slow burn) with runbook links.
5. Automate toil - replace recurring manual work (e.g. auto-remediation scripts that query Prometheus and restart deployments).
6. Test resilience - design chaos experiments and verify recovery meets RTO/RPO before marking them complete.
Writes blameless postmortems and balances reliability with feature velocity.
Full skill & source: https://github.com/Jeffallan/claude-skills/tree/e8be415bc94d8d6ebddc2fb50e5d03c6e27d4319/skills/sre-engineer