AI Prompt
role: >
You are a senior Site Reliability Engineer who designs Service Level Objectives
(SLOs) that match what users actually experience. You favor a few meaningful
objectives over many vanity metrics, and you turn every SLO into an error
budget policy and alerts a team can act on.
task: >
Design user-centric SLIs, SLOs, error budgets, burn-rate alerts, and an error
budget policy for the service described in the inputs. Then return the result
in the exact YAML output schema below.
inputs:
service_name: "${service_name:checkout-api}"
what_users_do: "${user_journeys:browse cart, apply coupon, place order, view order status}"
architecture: "${architecture:Node.js API behind a load balancer, PostgreSQL, Redis cache, Stripe for payments}"
traffic: "${traffic:about 40 requests per second at peak, strong evening and weekend peaks}"
telemetry_available: "${telemetry:load balancer access logs, Prometheus metrics, OpenTelemetry traces}"
current_pain: "${pain:slow order placement during sales, on-call fatigue from noisy CPU alerts}"
business_constraints: "${constraints:payments must never be double-charged; marketing runs flash sales monthly}"
compliance_window_days: ${window_days:28}
method:
- step: Map critical user journeys
detail: List 3-5 journeys from what_users_do and rank them by business impact. Ignore internal-only endpoints unless a user waits on them.
- step: Choose SLIs
detail: For each top journey, pick availability, latency, or correctness/freshness SLIs written as "good events / valid events". Name the exact data source and say where it is measured (load balancer, server, or client).
- step: Set targets
detail: Propose an SLO target and justify it against current pain, dependencies (a service cannot beat its critical dependencies), and cost. Prefer 99.5-99.9% unless there is strong evidence otherwise. Never 100%.
- step: Compute error budgets
detail: Convert each target into allowed bad events and allowed bad minutes for the compliance window, and show the arithmetic.
- step: Design burn-rate alerts
detail: Use multi-window, multi-burn-rate alerting (for example a 1h/5m window pair at 14.4x paging and a 6h/30m pair at 6x paging, plus a 3d/6h pair at 1x as a ticket). Write example PromQL or pseudo-queries using the telemetry the team actually has.
- step: Write the error budget policy
detail: Define concrete actions at 50%, 75%, and 100% budget consumed (for example a release freeze except fixes, or a reliability sprint), who decides, and how exceptions work.
- step: Retire noise
detail: Name the current alerts that should be downgraded or deleted once SLO alerts exist, such as raw CPU alerts.
output_format:
type: yaml
rules:
- Return only valid YAML that follows the schema below. No prose before or after it and no code fences.
- Use null for unknown values and record every guess in assumptions.
- Keep each description under 25 words.
schema:
service: string
compliance_window_days: integer
assumptions: [string]
user_journeys:
- name: string
business_impact: high | medium | low
slos:
- id: string (e.g. SLO-1)
journey: string
sli_type: availability | latency | correctness | freshness
sli_definition: "good events / valid events, in plain words"
good_event: string
valid_event: string
data_source: string
measured_at: load_balancer | server | client | synthetic
target_percent: number
latency_threshold_ms: integer or null
error_budget:
allowed_bad_ratio: number
allowed_bad_events_estimate: integer
allowed_bad_minutes: number
math: string
burn_rate_alerts:
- name: string
long_window: string
short_window: string
burn_rate: number
budget_consumed_at_trigger_percent: number
action: page | ticket
query_example: string
error_budget_policy:
thresholds:
- budget_consumed_percent: integer
actions: [string]
decision_owner: string
exception_process: string
alerts_to_retire:
- alert: string
reason: string
dashboards:
- panel: string
purpose: string
review_cadence: string
open_questions: [string]
constraints:
- Maximum 4 SLOs. If more seem necessary, explain the merge choice in assumptions.
- Every SLO must be measurable with the telemetry the team says it has; otherwise list the missing instrumentation in open_questions.
- Do not use CPU, memory, or other resource metrics as SLIs.
- Latency SLIs use a threshold ("requests faster than 800 ms"), not an average.
- Treat payment correctness (no double charge) as its own SLO or as an explicit invariant in assumptions.This prompt template is designed to help you get better results from AI models like ChatGPT, Claude, Gemini, and other large language models. Simply copy it and paste it into your preferred AI assistant to get started.
Browse our prompt library for more ready-to-use templates across a wide range of use cases, or compare AI models to find the best one for your workflow.