LLM Ops Automation Without Giving It the Keys
Why “LLM With Prod Credentials” Is the New Runbook-in-a-Loop
Wiring an LLM directly to production credentials is the modern version of while true; do …; done against a runbook. It will resolve incidents faster—until one hallucinated command, misread log line, or prompt injection turns “helpful automation” into an outage amplifier.
The core production-safe pattern is:
LLM = reasoning. System = enforcement.
The model proposes. Deterministic systems verify, authorize, and execute.
This separation gives you the benefits of LLMs (context synthesis, diagnosis, plan generation, summarization) without granting them the two things that make incidents unrecoverable: unbounded authority and un-audited execution.
Below is a concrete architecture that holds up in real production environments, with examples you can implement in a Kubernetes/SRE platform stack.
Reference Architecture: Advisor + Guarded Actuators
At a high level, the workflow looks like this:
- Ingest evidence: logs, traces, events, runbooks, tickets.
- LLM produces a structured plan: intent, constraints, ordered steps, expected signals, rollback.
- Translate steps into typed actions from a narrow catalog.
- Policy-as-code evaluates actions (risk, environment, approvals, error budget, change windows).
- Executor runs approved actions with least privilege and short-lived credentials.
- Verification is mandatory after each step; failures halt or rollback.
- LLM gets new evidence and can update the plan—without direct control.
The key is that every boundary is enforced by code you own: schemas, policies, and deterministic executors.
1) Keep the LLM as Planner + Diagnostician (No Tokens, No Shell)
Let the LLM read only the diagnostic universe:
- Centralized logs (e.g., Loki/ELK/Cloud Logging)
- Traces (Jaeger/Tempo)
- Metrics dashboards (Prometheus)
- Incident tickets and timelines
- Runbooks and known-issue docs
- Deployment diffs (from Git, not
kubectl)
Then force the model to output a structured plan. That plan should include:
- Intent: what outcome we want (“restore checkout latency to baseline”)
- Constraints: what must not happen (“no downtime”, “no scaling above budget”, “no changes to prod outside change window”)
- Ordered steps: each step references typed actions
- Expected signals: what metrics/logs should improve
- Rollback: what to do if signals don’t improve
A minimal JSON schema for a plan might look like:
{
"incident_id": "INC-10294",
"summary": "High 5xx on checkout",
"constraints": [
"prod only",
"no data migrations",
"must preserve error budget",
"no direct shell commands"
],
"steps": [
{
"id": "s1",
"action": "fetch_metrics",
"params": { "service": "checkout", "window": "15m" },
"expected": { "signal": "identify error spike correlation" }
},
{
"id": "s2",
"action": "rollback_deployment",
"params": { "namespace": "prod", "deployment": "checkout", "to_revision": 182 },
"expected": { "signal": "5xx rate drops below 0.5% in 10m" }
}
],
"rollback_plan": [
{
"action": "roll_forward_deployment",
"params": { "namespace": "prod", "deployment": "checkout", "to_revision": 183 }
}
]
}
Two practical notes:
- Prompt injection is real: if the LLM ingests tickets, chat logs, or even error pages, assume an attacker can influence that text. If the LLM has credentials, you’ve created a direct exploit path.
- Never grant “just one kubeconfig”: read-only kubeconfig is still a control plane foothold, and mistakes happen (contexts drift, RBAC changes, cluster-admin bindings appear).
2) Allow Only Typed Actions, Never Free-Form Commands
Free-form command generation is where boundaries die. The correct abstraction is a small action catalog with:
- A name:
restart_service,scale_deployment,rollback_deployment,drain_node - A strict schema for parameters
- Preconditions (what must be true before execution)
- Blast radius metadata (risk scoring, affected scope, rate limits)
- Default safety behavior (timeouts, idempotency)
Here’s an example action catalog entry (YAML) for Kubernetes deployment scaling:
name: scale_deployment
description: Scale a Kubernetes deployment to a target replica count.
schema:
type: object
required: [cluster, namespace, deployment, replicas]
properties:
cluster: { type: string }
namespace: { type: string, pattern: "^(prod|staging|dev)$" }
deployment: { type: string }
replicas: { type: integer, minimum: 0, maximum: 50 }
preconditions:
- "deployment_exists"
- "namespace_allowed"
metadata:
blast_radius:
scope: "single-deployment"
risk: 3 # 1-10
rate_limit:
max_per_10m: 3
requires_approval_if:
- "namespace == 'prod'"
- "replicas > 20"
If the model outputs something that looks like:
kubectl -n prod scale deploy/checkout --replicas=200
…you’re already in the danger zone. Typed actions prevent “creative” interpretations and make it possible to attach policy gates, approvals, and auditing uniformly.
3) Put Policy-as-Code in Front of Execution (LLMs Can’t Argue With Rego)
A policy layer is the line between “automation” and “accident.” The LLM should never be able to talk its way past a deny.
Common checks you should encode:
- Environment restrictions (prod vs staging)
- Change windows / freeze periods
- Required approvals (on-call, service owner)
- SLO and error budget state (avoid risky changes when budget is exhausted)
- Dependency health (don’t roll a service if its database is degraded)
- Risk score thresholds and blast radius caps
Using OPA (Open Policy Agent), you can model this as an admission decision for every proposed action.
Example Rego policy that blocks risky actions in prod outside a change window, and blocks actions when error budget is low:
package ops.guard
default allow := false
# Input shape:
# {
# "action": {"name":"scale_deployment","params":{...},"metadata":{"risk":3}},
# "context": {"env":"prod","change_window":false,"error_budget_remaining":0.08,"approvers":["alice"]}
# }
allow {
input.context.env != "prod"
}
allow {
input.context.env == "prod"
input.context.change_window == true
input.context.error_budget_remaining >= 0.10
approved_by_required_parties
input.action.metadata.risk <= 5
}
approved_by_required_parties {
# Require at least one approver in prod
count(input.context.approvers) >= 1
}
deny_reason["prod change requires active change window"] {
input.context.env == "prod"
input.context.change_window == false
}
deny_reason["insufficient error budget for risky change"] {
input.context.env == "prod"
input.context.error_budget_remaining < 0.10
input.action.metadata.risk >= 4
}
This policy is deterministic. It doesn’t care how persuasive the incident summary is, or whether the model claims “this is urgent.” If the policy says deny, the action does not run.
A good operational pattern is to return:
allow: true/falsedeny_reason[]required_approvals[](if you support conditional approval workflows)
…and to log the full decision input/output for auditability.
4) Execute Deterministically With Least Privilege (Short-Lived, Scoped Identity)
Once an action is authorized, execution must be:
- Boring (no surprises)
- Repeatable (idempotent where possible)
- Auditable (who/what/when/why)
- Least privilege (scoped roles, short-lived tokens)
A common trap is to make the executor a “god service” with broad cluster admin rights. Don’t. Instead:
- Use OIDC/workload identity to mint short-lived credentials
- Bind narrowly scoped RBAC roles per action category
- Separate executors by environment (prod executor is distinct from staging)
- Avoid shared keys and standing access
For Kubernetes, that often means:
- A dedicated service account per executor
- RBAC roles like “can patch deployment scale” but not “can read secrets”
- If using cloud-managed Kubernetes, rely on cloud IAM ↔ Kubernetes RBAC mapping (e.g., IRSA on EKS, Workload Identity on GKE)
Example: RBAC that only allows scaling deployments in a namespace:
apiVersion: rbac.authorization.k8s.io/v1
kind: Role
metadata:
name: deployment-scaler
namespace: prod
rules:
- apiGroups: ["apps"]
resources: ["deployments/scale"]
verbs: ["get", "update", "patch"]
---
apiVersion: rbac.authorization.k8s.io/v1
kind: RoleBinding
metadata:
name: deployment-scaler-binding
namespace: prod
subjects:
- kind: ServiceAccount
name: ops-executor
namespace: ops-system
roleRef:
kind: Role
name: deployment-scaler
apiGroup: rbac.authorization.k8s.io
Your executor then implements scale_deployment deterministically (using a Kubernetes client library), not by shelling out.
5) Make Verification Non-Optional (Every Step Proves It Worked)
Production automation without verification is just fast failure.
Each action step should define post-checks that can be evaluated automatically:
- Health probes (readiness/liveness, synthetic checks)
- Deployment status (rollout complete, replicas available)
- Metric checks (5xx rate, latency p95/p99)
- Diff checks (config drift, desired vs actual)
- Error budget impact (burn rate)
If verification fails:
- Halt and request human review, or
- Auto-rollback (if the action supports safe rollback)
A lightweight verification spec might look like:
{
"checks": [
{ "type": "promql", "query": "sum(rate(http_requests_total{service='checkout',code=~'5..'}[5m]))", "max": 1.0 },
{ "type": "k8s_rollout", "namespace": "prod", "deployment": "checkout", "timeout_sec": 300 }
],
"on_fail": "rollback_and_page"
}
The LLM can interpret the evidence (“5xx dropped but latency is still high; next step: check dependency db pool saturation”), but the pass/fail gate should be deterministic.
Concrete End-to-End Example: Roll Back a Bad Deploy (Safely)
Let’s walk the whole path.
Scenario: After deploying checkout revision 183, 5xx errors spike.
- The system collects evidence (Grafana panel snapshots, PromQL results, deployment history, recent commits).
- The LLM produces a plan proposing
rollback_deploymentto revision 182 and defines expected signals (5xx rate drops < 0.5% in 10 minutes). - The plan is converted to a typed action request:
{
"action": {
"name": "rollback_deployment",
"params": { "cluster": "prod-us", "namespace": "prod", "deployment": "checkout", "to_revision": 182 },
"metadata": { "risk": 4, "blast_radius": "single-deployment" }
},
"context": {
"env": "prod",
"change_window": true,
"error_budget_remaining": 0.22,
"approvers": ["oncall-sre"]
}
}
- OPA evaluates and returns
allow: true. - The executor performs rollback using a client library, with a short-lived identity limited to patching that deployment.
- Verification runs automatically:
- Kubernetes rollout status must be healthy
- PromQL must show 5xx below threshold
- If verification passes, the system records the action and outcome; the LLM generates a summary and suggests follow-ups (e.g., bisect commit, add regression test).
At no point did the model get the ability to run arbitrary commands or access production secrets.
Design Tips That Make This Work in Real Incident Response
A few practical lessons from implementing guarded automation:
- Treat the action catalog like an API product. Version it, document it, test it, and keep it small.
- Make policies environment-specific. Staging should be permissive enough for learning; prod should be strict.
- Log everything. Store: evidence inputs, LLM plan output, policy decision inputs/outputs, executor actions, verification results.
- Prefer “request approvals” over “break glass.” If you need break-glass, implement it as a separate audited flow with explicit human authentication—not a hidden model prompt.
- Use canaries by default. Typed actions can include rollout strategies (e.g., 10% traffic for 10 minutes) rather than full fleet changes.
The Takeaway: “Agents” Shouldn’t Be Agents in Production
In production operations, the most trustworthy LLM is one that:
- Diagnoses and proposes
- Explains tradeoffs and expected signals
- Updates plans based on new evidence
…but never directly executes. Advisors paired with guarded actuators is the pattern that scales across platforms, teams, and compliance regimes.
Separate reasoning from enforcement and you get automation you can actually trust—because the system can prove what happened, why it was allowed, and whether it worked.
