Skip to content

Guides

This section is for the moment something is wrong — or just was. Each guide starts from a symptom you might be staring at and walks the real investigation, command by command, to a diagnosis and a verified outcome. By the end of a guide you can run the same workflow on your own cluster, and you will know which commands answer which questions.

What you’re seeingGuide
Pods are crashing or won’t start, and you don’t know whyInvestigate a broken workload
You shipped a deploy and it never finished rolling outYour rollout is stuck
Memory or CPU keeps climbing and an OOM kill looks inevitableCatch resource exhaustion early
A node went down and everything on it is failing at onceA node just died
The incident is over and you need to know what changed before itWhat changed before the incident
You’d rather hit quota and capacity limits on your terms than in an outageCapacity & quota ahead of time

Each guide walks the real workflow, using output captured during live validation drills (abridged, never invented):

Every guide ends with a pointer to the matching agent skill in skills/ — the same workflows, packaged for the agents themselves.