Cloud Troubleshooting Methodology
Troubleshooting Methodology
A single repeatable process, run in order, every time — from spotting the problem to writing down what fixed it. Step through it below.
The Ten Steps
Documentation isn't its own step tucked at the end — it runs alongside every phase. The other nine steps are strictly ordered.
Determining the Scope
A web app is reachable from two branch offices but not a third. Run the check to see what that pattern tells you about where the problem actually lives.
Corporate Policies, Procedures & Impacts
Troubleshooting doesn't happen in a vacuum — standard operating procedures, access policies, and service-level agreements all shape what you're allowed to do and what it costs. Tap each to expand.
Standard Operating Procedures +
Govern how particular tasks get done and must be followed during troubleshooting — especially identity and access management policies, so you never grant users too much or too little access while fixing something else.
Escalation Policy +
Cloud policies govern case escalation — for example, to the CSP's technical support. Escalating can incur a cost, so administrators need a complete view of the environment before pushing a case upward.
Your SLA Exposure +
SLAs can enforce penalties on your organization for service outages — a reason to resolve issues quickly and within the terms you've committed to.
CSP's SLA to You +
CSPs carry their own SLAs backing service availability. If an outage falls within their responsibility, your organization may be awarded reduced fees or other compensation.
Reviewing the Process
Cloud troubleshooting runs on the same core methodology as on-premises troubleshooting always has — but the cloud adds variables worth watching for.
- Downtime spanning the user-to-cloud network connection, or an outage at the CSP itself
- A greater chance of security-specific access control misconfigurations blocking normal access
- Deployment issues tied to resource limits, feature deprecation, or regional service availability