"ITIL is for old data centres." It is a common view among cloud-native engineers, and it misses the point. ITIL's core ideas, such as restoring service quickly, fixing root causes and managing change risk, matter just as much in the cloud. What has changed is how you implement them.
This guide explains how ITIL applies to cloud operations today.
What is ITIL 4?
ITIL (originally the IT Infrastructure Library) is a widely used framework for IT service management. ITIL 4, introduced in 2019, centres on a service value system and 34 management practices, and explicitly integrates agile, DevOps and lean thinking. Its successor, ITIL (Version 5), was released in February 2026 and is being rolled out in phases, but the core practices below carry across.
Which ITIL practices matter most in cloud operations?
| Practice | Purpose | Cloud adaptation |
|---|---|---|
| Incident management | Restore normal service quickly | Automated detection, chat-based war rooms, runbooks |
| Problem management | Find and remove root causes | Postmortems, trend analysis, automation of fixes |
| Change enablement | Maximise successful changes | Pre-approved standard changes through CI/CD and IaC |
| Service request management | Fulfil routine requests | Self-service portals, automated provisioning |
| Monitoring and event management | Detect and respond to events | Observability platforms, SLO-based alerts |
| Continual improvement | Improve services over time | Regular service reviews, metrics-driven backlog |
How does incident management work in the cloud?
- Detect through monitoring and alerts tied to user impact.
- Log and prioritise (P1, P2, P3) using agreed definitions.
- Respond within SLA with clear roles: incident lead, communications, responders.
- Restore service with runbooks, rollback or failover.
- Communicate regularly with stakeholders.
- Close and review, handing off to problem management if needed.
See our guides to reducing MTTR and MSP SLAs.
How do you do change management without slowing DevOps?
Classify changes by risk:
Risk-based change management in the cloud
Illustration in progress
- Standard changes: low-risk, repeatable and pre-approved, such as deployments through a tested CI/CD pipeline or IaC modules. No change board needed.
- Normal changes: assessed and approved based on risk, ideally by peers through automated checks and reviews.
- Emergency changes: expedited with retrospective review.
Infrastructure as code, automated tests and pull-request reviews provide the evidence and control traditional change boards were designed for, at deployment speed.
Why is problem management so important?
Incident management restores service; problem management stops the incident from coming back. Track recurring incidents, run blameless postmortems, identify root causes and turn fixes into automation or architecture improvements. This is how a mature operations team reduces incident volume over time.
Incident management restores service; problem management stops the incident from coming back.
How does ITIL work with DevOps and SRE?
They are complementary:
How ITIL, DevOps and SRE work together
Illustration in progress
- DevOps speeds up delivery and shared ownership.
- SRE brings engineering practices like SLOs, error budgets and automation to reliability. See SLOs and error budgets.
- ITIL provides consistent service management processes, roles and accountability.
Many organisations use SRE practices within an ITIL-aligned service model, especially when working with a managed service provider.
How Crozaint applies ITIL
Crozaint runs managed services on an ITIL-based service model with a defined escalation matrix. Incident, problem, change and request management apply across all eight disciplines, from infrastructure and network to databases, DR and FinOps, with SLAs including a 30-minute response for P1 and P2 incidents.
After observability engagements, continuous improvement cycles can be aligned to ITIL 4 or SRE frameworks, depending on how each client operates. We automate standard changes through tools such as Terraform, Ansible, Jenkins, GitLab and GitHub Actions to keep processes fast.
Common mistakes to avoid
- Change boards for every routine deployment
- Closing incidents without problem management
- Copying ITIL processes without adapting to automation
- No agreed priority definitions
- Treating ITIL and DevOps as opposites
Conclusion
ITIL in the cloud is about outcomes, not paperwork. Restore fast, fix root causes, automate safe change and improve continuously.
Want mature operations without slowing your teams? Tell us what you're running in a 30-minute call.

