DevOps & Site Reliability · PPL

Site Reliability Engineering Certification

An advanced program designed to develop practical skills in system reliability, scalability, observability, incident management, automation, capacity planning, and resilient production operations.

  • 4 DaysDuration
  • PPLAccredited
  • 3 LanguagesArabic · English · Hindi
  • ₹15,999.00 Per delegate

This course is accredited by PPL

This is for all ppl accredited courses
2M+ Delegates trained worldwide
15,000+ Corporate clients
490+ Training locations
4.8 ★ Average learner rating
20% OFF Limited-time launch offer
— The journey

Course Outline

What the programme covers, module by module.

Module 1: Introduction to Site Reliability Engineering

  • SRE fundamentals
  • Evolution of SRE
  • SRE and DevOps
  • Reliability engineering
  • SRE responsibilities
  • Production operations

Module 2: SRE Principles & Practices

  • Reliability mindset
  • Engineering approach to operations
  • Automation
  • Risk management
  • Service ownership
  • Continuous improvement

Module 3: Understanding Service Reliability

  • Availability
  • Reliability
  • Durability
  • Performance
  • Service health
  • Reliability measurement

Module 4: Service Level Indicators

  • SLI fundamentals
  • Availability indicators
  • Latency indicators
  • Error indicators
  • Throughput indicators
  • SLI selection

Module 5: Service Level Objectives

  • SLO fundamentals
  • Reliability targets
  • SLO design
  • Measurement windows
  • SLO tracking
  • SLO evaluation

Module 6: Service Level Agreements

  • SLA concepts
  • SLO and SLA relationships
  • Service commitments
  • Availability commitments
  • Reliability expectations
  • Operational considerations

Module 7: Error Budgets

  • Error budget concepts
  • Budget calculation
  • Reliability trade-offs
  • Release decisions
  • Budget consumption
  • Error budget policies

Module 8: Monitoring Fundamentals

  • Infrastructure monitoring
  • Application monitoring
  • Service monitoring
  • Metrics
  • Monitoring strategies
  • Health indicators

Module 9: Observability

  • Observability principles
  • Metrics
  • Logs
  • Traces
  • Events
  • Telemetry correlation

Module 10: Alerting Strategy

  • Alert conditions
  • Actionable alerts
  • Alert severity
  • Alert routing
  • Alert fatigue
  • Alert optimisation

Module 11: Incident Management

  • Incident lifecycle
  • Incident detection
  • Incident classification
  • Response coordination
  • Escalation
  • Resolution

Module 12: Incident Response

  • On-call response
  • Incident roles
  • Communication
  • Triage
  • Mitigation
  • Service restoration

Module 13: Root Cause Analysis

  • Problem investigation
  • Event timelines
  • Root causes
  • Contributing factors
  • Corrective actions
  • Preventive actions

Module 14: Post-Incident Reviews

  • Review processes
  • Incident documentation
  • Learning culture
  • Action items
  • Follow-up
  • Reliability improvement

Module 15: Eliminating Toil

  • Toil concepts
  • Identifying repetitive work
  • Toil measurement
  • Automation opportunities
  • Operational efficiency
  • Toil reduction

Module 16: SRE Automation

  • Operational automation
  • Infrastructure automation
  • Deployment automation
  • Self-healing systems
  • Automated remediation
  • Automation workflows

Module 17: Capacity Planning

  • Capacity requirements
  • Resource utilisation
  • Demand forecasting
  • Growth planning
  • Capacity modelling
  • Resource allocation

Module 18: Scalability Engineering

  • Horizontal scaling
  • Vertical scaling
  • Load distribution
  • Autoscaling
  • Scalability bottlenecks
  • Scaling strategies

Module 19: Performance Engineering

  • Performance metrics
  • Latency
  • Throughput
  • Resource consumption
  • Bottleneck identification
  • Performance optimisation

Module 20: Distributed Systems Reliability

  • Distributed architecture
  • Network failures
  • Partial failures
  • Timeouts
  • Retries
  • Failure handling

Module 21: Resilience Engineering

  • Resilience principles
  • Fault tolerance
  • Redundancy
  • Graceful degradation
  • Failure isolation
  • Resilience patterns

Module 22: High Availability

  • Availability architecture
  • Redundancy
  • Load balancing
  • Failover
  • Replication
  • Availability strategies

Module 23: Disaster Recovery

  • Disaster recovery planning
  • Recovery objectives
  • Backup strategies
  • Recovery procedures
  • Failover planning
  • Recovery validation

Module 24: Kubernetes Reliability

  • Kubernetes availability
  • Pod health
  • Resource requests
  • Resource limits
  • Autoscaling
  • Workload resilience

Module 25: Reliable CI/CD

  • Deployment reliability
  • Automated validation
  • Progressive delivery
  • Canary releases
  • Blue-green deployments
  • Rollbacks

Module 26: Reliability Testing

  • Load testing
  • Stress testing
  • Failure testing
  • Resilience testing
  • Recovery testing
  • Reliability validation

Module 27: Chaos Engineering

  • Chaos engineering principles
  • Controlled experiments
  • Failure injection
  • Hypothesis development
  • Blast radius
  • Learning from failures

Module 28: On-Call Engineering

  • On-call responsibilities
  • Rotations
  • Escalation policies
  • Runbooks
  • Operational readiness
  • Handover practices

Module 29: SRE Metrics & Continuous Improvement

  • Reliability metrics
  • Operational metrics
  • Incident trends
  • Toil metrics
  • SLO performance
  • Improvement planning

Module 30: Enterprise SRE Implementation

  • SRE adoption strategy
  • SLI and SLO implementation
  • Observability design
  • Incident management workflow
  • Automation planning
  • Reliability improvement roadmap
— 01.2 · Is it right for you?

Who it's for & what's included

Pick a delivery method to see exactly who it suits and everything you receive.

Who it's for

Classroom

Best for learners who want face-to-face tuition and to network with peers in person.

What's included

Everything you get

  • Live instructor on-site
  • Printed workbook & materials
  • Group exercises & case studies
Who it's for

Online Instructor-Led

Best for learners who want a live instructor and a fixed schedule, without the travel.

What's included

Everything you get

  • Live instructor via video call
  • Digital workbook & resources
  • Session recordings
Who it's for

Self-Paced

Best for self-motivated learners who need maximum flexibility around work and life.

What's included

Everything you get

  • On-demand video lessons
  • Interactive quizzes
  • 24/7 access on any device
— What you will master

Course Objectives

01

Understand advanced Site Reliability Engineering principles and operational practices.

02

Define and manage SLIs, SLOs, SLAs, and error budgets for production services.

03

Implement effective monitoring, observability, and actionable alerting strategies.

04

Manage incidents systematically and conduct effective post-incident reviews.

05

Reduce operational toil through automation and engineering practices.

06

Design scalable, highly available, fault-tolerant, and resilient systems.

07

Apply capacity planning, performance engineering, reliability testing, and chaos engineering.

08

Establish sustainable SRE practices for continuously improving production reliability.

— Questions answered

Frequently Asked Questions

What is Site Reliability Engineering?
Site Reliability Engineering applies software engineering principles to IT operations to improve the reliability, scalability, performance, and maintainability of production systems.
Who should attend this course?
The course is suitable for SRE professionals, DevOps engineers, platform engineers, cloud engineers, infrastructure engineers, and experienced technical working professionals.
Do I need previous DevOps or cloud experience?
Yes. Familiarity with Linux, cloud infrastructure, networking, monitoring, containers, and DevOps concepts is recommended because this is an advanced-level program.
Which SRE topics are covered?
The course covers SLIs, SLOs, error budgets, monitoring, observability, incident management, toil reduction, automation, scalability, resilience, disaster recovery, Kubernetes reliability, and chaos engineering.
What practical skills will I develop?
You will develop skills in defining reliability targets, designing observability strategies, managing incidents, automating operational work, improving resilience, planning capacity, and maintaining reliable production services.
— Trusted by learners

What our delegates say

★★★★★

"The structure, the practice exams, the instructor — all top tier. Passed first try."

AS
Aarti SharmaSenior Project Manager · TCS
★★★★★

"Best training I have attended. The content is exactly what modern projects need."

JD
James DonovanProgramme Director · Capgemini
★★★★★

"24/7 support actually means 24/7 — got help on my mock exam at 2am. Worth every dollar."

MO
Maya OkaforPMO Lead · Standard Bank

★ 4.8 / 5 from 12,000+ verified learner reviews on Trustpilot & Google.

PPL Academy enquiry form

Get the course
that's right for you.

Our advisors respond within one business day.

Full name
Work email
Contact number
Message (optional)
Your details are never shared with third parties.
< 24h Response
Live & online Delivery
Certified Instructors