This job has been added to your Saved jobs.
You have reached the limit of 20 Saved Jobs. If you want to create a new one, please manage your Saved Jobs.
Site Reliability Engineer (SRE, AWS, Azure, Cloud)
Pizza Hut Digital & Technology
Job Expertise:
Job Domain:
Software Products and Web Services
Top 3 reasons to join us
- Flexible Friday afternoon
- 18 Annual Leave + 5 Recharge Days/ Year
- Hybrid working model
Job description
Incident Management and Reliability
- Independently lead complex incidents involving multiple systems, teams, or dependencies.
- Coordinate incident response activities, facilitate communication, and drive timely resolution.
- Lead or contribute to post-incident reviews and root cause analysis activities.
- Ensure corrective and preventive actions are identified, prioritized, tracked, and completed
Monitoring, Alerting, and Observability
- Design, implement, and continuously optimize monitoring, logging, alerting, and tracing solutions.
- Develop meaningful alerts based on service behavior, customer impact, and business priorities.
- Build and maintain dashboards that provide actionable insights into system performance and reliability.
Platform and Market Ownership
- Own SRE responsibilities for one or more markets, platforms, or critical services end-to-end.
- Establish and maintain operational excellence standards for assigned domains.
- Ensure monitoring coverage, dashboards, runbooks, and alerting configurations remain accurate, effective, and up to date.
- Continuously assess platform health, identify reliability risks, and drive improvements before incidents occur.
- Partner with engineering teams to ensure new features and services meet reliability requirements before production release.
- Define and track reliability metrics, including Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets.
Automation, and AI
- Develop and maintain tools, scripts, and automation solutions that reduce manual effort and improve operational efficiency.
- Identify and eliminate repetitive tasks through automation and self-service capabilities.
- Establish and promote best practices for the responsible use of AI within SRE workflows
Team Contribution and Mentoring
- Mentor and support Level 5 and Level 6 engineers in incident management, monitoring, automation, AI adoption, and operational best practices.
- Review monitoring configurations, dashboards, runbooks, and operational documentation to maintain quality standards.
- Share knowledge through training sessions, documentation, and post-incident learning activities.
- Contribute to the continuous improvement of team processes, standards, and ways of working.
Your skills and experience
- Own SRE responsibilities for one or more markets, platforms, or critical services end-to-end.
- Establish and maintain operational excellence standards for assigned domains.
- Ensure monitoring coverage, dashboards, runbooks, and alerting configurations remain accurate, effective, and up to date.
- Coordinate incident response activities, facilitate communication, and drive timely resolution.
- Design, implement, and continuously optimize monitoring, logging, alerting, and tracing solutions.
- Develop meaningful alerts based on service behaviours, customer impact, and business priorities.
- Review monitoring configurations, dashboards, runbooks, and operational documentation to maintain quality standards.
- Lead or contribute to post-incident reviews and root cause analysis activities.
- Influence technical decisions that improve platform stability, scalability, and operational efficiency
Why you'll love working here
Attractive Benefits:
- 100% salary during probation period
- Annual Leave: 18 days/ year
- Five “Recharge Days” – Extra days, in addition to company holidays.
- Flexible Friday afternoon
- Full salary insurance
- 13th-month bonus
- 1 day off for birthday
- Advanced health insurance (Generali)
- Regular engagement activities: sport clubs, internal event…
- Support Macbook and Monitor
Pizza Hut Digital & Technology
Company type
IT Product
Company industry
Software Products and Web Services
Company size
51-150
employees
Country
United Kingdom
Working days
Monday - Friday
Overtime policy
No OT
More jobs for you
Get similar jobs by email
Subscribe
NEW FOR YOU
Posted
21 hours ago
Senior Database Administrator (PostgreSQL, AWS, Agile)
At office
Ho Chi Minh
SUPER HOT
Posted
2 days ago
Senior Data / Platform Engineer (AI, AWS, Azure, GCP)
At office
Ho Chi Minh
Feedback