This job has been added to your Saved jobs.
You have reached the limit of 20 Saved Jobs. If you want to create a new one, please manage your Saved Jobs.
Job Expertise:
Job Domain:
Software Products and Web Services
Top 3 reasons to join us
- Flexible Friday afternoon
- 18 Annual Leave + 5 Recharge Days/ Year
- Hybrid working model
Job description
Roles & Responsibilities
Incident Management and Reliability
- Lead or coordinate incident response for assigned markets and services, with independent ownership of routine incidents and supported ownership of complex multi-system incidents
- Facilitate communication during incidents and drive timely resolution, including clean shift handoffs across global time zones
- Contribute to post-incident reviews and root cause analysis activities
- Ensure corrective and preventive actions are identified, tracked, and completed for owned services
Monitoring, Alerting, and Observability
- Implement and continuously optimize monitoring, logging, alerting, and tracing solutions
- Develop meaningful alerts based on service behavior, customer impact, and business priorities
- Build and maintain dashboards that provide actionable insights into system performance and reliability
Platform and Market Ownership
- Own day-to-day SRE responsibilities for one or more assigned markets, platforms, or services
- Ensure monitoring coverage, dashboards, runbooks, and alerting configurations remain accurate, effective, and up to date
- Assess platform health, identify reliability risks, and raise improvements before incidents occur
- Partner with engineering teams to ensure new features and services meet reliability requirements before production release
- Track reliability metrics for owned services, including Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets
Platform Engineering, Automation, and AI
- Develop and maintain tools, scripts, and automation that reduce manual effort and improve operational efficiency
- Identify and eliminate repetitive tasks through automation, with a bias toward self-service capabilities over ticket queues
- Contribute to auto-healing and auto-remediation capabilities: detection rules, remediation runbooks as code, and automated response workflows
- Build and extend internal platform tooling using infrastructure as code and CI/CD pipelines rather than manual configuration
- Apply team best practices for the responsible use of AI within SRE workflows, including AI-assisted diagnosis, triage, and remediation
Team Contribution
- Share knowledge through documentation, training sessions, and post-incident learning activities
- Review runbooks and operational documentation to maintain quality standards
- Contribute to the continuous improvement of team processes, standards, and ways of working
Your skills and experience
- 2+ years of experience in SRE, DevOps, production support, or infrastructure engineering roles
- Hands-on experience with monitoring and observability tooling (e.g., Datadog, Prometheus, Grafana, CloudWatch, or similar)
- Working knowledge of at least one major cloud provider (AWS preferred)
- Proficiency in at least one scripting or programming language (e.g., Python, Bash, Go) for automation, with demonstrated examples of automating away manual operational work
- Experience participating in incident response and on-call or shift-based operations
- Understanding of SLI/SLO concepts and reliability engineering fundamentals
- Ability to work follow-the-sun shift rotations, including structured handoffs with teams in other regions
- Strong written and verbal English communication skills for cross-region collaboration
- Experience with Kubernetes, Docker, and container orchestration in production
- Infrastructure as code experience (e.g., Terraform, CloudFormation)
- CI/CD and GitOps pipeline experience (e.g., GitLab CI, ArgoCD)
- Exposure to platform engineering concepts: internal developer platforms, self-service tooling, developer portals (e.g., Port, Backstage)
- Experience buildiExperience supporting distributed systems across multiple markets or regions
- Familiarity with AI-assisted operations tooling
- Relevant certifications (AWS, CKA, or similar)ng or contributing to auto-remediation or event-driven automation workflows.
Why you'll love working here
Attractive Benefits:
- 100% salary during probation period
- Annual Leave: 18 days/ year
- Five “Recharge Days” – Extra days, in addition to company holidays.
- Flexible Friday afternoon
- Full salary insurance
- 13th-month bonus
- Gift + 1 day off for birthday
- Advanced health insurance (Generali)
- Regular engagement activities: sport clubs, internal event…
- Support Macbook and Monitor
Pizza Hut Digital & Technology
Company type
IT Product
Company industry
Software Products and Web Services
Company size
51-150
employees
Country
United Kingdom
Working days
Monday - Friday
Overtime policy
No OT
More jobs for you
Get similar jobs by email
Subscribe
NEW FOR YOU
Posted
11 hours ago
Mid-Senior Java Backend Developer (Spring, English)
At office
Ho Chi Minh
SUPER HOT
Posted
1 day ago
Ass. Manager, Site Reliability Engineer (5yoe, English)
Hybrid
Ho Chi Minh
SUPER HOT
Posted
3 days ago
AI Support Engineer (≥2yoe, MLOPs, Excellent English)
Hybrid
Ho Chi Minh
HOT
Posted
4 days ago
Middle/ Senior DevOps Engineer (AWS, English) - BONUS
Hybrid
Ho Chi Minh - Da Nang
SUPER HOT
Posted
6 days ago
Mid/Sr Java Developer (English Required) - Up to 3200$
At office
Ho Chi Minh
Feedback