Việc làm này đã được thêm vào mục Việc làm đã lưu.
Bạn đã lưu tối đa 20 việc làm. Nếu bạn muốn lưu mới, hãy cập nhật Việc làm đã lưu.
3 Lý do để gia nhập công ty
- Build impactful products from the ground up
- Tackle high-scale, real-world tech challenges
- Grow with a multidisciplinary tech team
Mô tả công việc
DatVietVAC Group Holdings is looking for an SRE Engineer with strong DevOps expertise to operate and continuously improve the infrastructure and engineering tools of our Fan Commerce Platform.
The official position title is SRE Engineer. The role combines site reliability engineering with hands-on DevOps responsibilities, including cloud infrastructure, Kubernetes, CI/CD, monitoring, automation, and production incident response.
1. Cloud Infrastructure and Platform Operations
- Operate and maintain application infrastructure in cloud environments, including compute, databases, caching, message queues, and object storage.
- Manage and continuously improve Development, Staging, and Production environments.
- Operate and improve CI/CD pipelines that support application development, testing, and deployment.
- Provision and manage infrastructure as code using Terraform and Ansible.
- Manage network configurations, firewalls, routing, VPC connectivity, and cloud VPNs.
- Operate containerized workloads and Kubernetes clusters, particularly Google Kubernetes Engine.
- Perform routine system upgrades, security patching, configuration changes, and infrastructure maintenance.
- Maintain infrastructure documentation and configuration standards.
2. Monitoring, Reliability, and Incident Response
- Monitor platform availability, performance, capacity, and reliability against agreed service-level objectives.
- Build and maintain monitoring, logging, dashboard, and alerting systems using Prometheus, Grafana, Google Cloud Operations, Datadog, or equivalent tools.
- Participate in the production on-call rotation and respond to incidents in a timely manner.
- Troubleshoot application, infrastructure, network, database, cache, and message-queue issues.
- Support root-cause analysis and contribute to post-incident reports and corrective actions.
- Develop and maintain operational runbooks for platform services and recurring incident scenarios.
- Analyze operational data and system metrics to identify reliability and performance improvements.
3. Automation and Continuous Improvement
- Develop automation tools and scripts to reduce manual operational work and repetitive engineering tasks.
- Improve deployment automation, environment consistency, and release reliability.
- Support load and performance testing before high-traffic on-sale periods and major events.
- Assist with capacity preparation, scaling configuration, and infrastructure readiness.
- Identify technical debt, operational risks, and opportunities to improve platform resilience.
- Support backup, recovery, and system-security activities.
- Contribute to continuous improvements in infrastructure management, monitoring, and incident response.
4. Cross-functional Collaboration
- Collaborate closely with Backend Engineers, AI-Native Engineers, QC Engineers, and Product Owners to improve development and release processes.
- Support engineering teams in resolving environment, deployment, configuration, and production issues.
- Participate in release planning and production-readiness reviews.
- Communicate technical issues, operational risks, and dependencies clearly to relevant stakeholders.
- Work independently when required while contributing to the overall effectiveness of the engineering team.
- Maintain technical documentation, operational procedures, and service runbooks.
Yêu cầu công việc
Education
- Bachelor’s degree in information technology, Software Engineering, Computer Science, or a related field.
Experience
- At least three years of experience as a DevOps Engineer, Site Reliability Engineer, System Engineer, Cloud Engineer, or in a related role.
- Hands-on experience operating cloud infrastructure and production systems.
- Experience with containerized applications and Kubernetes in a production environment.
- Experience provisioning and managing infrastructure through Infrastructure as Code.
- Experience supporting production deployments and responding to system incidents.
- Experience with transactional, e-commerce, OTT, entertainment, or other high-traffic platforms is preferred.
Technical Knowledge and Skills
- Solid knowledge of TCP/IP, HTTP/1.1, HTTP/2, DNS, gRPC, VPC peering, Cloud VPN, and connectivity between cloud environments.
- Good understanding of clustering, replication, microservices, failover, HTTP/gRPC load balancing, and distributed-system concepts.
- Hands-on experience with Linux system administration.
- Experience with containers, Docker, Kubernetes, and preferably Google Kubernetes Engine.
- Experience provisioning and managing infrastructure using Terraform and Ansible.
- Experience implementing or maintaining CI/CD pipelines using Git, GitHub Actions, Argo CD, or equivalent tools.
- Experience with Google Cloud Platform or another major cloud provider.
- Experience with monitoring, logging, alerting, and performance-analysis tools.
- Familiarity with Prometheus, Grafana, Google Cloud Operations, Datadog, or equivalent solutions.
- Ability to monitor system performance, analyze bottlenecks, and troubleshoot production issues.
- Knowledge of system, cloud, network, and application security.
- Basic scripting or programming ability for infrastructure and operational automation.
- Familiarity with Agile/Scrum or Kanban development processes.
Soft Skills
- Strong troubleshooting, analytical, and problem-solving skills.
- Calm and structured when responding to incidents under pressure.
- Data-driven approach to operational and technical decisions.
- Strong teamwork and cross-functional collaboration skills.
- Proactive attitude and willingness to continuously develop technical capabilities.
- Ability to manage multiple responsibilities and organize work effectively.
- Strong ownership and attention to detail.
- Ability to work independently when required.
Language Skills
- Basic English communication skills.
- Ability to read and understand technical documentation in English.
Other Requirements
- Willingness to participate in the production on-call rotation.
- Willingness to provide operational support during major releases, on-sale periods, or live events when required.
Tại sao bạn sẽ yêu thích làm việc tại đây
• Full statutory insurance, including Social Insurance, Health Insurance and Unemployment Insurance, based on 100% of the official salary and in compliance with Vietnamese labor regulations.
• Working hours: Monday to Friday, from 8:30 AM to 5:30 PM, with a one-hour lunch break.
• 14 days of annual leave.
• Company-provided working equipment, including a laptop or desktop computer.
• Employee parking area.
• Annual PMP performance bonus, subject to individual KPI achievement and the Company’s business performance.
• Periodic health check-ups.
• Employee engagement programs and internal activities throughout the year.
30 years of DatVietVAC, 30 years of shaping trends, ...