Khám phá việc làm Cloud & Infrastructure nổi bật.
Xem ngay

Senior Site Reliability Engineer (SRE, GCP, Kubernetes)

DatVietVAC
+2
222 Pasteur, Phường Xuân Hòa, TP Hồ Chí Minh
Tại văn phòng
Đăng 1 giờ trước
Lĩnh vực:
Thương Mại Điện Tử
Truyền Thông, Quảng Cáo và Giải Trí
Sản Phẩm Phần Mềm và Dịch Vụ Web
Phần mềm và Dịch vụ Trí tuệ Nhân tạo

3 Lý do để gia nhập công ty

  • Build impactful products from the ground up
  • Tackle high-scale, real-world tech challenges
  • Grow with a multidisciplinary tech team

Mô tả công việc

DatVietVAC Group Holdings is looking for a Senior SRE Engineer to lead the infrastructure architecture, platform reliability, and production operations of our Fan Commerce Platform.

The platform operates across Google Cloud for computing and scaling and CMC Cloud for local data residency. You will be responsible for building a secure, scalable, and highly available infrastructure capable of supporting high-traffic on-sale periods and major entertainment events while maintaining operational efficiency and infrastructure costs within the approved budget.

1. Cloud Infrastructure Architecture

  • Design and implement multi-environment infrastructure, including Production, Staging, and Development, using Google Cloud services such as GKE, Cloud SQL, Memorystore, and Pub/Sub.
  • Design the hybrid-cloud architecture between Google Cloud and CMC Cloud, including private connectivity, VPN configuration, network segmentation, and data-flow separation.
  • Build and manage infrastructure as code using Terraform.
  • Develop and maintain CI/CD pipelines and automation tools that support engineering teams throughout the software development and release lifecycle.
  • Establish a comprehensive observability platform covering metrics, logs, traces, dashboards, and alerting.
  • Ensure that the infrastructure architecture supports scalability, maintainability, security, and local data-residency requirements.

2. Site Reliability and Incident Response

  • Define and manage Service Level Indicators, Service Level Objectives, and error budgets for core platform services.
  • Take ownership of platform availability, reliability, scalability, and operational readiness.
  • Lead capacity planning and load testing for high-traffic product launches, ticket or merchandise on-sales, and live events.
  • Design autoscaling strategies, traffic-management mechanisms, and overload-protection measures.
  • Participate in the 24/7 production on-call rotation and lead the response to high-severity incidents.
  • Lead blameless postmortems, identify root causes, and ensure that corrective and preventive actions are completed.
  • Develop incident-response procedures, operational runbooks, and disaster-recovery plans.
  • Coach and mentor engineers in production operations, incident management, and reliability practices.

3. Security and Compliance

  • Implement infrastructure security controls, including WAF, DDoS protection, and bot-management solutions using tools such as Cloudflare and Google Cloud Armor.
  • Manage secrets, encryption, identity, and access controls across cloud environments.
  • Establish backup, recovery, and disaster-recovery processes and conduct periodic recovery drills.
  • Collaborate with relevant teams on penetration testing, vulnerability remediation, security audits, and compliance requirements.
  • Ensure appropriate monitoring and audit logging for sensitive infrastructure and operational activities.

4. Performance and Cost Optimization

  • Monitor, analyze, and optimize cloud infrastructure costs through rightsizing, committed-use discounts, storage lifecycle management, and egress optimization.
  • Investigate performance bottlenecks and implement solutions to improve platform efficiency, scalability, and resilience.
  • Provide infrastructure and reliability recommendations to the Project Lead and engineering teams.
  • Evaluate and propose technologies that support the platform’s long-term architecture and business requirements.

Yêu cầu công việc

Education

  • Bachelor’s degree or higher in Information Technology, Software Engineering, Computer Science, or a related field.

Experience

  • At least five years of experience as a DevOps Engineer, Site Reliability Engineer, System Engineer, Cloud Engineer, or in a similar infrastructure role.
  • At least one year of experience at an equivalent senior level or in a technical leadership role.
  • Proven experience designing or operating production systems with high traffic or significant peak-load events, such as flash sales, on-sales, live events, entertainment platforms, or e-commerce platforms.
  • Hands-on experience operating Kubernetes and cloud infrastructure in a production environment.
  • Production experience with Google Cloud Platform is mandatory.

Technical Knowledge and Skills

  • Advanced knowledge of networking and security concepts, including TCP/IP, HTTP/1.1, HTTP/2, HTTP/3, DNS, gRPC, VPC peering, Cloud VPN, and on-premises-to-cloud connectivity.
  • Strong understanding of distributed systems, microservices, clustering, replication, failover, load balancing, and autoscaling.
  • Strong hands-on experience with Kubernetes, particularly Google Kubernetes Engine, and container technologies in production environments.
  • Proficiency in infrastructure as code and automation tools, particularly Terraform and Ansible.
  • Strong experience building and maintaining CI/CD pipelines using GitHub Actions, Jenkins, or equivalent tools.
  • Strong Linux administration skills, including experience with Ubuntu or CentOS.
  • Hands-on experience with Google Cloud services; additional AWS or Azure experience is an advantage.
  • Experience with performance testing, load testing, bottleneck analysis, capacity planning, and system optimization.
  • Knowledge of API gateways, reverse proxies, ingress controllers, and service mesh architecture.
  • Experience building observability systems using metrics, logs, traces, dashboards, and alerts.
  • Knowledge of SLI/SLO management, error budgets, incident response, postmortems, backup, and disaster recovery.
  • Experience with infrastructure security solutions such as WAF, DDoS protection, bot management, secrets management, and access control.

Soft Skills

  • Strong technical leadership and cross-functional collaboration skills.
  • Ability to mentor engineers and guide teams through incident response and production operations.
  • Calm, structured, and decisive when handling critical production incidents.
  • Strong analytical, troubleshooting, and risk-management skills.
  • Data-driven approach to reliability, performance, capacity, and cost optimization.
  • Strong attention to detail and a high sense of ownership.
  • Ability to explain technical decisions and trade-offs clearly to both technical and non-technical stakeholders.

Language Skills

  • Basic English communication skills.
  • Good ability to read and understand technical documentation in English.

Other Requirements

  • Willingness to participate in the production on-call rotation and support major on-sale periods or live events when required.
  • Ability to work under pressure during critical releases, peak-traffic events, and production incidents.

Tại sao bạn sẽ yêu thích làm việc tại đây

• Full statutory insurance, including Social Insurance, Health Insurance and Unemployment Insurance, based on 100% of the official salary and in compliance with Vietnamese labor regulations.
• Working hours: Monday to Friday, from 8:30 AM to 5:30 PM, with a one-hour lunch break.
• 14 days of annual leave.
• Company-provided working equipment, including a laptop or desktop computer.
• Employee parking area.
• Annual PMP performance bonus, subject to individual KPI achievement and the Company’s business performance.
• Periodic health check-ups.
• Employee engagement programs and internal activities throughout the year.

30 years of DatVietVAC, 30 years of shaping trends, ...

Mô hình công ty
Sản phẩm
Lĩnh vực công ty
Truyền Thông, Quảng Cáo và Giải Trí
Quy mô công ty
501-1000 nhân viên
Quốc gia
Vietnam
Thời gian làm việc
Thứ 2 - Thứ 6
Làm việc ngoài giờ
Không có OT

Việc làm tương tự dành cho bạn

Nhận các việc làm tương tự qua email Nhận thông báo