Explore outstanding Cloud & Infrastructure jobs.
See now

Middle Data Engineer (Apache Spark, Trino)

IMIP Technology And Solution Consultancy
+2
Tầng 5, Tòa nhà Lucky, số 81 Trần Thái Tông, Cau Giay, Ha Noi
At office
Posted 1 hour ago
Job Expertise:
Job Domain:
IT Services and IT Consulting

Top 3 reasons to join us

  • Attractive salary
  • Good working environment
  • Flexible management

Job description

Role Overview 

We are seeking a Middle Data Engineer to build and scale our Production Datahouse on-premise environment. This is a purely technical, hands-on role focused on implementing a high-performance, on-premise data architecture. 

You will own the end-to-end technical execution, managing the entire lifecycle of data as it moves from landing ingestion to the final serving layer. This role is not about high-level theory; it is about the hard engineering required to optimize on-premise hardware, tune distributed processing engines, and ensure a seamless data flow across our internal ecosystem. 

Key Responsibilities 

  • Data Pipeline Engineering: Design, build, and maintain automated data ingestion pipelines from several source systems using Airbyte and Airflow, ensuring reliability and scalability. 
  • Lakehouse Architecture Implementation: Develop and manage the Medallion Architecture (Bronze/Silver/Gold) using dbt-spark to enable structured, high-quality analytical datasets. 
  • Platform Performance Optimization: Optimize Spark and Trino workloads on Kubernetes and manage the full lifecycle of Apache Iceberg tables (compaction, snapshotting, and maintenance) to improve query performance, storage and reduce latency. 
  • Infrastructure & Reliability Management: Proactively monitor and maintain Data Platform resources using Grafana and Prometheus; analyze query patterns to enhance system stability and performance for production users. 
  • Data Security & Governance: Enforce row- and column-level security using Open Policy Agent (OPA) and develop foundational frameworks for Data Governance and Data Quality, embedding lineage, automated testing, and compliance directly into pipelines. 
  • Enterprise Data Integration: Architect and implement high-performance data bridges to deliver on-premise Lakehouse data into Microsoft Fabric, enabling seamless and secure hybrid-cloud connectivity.

Your skills and experience

  • 3+ years of hands-on experience in Data Engineering or Platform Engineering, building and operating production-grade data platforms. 
  • Strong expertise in distributed processing with Apache Spark and Trino, including performance tuning and optimization. 
  • Proven experience designing and maintaining reliable data pipelines using Airbyte and Apache Airflow. 
  • Solid understanding of modern Lakehouse architectures, dbt-based transformations, and Iceberg table management. 
  • Hands-on experience with Kubernetes, containerized deployments, and CI/CD for data workloads. 
  • Knowledge of data security, governance, and access control practices (row/column-level security, policy enforcement, IAM). 
  • Experience supporting BI/analytics use cases and integrating with enterprise reporting tools (e.g., Microsoft Fabric, Power BI). 
  • Strong troubleshooting, system reliability, and problem-solving skills in complex distributed environments. 
  • Ability to work independently in a highly technical, ownership-driven role. 

Why you'll love working here

**easy going and friendly environment 

**Remuneration Package: 
• Working hours: 5 days per week (in a professional and yet young, dynamic environment); 
• Salary: Competitive remuneration package (based on skills and experience); 
• 12 leave days per year 
 

IMIP specializes in productivity solutions for scanning, marketing automation and document capture.

Company type
IT Outsourcing
Company industry
IT Services and IT Consulting
Company size
51-150 employees
Country
Vietnam
Working days
Monday - Friday
Overtime policy
No OT

More jobs for you

Get similar jobs by email Subscribe