Việc làm này đã được thêm vào mục Việc làm đã lưu.
Bạn đã lưu tối đa 20 việc làm. Nếu bạn muốn lưu mới, hãy cập nhật Việc làm đã lưu.
Middle Data Engineer (Apache Spark, Trino)
3 Lý do để gia nhập công ty
- Attractive salary
- Good working environment
- Flexible management
Mô tả công việc
Role Overview
We are seeking a Middle Data Engineer to build and scale our Production Datahouse on-premise environment. This is a purely technical, hands-on role focused on implementing a high-performance, on-premise data architecture.
You will own the end-to-end technical execution, managing the entire lifecycle of data as it moves from landing ingestion to the final serving layer. This role is not about high-level theory; it is about the hard engineering required to optimize on-premise hardware, tune distributed processing engines, and ensure a seamless data flow across our internal ecosystem.
Key Responsibilities
- Data Pipeline Engineering: Design, build, and maintain automated data ingestion pipelines from several source systems using Airbyte and Airflow, ensuring reliability and scalability.
- Lakehouse Architecture Implementation: Develop and manage the Medallion Architecture (Bronze/Silver/Gold) using dbt-spark to enable structured, high-quality analytical datasets.
- Platform Performance Optimization: Optimize Spark and Trino workloads on Kubernetes and manage the full lifecycle of Apache Iceberg tables (compaction, snapshotting, and maintenance) to improve query performance, storage and reduce latency.
- Infrastructure & Reliability Management: Proactively monitor and maintain Data Platform resources using Grafana and Prometheus; analyze query patterns to enhance system stability and performance for production users.
- Data Security & Governance: Enforce row- and column-level security using Open Policy Agent (OPA) and develop foundational frameworks for Data Governance and Data Quality, embedding lineage, automated testing, and compliance directly into pipelines.
- Enterprise Data Integration: Architect and implement high-performance data bridges to deliver on-premise Lakehouse data into Microsoft Fabric, enabling seamless and secure hybrid-cloud connectivity.
Yêu cầu công việc
- 3+ years of hands-on experience in Data Engineering or Platform Engineering, building and operating production-grade data platforms.
- Strong expertise in distributed processing with Apache Spark and Trino, including performance tuning and optimization.
- Proven experience designing and maintaining reliable data pipelines using Airbyte and Apache Airflow.
- Solid understanding of modern Lakehouse architectures, dbt-based transformations, and Iceberg table management.
- Hands-on experience with Kubernetes, containerized deployments, and CI/CD for data workloads.
- Knowledge of data security, governance, and access control practices (row/column-level security, policy enforcement, IAM).
- Experience supporting BI/analytics use cases and integrating with enterprise reporting tools (e.g., Microsoft Fabric, Power BI).
- Strong troubleshooting, system reliability, and problem-solving skills in complex distributed environments.
- Ability to work independently in a highly technical, ownership-driven role.
Tại sao bạn sẽ yêu thích làm việc tại đây
**easy going and friendly environment
**Remuneration Package:
• Working hours: 5 days per week (in a professional and yet young, dynamic environment);
• Salary: Competitive remuneration package (based on skills and experience);
• 12 leave days per year
IMIP specializes in productivity solutions for scanning, marketing automation and document capture.