Build resilient data pipelines using Medallion architecture principles. Manage and optimize open storage formats including Parquet, Iceberg, and Delta Lake. Design, develop, and maintain distributed data processing workloads using Databricks and Apache Spark. Perform database operations, optimization, and maintenance for PostgreSQL and SQLite environments. Ensure data quality, reliability, scalability, and performance across data platforms. Collaborate with data consumers and stakeholders to support reporting, analytics, and operational data requirements. Monitor, troubleshoot, and improve data infrastructure and pipeline performance. Follow data engineering best practices, coding standards, and documentation processes.
Desired Candidate Profile
Bachelor's degree in Computer Science, Software Engineering, or a related field. Approximately 5 years of experience in Data Engineering or a related role. Advanced SQL skills with strong expertise in PostgreSQL. Strong proficiency in Python and DuckDB. Experience with data lakes and open-source query and storage layers. Hands-on experience working with Parquet, Iceberg, and Delta Lake. Experience with distributed computing frameworks, specifically Databricks and Apache Spark. Strong understanding of data modeling, ETL/ELT processes, and data pipeline development. Knowledge of database administration, performance tuning, and optimization techniques. Excellent problem-solving and analytical skills. Ability to work effectively both independently and within a collaborative team environment. Excellent English communication skills.