Tasks:
Develop and maintain data pipelines:
- Modify existing Lambda functions, SQL queries and processing jobs to add new features
- Refactor code to ingest/process external data
- Add quality checks to existing pipelines
Design and optimize data models:
- Design data models leveraging Iceberg format, for “silver” and “gold” layers
- Optimize queries for performance
Ensure data quality and governance:
- Apply data validation rules
- Build automated quality checks
Documentation and knowledge sharing:
- Document the data preparation process clearly, including transformations, code comments, and key decisions
- Store all code in GitLab with meaningful commit messages and proper versioning for code and datasets
What We're Looking For:
Data Engineering Design and Delivery
- 4–6 years of experience as Data Engineer
- Proven experience designing and building scalable, reliable, and maintainable data pipelines.
- Familiarity with CI/CD, automated testing, and Infrastructure as Code practices.
- Strong Python-based data pipeline development skills (or equivalent), with experience applying Test-Driven Development and AI-assisted coding practices.
- Strong data modeling and schema design expertise, including analytical models and scalable, maintainable data structures.
- Familiarity with DBT (Core or Cloud), including modular SQL transformations, data modeling, testing, documentation, and dependency management.