Top Selection Company**
🔥We are looking for a **Data Engineer for project-based employment**
**Grade: middle+|**Senior
**Rate:** from 272K to 285K
Citizenship/Location: RF
**Workload:** full-time
Term: long-term
Employment: only sole proprietor 📌
✅ **Mandatory requirements:**
- At least 4 years of experience as a Data Engineer;
- Experience in the full data lifecycle: from designing pipelines and models to production deployment and monitoring;
- Systems thinking, ability to design scalable and fault-tolerant solutions that consider data volume, velocity, and variety;
- Effective communication skills with analytics, business analytics, DevOps, and product development teams;
- Experience conducting load tests for DataLake platforms;
- Practical experience in writing efficient SQL queries for data analysis and transformation in StarRocks or similar OLAP systems (ClickHouse, Impala);
- Ability to create and maintain tables, partitions, views;
- Basic understanding of the StarRocks data model (duplicate/aggregate tables) for implementing ready-made solutions;
- Experience loading data (via files, INSERT, using simple connectors);
- Working with HMS via Spark or Hive to create/update tables, read metadata. Understanding the purpose of a metadata catalog;
- Confident work with Parquet, Iceberg, JSON, CSV. Understanding the advantages of columnar formats;
- Experience writing DAGs in Airflow (or similar) for scheduling regular ETL tasks. Understanding the principles of idempotency and task restarts;
- Integration of Data Ocean Nova with data sources (databases, BI tools). Understanding the architecture of such platforms (often microservice-based on K8s);
- Understanding how the listed components interact with each other in a unified platform. For example, how a query from StarRocks via HMS gets table metadata, and Ranger checks access rights;
- SQL (Advanced level);
- Complex JOINs, window functions, aggregations;
- Ability to read and analyze query execution plans (EXPLAIN) for basic optimization;
- Python (Intermediate level);
- Developing scripts for ETL, working with APIs, Pandas for processing moderate data volumes. Basic knowledge;
- PySpark: Ability to write and optimize Spark applications (DataFrame API) for batch data processing. Understanding the basics of transformations/actions, principles of data partitioning in Spark;
- Kubernetes: Basic understanding of concepts (Pod, Deployment, Service). Experience launching and monitoring your tasks (Spark, containers) in K8s. Ability to work with pod logs.
📆 **Tasks:**
- Independent development, implementation, and support of integration solutions using the technology stack adopted by the team (Java, Groovy, Apache Nifi, Airflow);
- Determining the technology stack for specific projects and tasks;
- Solving technically complex problems that other engineers in the team cannot solve;
- Promptly responding to information about problems in the area of responsibility, completing tasks within the set deadlines;
- Developing and maintaining the accuracy of documentation on the interaction of big data platform configuration units;
- Providing reports on your activities to the department head/manager as established by the management;
- Monitoring the quality of integration solutions with subsequent creation of tasks/defects for refactoring;
- Determining the technological strategy for the development of a project or product, working for the future;
- Building processes (e.g., CI/CD, code review), implementing and developing engineering practices..
@aliiS_a