Intern for the Causal AI team in Sberbank's "Finance" block to work on causal inference tasks, uplift modeling, and A/B test analysis. Seeking a final-year student with a strong mathematical background and Python skills.
Responsibilities:
- Collect and prepare data, participate in data mart assembly
- Conduct EDA and feature engineering, check data quality and completeness
- Train and validate ML models
- Assist in A/B test design and analysis: sample size calculation, split correctness verification, results breakdown
- Reproduce methods from articles and open-source libraries, compare with internal baselines on our data
- Write reusable code, tests, and documentation for the internal uplift platform
- Support current analytics: model bug fixes and ad hoc requests from the team and clients
Requirements:
- Incomplete higher education (final university years) in a technical field
- Solid mathematical foundation, especially in probability theory and statistics (understanding of p-value, confidence intervals, characteristics of basic distributions)
- Ability to write Python code, optimize it, and understand others' code
- Proficiency in core data analysis libraries (pandas, numpy, sklearn)
- Understanding of the core algorithms of classical ML (linear and tree-based models)
- Ability to turn a vague question into a testable hypothesis
Conditions:
- Comfortable modern office near "Leninsky Prospekt" metro station
- Option to choose a flexible schedule – office/hybrid
- Fixed-term employment contract for 3 months
- Corporate gym and recreation areas
- Over 400 educational programs from SberUniversity
- Adaptation program and mentor support at the start (for Junior level positions)
- Extended voluntary medical insurance, preferential family insurance, and corporate pension program
- Flexible mortgage discount equal to 1/3 of the Central Bank's key rate
- Prime subscription with the ability to share with three close individuals
- Referral bonus for recommending friends to the Sber team
Skills:
- Python
- scikit-learn
- CatBoost
- causal-learn
- DoWhy
- EconML
- Hypex
- Hadoop
- PySpark
- Hive
- HDFS
- pandas
- numpy
- Spark
- A/B Testing
- ML