Reach out directly about this role
We are developing post-training for large language models, including various online RL approaches, and are looking for someone who can not only implement ready-made ideas but also independently lead work from hypothesis to measurable model quality improvement.
What you will do:
We are looking for a strong engineer with good knowledge of Python and PyTorch, experience in LLM post-training / online RL / RLHF, and the ability to independently conduct ML experiments from hypothesis formulation to conclusions.
Experience with the following will be a plus:
Why join us:
If you are interested in developing online RL and turning research ideas into working solutions, message me directly.
Сбер
Russia2 months ago
Grade
Senior
Employment
Full-time
AI writes it from your resume. You only have to send it.
By company and country