Reach out directly about this role
We are developing post-training for large language models, including various online RL approaches, and are looking for someone who can not only implement existing ideas but also independently lead work from hypothesis to measurable model quality improvement.
What you will do:
We are looking for a strong engineer with good knowledge of Python and PyTorch, experience in LLM post-training / online RL / RLHF, and the ability to independently conduct ML experiments from hypothesis to conclusions.
Will be a plus: experience with GRPO and similar policy optimization approaches, reward models, LLM-as-a-judge, verifiers for general-domain tasks, as well as with vLLM, SGLang, verl, Megatron, or TRL.
Why join us:
If you are interested in developing online RL and turning research ideas into working solutions, write to me in direct messages.
Сбер
Russia2 months ago
Grade
Senior
Employment
Full-time
AI writes it from your resume. You only have to send it.
By company and country