Reach out directly about this role
By company and country
Full-time
Employment
Senior
Grade
Data Science & ML
Specialization
AI
Industry
Corporation
Company Type
We are developing post-training for large language models, including various online RL approaches, and are looking for someone who can not only implement existing ideas but also independently lead work from hypothesis to measurable model quality improvement.
What you will do:
We are looking for a strong engineer with good knowledge of Python and PyTorch, experience in LLM post-training / online RL / RLHF, and the ability to independently conduct ML experiments from hypothesis to conclusions.
Will be a plus: experience with GRPO and similar policy optimization approaches, reward models, LLM-as-a-judge, verifiers for general-domain tasks, as well as with vLLM, SGLang, verl, Megatron, or TRL.
Why join us:
If you are interested in developing online RL and turning research ideas into working solutions, write to me in direct messages.