Reach out directly about this role
We are developing post-training for large language models, including various online RL approaches, and are looking for someone who can not only implement ready-made ideas but also independently lead work from hypothesis to measurable model quality improvement.
What you will do:
We are looking for a strong engineer with good knowledge of Python and PyTorch, experience in LLM post-training / online RL / RLHF, and the ability to independently conduct ML experiments from hypothesis formulation to conclusions.
Experience with the following will be a plus:
Why join us:
If you are interested in developing online RL and turning research ideas into working solutions, message me directly.
Full-time
Employment
Senior
Grade
Data Science & ML
Specialization
AI
Industry
Corporation
Company Type
By company and country
Corporation
Company Type