Job responsibilities
- Utilize standard monitoring and processes to assess the health of Riot’s live services
- Lead the investigation and mitigation of live incidents as an Incident Commander, and participate in post-incident RCAs
- Use investigation and troubleshooting to return impacted systems to service quickly and keep our games green for players
- Identify improvements to the team’s tools, processes, and documentation, as well as broader improvements to prevent incidents and drive down incident impact for Riot’s live services
- Write tooling, automation, and services to implement these improvements
- Write and understand code in the team’s codebases, utilizing appropriate data structures, algorithms, and software testing best practices
What you’ll do
The standard working schedule consists of four 10-hour days, including one weekend day and three days off per week. This schedule supports our follow-the-sun operational model, ensuring continuous global coverage across timezones.
Requirements
- Bachelor’s Degree in Computer Science (or equivalent experience)
- 2+ years of industry experience or an advanced degree
- Hands-on experience programming in Java or Go
- Experience with technical processes such as code reviews and testing
- Experience debugging issues with production systems
- Experience with monitoring and event management platform
Preferred Qualifications
- Knowledge of cloud services (e.g. AWS and its common services)
- Knowledge of containerization technologies (e.g. Docker, Kubernetes)
- Knowledge of relational databases (e.g. MySQL)
- Knowledge of incident management processes (e.g. ITIL)
- Knowledge of Site Reliability Engineering (SRE) principles and best practices
- Experience deploying and operating services in a live environment