About the product
Examus is a proctoring system: ensuring the integrity of online exams. Universities and companies use our system for various scenarios, including admissions campaigns, internal and external exams, and certifications.
The load is seasonal and peak: during exam days, there are thousands of simultaneous sessions with video streams, and the cost of downtime is measured by disrupted exams for individuals. This imposes higher-than-market-average availability requirements.
Infrastructure to work with
- Cloud infrastructure in VK Cloud, object storage in Yandex Cloud
- Backend Python/Django + DRF, Celery (worker and beat), Node.js + Socket.IO for WebRTC signaling
- PostgreSQL (including managed), Redis
- Cloud provider load balancers, DNS, TLS, TURN/STUN (Coturn)
- Packaged product deployments on customer infrastructure: docker compose, installation scripts
Kubernetes clusters are managed by an external contractor. Kubernetes operations are not part of this role's responsibility — you need to be able to assign tasks to the contractor and verify the results, not administer the cluster.
Tasks
- Monitoring and Alerting. Achieve observability to a state where component degradation is visible before the client reports it: metrics, logs, traces, meaningful alerts, and prepared incident response scenarios.
- Traffic Ingestion Fault Tolerance. Eliminate single points of failure in the load balancing, DNS, and service publication scheme, ensure predictable failover and verifiable recovery.
- CI/CD. Development of build and deployment pipelines, reducing release time and risk, ensuring reproducible environments.
- Infrastructure as Code. Transitioning configurations to code, versioning, eliminating manual changes in production.
- Packaged Deployments. Packaging and supporting product installation in closed customer environments, assisting with implementation.
- Incident Management. Participating in incident analysis, conducting postmortems, driving actions to implementation rather than just documentation.
- Interaction with Cloud Providers and Contractors. Assigning tasks, monitoring deadlines and SLAs, escalating issues, verifying that contractor deliverables are actually completed.
What we expect
- 3+ years of experience in production infrastructure operations
- Proficient Linux and networking: TCP/IP, DNS, TLS, routing, ability to localize issues between application, cloud, and provider network
- Docker, docker compose
- Experience with monitoring systems (Prometheus, Grafana, or similar), setting up metrics and alerts from scratch
- CI/CD: GitLab CI, GitHub Actions, or similar
- Infrastructure as Code: Terraform, Ansible, or similar
- Scripting: Bash, Python
- Experience operating services in Russian clouds
- Ability to independently drive tasks to completion and communicate clearly: postmortems, diagrams, instructions
Will be a plus
- Experience working with external operations contractors: task assignment, acceptance, SLA disputes
- Basic understanding of Kubernetes at a user level: viewing logs, pod status, identifying problems, and correctly escalating them to the contractor
- PostgreSQL operation under load: replication, backups, recovery verification
- Experience with WebRTC, TURN/STUN, video traffic
- Experience deploying products in closed customer environments (on-premise, isolated networks)
- Understanding of regulatory requirements for data placement, experience with information security audits
- Participation in on-call rotations and building an on-call process
What we offer
- Influence on the operational architecture: the role is created to establish processes, not just maintain existing ones
- Direct contact with management, short decision-making paths
- A product with real-world load and a clear cost of failure — the results of your work are visible
- Official employment according to the Labor Code + VHI (Voluntary Medical Insurance) after probation
Application Question
Please answer in your cover letter — your response will help us understand your thinking process and speed up our conversation:
Imagine: during an exam, some users lose connection, but the service is generally "green" — all dashboards are normal. The client writes before an alert triggers. Where would you start, and how would you ensure that the system detects degradation before the client calls next time? Please answer in two to three paragraphs; your thought process is more important to us than a ready-made solution.
Selection Process
- Short introductory call (30 minutes)
- Technical interview with case studies (60-90 minutes)
- Final meeting with management, terms discussion