We are looking for a Senior DevOps Observability Engineer for an international product company in the iGaming sector.
We need a strong DevOps engineer with current production experience in Kubernetes and deep practical experience in observability, monitoring, and handling production incidents.
Responsibilities:
- Ensure the reliable operation of applications in Kubernetes in production
- Diagnose incidents and issues in Kubernetes clusters and distributed systems
- Develop the Observability platform for metrics, logs, and traces
- Work with VictoriaMetrics, Grafana, Loki, Tempo, Mimir, ELK, OpenTelemetry, and other observability tools
- Develop monitoring, alerting, and dashboards considering service architecture
- Maintain and develop the Observability strategy, standardize approaches for engineering teams
- Participate in the full cycle of handling production incidents - from OnCall and service recovery to RCA/Post-Mortem
- Implement reliability improvements based on incident outcomes and reduce the likelihood of their recurrence
- Interact with development/backend and other engineering teams
We want to see in your experience:
- 3+ years of experience as a DevOps Engineer
- Strong recent practical experience with Kubernetes in production
- Experience in troubleshooting and resolving incidents in Kubernetes clusters
- Experience in independently developing and maintaining Helm charts - templating, functions, custom charts of varying complexity
- Practical experience in building and developing monitoring/observability
- Experience with Grafana, VictoriaMetrics/Prometheus, Loki, ELK, OpenTelemetry, or a comparable stack
- Experience with production incidents, RCA/Post-Mortem, and subsequent reliability improvements
- CI/CD and GitOps experience - GitLab CI/CD, Argo CD, or comparable tools
- Confident work with Git
- Experience with Docker and container build tools
- Ability to independently troubleshoot complex production system issues and find root causes
- Experience interacting with cross-functional engineering teams
Will be a plus:
- OnCall experience and work with PagerDuty
- Experience with Sentry, Vector, Tempo, Mimir
- Experience with AWS/CloudWatch
- Terraform/Ansible and other Infrastructure as Code
- Experience with SLA/SLO and reliability engineering
What we offer:
- Work with a large-scale production infrastructure and a modern observability stack
- Tasks where you can influence observability and reliability approaches, not just maintain existing solutions
- An international product company and a strong engineering team that has been working together for a long time
- Office format 5/2 (hybrid is also possible)
- 8-hour workday + 30 minutes for lunch
- Flexible start of the workday from 8:00 to 10:00
- Breakfasts and lunches in the office at the company's expense
- Internal and external training
- Compensation for sports activities
- English and Greek language learning programs
- Regular corporate events for employees and their children
📩 Contact - @yuloby