TechWill is an IT company and a partner of VkusVill for the development of digital solutions.
We are responsible for the development of mobile and web applications, business process automation, artificial intelligence, DevOps, and information security for VkusVill.
Our solutions are used by over 1,000,000 VkusVill clients and employees.
We are building a cluster where GPU is the primary resource, and SRE must be able to operate it as proficiently as the agent part.
Responsibilities:
- Support of services and development teams from the infrastructure side;
- Ensuring the reliability and scalability of systems;
- Identification and elimination of performance bottlenecks;
- Configuration of monitoring, logging, and tracing systems;
- Prevention of potential failures;
- Optimization of CI/CD pipelines, implementation of Infrastructure as Code (IaC), and automation of routine tasks;
- Promotion of DevOps practices towards development: implementation of DevOps best practices such as SLA, SLO, SLI monitoring, incident analysis (postmortem), and change management;
- Participation in the creation and development of infrastructure platforms;
- Ensuring the security, reliability, fault tolerance, and rapid recovery of the platform after failures.
Requirements:
- Practical experience in administering and supporting Linux family information systems (Debian, Ubuntu, Rocky);
- Proficiency in bash or python as tools for automating routine activities;
- Practical experience with container orchestration systems (Kubernetes, Docker compose);
- Practical experience with containers (Docker), knowledge of Dockerfile basics and best practices in this area;
- Proficiency in configuration management systems (Ansible, Terraform, Pulumi) and practical experience applying such systems in IaC (Infrastructure as Code) construction processes;
- Application of GitLab CI and Jenkins tools in building build and delivery processes;
- Practical experience in applying and administering monitoring systems based on Prometheus, Zabbix, Grafana Stack, Alertmanager, VictoriaMetrics;
- Practical experience interacting with event streaming systems (Kafka, RabbitMQ);
- Knowledge of Agile methodologies, experience working with ticketing systems (Yandex Tracker, etc.) and documentation storage systems (Evawiki and other Wikis);
- Practical experience operating web servers and load balancers (Nginx, HAProxy, Traefik, APISIX);
- Practical experience administering relational database management systems (PostgreSQL, Greenplum, MySQL), clustering based on Patroni, as well as columnar DBMS ClickHouse;
- Practical experience applying NoSQL and Key-Value systems (Elasticsearch, OpenSearch, etcd, Redis, Memcached);
- Practical experience applying centralized log collection and storage systems based on stacks: Logstash, Vector, Graylog, Loki;
- Practical experience applying object storage systems based on S3 (MinIO), as well as tools for accessing them;
- Skills in working with cloud systems (Yandex Cloud, Cloud.ru) and their management systems;
- Practical experience in orchestrating data processing pipelines based on Apache Airflow;
- Experience deploying and supporting JupyterHub
- Proficiency in distributed data processing tools (Apache Spark, Spark Streaming);
- Experience working with Iceberg REST Catalog as a tabular data catalog;
- Practical experience applying MLflow for experiment tracking, model registration, and lifecycle management;
- Familiarity with vector databases (Qdrant) as part of an ML platform;
- Familiarity with the n8n process automation platform;
- Understanding/experience deploying and operating LLM models on GPUs (Triton Inference Server, vLLM, TensorRT-LLM or similar);
- Knowledge of GPU monitoring: DCGM, Prometheus, metrics on memory usage, cores, temperature, NVLink;
- Experience profiling inference: optimizing latency (TTFT, TPS), throughput, working with KV-cache;
- Managing GPU resources in K8s: HAMi, MIG, time-slicing, quotas;
- Understanding how to work with NVDEC/NVENC;
- Practical experience operating LLM agents: frameworks, orchestration of multi-step pipelines (LangGraph, LangChain or similar), LLM agent request lifecycle: planning → tool calling → validation;
- Experience with Model Context Protocol: MCP servers and gateways, tool registration/proxying, tool-calling routing between services.
Conditions:
- Work in an accredited IT company.
- Remote work: we value work results regardless of location;
- Official employment from the first day of work and curator support during adaptation.
- Transparent development system: clear grades, internal and external training, individual development plans, and competency matrices.
- Eco-friendly culture and adequate managers.
- Compensation for expenses on medical services, mental well-being, sports, team-building events, and the use of AI assistants.
- 15% bonus on purchases at VkusVill.
- Social responsibility: we encourage donation, provide financial assistance upon the birth of a child.
- Partner program “Green Light”: for referring acquaintances specialists, you can receive up to 50,000 rubles.