Reach out directly about this role
Company: Green Wave Palace ltd Location: RF (remote) Employment: Part-time Salary: Upon successful interview Employment Type: Self-employed / Individual Entrepreneur / Labor Code of RF
‼️ Can be combined with main job ‼️
• Design and implement production LLM services: from data ingestion and indexing to response generation and user feedback. • Build RAG pipelines: hybrid search (vector + BM25), context compression, re-ranking (cross-encoder/learning-to-rank), metadata filtering. • Orchestrate agentic workflows (LangChain / LangGraph): planning steps, tool invocation, error handling and fallbacks. • MCP (Model Context Protocol): ability to set up/connect MCP servers and publish tools/resources/prompts for hosts (e.g., Claude/ChatGPT/IDE), understanding MCP security and authentication. • Perform quality evaluation: automated and human-in-the-loop (groundedness, factuality, relevance, hallucination rate). • Develop and maintain REST/HTTP APIs (FastAPI, async/await), service integrations, and background processing queues. • Ensure reliability and security: PII control, guardrails, input data validation and sanitization.
• Python 3.x: asynchronous programming (asyncio/httpx), typing, Pydantic, FastAPI, SQLAlchemy; solid practical experience in production backend development. • Experience building RAG: embedding selection (OpenAI, e5, BGE, etc.), chunking/overlap strategy, index building and updating, vector databases (FAISS, Pinecone, Weaviate), hybrid search and re-ranking. • LangChain/LangGraph or any other agent framework; ability to build chains/graphs, connect tools, external APIs, and storage. • Working with multiple LLM providers (OpenAI, Anthropic, Mistral, Gemini, etc.), model routing and fallbacks; basic token parameter and system prompt configuration. • Practice in evaluation and observability: quality, latency, and cost metrics; ability to build a simple evaluation pipeline.
• LLMOps/Observability: Langfuse/Arize Phoenix, Promptfoo/Ragas, cost & latency dashboards, chain tracing. • Search: Elasticsearch/OpenSearch, hybrid (BM25 + dense), external reranker models (e.g., cross-encoder/Cohere ReRank). • Cloud and infrastructure: AWS/GCP/Azure (including Azure OpenAI/Bedrock/Vertex), Docker/K8s, queues (Celery/Kafka), Redis. Jobgether • Multimodality (VLM), OCR, structured fact extraction from documents.
Contacts: @irinaalekseevnashi
Part-time
Employment
Remote
Work Format
Senior
Grade
AI Engineering
Specialization
AI
Industry
By country
AI
Industry