MLOps Engineer
- Work arrangement: fully remote
- Employment type: Full Time
- Seniority: senior
- Posted:
Job description
122860 #MLOps Engineer Senior
Project Responsibilities Design and own the end-to-end MLOps architecture for production Machine Learning and Generative AI systems. Build and maintain ML training, validation, deployment, serving, monitoring, and retraining pipelines. Define and implement the model promotion lifecycle across development, staging, and production environments. Build and maintain model registry, experiment tracking, dataset/model versioning, and reproducible ML workflows. Design scalable training and inference infrastructure, including GPU-backed workloads. Build and maintain CI/CD pipelines for ML and AI workloads with automated testing, quality gates, approval controls, and rollback mechanisms. Implement progressive delivery approaches for ML models, including shadow testing, canary releases, and blue-green deployments. Collaborate closely with Platform Engineering to run AI workloads on shared Kubernetes infrastructure. Translate ML infrastructure requirements into technical requirements for platform, security, capacity, and architecture teams. Establish reusable MLOps project templates, shared pipeline components, and engineering standards. Implement automated ML testing, including data validation, model regression tests, training-serving consistency checks, and evaluation gates. Own production reliability of ML systems, including availability, latency, throughput, scalability, and operational stability. Build observability across infrastructure, data quality, model performance, and business impact. Configure monitoring, alerting, incident response, runbooks, rollback procedures, and post-incident reviews for AI systems. Implement automated model retraining based on schedules, events, and data/model drift. Manage the full model lifecycle, including deployment, monitoring, retraining, version promotion, and retirement. Track and optimize infrastructure, training, inference, and GPU-related costs. Build and operate the production layer for Generative AI and LLM-based applications. Implement LLM model gateways, routing, prompt management, prompt versioning, caching, and rate limiting. Operationalize RAG pipelines, including vector stores, embeddings, re-indexing, chunking strategies, and retrieval quality monitoring. Implement LLM guardrails, groundedness monitoring, sensitive-data protection, and prompt-injection mitigation. Build observability and permission controls for agentic AI systems. Implement token-level and request-level consumption metering across Generative AI applications. Build cost attribution mechanisms by business unit, use case, application, model, tenant, and environment. Optimize LLM workloads through model routing, prompt/context optimization, caching, batching, and selection of cost-efficient models. Apply hands-on Generative AI engineering practices, including prompt engineering, structured outputs, context management, embeddings, tool/function calling, and model adaptation where required. Evaluate emerging ML and Generative AI tools, models, and platforms and recommend technologies for adoption. Make and document architectural and build-vs-configure decisions for new AI capabilities. Produce architecture decision records, reference designs, technical documentation, and operational guidelines. Participate in architecture, capacity planning, technical roadmap, and platform strategy discussions. Perform code reviews, technical mentoring, pairing, and knowledge sharing to improve engineering standards across the team.
Skills
- ci/cd
- kubernetes
- ml
Languages
EN
Apply
Open this job in our interactive board to apply, save it, or sign up for matched alerts on similar roles.
View & apply