openqareer

MLOps / AI Operations Engineer

PwC · București

# MLOps / AI Operations Engineer **PwC** · București · `On-site` 🕒 **Статус:** *Опубликовано: 13 дней назад* · *Источник: Indeed* --- ### About the Role Job Description & Summary The opportunity Industrialize AI delivery through automated deployment, evaluation operations, observability, reliability engineering and transparent consumption management. What you will be doing - Build CI/CD pipelines for AI services, prompts, agent configurations, infrastructure and evaluation assets. - Automate environment provisioning, testing, deployment, rollback and release evidence. - Implement tracing, logging, model and agent monitoring, alerts and operational dashboards. - Operationalize evaluation thresholds, incident handling and continuous-improvement loops. - Monitor latency, capacity, token usage, infrastructure consumption and cost drivers. - Define runbooks, service ownership and production support handover. What we need from you - 4+ years in DevOps, platform engineering, ML engineering, SRE or cloud operations. - Strong automation, containers, cloud services, observability and Infrastructure as Code capability. - Experience deploying or operating ML, generative AI or distributed application workloads. - Understanding of release controls, reliability, security and cost optimization. Relevant AI technologies and tooling - Hands-on experience with GitHub Actions, Azure DevOps, GitLab CI or equivalent, plus Infrastructure as Code using Terraform, Bicep or comparable tooling. - Strong container and orchestration capability using Docker and Kubernetes, together with experience deploying AI or agent services across cloud and hybrid environments. - Experience operating model and prompt assets, agent configurations, evaluation datasets and release evidence using MLflow, platform-native registries or equivalent lifecycle tooling. - Practical implementation of agent tracing and observability using OpenTelemetry and tools such as LangSmith, MLflow, Langfuse, Azure Monitor, Prometheus or Grafana. - Ability to monitor model and agent quality, tool failures, retrieval performance, latency, token usage, cost, capacity and workflow-level service indicators. - Experience with progressive delivery, rollback, secrets management, vulnerability scanning, incident response and reliability practices for non-deterministic AI systems. Measures of success - Deployment frequency and success rate - Mean time to detect and restore - Evaluation and monitoring coverage - Service reliability and latency - Cost and consumption transparency Key interfaces - Other members of the AI Transformation & Agentic Systems Practice - PwC sector, functional, cloud, cyber, risk, Responsible AI and change specialists - Client business owners, product owners, technology teams and operational users - Technology alliance and implementation partners where relevant Contribution to the practice - Support proposals, client workshops and market development appropriate to seniority. - Contribute reusable methods, patterns, code, assets and lessons learned. - Coach colleagues and participate in the capability’s continuous learning agenda. - Uphold PwC quality, independence, confidentiality and risk-management requirements. #LI-BS1 #LI-Hybrid

Наблюдалась 2026-10-07, впервые 2026-09-23, источник — Indeed.

Открыть у работодателя