openqareer

Senior DevOps / SRE Engineer

Legend Holding Group · Dubai

# Senior DevOps / SRE Engineer **Legend Holding Group** · Dubai · `On-site` 🕒 **Статус:** *Опубликовано: 6 дней назад* · *Источник: Indeed* --- ### About the Role JD - Senior DevOps / SRE Engineer Experience: 5–6 Years Department: Technology / Engineering Employment Type: Full-Time Reporting To: Engineering Manager / IT Manager About the Role We are looking for a Senior DevOps / SRE Engineer with 5–6 years of experience to own the reliability, scalability, observability, security, automation, and cost efficiency of our technology platform. The role is highly hands-on and requires strong experience with AWS, Kubernetes/EKS, CI/CD, Infrastructure as Code, monitoring, incident management, and cloud cost optimization. Key Responsibilities DevOps & Cloud Infrastructure - Design, manage, and optimize production infrastructure on AWS. - Manage services including EKS, EC2, RDS, S3, ECR, ElastiCache, VPC, IAM, Load Balancers, Route 53, CloudFront, and CloudWatch. - Manage production and non-production environments. - Implement highly available, scalable, secure, and cost-efficient infrastructure. - Troubleshoot infrastructure, networking, and production issues. Kubernetes & CI/CD - Strong hands-on ownership of Kubernetes / AWS EKS. - Manage deployments, services, ingress, HPA, RBAC, secrets, namespaces, and node groups. - Troubleshoot pod, node, networking, resource, and scheduling issues. - Use Helm for application deployments. - Build and maintain CI/CD pipelines using GitHub Actions, Jenkins, or similar tools. - Implement automated deployments, rollbacks, and release processes. SRE & Reliability - Define and monitor SLIs, SLOs, and SLAs for critical services. - Improve system availability, resilience, performance, and scalability. - Implement proactive monitoring, alerting, health checks, and automated recovery. - Reduce operational toil through automation. - Participate in production incident management and on-call activities. - Lead RCA and postmortems for major incidents. - Drive corrective and preventive actions. Observability - Build and maintain monitoring, logging, and distributed tracing. - Work with tools such as Datadog, Prometheus, Grafana, CloudWatch, OpenTelemetry, Fluent Bit, or ELK/OpenSearch. - Create actionable dashboards and alerts. - Monitor latency, traffic, errors, saturation, infrastructure health, and service availability. Infrastructure as Code & Automation - Manage infrastructure using Terraform or equivalent IaC tools. - Automate infrastructure provisioning, deployments, scaling, backups, and operational processes. - Build reusable infrastructure and deployment components. - Minimize manual operational activities. Security & Disaster Recovery - Implement AWS and Kubernetes security best practices. - Manage IAM, secrets, access controls, and infrastructure security. - Implement backup, disaster recovery, RTO/RPO, and high-availability strategies. - Regularly test recovery and failover procedures. Cloud Cost Management / FinOps - Own BU-level cloud cost visibility, allocation, and optimization. - Implement resource tagging by BU, application, environment, and project. - Monitor AWS spend and budget vs actuals. - Identify cost anomalies and unnecessary/underutilized resources. - Optimize EKS/EC2, RDS, Redis, S3, logging, and data-transfer costs. - Build cost dashboards and reports for Business and Finance teams. - Drive measurable cloud cost savings without compromising reliability or performance. Required Skills - 5–6 years of hands-on DevOps / SRE / Cloud Engineering experience. - Strong AWS experience. - Strong Kubernetes / EKS experience. - Strong Docker and Linux experience. - Strong CI/CD experience. - Experience with Terraform and Helm. - Strong scripting skills in Bash, Python, or Go. - Strong monitoring and observability experience. - Experience handling production incidents and conducting RCA. - Good understanding of networking, security, and cloud architecture. - Understanding of SRE principles: SLI, SLO, SLA, Error Budgets, MTTR, capacity planning, and incident management. - Experience with cloud cost management / FinOps is highly preferred. Good to Have - Datadog - Prometheus / Grafana - OpenTelemetry - Kafka - Redis / ElastiCache - PostgreSQL / RDS - ArgoCD / GitOps - OpenCost - AWS certifications Key Success Metrics - Production availability and SLO achievement - Reduction in MTTR and recurring incidents - CI/CD reliability and deployment success - Infrastructure automation - Observability and alert coverage - Kubernetes health and resource utilization - Disaster recovery readiness - BU-level cloud cost visibility and optimization - Measurable cloud cost savings Ideal Candidate A hands-on engineer who thinks beyond deployment and focuses on: Reliability + Automation + Observability + Security + Scalability + Cost Efficiency Someone who can own the platform from AWS → Kubernetes → CI/CD → Monitoring → Incident Response → Automation → Cost Optimization. Pay: AED5,000.00 - AED6,000.00 per month Work Location: In person

Наблюдалась 2026-10-07, впервые 2026-09-30, источник — Indeed.

Открыть у работодателя