Senior Site Reliability Engineer – Data Platforms
# Senior Site Reliability Engineer – Data Platforms **City People Solutions** · Mandaluyong · `On-site` 🕒 **Статус:** *Опубликовано: сегодня* · *Источник: Indeed* --- ### About the Role We are looking for a Senior Data Platform Reliability Engineer to operate, maintain, and continuously improve data platforms running on Kubernetes, either on-premise or on AWS/GCP. Key Responsibilities: - Deploy releases and configuration changes through GitOps/DevOps. - Monitor platform and service health using logs, metrics, and observability tools. - Participate in incident response, root cause analysis, and 24x7 operational rotations. - Improve platform observability, automation, self-service capabilities, and reliability practices. - Troubleshoot system issues, integrations, errors, and misconfigurations. - Provide technical mentorship to junior engineers. - Promote platform standards, security best practices, and operational excellence. Qualifications: - 3+ years of experience supporting production data workloads/platforms such as Spark, Airflow, or Jupyter . - 5+ years of hands-on experience in ETL/ELT pipeline development and data transformation using Python/Java and SQL. - Strong hands-on experience with Kubernetes , including AWS EKS or GCP GKE. - Strong knowledge of Linux environments, microservices, and service communication . - Strong troubleshooting skills involving application crashes, resource contention, service latency, and scaling. - Experience analyzing logs, metrics, monitoring systems, and service KPIs . Nice to Have: - Experience with data/AI platforms such as Flink, Trino, Druid, or Ray.io . - Bash and Python automation/scripting experience. - Kubernetes or data certifications such as CKAD or AWS Certified Data Engineer . Work Location: Hybrid remote in Mandaluyong
Наблюдалась 2026-10-06, впервые 2026-10-06, источник — Indeed.