openqareer

Lead AI Platform Operations Engineer #AIDA

Sourceo · Singapore

Role Summary Responsible for ensuring the high availability, reliability, and performance of our OpenShift on-prem AI platform Lead proactive monitoring, outage detection, and incident response to minimize downtime and operational risk Design and maintain disaster recovery and business continuity processes to safeguard critical AI workloads Oversee cybersecurity operations, including vulnerability management, audits, and compliance with security standards for AIDA’s AI platform Collaborate closely with MLOps, LLMOps, and engineering teams to integrate automation, observability, and security best practices into platform operations How You will Make An Impact: Lead availability monitoring, outage detection, and performance optimization of our OpenShift on-prem AI platform Manage incident response, root cause analysis, and implement disaster recovery strategies to ensure business continuity Oversee cybersecurity operations including vulnerability management, threat detection, and access control enforcement Handle security audits, compliance reporting, and ensure alignment with Singtel policies, regulatory frameworks and industry best practices Collaborate with other developer teams to integrate monitoring, automation, and security best practices into AI/ML workflows Drive continuous improvement in platform operations through automation, observability, and operational excellence initiatives Lead AIDA AI platform operations function and coordinate distribution of work within team Skills for Success: Bachelor’s degree in Computer Science, Engineering, or a related field 10 years of experience in on-prem platform administration and/or operations Deep expertise in OpenShift platform operations and monitoring services including Elastic Search, Grafana and OpenTelemetry Strong background in incident management, SRE practices, and disaster recovery design Hands-on experience with security operations: IAM, SIEM/SOAR, vulnerability management, firewalls, endpoint detection Proficiency in infrastructure-as-code and automation scripting (Ansible) Familiarity with AI/ML infrastructure (GPU, model hosting) and their operational demands Knowledge of security compliance frameworks (ISO 27001, CIS, NIST) Excellent problem-solving, communication, and leadership skills, especially in high-pressure incident scenarios Forward thinking ability to identify possible failure scenarios and formulate effective response plans Are you ready to say hello to BIG Possibilities? Join Singtel to shape what's next and accelerate your career through meaningful work, continuous learning, and real impact.

Наблюдалась 2026-09-15, впервые 2026-09-04, источник — Indeed.

Открыть у работодателя