openqareer

Mid/Senior DevOps Engineer (Kubernetes, Linux)

Skylab · Thành phố Hồ Chí Minh

# Mid/Senior DevOps Engineer (Kubernetes, Linux) **Skylab** · Thành phố Hồ Chí Minh · `On-site` 🕒 **Статус:** *Опубликовано: 11 дней назад* · *Источник: Indeed* --- ### About the Role Top 3 reasons to join us - Hi-tech working environment - Young and dynamic colleagues - International experience Job description DevOps Engineer (Mid / Senior) Datacenter Infrastructure & Observability Location: Ho Chi Minh City, Vietnam (onsite) | Level: Mid / Senior | Type: Full-time | Experience: Mid 3+ years, Senior 5+ years About SkyLab SkyLab is a Singapore-headquartered AI infrastructure company operating across Southeast and North Asia. We design, build and operate GPU-as-a-Service (GPUaaS) environments on the latest NVIDIA platforms for AI model developers and enterprises, whose training and inference workloads depend on every node and link being healthy, 24 hours a day. About the role The DevOps team builds and runs the platform that keeps it that way. Our monitoring and automation stack watches hundreds of GPU servers across every site, covering hardware health, network fabric and alerting that gets the right people involved quickly. As new sites come online across the region, each one joins this platform from day one. You will join the team that owns it end to end. Everything is infrastructure-as-code, reviewed through merge requests, and runs on-premises rather than on public cloud. You will work closely with our Data Center Operations team and with engineers overseas, so clear written English is an important part of the job. Key responsibilities - Observability: operate and improve a multi-site monitoring platform, where each site keeps working on its own and a central NOC gives fleet-wide visibility. - Hardware and network health: use out-of-band server data and network telemetry to spot issues with GPUs, nodes and links early. - Automation: automate provisioning, configuration and deployment as code, and keep the CMDB accurate so it drives discovery and alerting. - AI-augmented operations: apply AI tools and agents to reduce manual work, such as alert triage, incident summaries and runbook automation, and use AI assistants to write and review code faster. - Alerting: tune alerts to be clear and actionable, and keep the runbooks behind them up to date. - Troubleshooting: investigate issues across hardware, network, OS and application layers, find the root cause, and share what was learned. - Reliability and security: help maintain highly available services, backup and recovery plans, secrets, and regular patching. - CI/CD and tooling: improve pipelines and build internal tools that make changes easier and safer for other teams. - Collaboration: document your work and partner with Data Center Operations and overseas engineers on design and problem-solving. - Our toolset: Prometheus/Grafana stack, Terraform, Ansible, Kubernetes (k3s), Proxmox, GitLab CI, NetBox (CMDB), Jira and PagerDuty. Your skills and experience Requirements (must-have) - Experience: Mid 3+ years in DevOps, SRE or infrastructure roles. Senior 5+ years, including running production systems. - Strong Linux administration and troubleshooting skills. - Hands-on experience with Ansible and Terraform. - Experience with Docker and Kubernetes in production. - Experience with Prometheus-based monitoring (PromQL, alert rules, exporters) and Grafana. - Solid networking fundamentals: TCP/IP, DNS, TLS, load balancing, VLANs and basic routing. - Scripting in Python and Bash. - Virtualization experience, ideally Proxmox or another KVM-based hypervisor. - Git-based workflow with code review and CI/CD, ideally GitLab CI. - Comfortable communicating in English, written and spoken. Nice to have You don't need all of these. Any of them will help you ramp up faster. - Server hardware and BMC management (Redfish, IPMI, iDRAC), ideally with GPU/HPC servers. - Datacenter networking, especially Juniper. - Thanos, Loki or Grafana at multi-site scale. - NetBox or another CMDB. - PostgreSQL high availability (Patroni, etcd) or object storage (MinIO, S3). - Incident tooling (PagerDuty, Jira Service Management) and on-call experience. - Secrets management (SOPS, Vault) and CVE remediation. - Experience building or using AI agents and LLM-based tools in operations (AIOps). - Public cloud (AWS or Azure), Go, or certifications such as CKA/CKAD. For Senior level, we also expect you to - Lead the design of new components and site rollouts. - Help set standards for alert quality, runbooks and code review. - Take a lead role in incident response and follow-up. - Mentor and support mid-level engineers. Soft skills - Calm and methodical when things go wrong. - Proactive, and comfortable working with teammates across time zones. - Enjoys improving how things are done. - Writes clearly, so others can follow your docs and runbooks. Why you'll love working here - Competitive package - Professional working environment - Opportunities to challenge and develop your career - Social insurance, health insurance, unemployment insurance as labor law stipulated - Premium Healthcare - Opportunity to participate in stock option program. - Public holiday and Annual leave in accordance with the Vietnamese labour law

Наблюдалась 2026-10-07, впервые 2026-09-25, источник — Indeed.

Открыть у работодателя