openqareer

Senior Network Reliability Engineer

Expleo Group · București 020276

# Senior Network Reliability Engineer **Expleo Group** · București 020276 · `On-site` 🕒 **Статус:** *Опубликовано: 26 дней назад* · *Источник: Indeed* --- ### About the Role Overview: Expleo is a global engineering, technology, and consulting service provider that partners with leading organizations to guide them through their business transformation, helping them achieve operational excellence and future-proof their businesses. Expleo benefits from more than 50 years of experience developing complex products in automotive and aerospace, optimizing manufacturing processes, and ensuring the quality of information systems. Leveraging its deep sector knowledge and wide-ranging expertise in fields including AI engineering, digitalization, automation, cybersecurity and data science, the group’s mission is to fast-track innovation through each step of the value chain. With a worldwide presence in 30 countries, our global footprint includes excellence centers around the world, including Romania since 1994. Responsibilities: The Network Reliability Engineering function is part of the Production Services organization and is responsible for the shared IT infrastructure used to run customer-facing services. The role is part of an international infrastructure team, working closely with colleagues across several European countries. The team is responsible for operating and improving reliable infrastructure services, with a strong focus on availability, resilience, automation and operational efficiency. The environment is complex and multidisciplinary, spanning Linux-based systems, shared infrastructure services, multi-data-center and hybrid environments, network services and security components. The role combines infrastructure operations with scripting, automation, infrastructure-as-code and observability practices to improve reliability and reduce repetitive manual work. Responsibilities: Operate and improve production infrastructure, troubleshoot incidents, perform RCA, and deploy controlled changes using tools such as Ansible, Terraform and scripting. Automate recurring infrastructure operations using scripting, infrastructure-as-code, and configuration-management tooling for provisioning, configuration, and controlled changes. Support and improve monitoring, alerting, dashboards and operational visibility for infrastructure services. Work with infrastructure, network, security, application and engineering teams to improve resilience, capacity, scalability and operational efficiency. Participate in out-of-office-hours changes and standby rotation when required, and contribute to post-incident reviews and continuous-improvement activities. Qualifications: Strong experience operating and troubleshooting production infrastructure in enterprise environments. Good Linux, Python/Bash scripting, Git and CI/CD knowledge. Experience with automation/IaC tools such as Ansible, Terraform or Puppet. Experience with monitoring and observability tools such as Zabbix, Prometheus or Grafana. Good networking knowledge: TCP/IP, DNS, routing, firewalls, load balancing and connectivity troubleshooting. Exposure to enterprise networking and infrastructure technologies such as Cisco Nexus/NX-OS, F5, Check Point and hybrid/cloud environments. Strong incident troubleshooting and Root Cause Analysis skills. Good English communication and cross-team collaboration skills. What do I need before I apply: Hybrid, a few days per month at the office. CIM only Benefits: Benefit Platform Holiday Voucher Private medical insurance Performance bonus Easter and Christmas bonus Employee referral bonus Bookster subscription Work from home options depending on project.

Наблюдалась 2026-09-15, впервые 2026-08-20, найдена на 2 площадках, источник — Indeed.

Открыть у работодателя