Senior Network Reliability Engineer

Job Locations
RO-B-Bucharest
Job area
IT & Digital
Employment type
Permanent
Workplace
Hybrid
Experience level
Management / Mid-Senior level

Overview

Expleo is a global engineering, technology, and consulting service provider that partners with leading organizations to guide them through their business transformation, helping them achieve operational excellence and future-proof their businesses.  
Expleo benefits from more than 50 years of experience developing complex products in automotive and aerospace, optimizing manufacturing processes, and ensuring the quality of information systems. Leveraging its deep sector knowledge and wide-ranging expertise in fields including AI engineering, digitalization, automation, cybersecurity and data science, the group’s mission is to fast-track innovation through each step of the value chain. 
With a worldwide presence in 30 countries, our global footprint includes excellence centers around the world, including Romania since 1994.

Responsibilities

The Network Reliability Engineering function is part of the Production Services organization and is responsible for the shared IT infrastructure used to run customer-facing services.

The role is part of an international infrastructure team, working closely with colleagues across several European countries. The team is responsible for operating and improving reliable infrastructure services, with a strong focus on availability, resilience, automation and operational efficiency.

The environment is complex and multidisciplinary, spanning Linux-based systems, shared infrastructure services, multi-data-center and hybrid environments, network services and security components. The role combines infrastructure operations with scripting, automation, infrastructure-as-code and observability practices to improve reliability and reduce repetitive manual work.

 

Responsibilities:

  • Operate, administer and continuously improve shared production infrastructure, with a focus on availability, resilience, performance and reliability.
  • Automate recurring operational and infrastructure-management tasks using scripting and automation tooling.
  • Use and improve infrastructure-as-code and configuration-management workflows for provisioning, configuration and controlled changes.
  • Troubleshoot infrastructure and connectivity issues across Linux systems, network services, security components and dependent platforms, and support complex production incidents.
  • Perform Root Cause Analysis for major incidents and translate findings into preventive improvements, automation and more resilient operational practices.
  • Prepare, test and deploy infrastructure changes using controlled, repeatable workflows, including implementation and rollback plans.
  • Support and improve monitoring, alerting, dashboards and operational visibility for infrastructure services.
  • Work with infrastructure, network, security, application and engineering teams to improve resilience, capacity, scalability and operational efficiency.
  • Participate in out-of-office-hours changes and standby rotation when required, and contribute to post-incident reviews and continuous-improvement activities.

 

Qualifications

  • Solid hands-on experience operating and troubleshooting production IT infrastructure in complex enterprise environments.
  • Good working knowledge of Linux and confidence working in command-line operational environments.
  • Practical scripting experience with Python and/or Bash for operational tasks, troubleshooting and automation.
  • Experience with at least one infrastructure automation, infrastructure-as-code or configuration-management tool, such as Ansible, Terraform or Puppet.
  • Familiarity with source control, preferably Git, and exposure to CI/CD or automated deployment workflows.
  • Experience with monitoring, alerting and observability tools or platforms, and the ability to use operational data to investigate incidents and improve reliability.
  • Good understanding of core infrastructure and networking concepts, including TCP/IP, DNS, DHCP, routing, load balancing, firewalls and connectivity troubleshooting.
  • Experience working in multi-site, multi-data-center, virtualized or hybrid infrastructure environments.
  • Familiarity with enterprise networking technologies such as Cisco IOS/NX-OS, Cisco Nexus or spine-leaf architectures.
  • Familiarity with or exposure to firewall, load-balancing or proxy technologies such as Check Point, F5 BIG-IP, HAProxy or similar solutions.
  • Exposure to cloud platforms such as AWS, Azure or GCP
  • Strong troubleshooting, incident management and Root Cause Analysis skills, with a focus on preventing recurrence and reducing repetitive operational work.
  • Good communication skills in English and the ability to collaborate effectively across infrastructure, security, application and engineering teams.

What do I need before I apply

  • Hybrid, a few days per month at the office.
  • CIM only

Benefits

  • Benefit Platform 
  • Holiday Voucher 
  • Private medical insurance  
  • Performance bonus 
  • Easter and Christmas bonus 
  • Employee referral bonus 
  • Bookster subscription  
  • Work from home options depending on project.

Options

Sorry the Share function is not working properly at this moment. Please refresh the page and try again later.
Share to social media

Can't find the job of your choice?
Upload your C.V. / Resume here for our recruiters to view.