- Location
- Bucharest, B, RO
- Type
- Full-time
- Department
- IT
- Closing date
- Today
- Source
- iCIMS
Description
Overview
Expleo is a global engineering, technology, and consulting service provider that partners with leading organizations to guide them through their business transformation, helping them achieve operational excellence and future-proof their businesses. Expleo benefits from more than 50 years of experience developing complex products in automotive and aerospace, optimizing manufacturing processes, and ensuring the quality of information systems. Leveraging its deep sector knowledge and wide-ranging expertise in fields including AI engineering, digitalization, automation, cybersecurity and data science, the group’s mission is to fast-track innovation through each step of the value chain. With a worldwide presence in 30 countries, our global footprint includes excellence centers around the world, including Romania since 1994.Responsibilities
We are looking for a Site Reliability Engineer to support an international client in the payments and financial technology sector. You will help design, deploy and operate services at scale, ensuring high availability, performance and security. Working closely with development teams, you will build and maintain the infrastructure and tooling that underpin platform reliability and operational excellence throughout the software development lifecycle.
Key responsibilities:
- Develop and improve shared services and tooling to increase delivery speed, availability, scalability and operational efficiency.
- Integrate community and open-source products into the technical ecosystem.
- Build tooling using scripting and programming languages, and collaborate with the team through knowledge sharing.
- Support business teams in migrating existing solutions to new infrastructure using Infrastructure as Code, automation and CI/CD.
- Work across teams to build fast, reliable and resilient production systems.
- Develop, configure and enhance the automation stack and CI/CD pipelines.
- Create and maintain technical and user documentation that colleagues can use, improve and share.
- Perform daily operational activities, including change management, application monitoring and application deployment.
- Embed security into service design and operations from the outset, in line with payment industry requirements.
Qualifications
- Academic or bachelor-level education, or equivalent practical experience.
- At least 3–5 years of Linux experience in high-availability environments.
- An autonomous, proactive approach and the ability to learn quickly and navigate complex automation stacks.
- A strong interest in open source, Infrastructure as Code, automation and self-healing systems that minimize manual support.
- A security-focused mindset and sound judgement when balancing stability with agility, operations with software engineering, and proactive with reactive work.
- Strong analytical, problem-solving, communication and collaboration skills, with the ability to engage technical and non-technical stakeholders at all levels.
Essential skills
- Excellent knowledge of Red Hat / CentOS Linux.
- Hands-on experience with Git, GitLab or Bitbucket, Git workflows and CI/CD practices.
- Hands-on experience with Puppet for configuration management and automation.
- Hands-on experience with Terraform to define and manage infrastructure as code.
- Comfortable with scripting and automation using Bash, Python, Ruby and related tooling.
- Proficiency in at least one programming language, such as Python, Go, Java, Ruby or Perl.
- Understanding of networking concepts and protocols, including IP, DHCP, DNS, BGP and load balancing.
- Experience with multi-datacenter environments spanning different countries.
- Experience with hybrid infrastructure combining on-premises systems and cloud platforms.
- Experience with security standards such as PCI DSS.
- Excellent communication skills in English.
Desired skills
- Familiarity with service discovery in a micro-frontend architecture.
- Previous experience as an SRE or in a similar operational role.
- Cloud technology certifications covering AWS, GCP or Azure.
- Experience with monitoring and alerting tools such as Prometheus, Grafana, Datadog, PagerDuty or New Relic.
- Understanding of cloud compliance requirements and security best practices.
What do I need before I apply
- Hybrid, a few days per month at the office.
- CIM only
Benefits
- Benefit Platform
- Holiday Voucher
- Private medical insurance
- Performance bonus
- Easter and Christmas bonus
- Employee referral bonus
- Bookster subscription
- Work from home options depending on project