- Location
- 34TH ST BONIFACIO GLOBAL CITY TAGUIG, Philippines
- Workplace
- Hybrid
- Type
- Full-time
- Department
- Engineering
- Seniority
- Lead
- Experience
- 8+ years
- Source
- Workday
Description
Job Title: Senior Network SRE Automation Engineer
Job Summary
We are seeking a highly skilled and experienced Senior Network SRE Automation Engineer to join our team. In this role, you will work closely with cross-functional engineering teams across various domains to bridge the gap between traditional network operations and modern DevOps practices. You will play a critical role in ensuring the safety, soundness, and reliability of Day 2 Operations workflows, with a focus on automation, recovery, and operational excellence. Additionally, you will provide occasional network escalation support during major production incidents, conduct complex deep dives, and lead root cause analysis (RCA) efforts on the SRE-OPS front.
Key Responsibilities
- Collaborate with cross-functional engineering teams to design and implement automated workflows for Day 2 Operations, ensuring reliability, scalability, and built-in recovery mechanisms.
- Provide escalation support for major production incidents, performing in-depth troubleshooting and resolution of complex network issues.
- Conduct detailed root cause analysis (RCA) and post-incident reviews to drive continuous improvement in operational processes.
- Develop and maintain CI/CD pipelines for network automation, ensuring seamless integration and deployment of changes.
- Leverage expertise in Python to build and maintain automation scripts and tools for network operations.
- Integrate and manage network automation platforms such as Cisco NSO and Nautobot to streamline network management and provisioning.
- Ensure operational workflows adhere to enterprise standards for safety, soundness, and reliability.
- Work with GitHub for version control and collaboration on automation projects.
- Partner with financial enterprise teams to ensure compliance with industry standards and best practices.
- Drive continuous improvement in operational processes and workflows.
Required Qualifications
- Education: Bachelor’s degree in computer science, Information Technology, Network Engineering, or a related field (or equivalent experience).
- Experience:
- 8+ years of experience in network engineering, site reliability engineering, or network automation.
- Strong background in network observability, monitoring tools, and troubleshooting complex network issues.
- Proven experience working in financial enterprise environments with a focus on regulatory compliance and operational excellence.
- Technical Skills:
- Strong proficiency in Python, GraphQL, or other relevant programming languages, with the ability to develop scripts and tools for automation and observability.
- Experience with CI/CD pipelines and tools such as GitHub, Jenkins, or similar.
- Expertise in network automation platforms like Cisco NSO and Nautobot.
- Strong understanding of network protocols (TCP/IP, BGP, OSPF, DNS, etc.).
- Knowledge of Infrastructure as Code (IaC) tools like Terraform or Ansible.
- Familiarity with observability tools such as Prometheus, Grafana, Splunk, ELK Stack, or Datadog.
- Understanding of cloud networking (AWS, Azure, GCP) and hybrid environments.
- Basic familiarity with container orchestration tools (e.g., Kubernetes) and service meshes.
- Soft Skills:
- Strong problem-solving and analytical skills, with the ability to perform deep dives and root cause analysis (RCA).
- Excellent communication and collaboration abilities to work with cross-functional engineering teams.
- Ability to work in a fast-paced, dynamic environment and handle production escalations.
Preferred Qualifications
- Strong proficiency in Python, GraphQL, or other relevant programming languages, with the ability to develop scripts and tools for automation and observability.
- Extensive experience with observability tools, including setting up monitoring, alerting, and visualization workflows.
- Experience with AI/ML-based network monitoring tools.
- Certifications such as CCNP, CCIE, AWS Advanced Networking, or Kubernetes certifications (CKA/CKAD).
- Familiarity with chaos engineering practices to test network resilience.
- Experience with container orchestration tools (e.g., Kubernetes) and service meshes.
This role offers the opportunity to work in a dynamic, global environment, driving innovation and operational excellence in critical network domains. You will collaborate with cross-functional teams, contribute to building reliable and automated workflows, and ensure the safety, soundness, and reliability of Day 2 Operations. If you are passionate about network reliability, automation, and observability, we encourage you to apply!
------------------------------------------------------
Job Family Group:
Technology------------------------------------------------------
Job Family:
Systems & Engineering------------------------------------------------------
Time Type:
Full time------------------------------------------------------
Most Relevant Skills
Please see the requirements listed above.------------------------------------------------------
Other Relevant Skills
For complementary skills, please see above and/or contact the recruiter.------------------------------------------------------
Citi is an equal opportunity employer, and qualified candidates will receive consideration without regard to their race, color, religion, sex, sexual orientation, gender identity, national origin, disability, status as a protected veteran, or any other characteristic protected by law.
If you are a person with a disability and need a reasonable accommodation to use our search tools and/or apply for a career opportunity review Accessibility at Citi.
View Citi’s EEO Policy Statement and the Know Your Rights poster.