- Location
- APAC - India - Pune · APAC-India-Pune
- Department
- Engineering
- Experience
- 3+ years
Description
Company overview:
TraceLink is the world’s largest Agentic Business Network, enabling life sciences and healthcare companies to build and manage a scalable digital workforce of governed, no-code AI agents that execute and coordinate mission-critical supply chain operations alongside human teams. Powered by the Integrate-Once™ OPUS platform, TraceLink links more than 300,000 network participants, enabling multi-enterprise processes at global scale.
Founded in 2009 with the simple mission of protecting patients, today Tracelink has 5 global offices, over 800 employees and more than 1700 customers in over 60 countries around the world. Our expanding product suite continues to protect patients and now also enhances multi-enterprise collaboration through innovative new applications such as MINT.
Tracelink is recognized as an industry leader by Gartner and IDC, and for having a great company culture by Comparably.
Role summary
We're looking for an experienced, driven and passionate engineering team member with backgrounds in programming, distributed systems and Kubernetes to help our SRE team improve its Service Mesh and Kubernetes architecture. The SRE group is building and expanding on the critical need to maintain visibility and provide scalability of the TraceLink global platform. Within SRE, you'll have plenty of opportunities to share your strengths, guide us on how to build a scalable platform and collaborate closely with various engineering stakeholders.
You will work in a global team, in an inclusive environment with AWS cloud-based deployments and focus on ensuring services are running smoothly, continuously assess opportunities to reduce toil and help improve service availability and reliability, optimise AWS resources usage across multiple environments to deliver cost effective services to the engineering organisation.
Responsibilities
As a member of the SRE core team, ensure high availability, performance and reliability expected by our customers and delivery to defined OKRs
Design, build, document, test new tools and technologies as part of an Agile development team. Maintain and improve these to eliminate bugs, increase performance/efficiency, or extend capabilities
Play an active role in the development process, deliver on commitments, communicate issues, work with others both in the team and in other teams
Gather and analyse metrics from systems to assist in performance tuning and fault finding.
Motivated, self-organized and have good time & work management skills.
Implementing and monitoring systems to proactively detect and address issues
Qualifications
3+ years of experience with increasing responsibility as an SRE/Cloud Engineer
Strong understanding of cloud deployment and management practices in Multi Cloud Environments (AWS and Azure)
Hands-on experience with Terraform Helm and Kubernetes
Hands-on experience with tools and techniques to diagnose and uncover container and overall system performance
Proactive approach to identifying problems, performance bottlenecks, and areas for improvement
Skilled in AWS services both from technology and cost optimisation perspectives
Multi Cloud skills with AWS and Azure are preferred.
Experience working with mature development practices and tools for source control, security, and deployment
Hands on experience with Python/Shell
Excellent communication skills, written and verbal
Strong analytical and problem-solving skills
Nice to have: Experience in Observability, Istio and Agentic AI capabilities.
Please see the Tracelink Privacy Policy for more information on how Tracelink processes your personal information during the recruitment process and, if applicable based on your location, how you can exercise your privacy rights. If you have questions about this privacy notice or need to contact us in connection with your personal data, including any requests to exercise your legal rights referred to at the end of this notice, please contact [email protected].