- Location
- Sysco LABS - Sri Lanka
- Type
- Full-time
- Department
- Engineering
- Seniority
- Senior
- Source
- Workday
Description
JOB DESCRIPTION
Senior Technical Operations Engineer
About Sysco LABS:
Sysco LABS is the Global In-House Center of Sysco Corporation (NYSE: SYY), the world’s largest foodservice company. Sysco ranks 55th in the Fortune 500 list and is the global leader in the trillion-dollar foodservice industry.
Sysco operates 333 distribution centers across 10 countries, with 75,000 colleagues serving approximately 670,000 customer locations, including restaurants, healthcare and educational facilities, lodging establishments, entertainment venues and more. For fiscal year 2026, which ended June 27, 2026, the company generated sales of more than $84 billion.
Sysco LABS Sri Lanka delivers the technology that powers Sysco’s end-to-end operations.
Sysco LABS’ enterprise technology is present in the end-to-end foodservice journey, enabling the sourcing of food products, merchandising, storage and warehouse operations, order placement and pricing algorithms, the delivery of food and supplies to Sysco’s global network and the in-restaurant dining experience of the end-customer.
About the Team & Scope
Technical Operations safeguards the business continuity, availability, stability, and digital experience of Sysco’s critical technology platforms.
Through 24/7 operational coverage, Technical Operations provides proactive production monitoring, rapid incident detection and service restoration, application and production support, problem management, release and change support, operational readiness, hypercare, stakeholder communication, service governance, and continuous operational improvement to minimize business disruption and protect Sysco’s customer, business, and brand experience.
Technical Operations works closely with Engineering, SRE, DevOps, Platform Engineering, Product, Business, Security, Infrastructure, vendors, and other technology functions. Technical Operations primarily owns the operational management and support of production applications, while partnering with engineering functions on deeper reliability engineering, architectural remediation, platform engineering, and systemic reliability improvements.
Role Focus - Serve as the senior technical escalation point and lead complex production troubleshooting.
Responsibilities:
- Provide senior-level support for business-critical applications, managing incidents, defects, and alerts within agreed SLAs to maintain availability and business continuity.
- Act as the technical lead during P1/P2 and Major Incidents, leading bridge calls, triage, troubleshooting, escalation, service restoration, recovery validation, stakeholder communication, and closure while clearly communicating customer and business impact.
- Lead complex investigations across front-end and back-end applications, APIs, databases, integrations, and upstream and downstream services using logs, metrics, traces, and errors.
- Partner with Engineering, SRE, DevOps, Infrastructure, Security, Product, Business, vendors, and resolver teams to mitigate production risks and resolve escalated issues.
- Monitor applications, infrastructure, databases, ETL/data pipelines, APIs, integrations, and service dependencies using Datadog dashboards and alerts.
- Lead Problem Management and blameless postmortems, driving corrective, preventive, and long-term remediation actions.
- Analyze availability, SLA, MTTA, MTTR, backlog, performance, defects, alerts, and recurring defect trends to identify risks and improve services.
- Improve monitoring, observability, dashboards, alerts, and logging by addressing coverage gaps, noisy alerts, anomalies, and degradation trends to enable earlier detection and faster troubleshooting.
- Support production releases, migrations, and changes through validation, monitoring, hypercare, and coordination of release-related defect resolution.
- Implement automation, self-service, and AI-assisted solutions to automate repetitive activities, reduce operational toil, accelerate troubleshooting, and improve MTTR and operational efficiency.
- Drive operational and service-management improvements by improving runbooks, SOPs, workflows, knowledge management, incident practices, troubleshooting standards, and automation.
- Mentor Associate and Technical Operations Engineers in troubleshooting, incident management, RCA, operational practices, and stakeholder communication.
- Support cross-functional activities, Security, Legal, Audit, and Compliance activities as a senior production and operational point of contact.
Requirements:
- Bachelor’s degree in Software Engineering.
- 3+ years of relevant experience in Technical Operations, Production Support, or a similar capacity.
- Strong working knowledge of ITIL/ITSM and production-support practices, including Incident, Problem, Change, Release, Service Request, Major Incident Management, postmortems, ITSM KPIs.
- Familiarity with ServiceNow, Jira, Datadog, or equivalent tools.
- Demonstrated experience leading technical investigations during P1/P2 Major Incidents, including bridge coordination, escalation, communication, and resolution.
- Advanced hands-on experience with monitoring, logging, observability, and APM tools, including dashboards, monitors, log analysis, anomaly investigation, and alert tuning.
- Strong experience troubleshooting production defects, including front-end and back-end application issues and upstream and downstream service dependencies.
- Strong understanding of modern web and mobile applications, client-server architectures, microservices, and REST APIs.
- Experience with Swagger/OpenAPI or equivalent API-testing technologies.
- Working knowledge of cloud and container technologies, such as AWS and Docker; exposure to Kubernetes and cloud-native application environments is advantageous.
- AWS Certified Solutions Architect certification is required, and ITIL 4 certification is preferred.
- Strong scripting and automation capabilities using Python, Bash/Shell, PowerShell for workflow automation, self-service capabilities, and toil reduction.
- Familiarity with modern web technologies and application frameworks, such as Node.js, React.js, JavaScript for front-end and back-end troubleshooting.
- Working knowledge of Release Management, Change Activities and hypercare practices in general.
- Strong analytical, troubleshooting, problem-solving, and decision-making skills, with the ability to lead complex production investigations and make sound decisions under pressure.
- Strong verbal and written communication skills, with the ability to explain technical issues clearly to technical, non-technical, senior, and global stakeholders.
- Strong stakeholder management and collaboration skills, with the ability to partner with Engineering/L3, SRE, DevOps, Infrastructure, Security, Product, Business, vendors, and other resolver teams to drive issues to resolution.
- Demonstrated ability to mentor junior engineers, share technical knowledge, and improve team-level troubleshooting and operational practices.
- Experience working in a 24/7 production-support environment, including shift-based operations, incident response, and on-call support.
Benefits:
- Performance-based annual bonus
- Performance rewards and recognition
- Agile Benefits - special allowances for Health, Wellness & Academic purposes
- Paid birthday leave
- Team engagement allowance
- Comprehensive Health & Life Insurance Cover - extendable to parents and in-laws
- Hybrid work arrangement
Sysco LABS is an Equal Opportunity Employer.