Hiring.Camp

HS Hybrid Cloud Senior Staff Engineer

TJX

·

Today

Location
APAC Home Office Hyderabad IN, India
Workplace
Hybrid
Type
Full-time
Department
Engineering
Seniority
Senior
Education
Bachelor
Source
Workday

Description

TJX Companies

At TJX Companies, every day brings new opportunities for growth, exploration, and achievement. You’ll be part of our vibrant team that embraces diversity, fosters collaboration, and prioritizes your development. Whether you’re working in our four global Home Offices, Distribution Centers or Retail Stores—TJ Maxx, Marshalls, Homegoods, Homesense, Sierra, Winners, and TK Maxx, you’ll find abundant opportunities to learn, thrive, and make an impact. Come join our TJX family—a Fortune 100 company and the world’s leading off-price retailer. 

Job Description:

Job Description 

Job Title: Senior Staff Engineer – Cloud Engineering and Enablement – Hosting & Operations 
Job Family / Role Category: IT Infrastructure & Networking / Cloud Engineering 
Job Level: Senior Staff Engineer 
Location: India 
Work Model: Hybrid 
Team: Cloud Engineering and Enablement – Hosting & Operations 
Reports To: IT Engineering Manager – Cloud Operations 

Role Overview 

At TJX, our technology teams support a global off-price retail business by delivering reliable, secure, and scalable technology platforms. The Cloud Engineering and Enablement team operates and improves the enterprise cloud foundation that enables teams to build, deploy, and support modern technology services. 

We are seeking a Senior Staff Engineer – Cloud Operations to provide technical leadership across our enterprise cloud operating model, with a primary focus on Microsoft Azure and additional support for Google Cloud Platform, Oracle Cloud Infrastructure, and related cloud services. 

This role is responsible for improving platform stability, operational readiness, incident response, observability, automation, AI-enabled operations, and continuous improvement across the multi-cloud environment. The role requires hands-on technical depth, strong operational judgment, and the ability to influence engineering, infrastructure, security, product, and vendor teams. 

What You’ll Do 

The Senior Staff Engineer – Cloud Operations will lead complex cloud operations initiatives, improve reliability, reduce operational risk, and help mature how cloud services are supported at scale. 

Key responsibilities include: 

What You’ll Need 

  • Lead medium-to-very-high complexity cloud operations initiatives across Azure-first platforms, shared services, multi-cloud integrations, and production support capabilities. 

  • Own operational readiness for cloud services, including monitoring, alerting, incident response, runbooks, recovery patterns, release readiness, and support models. 

  • Improve reliability by analyzing incidents, telemetry, service health, capacity signals, and recurring operational issues across Azure, GCP, OCI, and related services. 

  • Use AI-enabled operations tooling, including SRE agents where appropriate, to improve incident triage, correlation, response, remediation guidance, and resolution time. 

  • Lead complex incident response and troubleshooting across cloud infrastructure, AKS, networking, identity, security tooling, application platforms, vendor-managed services, and multi-cloud dependencies. 

  • Improve incident, problem, and change management practices, including post-incident reviews, root cause analysis, corrective actions, and knowledge management. 

  • Design and implement automation to reduce toil, validate platform health, accelerate recovery, and improve repeatable operational execution across cloud environments. 

  • Define and influence cloud operations standards, reference patterns, and roadmap priorities that improve reliability, reduce incident recurrence, improve mean time to resolution, and mature operational practices across Azure-first multi-cloud platforms. 

  • Partner with engineering teams to define production readiness, operational acceptance criteria, support handoffs, and environment validation practices. 

  • Champion DevSecOps, site reliability engineering, infrastructure as code, policy-as-code, AI-assisted operations, and operational excellence practices. 

  • Maintain operational standards, runbooks, playbooks, diagrams, reusable patterns, and service support documentation. 

  • Mentor engineers, lead technical discussions, and contribute to onboarding, interviews, and knowledge-sharing activities. 

We are looking for a hands-on technical leader with strong cloud operations experience, sound engineering practices, and a customer-focused mindset. 

The ideal candidate brings deep Azure operations experience, understands multi-cloud operating models across GCP and OCI, and is comfortable leading complex troubleshooting, creating durable fixes, mentoring engineers, and raising operational maturity across the cloud platform. 

Required qualifications and experience include: 

  • 8+ years of engineering experience in cloud, infrastructure, platform engineering, site reliability engineering, or a related technical domain. 

  • Hands-on experience operating enterprise-scale Microsoft Azure environments, including compute, networking, identity, storage, monitoring, security, and platform services. 

  • Baseline familiarity with Google Cloud Platform and Oracle Cloud Infrastructure concepts, with enough operational exposure to collaborate in a multi-cloud environment and apply consistent reliability, support, and operational practices. 

  • Experience with incident, problem, and change management, including production readiness, operational support models, post-incident reviews, and corrective actions. 

  • Experience operating and troubleshooting Azure Kubernetes Service, including cluster health, node pools, ingress, networking, identity integration, observability, and platform upgrades. 

  • Strong knowledge of Azure networking, including virtual networks, network security groups, route tables, Azure Firewall, private endpoints, DNS, load balancing, and hybrid connectivity troubleshooting. 

  • Experience with monitoring, alerting, logging, dashboards, service health checks, operational reporting, and AI-enabled operations tools used to support incident triage and resolution. 

  • Hands-on automation experience using PowerShell, Python, Bash, Terraform, Bicep, ARM, GitHub Actions, Azure DevOps pipelines, or similar tools. 

  • Strong understanding of DevSecOps, CI/CD, infrastructure as code, policy-as-code, release controls, security, compliance, RBAC, least privilege, secrets management, and privileged access practices. 

  • Ability to lead complex troubleshooting, translate operational data into engineering fixes, and drive measurable reliability improvements. 

  • Experience creating technical documentation, runbooks, support models, reusable patterns, and operational standards. 

  • Demonstrated ability to mentor engineers, influence technical direction, communicate clearly with stakeholders, and lead complex work independently in an Agile environment. 

Preferred Qualifications 

Great to have: 

  • Experience applying site reliability engineering practices, including service level indicators, service level objectives, error budgets, toil reduction, reliability reviews, and operational maturity assessments. 

  • Experience supporting Azure Landing Zones, enterprise-scale cloud operating models, subscription management, policy enforcement, and shared platform services. 

  • Hands-on experience supporting or improving GCP and/or OCI operations, including platform monitoring, identity and access management, networking, security controls, landing zone patterns, automation, and operational support models. 

  • Experience with Kubernetes operations and security, including network policies, workload identity, pod security, image scanning, certificate rotation, and cluster lifecycle management. 

  • Experience using or integrating AI-enabled SRE agents, AIOps platforms, or intelligent automation to improve incident detection, correlation, remediation recommendations, and mean time to resolution. 

  • Experience improving production readiness through release gates, operational acceptance criteria, environment validation, support handoffs, and service transition practices. 

  • Understanding of cloud resiliency practices, including backup and restore, disaster recovery, high availability patterns, capacity management, and dependency mapping. 

  • Exposure to cost management, tagging, capacity planning, utilization optimization, quota management, and operational reporting. 

  • Experience supporting global platforms, regulated environments, or business-critical technology services. 

Key Competencies 

  • Technical leadership and engineering judgment 

  • Cloud operations and reliability mindset 

  • Ownership, urgency, and operational discipline 

  • Incident leadership and problem-solving 

  • Automation and continuous improvement 

  • DevSecOps and secure engineering practices 

  • Stakeholder communication and influence 

  • Mentoring and knowledge sharing 

  • Collaboration across global and cross-functional teams 

  • Agile ways of working 

Minimum Education 

Bachelor's degree in computer science, Information Technology, Engineering, or a related discipline, or equivalent practical experience. 

Minimum Experience 

  • 8+ years of relevant engineering experience in cloud, infrastructure, platform engineering, site reliability engineering, technology operations, or related technical areas. 

  • Experience working in large-scale enterprise environments with production-critical technology services. 

Hiring / Intake Notes 

  • Role type: Individual Contributor 

  • Role capacity: Senior technical lead / hands-on engineering leader 

  • Suggested interview focus areas: 

  • Azure cloud operations and troubleshooting 

  • AKS and container platform operations 

  • Incident, problem, and change management 

  • Automation and infrastructure as code. 

  • Observability and production readiness 

  • DevSecOps and secure operations 

  • Multi-cloud awareness across GCP and OCI 

  • Technical leadership, mentoring, and stakeholder communication 

In addition to our open door policy and supportive work environment, we also strive to provide a competitive salary and benefits package. TJX considers all applicants for employment without regard to race, color, religion, gender, sexual orientation, national origin, age, disability, gender identity and expression, marital or military status, or based on any individual's status in any group or class protected by applicable federal, state, or local law. TJX also provides reasonable accommodations to qualified individuals with disabilities in accordance with the Americans with Disabilities Act and applicable state and local law.

Address:

Salarpuria Sattva Knowledge City, Inorbit Road

Location:

APAC Home Office Hyderabad IN

Skills

PythonAzureGCPKubernetesTerraformCI/CDOracleGitHubAgileDevOpsSREComplianceChange Management