Hiring.Camp

Lead DevOps Engineer

Sphera colleague

·

Today

Location
CA Remote, Canada
Workplace
Remote
Type
Full-time
Department
Engineering
Seniority
Lead
Source
Workday

Description

Sphera is a leading global provider of enterprise software and services that enables companies to manage and optimize their environmental, health, safety and sustainability. Our mission is to create a safer, more sustainable and productive world.

Sphera is a portfolio company of Blackstone, a U.S.-based alternative asset investment company that focuses on private equity, technology and innovation, and more. Blackstone businesses succeed through strong partnerships, a personalized approach and a commitment to exceptional performance with uncompromising integrity. Sphera and Blackstone are leaders in the Environmental, Social and Governance (ESG) space.

We are guided by our core values of Customer Centricity, Accountability, Bias to Action, Innovation, and Collaboration. These values help us recruit the right talent to join our rapidly expanding team of around the globe. It is important to us that each and every Spherion is not only eager to challenge themselves and knows how to get work done but is an awesome addition to our company culture.

Lead DevOps Engineer Job Description

As the Lead DevOps Engineer, you will be a key technical leader responsible for the reliability, scalability, and automation that power all of the products our team supports across the SpheraCloud SaaS platform, with a particular focus on our core platform and our AI initiatives. From designing and maintaining the automation behind our cloud-based deployments, to enabling auto-scaling and safe, one-click rollbacks, to building and securing the CI/CD pipelines that make our engineering teams faster, you will make a meaningful, day-to-day impact on the business. You will also lead the buildout and operation of the infrastructure that supports our AI and data workloads, and partner closely with our Engineering and Security teams on the network and firewall changes that keep our environments secure and compliant.

Key Responsibilities

  • Provide technical leadership for DevOps and platform engineering across all of the products the team supports, prioritizing our core platform and AI initiatives.
  • Own and continuously improve our CI/CD pipelines, including integrated code-quality and security scanning with SonarCloud and Black Duck.
  • Configure and enforce permissions across repositories and pipelines so that scans run reliably and that access, approval, and enforcement controls are correctly applied.
  • Design, build, and maintain the automation that powers cloud-based deployments, including auto-scaling, progressive delivery (blue/green and canary), and reliable, one-click rollback.
  • Implement and manage infrastructure primarily as code with Terraform, using the Azure CLI for emergency and ad-hoc operations.
  • Build, deploy, and manage workloads using Docker on Azure Container Apps, as well as Azure Web Apps and Function Apps.
  • Own and operate key platform resources, including Azure API Management, Elasticsearch, and Databricks, ensuring they are reliable, secure, performant, and cost-efficient.
  • Build and operate the infrastructure and pipelines that support our AI, machine learning, and data workloads, covering model and data-pipeline deployment, serving, and monitoring (MLOps).
  • Partner daily with our Engineering and Security teams on network and firewall rule change requests, ensuring changes are well-documented, reviewed, and compliant.
  • Develop and evolve self-service platform capabilities and internal developer tooling (“golden paths”) that improve developer velocity, consistency, and safety across teams.
  • Perform and automate system administration, including provisioning, configuration, maintenance, and disaster recovery.
  • Work with our NOC team to build observability into the platform (metrics, logs, traces, and alerting) to ensure performance and reliability, and to monitor the behavior, quality, and cost of AI and data workloads in production.
  • Embed security and compliance into pipelines and infrastructure (DevSecOps), including automation, auditing, and tooling for security, compliance, and resource usage.
  • Work side by side with engineers, guiding critical projects and using your subject-matter expertise to solve complex problems, while mentoring and growing others on the team.
  • Partner with our FinOps team and own the actions that support their cost-optimization goals across our cloud and AI resources.
  • Take ownership of operational excellence to establish and maintain business-focused KPIs and SLOs, using these metrics both as a basis for positive transformation and as a visible representation of the team’s success.

Required Skills and Experience

  • Excellent communication and collaboration skills, with the ability to communicate effectively across multiple disciplines and levels.
  • Demonstrated technical leadership or mentorship of engineers on complex, cross-team initiatives.
  • Strong analytical skills with a proven ability to drive reliability and performance improvements.
  • Proven success building and optimizing CI/CD for an Azure-based SaaS application, including integrated code-quality and security scanning (e.g., SonarCloud and Black Duck).
  • Experience configuring and enforcing repository and pipeline permissions to ensure scans run reliably and that access and approval controls are correctly applied.
  • Hands-on experience with Azure DevOps (and/or GitHub Actions).
  • Strong infrastructure-as-code experience with Terraform.
  • Proficiency with the Azure CLI for operational and emergency tasks.
  • Strong scripting skills in at least one of PowerShell, Bash, or Python.
  • Experience building and running workloads with Docker on Azure Container Apps, Azure Web Apps, and Function Apps.
  • Experience operating platform services such as Azure API Management, Databricks, and Elasticsearch.
  • Experience managing network and firewall configurations and change-request processes.
  • Strong knowledge of both Microsoft/Windows and Linux server environments.
  • Working knowledge of version control and Git-based workflows.
  • Experience working within an agile software development lifecycle.
  • Strong understanding of security principles and secure-by-design practices.
  • Working knowledge of Microsoft SQL Server and of IIS configuration and management.
  • Working knowledge of local and wide-area networking and associated technologies (switches, routers, firewalls, VPNs, VNets, application gateways).
  • Bachelor’s degree in Computer Science or a similar area of study, or equivalent practical experience.

Preferred Experience

  • Experience supporting AI/ML or generative-AI workloads in production, including MLOps pipelines, model deployment and serving, and monitoring model performance and cost.
  • Familiarity with Azure AI services such as Azure Machine Learning and Azure OpenAI, and/or common ML frameworks.
  • Advanced experience tuning and scaling Databricks (Spark) and Elasticsearch clusters.
  • Experience provisioning and scaling GPU-based compute for AI workloads.
  • Experience with vector databases and retrieval-augmented generation (RAG) architectures.
  • Experience building internal developer platforms / platform engineering.
  • Deeper experience with container orchestration concepts, including Kubernetes (which underpins Azure Container Apps).
  • Observability stacks such as Grafana, Prometheus, OpenTelemetry, New Relic, Datadog, or Rapid7.
  • Additional PaaS services (App Service Environment, Azure SQL, etc.).
  • Additional automation and delivery tooling: Jenkins, Octopus, Ansible, or similar.
  • Data-platform experience: data pipelines, data lakes, and query engines such as Trino/Presto; Hadoop/HDFS a plus.
  • Relevant Azure certifications (e.g., Azure DevOps Engineer Expert or Azure Solutions Architect Expert).


Sphera is proud to be an Equal Opportunity Employer. We celebrate diversity and are committed to creating an inclusive environment for all colleagues.

This job description is intended to convey information essential to understanding the scope of the job and the general nature and level of work performed by job holders within this job. This job description is not intended to be an exhaustive list of qualifications, skills, efforts, duties, responsibilities or working conditions associated with the position.

Skills

PythonAzureDockerKubernetesTerraformAnsibleJenkinsCI/CDLinuxSQLElasticsearchSQL ServerMachine LearningSparkHadoopDatabricksGitGitHubDevOpsCompliance

Similar Jobs

30

Devops Lead Engineer

Career Site · Mumbai - India

Yesterday

Lead Site Reliability Engineer

Boeing · USA - Berkeley, MO, United States of America · Onsite

4 days ago

Lead Site Reliability Engineer

Boeing · USA - Berkeley, MO, United States of America · Onsite

4 days ago

Lead DevOps Engineer

Particle41Llc · USA - Remote · Remote

5 days ago

Lead Site Reliability Engineer

Disney · IND - Bangalore - Ecoworld 6B - 7th Floor, India · Onsite

5 days ago

Lead Site Reliability Engineer

Disney · IND - Bangalore - Ecoworld 6B - 7th Floor, India · Onsite

5 days ago

Lead Site Reliability Engineer

Disney · IND - Bangalore - Ecoworld 6B - 7th Floor, India · Onsite

5 days ago

Platform Integration Lead Engineer

Cat · Bangalore, Karnataka, India

5 days ago

Lead Site Reliability Engineer

Fis · IND BNGL FL2-3 TWR 3, India

6 days ago

Lead Mac Platform Engineer

Capitalone · Richmond, VA, United States of America +2

6 days ago

Lead Site Reliability Engineer

"Intellum, Inc." · Remote, United States · Remote

1 week ago

Site Reliability Engineer Lead

Td · TD Centre - South - 79 Wellington Street West, Toronto, Ontario, Canada · Onsite

1 week ago

Lead Data Platform Engineer

Collegeboard · Remote - Virginia, United States of America · Remote

1 week ago

Lead Site Reliability Engineer

Cat · Chicago, Illinois, United States of America +1

1 week ago

Lead DevOps Engineer

Eaton · Peachtree City, GA,US, US +1 · Hybrid

1 week ago

Atlassian Lead Platform Engineer

Koniag Government Services · Washington, DC, USA

1 week ago

Lead Site Reliability Engineer

Rb · San Francisco, CA, United States of America +1

1 week ago

Lead Site Reliability Engineer

Mastercard · Dublin, Ireland

1 week ago

Lead DevOps Engineer

Rbc · RBC WATERPARK PLACE, 88 QUEENS QUAY W:TORONTO, Canada

1 week ago

Lead DevOps Engineer

Rbc · RBC WATERPARK PLACE, 88 QUEENS QUAY W:TORONTO, Canada

1 week ago

Site Reliability Engineer Lead

Our latest jobs · Auckland - Te Kupenga, New Zealand

1 week ago

Lead Data Platform Engineer

Job Listings · Bangalore, India

1 week ago

Lead Test Platform Engineer

Loft Orbital · Abu Dhabi · Onsite

2 weeks ago

Lead AI Platform Engineer

Mastercard · Dublin, Ireland

2 weeks ago

Lead DevOps Engineer

Veriday Inc. · Toronto, Ontario, Canada

2 weeks ago

Platform Engineer (Senior/Lead)

Scott Logic · Leeds · Hybrid

2 weeks ago

Lead Data Platform Engineer

Urban Connect · Bucharest · Hybrid

2 weeks ago

Lead Engineer – AI Platform

Waters · Bangalore, IN

2 weeks ago

Lead Engineer – AI Platform

International Waters · Bangalore, IN

2 weeks ago

Lead Site Reliability Engineer

Draftkings · Boston, MA, United States of America

2 weeks ago
Remote Lead DevOps Engineer at Sphera colleague | Hiring.Camp