Hiring.Camp

Senior Dev Ops Engineer

Uchicago

·

Today

Salary
$150k – $170k
Location
NBC Tower, United States of America
Workplace
Hybrid
Type
Full-time
Department
Engineering
Seniority
Senior
Closing date
Today
Source
Workday

Description

Department

Globus Systems Operations


About the Department

Globus (www.globus.org) is a sustainable, non-profit unit within The University of Chicago delivering solutions to the research community worldwide. Globus develops and provides critical services that support scientific research for governmental, academic, and commercial organizations in a wide range of disciplines including life sciences, physics, and astronomy. We develop and operate commercial-quality, cloud-based software application and platform services used by 10s of thousands of researchers to manage their large–and growing–data management challenges. We have offices located at the NBC Tower in the heart of downtown Chicago and remote employees who work-from-home. Globus, together with Globus Labs, a research group within the University of Chicago, and part of the Data Science and Learning Division at Argonne National Labs, develop and deploy cutting edge technologies to solve new challenges facing the scientific community and enable break-through scientific discoveries.


Job Summary

The Globus Operations team is a 4-5 person group that designs, builds, and operates the cloud infrastructure behind a research-computing platform used by hundreds of institutions worldwide. Globus is a hybrid solution combining AWS-hosted orchestration services with installable applications, delivering identity and access management, data transfer and sharing, and task automation as both software-as-a-service (SaaS) and platform-as-a-service (PaaS).

As a senior member of this small, high-autonomy team, you will ensure software development and operational best practices are effectively integrated across our services and AWS infrastructure. The team owns an unusually broad estate — an AWS Organization of roughly thirty production accounts, a golden-image pipeline producing around thirty machine images, a self-hosted monitoring platform, a centralized security-logging pipeline, and in-house compliance and security automation — all managed as code in Python, Bash, Terraform, CloudFormation, and Ansible. You will work both collaboratively and independently on complex issues and projects, in some cases making progress under minimal guidance. You will also embed with select software engineering teams to provide operational and infrastructure guidance on best practices.

We are looking for a senior engineer equally comfortable improving a build pipeline and debugging a production incident, and who will be a leader and educator both within and across teams.

Responsibilities

  • Architecture and Design: Participate in the definition and documentation of cloud infrastructure architecture — networking, monitoring, logging, security, backup, and deployment — across production and development environments. Design new systems and tools, identify opportunities for technical improvement, and review and test solutions to ensure standards are met.
  • Software Development: Develop, test, document, and maintain high-quality software and infrastructure as code for deploying and operating Globus cloud services, in Python, Bash, Terraform, and CloudFormation. Contribute to shared internal libraries and expand automated test coverage of operational tooling.
  • Build and Release Automation: Maintain and evolve the machine-image and deployment supply chain — Packer-built AMIs produced through AWS CodeBuild, published via SSM Parameter Store, KMS-encrypted and shared cross-account.
  • Monitoring and Observability: Maintain the current monitoring and alerting systems (self-hosted Nagios/NRPE with configuration generated from live AWS inventory, CloudWatch alarms, PagerDuty routing) and help in the migration to Prometheus.
  • Security Operations and Compliance: Contribute to Globus's security plan, controls, monitoring, and reporting. Extend in-house continuous-compliance automation, operate intrusion detection and endpoint protection tooling, maintain hardened machine images.
  • Identity, Access, and Trust: Administer IAM roles, policies, and cross-account trust across the AWS Organization.
  • SRE/Operations: Deploy, operate, and monitor production Globus services for high availability. Participate in an on-call rotation, lead incident response, and author root-cause analyses. Identify recurring operational toil and eliminate it through automation and better defaults.
  • Support: Act as a technical consultant and resource for other team members, including the engineering and user-support teams, assisting with operational issues and troubleshooting.
  • Designs new systems, features, and tools. Solve complex problems and identify opportunities for technical improvement and performance optimization. Review and test code to ensure appropriate standards are met.
  • Utilize technical knowledge of existing and emerging technologies, including public cloud offerings from Amazon Web Services, Microsoft Azure, and Google Cloud.
  • Performs other related work as needed.


Minimum Qualifications

Education:

Minimum requirements include a college or university degree in related field.


Work Experience:

Minimum requirements include knowledge and skills developed through 5-7 years of work experience in a related job discipline.


Certifications:

---

Preferred Qualifications


Experience:


  • Strong experience operating AWS at scale — IAM, VPC (peering, endpoints, security groups), EC2 and Auto Scaling Groups, ECS (EC2 and Fargate), S3, RDS, Route 53, Lambda, Systems Manager, CloudWatch — including cross-account work in a multi-account AWS Organization.
  • Strong experience building Infrastructure as Code with Terraform (modules, remote state, provider upgrades, reconciling drift against live environments), the ability to maintain and incrementally retire legacy CloudFormation is a plus.
  • Experience automating infrastructure in Python — maintainable, reusable libraries rather than only scripts, using boto3 — and in Bash.
  • Building Amazon machine images and the instances they run on, plus Docker containers, and CI/CD tooling (AWS CodeBuild, GitHub Actions).
  • Linux administration and troubleshooting across Ubuntu and RPM-based distributions, including systemd, syslog/rsyslog, storage, and SSH/sudo configuration.
  • Networking, firewalls, routing, and DNS; web servers and TLS.
  • Experience with check-based monitoring, particularly self-hosted Nagios/NRPE, to operate, extend, and update.
  • Incident management and on-call practice using a platform such as PagerDuty, and writing root-cause analyses.

Preferred Competencies


  • Excellent verbal and written communication skills, including clear technical documentation and post-incident analyses.
  • Strong analytical and problem solving skills, including systematic troubleshooting of distributed systems.
  • Excellent organizational skills, constant attention to detail, and the ability to prioritize and manage workload to meet critical project milestones and deadlines.
  • Work both independently and as a team member, taking end-to-end ownership of services with minimal supervision and switching comfortably between long-horizon engineering projects and time-sensitive operational work.
  • Receptive to feedback; willing to learn and embrace continuous improvement.
  • Experience with software development fundamentals — code review, automated testing, static analysis — and experience integrating development and deployment frameworks with the monitoring, operations, and orchestration needed to run applications securely and at scale on public cloud platforms.

Working Conditions


  • Occasional evening or weekend hours.
  • Participates in an on-call support rotation.
  • Option available for remote work with occasional required attendance at in-person meetings.

Application Documents


  • Resume (required)
  • Candidates may be asked to complete a technical coding challenge
  • Finalists will be required to provide professional references. Reference checks will be conducted prior to an offer being extended. Candidates may also be asked to complete an in-person interview as part of the selection process


The University of Chicago uses AI-assisted tools to streamline and augment some recruitment processes; however, AI is not used to make hiring decisions.

When applying, the document(s) MUST be uploaded via the My Experience page, in the section titled Application Documents of the application.


Job Family

Information Technology


Role Impact

Individual Contributor


Scheduled Weekly Hours

37.5


Drug Test Required

No


Health Screen Required

No


Motor Vehicle Record Inquiry Required

No


Pay Rate Type

Salary


FLSA Status

Exempt


Pay Range

$150,000.00 - $170,000.00

The included pay rate or range represents the University’s good faith estimate of the possible compensation offer for this role at the time of posting.


Benefits Eligible

Yes

The University of Chicago offers a wide range of benefits programs and resources for eligible employees, including health, retirement, and paid time off. Information about the benefit offerings can be found in the Benefits Guidebook.


Posting Statement

The University of Chicago is an equal opportunity employer and does not discriminate on the basis of race, color, religion, sex, sexual orientation, gender, gender identity, or expression, national or ethnic origin, shared ancestry, age, status as an individual with a disability, military or veteran status, genetic information, or other protected classes under the law. For additional information please see the University's Notice of Nondiscrimination.

 

Job seekers in need of a reasonable accommodation to complete the application process should call 773-702-5800 or submit a request via Applicant Inquiry Form.

 

All offers of employment are contingent upon a background check that includes a review of conviction history.  A conviction does not automatically preclude University employment.  Rather, the University considers conviction information on a case-by-case basis and assesses the nature of the offense, the circumstances surrounding it, the proximity in time of the conviction, and its relevance to the position.

 

The University of Chicago's Annual Security & Fire Safety Report (Report) provides information about University offices and programs that provide safety support, crime and fire statistics, emergency response and communications plans, and other policies and information. The Report can be accessed online at: http://securityreport.uchicago.edu. Paper copies of the Report are available, upon request, from the University of Chicago Police Department, 850 E. 61st Street, Chicago, IL 60637.

Skills

PythonAWSAzureDockerTerraformAnsibleCI/CDLinuxData ScienceGitHubSRECompliance

Similar Jobs

30

Senior Network Security Engineer, Network Security Engineering Group - Security Engineering Section (RMI Security Eng. & Ops Dep)

Rakuten·Rakuten Crimson House, Japan +1

Today

Senior Network Security Engineer, System Security Engineering Group - Security Engineering Section (RMI Security Eng. & Ops Dep)

Rakuten·Rakuten Crimson House, Japan +1

Today

Senior Production Engineer (Cloud and App Ops) - freelance

Netcompany·Strasbourg, Grand Est·Hybrid, Onsite

2d ago

London - Senior ML Ops Engineer (Experiences)

Tripadvisor·UK +1·Hybrid

2d ago

Senior Dev Ops Engineer

Hermeus·Los Angeles, CA·Onsite

4d ago

Sr. System Administrator / Dev Ops - Linux Kubernetes

Red Arch Solutions·Linthicum, MD

5d ago

Senior Platform Engineer - ML Ops

Medtronic·IRL-G Galway Parkmore East 1, Ireland

5d ago

Senior Platform Engineer - ML Ops

Medtronic·IRL-G Galway Parkmore East 1, Ireland

5d ago

Ground Senior Ops Training Engineer - Level 4

Northrop Grumman·COAU09, US

6d ago

Ground Senior Ops Training Engineer - Level 4

Northrop Grumman·Aurora, CO

6d ago

Senior SRE/ML Ops Engineer - Performance

Vibe·Paris·Hybrid

1w ago

Senior Security Engineer - Sec Ops

Relx·Home based-Pennsylvania, US +5·Remote

1w ago

Senior Security Engineer - Sec Ops

RELX Jobs·Home based-Pennsylvania, US +5·Remote

1w ago

Senior Engineer (Payment & Ops) - Leaply

Skelar·Warsaw +2·Remote

1w ago

Sr Dev Ops Engineer I

Jeppesen ForeFlight·Bengaluru, Karnataka

1w ago

Senior AI Ops Engineer - ITS

Civil Group·Minneapolis, MN·Hybrid

2w ago

Senior Engineer, Dev Ops

Presidio·Irving, TX

3w ago

Senior Staff Engineer, ML Ops (R4941)

Shieldai·San Mateo, California·Onsite

3w ago

Senior Principal Engineer, Central Ops & Systems

Docker·Seattle, WA·Remote

3w ago

Senior Data Ops engineer (x/f/m)

Doctolib·Paris, France

4w ago

Senior Software Engineer, Data Ops - PDEGO

Sonymusicentertainment·New York, US

1mo ago

Senior AI Ops Engineer

Bottomlinetechnologies·India

1mo ago

Senior Software Engineer, Ops Platform

Lime·Canada·Remote

1mo ago

Senior Engineer - AI/ML -Ops (6-8 Yrs)

Aptiv·IND - Technical Center India, Chennai

1mo ago

Autonomy Engineer, Ops Research (Senior - Principal)

Trueanomalyinc·Denver, CA +2

1mo ago

Senior ML Ops Engineer

KAYAK·Berlin Office·Hybrid

1mo ago

Sr Dev Ops Engineer

Tailored Brands·Houston, TX·Remote, Hybrid

1mo ago

Assistant VP, Senior Technical Developer, Group Technology & Ops

Uobgroup·Central Region, Singapore·Hybrid

1mo ago

Sr Dev-Ops Engineer (Altium)

Renesas Electronics·Shanghai, China

1mo ago

Senior Gateway Ops Engineer

Tencent·US-California-Palo Alto, US·Onsite

1mo ago