- Location
- Atlanta Support Center, United States of America
- Type
- Full-time
- Department
- Engineering
- Seniority
- Senior
- Experience
- 5+ years
- Education
- Bachelor
- Source
- Workday
Description
Inspire Brands is hiring two Senior Site Reliability Engineers to help build and scale reliable, resilient, and observable systems supporting high-traffic, customer-facing digital platforms. These role blends software engineering, systems thinking, and operational excellence to reduce toil, prevent incidents, and improve system reliability at scale.
The ideal candidate has hands-on experience applying and implementing SRE principles — not just supporting production systems, but engineering reliability into them.
RESPONSIBILITIES
Reliability Engineering
- Define and manage SLIs, SLOs, and Error Budgets for critical services
- Drive production readiness reviews and reliability requirements into architecture and design
- Perform capacity planning, failure mode analysis, and dependency risk assessments
- Identify systemic reliability risks and drive remediation before they cause customer impact
Observability
- Design monitoring, alerting, logging, and tracing solutions using modern observability tooling
- Improve signal-to-noise ratio and reduce alert fatigue
- Build dashboards and telemetry that reflect true service health, not just infrastructure metrics
Incident Management
- Lead technical response for high-severity incidents
- Drive blameless postmortems and root cause analysis focused on systemic fixes
- Continuously improve detection, response, and recovery processes
- Participate in an on-call rotation
Automation & Toil Reduction
- Identify and eliminate manual, repetitive operational work through automation
- Build self-healing systems, tooling, and scripts to reduce human intervention
- Improve CI/CD pipelines and deployment safety (canary, rollback, blue-green)
- Support Infrastructure as Code (Terraform, Bicep, or similar)
Performance & Scalability
- Conduct load testing, performance benchmarking, and bottleneck analysis
- Partner with engineering to design systems for horizontal scalability and fault tolerance
Collaboration & Culture
- Partner with engineering teams to implement resiliency patterns (circuit breakers, retries, graceful degradation, rate limiting)
- Mentor engineers on SRE best practices
- Promote a culture of engineering-driven reliability over reactive operations
EDUCATION AND EXPERIENCE QUALIFICATIONS
Required Qualifications
- 5+ years experience in Site Reliability Engineering, Software Engineering, or Platform Engineering
- 2+ years experience with Kubernetes and containerized workloads
- 4-year degree in Computer Science or related field
Preferred Qualifications
- Experience with chaos engineering or resiliency testing
- Experience with high-volume, high-availability transactional systems
- Experience with AI-assisted observability or operational automation
- Experience making meaningful contributions to internal SRE tooling, frameworks, or platforms
REQUIRED KNOWLEDGE, SKILLS, OR ABILITIES
- Strong programming/scripting skills (Python, Go, Java, or Node.js)
- Demonstrated experience defining and operating against SLOs/Error Budgets
- Strong skills in leading incident response and root cause analysis for production systems
- Solid understanding of distributed systems and microservices architecture
- Deep knowledge and expertise in at least one major cloud platform (Azure, AWS, or GCP)
- Expertise with observability platforms and monitoring strategy
This position is based in our Atlanta Support Center, with an expected on-site presence of 80%.