- Salary
- $122k – $203k
- Location
- North Hills, NY - 3400 New Hyde Park Rd, United States of America
- Workplace
- Remote, Hybrid
- Type
- Full-time
- Department
- Engineering
- Experience
- 1+ years
- Education
- PhD
- Source
- Workday
Description
Company
Cox Automotive - USAJob Family Group
Job Profile
Management Level
Flexible Work Option
Travel %
Work Shift
Compensation
Compensation includes a base salary in the range of $121,800.00 - $203,000.00. The base salary may vary within the anticipated base pay range based on factors such as the ultimate location of the position and the selected candidate’s knowledge, skills, and abilities. Position may be eligible for additional compensation that may include an incentive program.Job Description
We're hiring a Sr. Software Engineer - Reliability Engineer who can code across the stack and cares deeply about reliability. You'll design and build infrastructure, observability tooling, and operational systems—treating resilience and debuggability as first-class concerns. You'll own projects end-to-end: from architecture → code → deployment → production. You'll split time between infrastructure-as-code, incident response, system improvements, and mentoring. You'll work on a team that ships quality systems while maintaining operational excellence across a platform serving millions of dealership transactions daily.
What You'll Do:
SRE Best Practices & System Health Management
- Design and implement resilience initiatives: redundancy, failover, disaster recovery, data protection.
- Write infrastructure-as-code (Terraform); manage 50+ AWS accounts with infrastructure patterns.
- Own system health: proactively maintain application performance, minimize downtime, ensure consistent user experience.
- Evolve team's SRE standards and practices.
Application Monitoring & Observability
- Build observability into systems: logging, metrics, distributed tracing, alert design.
- Improve monitoring frameworks; enable faster incident detection and resolution.
- Design dashboards and alerts that help teams understand system behavior.
- Partner with application teams on service instrumentation.
AWS Cost Optimization
- Drive significant reductions in cloud spend through architectural improvements and resource utilization.
- Review infrastructure for efficiency; identify and eliminate waste.
- Balance cost, performance, and reliability in design decisions.
Software Development & Architecture
- Build production systems, APIs, internal tools, and automation with clean, well-tested code.
- Design for maintainability, operational simplicity, and reliability.
- Participate in code review and technical design discussions.
- Mentor junior engineers on code quality and architectural thinking.
Operations & Incident Response
- Participate in on-call rotations; debug and resolve production incidents.
- Conduct postmortem analysis; drive systemic improvements.
- Develop operational procedures and runbooks.
Qualifications:
- 5+ years software engineering, platform engineering, or infrastructure engineering experience.
- Strong coding in Python, Go, Java, or equivalent; writes clean, testable code.
- AWS hands-on: EC2, RDS, DynamoDB, S3, Aurora, Lambda, VPCs, Athena.
- Terraform or equivalent infrastructure-as-code experience.
- Docker and container orchestration (Kubernetes or similar).
- Debugging on Linux and Windows platforms; able to troubleshoot complex systems using logs, metrics, and architectural knowledge.
- System design thinking: can architect scalable systems and reason about trade-offs.
Bachelor’s degree in a related discipline and 4 years’ experience in a related field. The right candidate could also have a different combination, such as a master’s degree and 2 years’ experience; a Ph.D. and up to 1 year of experience; or 16 years’ experience in a related field.
Highly Valued:
- Experience with observability tools (New Relic, Splunk, Prometheus).
- Incident response experience; familiar with postmortem practices.
- Interest in or hands-on experience with SRE concepts (SLOs, resilience, failure modes).
- Cost optimization mindset; has identified and eliminated cloud waste.
- Windows and Linux system troubleshooting and performance analysis.
- Experience with CI/CD pipelines and deployment automation.
Drug Testing
Benefits
About Us