- Location
- (USA) VIZIO SERVICES IRVINE CA Irvine Home Office, United States of America
- Type
- Full-time
- Department
- Engineering
- Seniority
- Senior
- Experience
- 1+ years
- Education
- Master
- Closing date
- Today
- Source
- Workday
Description
What you'll do...
Position: Senior Site Reliability Engineer
Job Location: 39 Tesla, Irvine, CA 92618
Duties: Drive the design and evolution of monitoring and observability frameworks that enable proactive detection, root cause analysis, and rapid resolution of customer-impacting incidents. Lead the development and integration of automation tools to streamline operational workflows, reduce toil, and enhance the reliability of customer service platforms. Participate in on-call rotations, applying deep technical expertise to swiftly diagnose and mitigate production issues, ensuring high availability and minimal disruption to customer support experiences. Collaborate with engineering teams to embed reliability into the software development lifecycle, championing a culture of shared ownership and “you build it, you run it.” Define and manage SLIs, SLOs, and SLAs to align service reliability with business expectations and continuously improve system performance. Apply proven reliability patterns and practices, leveraging hands-on experience to architect resilient systems that scale with customer demand. Lead post-incident reviews and blameless retrospectives, identifying systemic improvements and fostering a culture of continuous learning and operational excellence. Analyze system performance and advocate for cost-effective optimizations, balancing infrastructure efficiency with world-class service reliability. Identify repetitive and routine tasks in (Continuous Integration/Continuous Delivery) CI/CD, testing, or any other process that can be automated. Implement telemetry features as required under guidance. Apply security policy requirements to component/module during code development/configuration. Detect and document defects, bugs, and errors for assigned component/module and conduct analysis to determine the sources under guidance. Troubleshoot performance and availability bottlenecks for assigned application under guidance. Work with business partners to identify and document critical applications. Interpret and follow procedures in contingency plans. Explain the contingency and disaster recovery plans for assigned environment. Execute established procedures necessary to continue operations in an emergency. Participate in the design of a minimum operating environment for a computer-based facility. Utilize established criteria (for example, probability of failure, frequency of failure) to measure site reliability. Monitor site reliability conditions and new reliability requirements. Assist in the design and development of a reliability program plan for a specific site environment. Apply appropriate tools, services, or applications for reliability prediction and other site improvements. Research and assess various reliability models for different site environments. Suggest metrics to monitor software or system performance. Monitor current performance data to ensure compliance with defined SLOs for multiple applications/systems. Determine thresholds for monitoring metrics and triggers alerts based on thresholds. Help with specific procedures to proactively check the health of applications and infrastructure, including a variety of operating systems, hardware, and software. Make recommendations regarding situational awareness and alerting. Make recommendations regarding instrumentation gaps and alerting logic, including a variety of operating systems, hardware, and software.
Minimum education and experience required: Master's degree or the equivalent in Computer Science, Computer Engineering, Computer Information Systems, Software Engineering, Electrical Engineering, Information Systems Security, or related area and 1 year of experience in site reliability engineering, site and system administration, infrastructure management, or related area; OR Bachelor's degree or the equivalent in Computer Science, Computer Engineering, Computer Information Systems, Software Engineering, Electrical Engineering, Information Systems Security, or related area and 3 years of experience in site reliability engineering, site and system administration, infrastructure management, or related area.
Skills required: Experience building, supporting, and maintaining Databricks Platforms and Databricks Infrastructure. Experience creating and managing AWS services, including s3 buckets, IAM roles and policies, VPC, Subnets, and VPCE Endpoints. Experience providing production support for applications running on data bricks and addressing all infrastructure requests from Data Engineering and Data Science teams. Experience working with Databricks and AWS service teams to understand new features, implement updates, and resolve vendor related issues. Experience investigating, troubleshooting, and optimizing Databricks spark jobs and providing resolutions. Experience supporting applications running on Elastic Kubernetes Service clusters and Kafka clusters on AWS and GCP. Experience building Infrastructure as Code using Terraform to automate the creation and management of Databricks resources. Experience using PagerDuty, Monte Carlo, and Cloud Watch for alerting, incident response, and monitoring. Experience using Airflow for workflow orchestration and job scheduling. Experience with migrations from AWS to GCP, setting up infrastructure, and enabling data movement from AWS s3 to Google Cloud Storage. Experience working with offshore, Central Ops, and QA teams to improve system reliability, reduce operational overhead, enhance observability, ensure safe deployments, reduce incident response time, and align engineering practices with business goals. Employer will accept any amount of experience with the required skills.
Salary Range: $108,000/year to $216,000/year. Additional compensation includes annual or quarterly performance incentives.
Benefits: At Walmart, we offer competitive pay as well as performance-based incentive awards and other great benefits for a happier mind, body, and wallet. Health benefits include medical, vision and dental coverage. Financial benefits include 401(k), stock purchase and company-paid life insurance. Paid time off benefits include PTO (including sick leave), parental leave, family care leave, bereavement, jury duty and voting. Other benefits include short-term and long-term disability, education assistance with 100% company paid college degrees, company discounts, military service pay, adoption expense reimbursement, and more.
Eligibility requirements apply to some benefits and may depend on your job classification and length of employment. Benefits are subject to change and may be subject to a specific plan or program terms. For information about benefits and eligibility, see One.Walmart.com.
Wal-Mart is an Equal Opportunity Employer.
#LI-DNI #LI-DNP
Walmart and its subsidiaries are committed to maintaining a drug-free workplace and has a no tolerance policy regarding the use of illegal drugs and alcohol on the job. This policy applies to all employees and aims to create a safe and productive work environment.