Hiring.Camp

Site Reliability Engineer

HighRadius

Location
Hyderabad, Telangana, India · Hyderabad
Department
Engineering

Description

Job Summary:
We are looking for a highly skilled and adaptable Site Reliability Engineer (7+ Years) to become a key member of our Cloud Engineering team. In this crucial role, you will be instrumental in designing and refining our cloud infrastructure with a strong focus on reliability, security, and scalability. As an SRE, you'll apply software engineering principles to solve operational challenges, ensuring the overall operational resilience and continuous stability of our systems. This position requires a blend of managing live production environments and contributing to engineering efforts such as automation and system improvements.


Key Responsibilities:
● Cloud Infrastructure Architecture and Management: Design, build, and maintain resilient cloud infrastructure solutions to support the development and deployment of scalable and reliable applications. This includes managing and optimizing cloud platforms for high availability, performance, and cost efficiency.
● Enhancing Service Reliability: Lead reliability best practices by establishing and managing monitoring and alerting systems to proactively detect and respond to anomalies and performance issues. Utilize SLI, SLO, and SLA concepts to measure and improve reliability. Identify and resolve potential bottlenecks and areas for enhancement.
● Driving Automation and Efficiency: Contribute to the automation, provisioning, and standardization of infrastructure resources and system configurations. Identify and implement automation for repetitive tasks to significantly reduce operational overhead. Develop Standard Operating Procedures (SOPs) and automate workflows using tools like Rundeck or Jenkins.
● Incident Response and Resolution: Participate in and help resolve major incidents, conduct thorough root cause analyses, and implement permanent solutions. Effectively manage incidents within the production environment using a systematic problem-solving approach.
● Collaboration and Innovation: Work closely with diverse stakeholders and cross-functional teams, including software engineers, to integrate cloud solutions, gather requirements, and execute Proof of Concepts (POCs). Foster strong collaboration and communication. Guide designs and processes with a focus on resilience and minimizing manual effort. Promote the adoption of common tooling and components, and implement software and tools to enhance resilience and automate operations. Be open to adopting new tools and approaches as needed.


Required Skills and Experience:
● Cloud Platforms: Demonstrated expertise in at least one major cloud platform (AWS, Azure, or GCP). Extensive experience with containerization (Docker) and orchestration (Kubernetes) technologies.
● Automation & IaC: Proficiency in scripting languages (shell and Python). Experience with configuration management tools (Ansible or Puppet). Must have exposure to Infrastructure as Code (IaC) tools (Terraform or CloudFormation).
● Monitoring & Observability: Experience setting up and configuring monitoring tools (Prometheus, Grafana, or the ELK stack). Hands-on experience implementing OpenTelemetry for observability. Familiarity with monitoring and logging tools for cloud-based applications.
● Service Reliability Concepts: A strong understanding of SLI, SLO, SLA, and error budgeting.
● Soft Skills & Mindset: Excellent communication and interpersonal skills for effective teamwork. We value proactive individuals who are eager to learn and adapt in a dynamic environment. Must possess a pragmatic and adaptable mindset, with a willingness to step outside comfort zones and acquire new skills. Ability to consider the broader system impact of your work. Must be a change advocate for reliability initiatives.


Desired/Bonus Skills:
● Experience with DevOps toolchain elements like Git, Jenkins, Rundeck, ArgoCD, or Crossplane.
● Experience with database management, particularly MySQL and Hadoop.
● Knowledge of cloud cost management and optimization strategies.
● Exposure to Gen AI.
● Understanding of cloud security best practices, including data encryption, access controls, and identity management.
● Experience implementing disaster recovery and business continuity plans.
● Familiarity with ITIL (Information Technology Infrastructure Library) processes

Skills

PythonAWSAzureGCPDockerKubernetesTerraformAnsibleJenkinsMySQLHadoopGitDevOpsSREITIL

Similar Jobs

30

Staff Platform Engineer

Robotsandpencils · US Remote +1 · Remote

Today

Staff Platform Engineer

Vodafone · London, England,GB, GB

Today

Enablement Platform Engineer

Endava · Warsaw, Mazovia, Poland · Hybrid

Today

DevOps Engineer

Verint · Herzliya, Israel, IL

Today

Software Engineer, Platform

Aerovect · Atlanta - Hybrid · Hybrid

Today

Data Platform Engineer

Accenture · Bengaluru, BDC7B, India

Today

Data Platform Engineer

Accenture · Bengaluru, BDC7B, India

Today

Data Platform Engineer

Accenture · Bengaluru, BDC7B, India

Today

DevOps Engineer

Netcompany · London, England, United Kingdom · Hybrid

Yesterday

DevOps Engineer

Talan · Málaga, AN, Spain · Remote

Yesterday

DevOps Engineer

CFS · Sydney, NSW, Australia

Yesterday

Senior Platform Engineer

Mmc · Cluj-Napoca - Decembrie, Romania · Hybrid

Yesterday

Cloud platform engineer

Professional Kyndryl · PELML Lima (PELML) La Molina, Peru · Remote

Yesterday

Cloud platform engineer

Professional Kyndryl · PELML Lima (PELML) La Molina, Peru · Remote

Yesterday

DevOps Engineer

OpenObserve Inc. · India · Remote

Yesterday

AI Platform Engineer

SatoshiLabs · Prague · Hybrid

Yesterday

DevOps Engineer

Ncr · Belgrade Campus, Serbia

Yesterday

Platform AI Engineer

Hewlett Packard Enterprise (HP) · Bengaluru, Karnātaka, India · Hybrid, Onsite

Yesterday

DevOps Engineer

KLA · IND-Tamil Nadu-Chennai-KLA, India

Yesterday

Platform AI Engineer

Hewlett Packard Enterprise (HP) · Bengaluru, Karnātaka, India · Hybrid, Onsite

Yesterday

Platform AI Engineer

Hewlett Packard Enterprise (HP) · Bengaluru, Karnātaka, India · Hybrid, Onsite

Yesterday

DevOps Engineer

Teliacompany · Vilnius, Lithuania

Yesterday

AI Platform Engineer

Teamsystem · MILANO_P.ZZA LUIGI EINAUDI, Italy +1

Yesterday

DevOps Engineer

Accenture · Santiago, Andes, Chile · Hybrid

2 days ago

ServiceNow Platform Engineer

Lightfeatheriollc · Washington, DC +1 · Hybrid

4 days ago

Platform Product Engineer

Zonecompanysoftwareconsultingllc · Spain

4 days ago

Platform Product Engineer

Zonecompanysoftwareconsultingllc · United Kingdom

4 days ago

Platform Product Engineer

Zonecompanysoftwareconsultingllc · Czech Republic

4 days ago

DevOps Engineer

InfoImage · Brisbane, CA

4 days ago

Power Platform Engineer

Openings - Red Thread · E. Hartford, CT

4 days ago
Site Reliability Engineer at HighRadius | Hiring.Camp