Hiring.Camp

Site Reliability Engineer, Incident management

Qualys Careers

·

Yesterday

Location
Pune, India
Type
Full-time
Department
Engineering
Source
Workday

Description

Come work at a place where innovation and teamwork come together to support the most exciting missions in the world!

The Site Reliability Engineer – Incident management has the responsibility of monitoring, maintaining and managing entire Qualys infrastructure and services installed at different datacenters. When there is any malfunction in Product/Services, the Site Reliability Engineer- Incident Management technician Monitor, troubleshoots, repairs and gets the service/system back up as quickly as possible. Ensure maximum possible service availability and performance, provide support services for Engineering and other technical teams and to collaborate for quicker resolution. End to end Incident management, documentation and task automation are also part of responsibility.

WHERE THIS ROLE SITS

The SRE - incident management role partners closely with SRE, Engineering, DevOps, Infrastructure, and Security teams to ensure production systems remain resilient and operationally efficient.

WHAT YOU WILL DO

Reliability and Operations

· Maintain highly available and scalable applications and services.

· Monitor application and infrastructure health using observability tools.

· Respond to incidents, troubleshoot production issues, and perform root cause analysis.

· Participate in on-call rotations and incident response processes.

· Track and improve service reliability, latency, performance, and efficiency.

· Responsible for basic troubleshooting platform/product issues to isolate the problems and take appropriate action to resolve

· Ensure creation and timely resolution to incident tickets tracking and resolution of the incident

Automation and Platform Engineering

· Create automation scripts and tooling to reduce manual operational effort.

Collaboration and Continuous Improvement

· Site Reliability Engineer- Incident Management must carefully track and document all issues and resolutions in detail on the ticketing tool / documentation tools.

· Escalate the issue as needed to management, other IT resources or 3rd party vendors for assistance in reaching a resolution

WHAT GOOD LOOKS LIKE

· Proactively identifies reliability risks before they impact customers.

· Uses automation to eliminate repetitive operational tasks.

· Responds effectively to incidents and drives meaningful root cause analysis.

· Collaborates effectively across engineering and infrastructure teams.

DISTINGUISHING EXPECTATION

The SRE – Incident management is measured by improvements in reliability, automation, operational efficiency, and service availability. Success is demonstrated through proactive monitoring problem prevention, reduced operational burden, and continuous improvement of production systems.

REQUIRED QUALIFICATIONS

· Bachelor’s degree in Computer Science, Engineering, or related field (or equivalent experience).

· Experience in Site Reliability Engineering, NOC operations, or Cloud Infrastructure roles.

· Knowledge of Linux/Unix systems administration.

· Experience with cloud platforms such as AWS, OCI, or Google Cloud Platform.

· Proficiency in Python, Go, Bash, Java, or similar scripting/programming languages.

· Familiarity with Docker and Kubernetes.

· Familiarity with Terraform, CloudFormation, or other IaC tools.

· Experience with monitoring and logging tools such as Prometheus, Grafana,

AppDynamics, ELK, or Splunk.

· Understanding of networking, security, and distributed systems concepts.

PREFERRED QUALIFICATIONS

· Experience managing large-scale distributed systems.

· Knowledge of reliability engineering principles and automation best practices.

· Cloud certifications.

· Experience with incident management and postmortem culture.

WORK ENVIRONMENT

· Full-time role.

· On-call participation required.

· Hybrid or remote flexibility depending on company policy.

· Cross-functional collaboration with engineering and operations teams.

· Site Reliability Engineer- Incident Management team will operate 24*7*365 days. Monthly shift rotation basis

ABOUT QUALYS

Qualys, Inc. is a pioneer and leading provider of cloud-based IT, security, and compliance solutions, serving more than 10,000 customers across over 130 countries. Qualys delivers innovative security and compliance solutions that help organizations simplify security operations and reduce risk.

Skills

PythonJavaAWSDockerKubernetesTerraformLinuxSplunkDevOpsSREComplianceGo

Similar Jobs

30

AI Platform Engineer

Job Openings · Highland Heights, OH

Today

Staff Platform Engineer

Coherehealth · Hyderabad, Telangana, India

Today

DevOps Engineer

Endava · Ho Chi Minh City, Ho Chi Minh City, Vietnam · Hybrid

Today

Security Platform Engineer

Ericsson

Today

Site Reliability Engineer

Yuno · Europe +13 · Remote

Today

Senior Platform Engineer

Viavisolutions · Chennai, IND, India

Yesterday

Senior Platform Engineer

Ig · Bangalore, India

Yesterday

Site Reliability Engineer

Fis · AUS SYDN 55, Australia · Hybrid

Yesterday

Data Platform Engineer

Property Me · Sydney, New South Wales, Australia

Yesterday

SRE engineer

P2P.Org · Remote - Europe +3 · Remote

Yesterday

Site Reliability Engineer

Ontrac Solutions · Karachi, PK · Remote

Yesterday

Site Reliability Engineer

Ontrac Solutions · New Delhi, IN · Remote

Yesterday

Site Reliability Engineer

Ontrac Solutions · Vilnius, LT

Yesterday

Cloud Platform Engineer

Accenture · Pune, PDC3C, India

Yesterday

DevOps Engineer

Accenture · Bengaluru, BDC7B, India

Yesterday

DevOps Engineer

Accenture · Bengaluru, BDC7B, India

Yesterday

DevOps Engineer

Accenture · Bengaluru, BDC7B, India

Yesterday

DevOps Engineer

Accenture · Bengaluru, BDC7B, India

Yesterday

DevOps Engineer

Accenture · Bengaluru, BDC7B, India

Yesterday

DevOps Engineer

Accenture · Bengaluru, BDC7B, India

Yesterday

Cloud Platform Engineer

Accenture · Chennai, CDC2E, India

Yesterday

Cloud Platform Engineer

Accenture · Bengaluru, BDC7A, India

Yesterday

Data Platform Engineer

Accenture · Bengaluru, BDC7C, India

Yesterday

DevOps Engineer

Accenture · Hyderabad, HDC4A, India

Yesterday

Data Platform Engineer

Accenture · Chennai, CDC2A, India

Yesterday

Data Platform Engineer

Accenture · Hyderabad, HDC4A, India

Yesterday

DevOps Engineer

Accenture · Chennai, CDC2A, India

Yesterday

Data Platform Engineer

Accenture · Kolkata, KDC1A, India

Yesterday

DevOps Engineer

Barclays · Pune, Gera Commerzone SEZ, India

Yesterday

DevOps Engineer

Barclays · Pune, Gera Commerzone SEZ, India

Yesterday