Senior Site Reliability Engineer (SRE) – Application Observability & Readiness (Azure)
Encora
·Today
- Location
- Peru
- Department
- CSA Billable
- Seniority
- Senior
- Experience
- 5+ years
- Source
- Greenhouse
Description
Main Responsibilities
- Collaborate with development teams to design and implement monitoring, alerting, dashboards, and APM instrumentation across applications and services.
- Lead the implementation, configuration, and optimization of Application Performance Monitoring (APM) solutions.
- Apply observability best practices using tools such as Azure Monitor, Application Insights, New Relic, and Log Analytics (KQL).
- Enable code-level instrumentation, distributed tracing, and structured logging to improve application visibility and reliability.
- Design and maintain application-level monitoring dashboards and operational health metrics.
- Define and implement Service Level Indicators (SLIs), Service Level Objectives (SLOs), and effective alerting strategies based on latency, error rates, traffic, and resource saturation.
- Continuously improve monitoring and alerting mechanisms through production insights and incident learnings.
- Participate in production readiness reviews, identifying operational risks, observability gaps, and potential failure scenarios before deployment.
- Support incident analysis and post-incident improvements through enhanced telemetry and monitoring practices.
- Partner with engineering teams to ensure applications are reliable, scalable, and production-ready.
Mandatory Requirements
- Strong experience supporting and operating applications in Microsoft Azure IaaS environments.
- Hands-on experience with application observability, monitoring, and reliability engineering practices.
- Mandatory experience with DBT, Databricks, and SQL (minimum 1 year of experience).
- Experience implementing and managing APM solutions such as Application Insights, New Relic, or similar platforms.
- Experience designing dashboards and monitoring solutions using Azure Monitor, Application Insights, and Log Analytics (KQL).
- Familiarity with CI/CD environments including Azure DevOps and GitHub Actions.
- Solid understanding of cloud-native architectures and distributed application systems.
- Practical SRE mindset with experience in incident analysis, root cause investigation, and proactive problem prevention.
- Strong verbal and written English communication skills, with the ability to collaborate effectively with global teams.
Preferred Requirements
- Experience with scripting and automation using PowerShell and/or Bash.
- Knowledge of scalability, availability, and resilience patterns in modern cloud environments.
- Experience driving production readiness and operational excellence initiatives.
- Exposure to reliability engineering best practices in enterprise-scale environments.