- Location
- Bolivia; Colombia; Costa Rica; Peru
- Department
- CSA Billable
- Experience
- 6+ years
- Source
- Greenhouse
Description
Job Title: Site Reliability Engineer (SRE)
Key Skills: Kubernetes, AWS/Azure/GCP, Terraform, Python, Observability, CI/CD
Experience: +6 YOE.
Location: Costa Rica, Peru, Colombia, and Bolivia.
Mode: Remote.
We at Coforge are hiring Site Reliability Engineer (SRE) (#22323) with the following skill set.
Key Responsibilities
· Design, build, and operate scalable and highly available cloud platforms.
· Ensure reliability, performance, and stability of distributed production systems.
· Implement and maintain Infrastructure as Code using Terraform or similar tools.
· Manage Kubernetes-based and containerized environments.
· Define and operate SLOs, SLIs, error budgets, dashboards, runbooks, and alerting standards.
· Implement observability, monitoring, and incident response practices.
· Participate in on-call rotations and respond to production incidents.
· Collaborate with engineering teams to improve automation, scalability, and platform resilience.
· Conduct postmortem reviews and drive continuous reliability improvements.
Required Skills & Qualifications
· Bachelor’s degree in Computer Science, Engineering, Information Systems, Software Engineering, or a related technical field, or equivalent practical experience.
· 6+ years of experience in Site Reliability Engineering, Platform Engineering, Infrastructure Engineering, DevOps Engineering, Backend Engineering, or Production Engineering.
· Strong software engineering skills in at least one language such as Python, Go, Java, TypeScript, or C#.
· Strong understanding of distributed systems, microservices, APIs, asynchronous processing, queues, databases, caching, retries, idempotency, and failure modes.
· Experience with cloud infrastructure on AWS, Azure, or GCP.
· Experience with Kubernetes, containers, Terraform or similar IaC tooling, CI/CD pipelines, and Linux-based systems.
· Experience with observability tools such as Datadog, Prometheus, Grafana, OpenTelemetry, CloudWatch, New Relic, Splunk, or Sentry.
· Experience defining and operating SLOs, SLIs, error budgets, alerting standards, dashboards, runbooks, and incident response practices.
· Strong communication skills and experience working across cross-functional teams.
Preferred Skills
· Cloud, Kubernetes, Infrastructure, Reliability Engineering, Security, or DevOps certifications.
· Experience in logistics, transportation, final-mile delivery, field-service software, routing, dispatch, or fleet operations.
· Experience working with operational SaaS or marketplace platforms.
· Experience driving automation, platform reliability, and operational excellence initiatives.
Posted On: 14-08-2026
At Coforge, we hire professionals based solely on their skills and qualifications and do not discriminate based on age, disability, religion, gender, sexual orientation, socioeconomic status, or nationality.