Hiring.Camp

Site Reliability Engineer Senior

Cobre

·

Jul 3, 2026

Location
Colombia · Latin america
Department
Tech
Seniority
Senior
Experience
5+ years
Source
Greenhouse

Description

What is Cobre, and what do we do?
Cobre is Latin America’s leading instant b2b payments platform. We solve the region’s most complex money movement challenges by building advanced financial
infrastructure that enables companies to move money faster, safer, and more efficiently.
We enable instant business payments—local or international, direct or via API—all from a single platform.
Built for fintechs, PSPs, banks, and finance teams that demand speed, control, and efficiency. From real-time payments to automated treasury, we turn complex financial processes into simple experiences.
Cobre is the first platform in Colombia to enable companies to pay both banked and unbanked beneficiaries within the same payment cycle and through a single interface.
We are building the enterprise payments infrastructure of Latin America!
What we are looking for:
The Cobre Infrastructure team and their SRE engineers are professionals who face the daily challenges that allow us to improve the technological level of our products. We enjoy
each project or task, giving 100% and learning from each other. Our main goal is to maintain the reliability of our systems. To achieve this, we collaborate with other teams to find the most effective solutions, maintain high-reliability processes, adopt the necessary safety measures and optimize time and cost in every decision.
What would you be doing:
● Support teams in defining the infrastructure that will support the solution architecture.
● Support all the infrastructure (aws services and k8s clusters) and company products Culture zero-downtime deployments.
● Assisting with troubleshooting application issues and incidents related with infrastructure services.
● Review code instrumentation with development teams and ensure necessary dashboards are created to monitor.
● Document and maintain runbooks and procedures, automate as much as possible (AI-driven auto-remediation for incidents) to reduce MTTR.
● Perform periodic load and scalability testing to establish baselines, drift, and capacity planning.
● Design and implement peak readiness reviews for anticipated high-volume times.
● Lead weekly operational state reviews covering performance trends, anomalies, errors and other availability events with SREs, product owners, and development teams.
● Contribute to incident management, postmortems, and reliability reviews.
● Socialize SRE culture across teams within the organization to publicize the value of SRE, mentor and train other engineers around proactive reliability decision making and planning.

What do you need:
● Proven experience of at least 5 years as SRE or DevOps, with a strong focus on highly available and scalable environments, cloud infrastructure, observability, and incident management.
● In-depth technical knowledge of microservices architecture and cloud platforms (e.g., AWS, Kubernetes), along with proficiency in Infrastructure as Code (IaC) tools (e.g., Terraform, Pulumi).
● Strong mindset for automation and continuous improvement with a huge interest in AIOps / AI-driven auto-remediation (n8n, aws bedrock, python scripting …)
● Understanding of secure-by-design infrastructure principles.
● Exposure to GitOps and declarative configuration patterns.
● Basic knowledge in Port.io or any Internal Developer Platform (Backstage, Cortex).
● Strong understanding of monitoring, logging, and alerting tools, with a track record of improving system reliability and performance. (e.g., NewRelic, Datadog, Cloudwatch…)
● Proven experience troubleshooting, mitigating, and resolving issues in a distributed system.
● Ability to define and execute the SRE strategy, aligning it with company goals and driving the adoption of SRE practices across multiple teams.
● Resilience in facing challenges and promoting a fail-fast, learn-fast culture that embraces innovation and experimentation.
● Exceptional communication skills to effectively convey complex technical concepts to both technical and non-technical stakeholders.
● Ability to actively listen and understand diverse team and stakeholder needs,demonstrating empathy in decision-making and conflict resolution

Skills

PythonAWSKubernetesTerraformDevOpsSREMicroservices

Similar Jobs

30

DevOps Engineer

BeVera Solutions LLC · Atlanta, GA

Today

Senior Platform Engineer

Natera · US Remote · Remote

Today

Senior Platform Engineer

Focused · Chicago, Illinois, United States +1

Today

Site Reliability Engineer

Bay Systems Consulting · Berkeley, CA

Today

Site Reliability Engineer

OnBoard · United States

Today

DevOps Engineer

Sprezzatura Management Consulting · Remote, US · Remote

Today

Devops Engineer

Miratech · Bengaluru, KA, India · Remote

Today

DevOps Engineer

Inetum · Lisbon, Lisbon, Portugal · Hybrid

Today

Release Platform Engineer

bet365 · Manchester, England, United Kingdom · Hybrid

Today

Release Platform Engineer

bet365 · Stoke-on-Trent, England, United Kingdom · Hybrid

Today

Site Reliability Engineer

Fortive · IN +1 · Remote

Today

Site Reliability Engineer

Fluke · Remote, India · Remote

Today

DevOps Engineer

Capital Markets Gateway · Brno · Hybrid

Today

DevOps Engineer

Agdata · Pune, Maharashtra

Yesterday

Site Reliability Engineer

Andurilindustries · Waltham, Massachusetts, United States

Yesterday

Site Reliability Engineer

TJX · CAN Home Office Mississauga ON, Canada

Yesterday

DevOps Engineer

Swbc · SWBC Headquarters, United States of America

Yesterday

Engineer - DevOps

Allegion · Bangalore, India

Yesterday

Site Reliability Engineer

Truist · Atlanta GA - 303 Peachtree Center Avenue - Garden Offices, United States of America +2 · Remote, Hybrid

Yesterday

Site Reliability Engineer

Truist · Atlanta GA - 303 Peachtree Center Avenue - Garden Offices, United States of America +2 · Remote, Hybrid

Yesterday

Power Platform Engineer

RSM · SLV-San Salvador-Calle Cortez Blanco #8 Urb. Madreselva, El Salvador

Yesterday

Devops Engineer

Playpower Labs · Fully Remote - India Only · Remote

Yesterday

Staff Platform Engineer

Ftmo · Prague office · Onsite

Yesterday

DevOps Engineer

G2I · USA

Yesterday

Technology Platform Engineer

Accenture · Bengaluru, BDC7C, India

Yesterday

Technology Platform Engineer

Accenture · Chennai, CDC2F, India

Yesterday

Senior Platform Engineer

Amadeus · Lisbon - Carnaxide, Portugal · Hybrid

Yesterday

Cloud Platform Engineer

Accenture · Pune, PDC5A, India

Yesterday

Cloud Platform Engineer

Accenture · Bhubaneswar, BBDC1A, India

Yesterday

Data Platform Engineer

Accenture · Bengaluru, BDC9A, India

Yesterday
Site Reliability Engineer Senior at Cobre | Hiring.Camp