Hiring.Camp

Site Reliability Engineer

Razorpaysoftwareprivatelimited

·

Today

Location
Bengaluru
Department
Engineering
Experience
5+ years
Source
Greenhouse

Description

Razorpay is one of India’s leading full-stack financial technology companies, powering the way businesses move, manage, and grow money. Founded in 2014 by Harshil Mathur and Shashank Kumar with a simple vision - to simplify payments for Indian businesses - we’ve since grown into a fintech powerhouse driving India’s digital payment revolution.

Razorpay powers millions of businesses with a smarter, scalable stack that goes beyond transactions to help them truly build and grow.

From building AI-native agentic payments, to AI-assisted fraud detection and real-time risk intelligence to automated reconciliation, smart payouts, and predictive financial insights, we are embedding intelligence across our stack to make money movement faster, safer, and more efficient. In close collaboration with ecosystem partners - including banks, networks, regulators - we are pioneering industry-first solutions that are shaping the next era of fintech

Across India, Singapore and Malaysia, our products span everything from seamless checkouts to payroll automation - powering a fintech ecosystem that’s redefining how money moves across Asia.

Today, that ecosystem supports everyone from early-stage startups to some of India’s largest enterprises, enabling them to accept, process, and disburse payments at scale while expanding into new ways of managing money more efficiently.

Our scale speaks volumes: Razorpay processes $180+ billion in annualized transactions, powering leading businesses like Airbnb, Facebook, WhatsApp, Airtel, CRED, BookmyShow, Zomato, Swiggy, Lenskart, Mirae Asset Capital markets, Indian Oil, National Pension Scheme - and over 100 of India’s unicorns. With strong roots in India and growing operations in Southeast Asia, we are shaping the next chapter of financial technology across the region.

We are backed by global investors including GIC, Peak XV Partners (formerly Sequoia Capital India & SEA), Tiger Global, Ribbit Capital, Matrix Partners, MasterCard, and Salesforce Ventures, having raised over $740 million to date. Strategic acquisitions - including Ezetap (POS and offline payments), Curlec (Malaysia expansion), BillMe (digital invoicing), and POP (rewards-first UPI) - along with earlier moves in fraud prevention, payroll, and lending, have further strengthened our platform and widened our footprint across Asia.

But what truly sets Razorpay apart is our culture. At Razorpay, ownership is our oxygen - you own what you build, with no micromanagement or red tape, just the runway to make your ideas fly. Learning is a lifestyle - if you’re curious, you’ll feel at home here. People > Pedigree - we hire for attitude, hustle, and hunger more than degrees. Transparency thrives over titles - this is where interns question CXOs and CXOs say “thank you.” Guided by our values of Customer First, Autonomy & Ownership, Agility with Integrity, Transparency, Challenging the status quo and a strong belief that Razorpay grows with Razors,  you’ll be part of a 3000+ strong team building not just products, but the financial infrastructure of the future.

About the role

You will be one of the founding SREs at Razorpay, embedded with the payment platform teams that move money for millions of businesses. Your mandate is to take our payment flows from three nines to four and five nines of availability. In payments, a failed request is not a retry, it is a customer's money in limbo. You will define what reliability means here, build the systems that enforce it, and set the standard every future SRE is measured against.

What you will do

  • Define SLIs, SLOs, and error budgets for critical payment flows (authorization, capture, refunds, settlements, webhooks) and make them the shared language between product and platform teams.
  • Own the release lifecycle for payment services: design progressive rollout pipelines (canary, staged, feature-flagged), automated rollback triggers, and make "can we roll back in under 5 minutes" a launch-blocking question.
  • Carry the pager for payment-critical services, lead incident command during outages, and drive blameless postmortems where action items actually ship.
  • Eliminate toil through software: build automation for failover, capacity management, load shedding, and degradation so that known failure classes cannot recur.
  • Harden payment flows against distributed systems failure modes: retry storms, thundering herds, cascading failures, partial outages of banks and network partners, idempotency violations, and reconciliation gaps.
  • Run production readiness reviews for new payment services and hold the line on launch gates using error budget data, not opinion.
  • Instrument what matters: design alerting that pages on customer-facing symptoms, not noise, and cut mean time to detection and recovery quarter over quarter.
  • Practice failure on purpose: game days, chaos experiments, and failure injection against payment-critical paths.

What we are looking for

  • 10+ years of engineering experience, with at least 5 years operating large-scale distributed systems in production (high QPS, multi-region, or systems where sub-1 percent error rates were business-critical).
  • Strong software engineering skills in at least one of Go, Java, or Python. You have built tools and services, not just configured them.
  • Deep understanding of distributed systems failure modes and the patterns that contain them: circuit breakers, backpressure, bulkheading, graceful degradation, idempotency.
  • Solid fundamentals in Linux internals, networking, and databases under load (replication, failover, connection pool exhaustion, lock contention).
  • Hands-on experience designing or significantly improving deployment pipelines: canary analysis, automated rollback, feature flags.
  • Genuine on-call ownership: you have carried a pager for systems that mattered, led incidents, and can walk us through a specific outage you handled and what you changed afterward.
  • Fluency with modern observability (metrics, tracing, structured logging; e.g. Prometheus, Grafana, OpenTelemetry, Datadog, Coralogix, Clickhouse or similar) and experience reducing alert noise.
  • Experience defining SLOs and error budget policies from scratch. Contributions to reliability tooling, open source or internal, that other teams adopted.
  • The judgment and communication skills to tell a product team "not yet" with data, and the pragmatism to help them get to "yes" quickly.

Nice to have

  • Experience in payments, fintech, banking, trading, or another domain where correctness and money are coupled (transactional consistency, exactly-once semantics, reconciliation).
  • Experience with Kubernetes at scale, service mesh, and traffic management.
Razorpay believes in and follows an equal employment opportunity policy that doesn't discriminate on gender, religion, sexual orientation, colour, nationality, age, etc. We welcome interests and applications from all groups and communities across the globe.
 
Follow us on LinkedIn & Twitter

Skills

PythonJavaKubernetesLinuxSalesforceSRE

Similar Jobs

30

DevOps Engineer

Accenture·Taguig, Uptown Bonifacio Tower 2 +1

Today

DevOps Engineer

Tetrad Digital Integrity LLC·Washington, DC·Remote, Onsite

Today

DevOps Engineer

Booz Allen Hamilton·McLean, VA +1

Today

DevOps Engineer

Booz Allen Hamilton·Undisclosed Location - USA, VA

Today

DevOps Engineer

Study Now

Today

DevOps Engineer

Securonix·Pune, Maharashtra

Today

DevOps Engineer

Maropost·Mohali, Punjab

Today

Platform Support Engineer

Synechron·Bengaluru - GTP, India +1

Today

Senior Platform Engineer

Damia Group·Porto·Remote, Hybrid, Onsite

Today

DevOps Engineer

Damia Group·Porto·Remote, Hybrid, Onsite

Today

Collibra Platform Engineer

HarmonyTech·Washington, DC

Today

Staff Platform Engineer

Veeamsoftware·Remote, US·Onsite, Remote

1d ago

DevOps Engineer

" MAXISIQ, Inc."·Chantilly, VA

1d ago

Autonomy Platform Engineer

Doordash USA·San Francisco, CA·Remote, Hybrid, Onsite

1d ago

Platform Support Engineer

Marketing·Remote, OR·Remote

1d ago

Fullstack Platform Engineer

Songpush·Germany·Onsite

1d ago

Site Reliability Engineer

Andurilindustries·Waltham, Massachusetts

1d ago

DevOps Engineer

Wabtec·Bengaluru, KA

1d ago

Site Reliability Engineer

Momentum Financial Services Group·Hyderabad·Hybrid

1d ago

DevOps Engineer

Brio Digital·London, England·Hybrid

1d ago

DevOps Engineer

ACI Worldwide·Pune, Maharashtra·Hybrid

1d ago

Site Reliability Engineer

H&M Group·Bangalore, India

1d ago

Lead Engineer - Platform

Caci·1BQ WASHINGTON DC, US +1·Hybrid, Remote

1d ago

SRE Platform Engineer

Caci·1BQ WASHINGTON DC, US +1·Hybrid, Remote

1d ago

DevOps Engineer

Leidos·9358 Undisclosed DC Customer Site, US·Remote, Onsite

1d ago

DevOps Engineer

KBR Careers·Colorado Springs, 2424 Garden of the Gods Rd.·Onsite

1d ago

Senior Platform Engineer

Drawbridge Partners·Remote

1d ago

DevOps Engineer

Canals·Santiago +2·Remote

1d ago

DevOps Engineer

Proofpoint·Toronto, Canada +1

1d ago

Site Reliability Engineer

Securonix·Pune, Maharashtra

1d ago
Site Reliability Engineer at Razorpaysoftwareprivatelimited | Hiring.Camp