Hiring.Camp

Platform Reliability Engineer, Azure

Wellfit

·

May 19, 2026

Salary
$130k – $150k/yr
Location
Irving, TX
Workplace
Hybrid
Department
Engineering
Source
Lever

Description

Wellfit is the dental industry’s fintech solution, breaking down financial barriers so patients, providers, employers, and payors can all access better care. As a healthcare fintech innovator, we’re transforming the patient journey and redefining what’s possible in dental care.

About Wellfit

Wellfit is the dental industry’s fintech solution, breaking down financial barriers so patients, providers, employers, and payors can access better care. As a healthcare fintech innovator, we are transforming the patient journey and redefining what is possible in dental care. Today, Wellfit supports a growing production platform serving 1,100+ offices and processing $1.5B+ in annual transactions. As we continue to scale, reliability, observability, alerting, and production readiness are critical to how we support our customers and deliver with confidence.


About the Role:
We are seeking a hands-on Platform Reliability Engineer, Azure to help strengthen the reliability, visibility, and operational maturity of our Azure-based platforms.

This role is ideal for someone who enjoys working directly in Azure, improving production systems, troubleshooting issues across infrastructure and application layers, and building practical monitoring and alerting solutions that help teams respond faster and operate more confidently.

You do not need to be an expert in every part of the stack on day one. We are looking for someone with strong Azure experience, solid troubleshooting instincts, a DevOps/reliability mindset, and the ability to collaborate closely with engineering teams across systems, services, and applications.

What You’ll Do

• Own and improve monitoring, alerting, and observability across Azure-based production systems.
•Work directly in Azure Monitor, Application Insights, App Services, logs, metrics, traces, and related Azure tooling to troubleshoot reliability and performance issues.
• Build and refine practical alerting workflows, including Sev0/Sev1 alert routing, escalation paths, and runbook integration.
• Create and maintain clear, actionable runbooks that help on-call engineers respond confidently to production incidents.
• Partner with engineering teams to investigate issues across infrastructure, configuration, deployments, services, and application behavior.
• Support release readiness by improving visibility into critical Azure resources before, during, and after production deployments.
• Build and maintain dashboards in tools such as Grafana, Azure Monitor, Application Insights, or similar observability platforms.
• Help configure incident routing integrations, including Slack/webhook-based alert delivery to the appropriate team channels.
• Automate repeatable operational tasks using PowerShell, Logic Apps, Azure tooling, or similar workflow automation methods.
• Contribute to RCA documentation, incident follow-up, reliability improvements, and operational playbook development.

What We’re Looking For
• Hands-on experience supporting production systems in Azure.
• Strong working knowledge of Azure App Services, Azure Monitor, Application Insights, and Azure production troubleshooting.
• Experience with DevOps, cloud operations, site reliability, platform engineering, or production support in a hands-on environment.
• Strong troubleshooting instincts and the ability to work through ambiguous production issues.
• Comfort working across logs, metrics, traces, alerts, configurations, deployments, and service dependencies.
• Ability to collaborate with software engineering teams across the stack, including .NET, Angular, SQL, APIs, and cloud services.
• Experience building or improving dashboards, alerts, runbooks, incident workflows, or operational playbooks.
• Working knowledge of scripting or automation, preferably with PowerShell, Logic Apps, CLI tooling, or similar technologies.
• Clear communication skills with the ability to document findings, explain issues, and drive follow-through after incidents.
• A high-ownership mindset with the ability to create structure, improve processes, and operate effectively in a fast-moving environment.

Preferred Experience
• Azure certifications.
• Grafana, Prometheus, DataDog, Dynatrace, or similar observability/APM tools.
• Slack integrations, webhooks, Logic Apps, or incident routing workflows.
• Azure Front Door, CDN, Function Apps, WebJobs, Service Bus, Event Hub, Event Grid, SQL Pools, App Service Plans, or related Azure services.
• Experience in healthcare, fintech, payments, or other high-availability environments.
• Experience in startup, SMB, or scale-up environments where ownership is broad and hands-on.

What Success Looks Like 
• You understand how production systems are monitored, where alerting gaps exist, and how to improve them.
• You can work directly in Azure to investigate issues, improve visibility, and support reliable operations.
• You build practical runbooks, dashboards, and alerting workflows that teams actually use.
• You collaborate well with engineers, ask strong troubleshooting questions, and help drive issues to resolution.
• You bring ownership, curiosity, and a builder mindset to a growing platform environment.

Why Wellfit
• Make an Impact: Your work will directly strengthen the reliability of a fast-growing healthcare fintech platform supporting 1,100+ offices and $1.5B+ in annual transactions.
• Build and Own: This is a high-impact role where you will help shape how we monitor, operate, and scale production systems.
• Work Flexibly: Hybrid model based in Dallas with 3 days per week in office.
• Comprehensive Benefits: Full medical, dental, vision, generous PTO, bonus eligibility, and 401(k) matching.
• Fast-Growth Environment: A rare opportunity to grow with a profitable startup on a national trajectory.

 

Skills

AngularAzureSQLDevOps

Similar Jobs

30

Platform Reliability Engineer

Appnovation·New York, Austin +2

1d ago

Platform Reliability Engineer

Appnovation·Toronto, Montreal +2

1d ago

Platform Reliability Engineer

ConocoPhillips·London Office, UK·Onsite

2w ago

Platform & Reliability Engineer

Volka·Cyprus, Limassol·Onsite

1mo ago

Platform Reliability Engineer

Worldquant·Montevideo +1·Remote

10mo ago

Senior Site Reliability Engineer, Platform Infrastructure

Cricut·South Jordan, UT·Hybrid

1d ago

Site Reliability Engineer III (SRE) - Guidewire Cloud Platform (Application)

Guidewire·Krakow Office, Poland·Hybrid

2d ago

Senior Site Reliability Engineer (SRE) - Guidewire Cloud Platform (Application)

Guidewire·Dublin Office, Ireland·Hybrid

2d ago

Site Reliability Engineer II (AI Platform)

Opentable·Toronto, Canada·Remote, Hybrid

2d ago

Site Reliability Engineer (SRE) – Cloud Platform

Fis·IND PUNE FL7, India

3d ago

Engineer, Platform Engineering & Reliability Specialist (Azure)

Nasdaq·GA - Glenridge Point, US·Hybrid, Onsite

3d ago

Senior Platform Reliability Engineer - Azure

LSEG·IND-Hyderabad-CapitaLand, India +1

6d ago

Associate - AI Tooling Ops - Platform Reliability Engineer

Jefferies·Pune, India

1w ago

Site Reliability Engineer, Kubernetes Platform (Top Secret Clearance)

Spacex·Hawthorne, CA +1

1w ago

Principal Site Reliability Engineer, Platform Engineering: Dedicated

Gitlab·Remote, Canada·Remote

1w ago

Senior Platform Reliability Engineer (Fabric and Interconnect)

Firmus Technologies·Sydney, New South Wales

1w ago

Senior Platform & Reliability Engineer (all genders)

Contabo·München, Bayern·Remote, Hybrid, Onsite

1w ago

Site Reliability Engineer, AI Platform

Algolia·Paris, France

1w ago

Senior Site Reliability Engineer, AI Platform

Algolia·Paris, France

1w ago

Site Reliability Engineer – Cloud Native Platform (OpenShift)

KPN·Amersfoort, UT·Hybrid

2w ago

MTS 2, Platform Reliability Engineer

Ebay·Bangalore, India·Hybrid

2w ago

Site Reliability Engineer - Data Platform

IMC·Amsterdam, Netherlands

2w ago

Site Reliability Engineer (Managed Patching & Platform Automation)

Swift·Kuala Lumpur, Malaysia

2w ago

Software Engineer II Platform Data Reliability

Sony Interactive Entertainment·San Mateo, CA·Hybrid

2w ago

Senior Platform Reliability Engineer

Firmus Technologies·Melbourne, Victoria

3w ago

Senior Lead Engineer – Cloud Platform, DevOps & Site Reliability

Qualcomm·Hyderabad, TS

3w ago

Staff Site Reliability Engineer - AI Platform Runtime

Nvidia·Santa Clara, CA

3w ago

Staff Site Reliability Engineer - AI Platform Runtime

Nvidia·Santa Clara, CA

3w ago

Platform Site Reliability Engineer (SRE)

Broadridge Financial Solutions·Manila - 6805 Ayala Ave, Philippines

3w ago

Lead Engineer, Platform Engineering & Reliability

Nasdaq·GA - Glenridge Point, US·Hybrid, Onsite

4w ago