Senior Data Engineer - Databricks & Streaming - Healthcare AI (Onsite, Evening Shift, Lahore, PKR Salary)
Hr Pod Hiring Talent Globally
·Today
- Location
- Lahore
- Type
- Full-time
- Department
- Engineering
- Seniority
- Senior
- Experience
- 4+ years
- Closing date
- Today
- Source
- CareersPage
Description
Requirements:
- 4+ years of experience in data engineering, with substantial production experience in Databricks.
- Strong experience with Spark SQL, PySpark, Delta Lake, Medallion Architecture, and Delta Live Tables (DLT).
- Hands-on experience with Structured Streaming or equivalent production-grade streaming ingestion using Azure Event Hubs, Kafka, or Kinesis.
- Strong understanding of checkpoint recovery, watermarking, and deduplication strategies.
- Demonstrated experience debugging source-to-warehouse data discrepancies.
- Ability to walk through a real-world incident involving mismatched record counts and explain how the root cause was identified and resolved.
- Proven experience integrating third-party REST APIs in production.
- Experience handling pagination edge cases, rate and row limits, retries, and schema drift.
- Experience with entity resolution or data matching involving messy, real-world text data.
- Experience with a metrics or semantic layer such as Holistics AML/AQL, dbt Metrics, or LookML.
- Working understanding of why non-additive measures cannot be reliably calculated from pre-aggregated rollups.
- Strong SQL and Python skills, with the ability to own data pipelines end-to-end with minimal oversight.
- Strong written and spoken English, with the ability to collaborate effectively with a US-based team asynchronously.
- Healthcare data experience, including referrals, payer taxonomy, claims/eligibility, or other PHI-adjacent datasets.
- Familiarity with HIPAA handling expectations.
- Experience with voice-agent, call-center, or telephony/conversation data.
- Familiarity with call transcripts and containment or outcome metrics.
- Hands-on experience with Holistics, specifically AML/AQL modeling.
- Experience with the broader Azure ecosystem beyond Event Hubs, including ADLS, ADF, and Key Vault.
Responsibilities:
- Build and maintain resilient ingestion pipelines for third-party vendor REST APIs.
- Work primarily with voice-AI observability and telephony platforms.
- Handle different pagination schemes, including offset/limit and page/cursor models.
- Manage row and rate limits through time-windowing and adaptive bisection.
- Implement robust schema-drift handling through contract and column-presence checks.
- Ensure alerts are triggered when fields are renamed, moved, or removed rather than silently propagating null values.
- Own streaming ingestion from Azure Event Hubs into Databricks using Structured Streaming and/or Auto Loader.
- Manage checkpoints and offsets, watermarking, and at-least-once deduplication.
- Perform source-parity reconciliation across ingested and production data.
- Investigate row counts, dropped or duplicated events, late-arriving data, and schema mismatches.
- Identify and resolve the root cause when ingested data does not match production sources.
- Develop and maintain Delta Lake pipelines using a Medallion Architecture (Bronze Silver Gold).
- Use Spark SQL and PySpark to build and maintain production data pipelines.
- Implement idempotent MERGE upserts.
- Work with Delta Live Tables and materialized-view constraints, including CREATE OR REFRESH and LIVE references.
- Understand and manage differences between DLT and job execution contexts.
- Build entity-resolution pipelines for dirty, free-text data.
- Normalize practice, provider, and payer names using regex, canonical dictionaries, fuzzy matching, confidence-scored crosswalks, and override tables.
- Maintain the semantic and metrics layer with rigorous metric definitions.
- Define and maintain accurate denominators, data grain, and cohort boundaries.
- Ensure the correct handling of non-additive aggregates, including medians and percentiles that cannot be reliably supported through aggregate-aware pre-aggregation.
- Ensure every metric remains accurate and reproducible.
- Instrument data quality across the entire pipeline.
- Monitor data freshness, source parity, data contracts, and other critical quality checks.
- Build alerting mechanisms that identify data issues before they reach dashboards.