- Location
- Gurugram, Haryana
- Type
- Full-time
- Department
- Engineering
- Education
- Master
- Closing date
- Today
- Source
- ApplyToJob
Description
| NEOLYTIX | PROVIDER ANALYTICS Data Engineer — Healthcare Data Platform Reports To: Lead Architect, Provider Analytics |
| FUNCTION Provider Analytics | EXPERIENCE 4–7 Years | LOCATION Gurgaon (Office-based) | LEVEL Mid-Senior IC |
Why This Role Exists
Platform architecture is owned in-house by a hands-on data architect who sets the design — medallion lakehouse structure, modeling standards, and the validation framework. This role builds the platform inside that design, owning pipelines end-to-end from ingestion through production operation, and is explicitly expected to grow into independent ownership of the platform within 6–12 months. We are hiring an inheritor, not a ticket-taker.
What this role is not: it is not a Power BI report developer, and it is not a passive ETL operator who waits for a mapping document and a ticket queue. You should be able to point to at least one pipeline or data product you owned end-to-end — through design input, build, and production operation — and bring the appetite to take over a platform, not just work inside one.
- The medallion lakehouse (bronze / silver / gold) on the Microsoft data stack, implemented to the architect’s design and documented so every decision is reproducible.
- Automated extraction connectors for EHR / practice-management systems (target: 4 systems in year one), plus ingestion of clearinghouse 835 / 837 files and client-supplied feeds — replacing manual monthly pulls.
- A claim lineage database linking charges, claims, remittances, and payments across sources (target: ≥90% linkage rate), enabling true first-pass-resolution and denial analytics.
- Dimensional (star-schema) models for core RCM KPIs — first-pass resolution, days in AR, denial rates, net collection rate — serving certified Power BI semantic models.
- The automated validation harness inside the pipelines: row counts, reconciliation-to-source totals, period-over-period variance thresholds, and alerting — so errors are caught before a client sees them.
- A bounded InCredibly workstream: support for an in-flight client data migration and integration groundwork (API / FHIR) for the platform’s reporting needs.
- Own pipeline reliability end-to-end: orchestration, monitoring, failure recovery, schema-drift handling, and refresh SLAs.
- Absorb and document the architecture as it is built — decision records, runbooks, and standards — with the explicit goal of independent platform ownership within 6–12 months.
- Implement the technical half of data governance: lineage tracking, enforcement of certified metric definitions in the transformation layer, and role-based access to data assets.
- Integrate new data sources as clients onboard — API-based where available, structured file exchange (SFTP) where not.
- Handle PHI to HIPAA standards: minimum-necessary access, encryption in transit and at rest, and audit-ready handling in every pipeline.
- Work day-to-day with the associate data engineer; grow into a mentoring role as ownership expands.
- Strong SQL and solid Python for data engineering.
- Hands-on Azure data stack experience — Azure Data Factory plus Synapse, Fabric, or Databricks (equivalent lakehouse experience on another cloud considered with strong fundamentals).
- Has owned at least one pipeline or data product end-to-end — contributed to its design, built it, and operated it in production — as opposed to executing assigned tickets inside someone else’s build.
- Working knowledge of dimensional modeling applied in production.
- A demonstrated validation and reconciliation mindset: you can explain, concretely, how you knew your numbers were right. This is non-negotiable at any level.
- Disciplined handling of sensitive / regulated data; HIPAA or PHI exposure preferred.
- Evidence of fast learning and increasing scope — this role is a succession seat, and trajectory matters as much as current state.
- US healthcare claims and remittance data (835 / 837), EHR / practice-management exports, or RCM metrics exposure.
- FHIR / HL7 integration experience; REST API development.
- Power BI semantic-model / dataset design (the front end sits on your models).
6 Months: Connectors and the claim lineage database in production; operating pipelines with review-only oversight; manual extraction hours measurably reduced.
12 Months: 4 EMR/PM connectors live with a repeatable playbook; ≥90% claim linkage; ≥97% scored data accuracy on certified reports; independently owning platform run and extension.