- Location
- Louisville, KY, US
- Workplace
- Remote
- Type
- Contract
- Department
- IT
- Experience
- 4+ years
- Education
- Master
- Source
- Breezy HR
Description
We are looking for a highly skilled Databricks Developer with expertise in building and managing modern data platforms using the Databricks Lakehouse architecture. The ideal candidate will have strong experience in PySpark, Python, SQL, Delta Lake, data modeling, data quality frameworks, and enterprise-scale data engineering solutions. The role involves designing, developing, and optimizing scalable data pipelines that support analytics, reporting, and AI/ML initiatives.
Key Responsibilities
Design and implement scalable data solutions using Databricks Lakehouse Architecture.
Develop and maintain data pipelines using PySpark, Python, and SQL.
Build and optimize ETL/ELT workflows for batch and near real-time data processing.
Implement Delta Lake features including ACID transactions, time travel, schema evolution, and data versioning.
Design and maintain enterprise data models to support reporting and analytics requirements.
Ensure data quality through validation, monitoring, reconciliation, and governance controls.
Develop and manage data catalogs, metadata management, and data lineage processes.
Collaborate with business stakeholders, architects, and analytics teams to gather and translate requirements into technical solutions.
Optimize Databricks workloads for performance, scalability, and cost efficiency.
Implement security, access controls, and governance best practices within the Databricks ecosystem.
Support troubleshooting, root cause analysis, and production issue resolution.
Contribute to data platform modernization and cloud migration initiatives.
Required Technical Skills
Databricks
Strong experience with Databricks Architecture and platform administration.
Hands-on expertise in Databricks Lakehouse Architecture.
Deep understanding of Delta Lake concepts and implementation.
Experience with Unity Catalog / Data Catalog and metadata management.
Knowledge of Databricks Workflows, Jobs, Clusters, and Performance Tuning.
Data Engineering
Strong proficiency in PySpark for large-scale data processing.
Advanced Python programming skills.
Expert-level SQL development and query optimization.
Experience in building robust ETL/ELT pipelines.
Strong understanding of data modeling techniques including:
Star Schema
Snowflake Schema
Dimensional Modeling
Data Vault (preferred)
Data Governance & Quality
Experience implementing data quality frameworks and validation checks.
Knowledge of data lineage, metadata management, and governance processes.
Experience with data reconciliation, profiling, and monitoring tools.
Cloud & Platform Experience (Preferred)
Azure Databricks
Azure Data Lake Storage (ADLS)
Azure Data Factory
Azure Synapse Analytics
CI/CD pipelines (Azure DevOps, GitHub Actions)
Qualifications
Bachelor's or Master's degree in Computer Science, Information Technology, Data Engineering, or a related field.
4-8 years of experience in Data Engineering and Analytics.
Minimum 3+ years of hands-on experience with Databricks and PySpark.
Experience working in Agile development environments.