- Location
- Bangalore, India
- Type
- Full-time
- Department
- Engineering
- Seniority
- Senior
- Source
- Workday
Description
Position Summary:
We are seeking a hands-on Data Engineer building on Databricks who is growing their Lakehouse and performance-engineering depth, with a builder's mindset for AI-assisted operations.
Key Responsibilities:
- Design and develop scalable data pipelines and Lakehouse solutions on Databricks.
- Build and tune Databricks workloads for performance and cost, including cluster sizing, query optimization, and Delta Lake table design.
- Implement and utilize best practices for partitioning, clustering, and workload isolation.
- Track performance trends, identify high-cost queries, and partner with source teams and end users to resolve long-running loads.
- Design and operationalize Unity Catalog for data governance — access control, lineage, and security.
- Build monitoring and self-healing automation using Databricks-native AI and agentic capabilities.
- Contribute to CI/CD workflows for Databricks assets, applying DevOps best practices for deployment and release management.
- Deliver assigned pipelines and workloads with guidance from senior engineers, growing toward independent ownership.
What Success Looks Like (First 6–12 Months)
- In your first 6–12 months, you'll independently build and tune production pipelines, implement Unity Catalog access controls as designed, and contribute to monitoring automation.
Required Qualifications:
- Bachelor or Master’s degree in Computer Science, Information Technology or equivalent years of relevant experience.
- 3+ years of data engineering experience (with a focus on data integration), including 1+ years hands-on Databricks in enterprise settings.
- Solid understanding of Databricks Lakehouse architecture, Delta Lake, Unity Catalog, and Workflow orchestration.
- Working ability to tune Spark workloads for cost and performance.
- Strong Python (PySpark) and SQL skills.
- Working knowledge of CI/CD practices and DevOps principles applied to data workloads.
- Experience with observability tooling for Databricks.
Preferred Qualifications:
- Experience with Databricks-native AI capabilities and agentic frameworks.
- Familiarity with Databricks Serverless Compute and DBSQL performance tuning.
- A Databricks Certified Professional is nice to have.
- Exposure to Infrastructure-as-Code is a plus.
Competencies:
- Performance-engineering mindset — measures, tunes, and re-measures.
- Curiosity for AI-native operations and continuous automation.
- Strong sense of platform ownership — quality, cost, and reliability.
- Effective communication with engineering peers, vendors, and business stakeholders.