- Location
- Bengaluru, KA,IN, IN
- Type
- Full-time
- Department
- Engineering
- Experience
- 2+ years
- Education
- Bachelor
- Source
- Eightfold
Description
Develop and maintain ETL/ELT data pipelines for batch data processing and analytics workloads. Build data transformation solutions using Python, PySpark, and SQL. Work with Databricks to ingest, transform, process, and curate enterprise datasets. Integrate data from relational databases, APIs, files, and cloud storage. Support migration of datasets and pipelines from legacy or existing platforms to modern cloud-based data platforms. Perform data validation, reconciliation, and quality checks to ensure accuracy and completeness of migrated and processed data. Develop and maintain datasets and data models used by reporting, analytics, and downstream applications. Work with Azure Data Lake and related Azure data services for storing and processing enterprise data. Troubleshoot data pipeline failures, performance issues, and data quality problems. Participate in code reviews and follow standard development, testing, deployment, and CI/CD practices. Collaborate with Data Engineers, Analysts, Data Scientists, and business teams to understand data requirements. Support AI-related data requirements such as preparing and processing structured and unstructured datasets. Gain hands-on exposure to Generative AI concepts, LLMs and Retrieval-Augmented Generation (RAG) as part of AI-enabled data initiatives. 2-5 years of experience in Data Engineering, Data Analytics Engineering, or a related role. Experience with Spark / PySpark for data processing. Hands-on experience with Databricks or a similar cloud data processing platform. Understanding of Data Lake / Lakehouse concepts. Experience working with structured and semi-structured data such as CSV, JSON, and Parquet. Understanding of data quality, data validation, and reconciliation techniques. Experience with at least one cloud platform, preferably Microsoft Azure. Familiarity with Git and CI/CD / DevOps practices. Candidates should be willing to learn and work with: Generative AI and Large Language Models (LLMs) Preparing enterprise data for AI use cases Embeddings and semantic search Vector databases / vector search Retrieval-Augmented Generation (RAG) AI-assisted data engineering and automation Bachelor's degree in Computer Science, Information Technology, Engineering, or a related field. Strong willingness to learn new data and AI technologies.