- Location
- Frankfurt
- Type
- Full-time
- Department
- Engineering
- Closing date
- Today
- Source
- CareersPage
Description
Data Engineer
Location: Frankfurt, Germany / EU
Start: ASAP
We are looking for a strong, hands-on Data Engineer with deep Databricks and Apache Spark experience to support a Frankfurt-based banking client.
The ideal candidate will have extensive experience designing and implementing scalable data engineering solutions using Databricks, Spark/PySpark, with a strong understanding of data modeling and modern data architectures. Experience working with unstructured and semi-structured data for AI/ML and RAG use cases is highly desirable.
Key Responsibilities:
- Design, develop, and optimize data engineering pipelines and data processing solutions using Databricks and Apache Spark.
- Build scalable and reliable data pipelines using PySpark and related Spark technologies.
- Work with large and complex datasets across structured, semi-structured, and unstructured data sources.
- Design and implement effective data models to support analytics, AI/ML, and downstream data consumption.
- Process and transform unstructured and semi-structured data for AI-driven use cases, including RAG (Retrieval-Augmented Generation).
- Develop data ingestion, transformation, cleansing, and enrichment workflows.
- Optimize Spark jobs and Databricks workloads for performance, scalability, and reliability.
- Work closely with data scientists, ML/AI engineers, architects, and business stakeholders to deliver production-ready data solutions.
- Apply strong engineering practices around data quality, testing, monitoring, and operational reliability.
- Contribute to the design and evolution of modern cloud-based data platforms.
Must-Have Requirements:
- Strong hands-on experience with Databricks in production environments.
- In-depth knowledge of Apache Spark and PySpark, including performance tuning and optimization.
- Strong Data Engineering background, with experience building production-grade data pipelines.
- Solid understanding of data modeling, data structures, and modern data architectures.
- Proven experience processing large-scale datasets.
- Experience working with unstructured and semi-structured data.
- Practical experience preparing and transforming data for AI/ML and RAG use cases.
- Strong Python skills, particularly for data engineering and PySpark development.
- Experience with data ingestion, transformation, orchestration, and pipeline automation.
- Ability to work independently in a fast-paced banking/enterprise environment.
- Candidate must be located within the EU.
Nice-to-Have:
- Experience with Generative AI / LLM / RAG architectures.
- Knowledge of vector search, embeddings, chunking, and document-processing pipelines.
- Experience with Delta Lake / Delta tables and modern lakehouse architectures.
- Experience with cloud platforms such as Azure, AWS, or GCP.
- Experience in banking or other regulated financial-services environments.
- Knowledge of data governance, security, lineage, and compliance requirements.
- Experience with CI/CD and DevOps practices for data platforms.