- Location
- Hamburg, Deutschland, Hamburg
- Type
- Full-time
- Department
- Engineering
- Seniority
- Senior
- Experience
- 3+ years
- Education
- Master
- Source
- Personio
Description
Your tasks
Your role
As a Senior Data Engineer, you will own key parts of our data ecosystem and platform capabilities that enable advanced analytics, machine learning, and NLP/LLM applications. Your focus is to make data usable at scale: well-structured, traceable, governed, and accessible for downstream AI/ML use cases and enterprise applications. You will work closely with the materials experts, simulation teams, lab stakeholders, and AI/ML colleagues to translate product goals into robust data products and production-ready solutions.
Key responsibilities
- Design, implement, and operate scalable data pipelines for structured and unstructured data, including batch processing and event-driven or streaming workflows where needed.
- Develop cloud-native and on-premises architectures for data and AI workloads (primarily on Microsoft Azure).
- Build and evolve a materials data ecosystem linking physics-based modeling/simulation data and experimental laboratory data
- Handle large-scale scientific datasets (e.g., atomistic simulations, DFT/MD, high-throughput campaigns), including efficient storage, metadata, and performant access patterns
- Integrate data from HPC/simulation workflows and laboratory systems (instrument exports, LIMS/ELN where applicable) into curated, analysis-ready datasets
- Define and implement data models, metadata standards, and provenance to ensure traceability, reproducibility, and auditability across simulations and experiments
- Establish robust data quality practices (validation rules, unit consistency, schema controls) and data quality monitoring aligned with operational SLAs/SLOs
- Implement data governance foundations (cataloging, access control, lineage) and enable policy-driven data sharing across teams
- Work with the DevOps team to implement and improve CI/CD pipelines, deployment automation, and infrastructure requirements for data and AI workloads.
- Ensure reliability, security, GDPR compliance, monitoring/observability (logging, metrics, alerting), and cost efficiency of cloud platforms
- Provide technical leadership through design reviews, documentation of standards, and mentoring where appropriate
Deine Aufgaben
Your qualification
Your profile
Education & experience
- Degree in Computer Science, Data Science & Engineering, Mathematics, Natural Sciences, or a comparable field
- 3+ years of professional experience in data engineering.
- Experience designing or operating data ecosystems that unify multiple data domains with strong governance and provenance
- Strong Python skills and experience with data processing frameworks.
- Strong SQL skills and experience with data modeling for analytics and production use cases
- Familiarity with Vector databases, semantic search, text chunking strategies, LLM workflows and RAG architectures.
- Experience with message brokers and asynchronous processing patterns; practical experience with RabbitMQ is a strong plus
- Proven experience with Microsoft Azure (Data Services, Compute, Storage, Azure OpenAI); AWS/GCP experience is also valued
- Solid understanding of DataOps practices, CI/CD pipelines, and automation
- Familiarity with cloud databases, security concepts, and enterprise integration patterns
- Experience with orchestration tools and operational reliability practices
Practical knowledge of distributed data processing and scaling patterns for ingestion/transform/query of very large datasets - Bonus: familiarity with computational materials science/materials informatics, simulation pipelines, or lab data management (LIMS/ELN/instrumentation exports)
- Structured, pragmatic, and hands-on with strong ownership
- Able to communicate clearly across technical and non-technical stakeholders
- Curious about new technologies and able to translate them into reliable production systems
- Excellent communication skills in English; German is a plus
Deine Qualifikation
We offer
What we offer
- Innovative, fast-growing environment with a long-term perspective
- Flat hierarchies, fast decisions, and direct collaboration with management
- Permanent employment with flexible working hours
- Hybrid setup with remote work up to 2 days/week
- Pension scheme, corporate benefits, team events, and a well-connected office location
Wir bieten Dir
Are you interested?
Are you interested?
Send your CV, LinkedIn profile and a short note highlighting relevant experience in data platforms, data pipelines, and scientific data (simulation/experimental), etc. to: [email protected]