Hiring.Camp

Data Engineer with strong Python Spark

Principal33

·

Today

Location
Remote
Workplace
Remote
Type
Full-time
Department
Engineering
Experience
5+ years
Source
Personio

Description

Your mission

Data pipelines and integration
•  Design, build and operate scalable batch, micro-batch and streaming data pipelines using Python, PySpark and Azure Databricks.
•  Integrate internal and external data sources, including REST APIs, GraphQL, WebSocket and gRPC interfaces, databases, files, event streams and third-party data feeds.
•  Develop robust ingestion solutions for structured, semi-structured and unstructured data, including JSON, CSV, Parquet, Delta and API-based payloads.
•  Build reliable web scraping and data acquisition components where APIs or managed integration mechanisms are unavailable.
•  Implement pagination, throttling, retries, exponential backoff, checkpointing, schema evolution and recovery patterns for external integrations.
•  Design pipelines that support idempotent processing, reprocessing, controlled backfills and graceful recovery from partial failures.
•  Develop and maintain batch and streaming patterns for market, weather, fundamental and time-series data.
 
Software engineering
•  Develop modular, reusable and testable Python and PySpark components rather than relying on monolithic notebooks.
•  Apply object-oriented and functional design principles appropriately to data-focused development.
•  Structure solutions as maintainable software projects with clear separation between source code, configuration, tests, deployment assets and notebooks.
•  Write clean, readable and well-documented code using type hints, meaningful interfaces and appropriate design patterns.
•  Build automated unit, integration, contract and data-quality tests and incorporate them into delivery pipelines.
•  Conduct code reviews and promote engineering standards covering readability, testability, security, performance and maintainability.
•  Package reusable functionality as Python modules or wheels where appropriate.
•  Troubleshoot complex issues across source systems, APIs, processing logic, infrastructure and production runtime environments.
 
Data architecture and modelling
•  Design maintainable data models that support analysts, traders, reporting solutions and downstream data products.
•  Implement Lakehouse and Medallion architecture patterns across Bronze, Silver and Gold layers.
•  Preserve raw data appropriately while applying cleansing, validation, standardisation and business transformations in downstream layers.
•  Design solutions for schema evolution, data retention, lineage and reproducible processing.
•  Apply sound data architecture principles across operational, analytical, event-based and time-series workloads.
•  Optimise data layouts, partitioning, joins, file sizes, caching and Spark execution plans for performance and cost.
 
Orchestration and DataOps
•  Design, schedule and operate workflows using Databricks Workflows and Astronomer.
•  Implement dependency management, parameterisation, environment-specific configuration and controlled promotion across development, test and production environments.
•  Define operational runbooks and support effective diagnosis, recovery and problem management.
•  Monitor pipeline health, freshness, completeness, performance and data-quality indicators.
•  Use production-safe release patterns, including controlled rollouts, rollback and validation where appropriate.
 
GitOps, CI/CD and Infrastructure as Code
•  Manage all production code through Git using clear branching, pull-request and review practices.
•  Build and maintain automated CI/CD pipelines using GitHub Actions and/or Azure DevOps.
•  Deploy Databricks jobs, pipelines and application artefacts using Databricks Asset Bundles or equivalent approved mechanisms.
•  Provision and configure relevant cloud and Databricks resources through Terraform.
•  Treat application code, infrastructure, data pipeline definitions and operational configuration as version-controlled artefacts.
•  Apply automated validation, security scanning and testing before production deployment.
•  Contribute to reusable pipeline templates, engineering standards and platform automation.
 
Reliability, performance and cost efficiency
•  Engineer solutions for availability, recoverability, scalability and predictable operational behaviour.
•  Optimise Spark workloads through appropriate partitioning, built-in Spark functions, efficient joins, adaptive execution and avoidance of unnecessary shuffles or UDFs.
•  Select suitable compute models and cluster configurations based on workload characteristics.
•  Apply cost-awareness to pipeline design, compute sizing, scheduling, storage and data-retention decisions.
•  Monitor resource consumption and identify opportunities to reduce processing times and cloud costs without compromising reliability or data quality.
•  Balance immediate delivery requirements with sustainable architecture and long-term maintainability.
 
Data quality, governance and security
•  Implement automated data validation, schema checks, null checks, referential-integrity controls and business quality rules.
•  Detect and manage schema drift and unexpected changes in source data.
•  Use Delta Lake and Unity Catalog capabilities to support data lineage, access control, metadata and governance.
•  Ensure secrets and credentials are handled securely using approved secret-management mechanisms and managed identities.
•  Maintain technical documentation, metadata and operational information for assigned data products.
•  Collaborate with data governance, architecture, security and platform teams to ensure alignment with enterprise standards.
 
Collaboration and delivery
•  Work closely with traders, analysts, data scientists, software engineers, product owners and platform teams to translate business requirements into robust technical solutions.
•  Communicate design decisions, risks, dependencies and technical trade-offs clearly to technical and non-technical stakeholders.
•  Contribute reusable components, templates, documentation and engineering guidelines for the wider data community.
•  Work effectively in a distributed, international and cross-functional environment.


Your profile

An experienced data engineer with at elast 5 years of experience, ideally in the energy sector and/or trading.
Be able to operate fundamental power-price forecasting models for short- to mid-term trading and large amounts of data.
Be able to work fully remote in a collaborative environment with interdisciplinary teams.
Good communication skills and professional behaviour.


What we offer

  • Medical insurance: Your health, and your family's, is our top priority. You're fully covered.
  • Holiday flat in Valencia: Dreaming of sunny days in Spain? Our company flat is ready for your next getaway.
  • Gifts for special occasions: We love celebrating you. Expect thoughtful surprises on Easter, Women's Day, Father's Day, and more.
  • Anniversary gifts: Your 1st, 5th, and 10th work anniversaries are milestones worth celebrating, and we mark each of them with a special gift.
  • Team-building events: We value connection beyond work, offering engaging team experiences in great locations to inspire collaboration and fun.
  • End of year celebrations: We wrap up each year with a special celebration, filled with joy, laughter, and unforgettable moments.
  • Day off on your birthday: Your special day is yours to enjoy, no work required.
  • Community and social initiatives: We bring people together through activities like Bring Your Kids to the Office Day, donation drives, tree planting, sporting events, decorating the Christmas tree with your colleagues, celebrating new office openings, Principal33's anniversary, and more.
  • Access to Udemy: Growth matters, so you get unlimited access to thousands of courses. Develop your skills anytime, anywhere.
  • German language courses: Dedicated courses to help you build your German skills, supporting both personal and professional growth.
  • Personal and professional growth: We support your development through masterclasses with experienced trainers, and we're open to investing in courses and training that help you grow.


Email address

Skills

PythonAzureTerraformCI/CDSparkDatabricksGitGitHubRESTGraphQLgRPCWebSocketDevOps

Similar Jobs

30

Senior Data Engineer – Data Analytics Platform 80-100% (f/m/d) - (Contract through our external payroll partner with immediate start for 12 months with possible extension)

Juliusbaer · Zurich, Switzerland

4 days ago

Junior Test Data Engineer - CLM 100% (f/m/d) (Contract through our external payroll partner with immediate start until 31.03.2027, possible extension)

Juliusbaer · Zurich, Switzerland

4 days ago

SAP BRIM Data Analyst / Developer with BODS and Emigal

"Data-Core System, Inc." · Middletown, Pennsylvania

6 days ago

Data Engineer With Python

Abb · Bangalore, Karnataka, India

6 days ago

Data Engineer with AI

AutoDesk · EMEA - Poland - Kraków - Lubomirskiego

2 weeks ago

AI/ML Data Engineer (TS/SCI with Poly Required)

GCI · Chantilly, VA

2 weeks ago

Data Engineer with German

DXC Technology · WSW03 - Warsaw Skyliner, DXC Technology Polska (WSW03), Poland +3

2 weeks ago

AWS Data Engineer with BI Connector To snowflake migration Exp(6+ Yrs of Exp & Immediate Joiner looking for)

3Pillarglobal · India · Remote

4 weeks ago

Data & AI Engineer Full Stack Developer with Italian

DXC Technology · WSW03 - Warsaw Skyliner, DXC Technology Polska (WSW03), Poland · Hybrid

1 month ago

Threat Detection & Automation Engineer (with focus on Data Engineering)

Northwestern Mutual · Milwaukee, WI Corporate, United States of America +1

1 month ago

(Junior or Mid) Data Engineer with Snowflake

Aviva the most attractive choice · Poland - Warsaw - ASEC

1 month ago

Senior Data Engineer with Credit Risk

Synechron · New York, NY, United States of America

1 month ago

Data Engineer | Pipelines With a Mission (AI Track Inside)

Thales · Bucharest Orhideea, Romania · Hybrid

1 month ago

Senior Data Engineer with ETL

Damia Group · Lisbon, Porto, , Porto · Remote, Hybrid, Onsite

1 month ago

Data Engineer with BigData

Damia Group · Porto, Porto · Remote, Hybrid, Onsite

1 month ago

Software Engineer- 8-10 years of experience in Data Warehousing (DWH) with expertise in cloud data platforms and database

Cisco · IND-HYDERABAD, India +1 · Hybrid

1 month ago

Data Center Facilities Engineer, (Clearance Required TS/SCI with Poly) Manassas, VA Onsite

Hewlett Packard Enterprise (HP) · Chantilly, Virginia, United States of America · Remote, Hybrid, Onsite

2 months ago

Data Center Facilities Engineer, (Clearance Required TS/SCI with Poly) Manassas, VA Onsite

Hewlett Packard Enterprise (HP) · Chantilly, Virginia, United States of America · Remote, Hybrid, Onsite

2 months ago

Data Center Facilities Engineer, (Clearance Required TS/SCI with Poly) Manassas, VA Onsite

Hewlett Packard Enterprise (HP) · Chantilly, Virginia, United States of America · Remote, Hybrid, Onsite

2 months ago

AI/ML Data Engineer (TS/SCI with Poly Required)

GCI · Chantilly, VA

2 months ago

Senior Data Engineer with Python and Databricks

Endava · Cluj-Napoca, CJ, Romania · Hybrid

2 months ago

Data Center Systems Engineer - TS/SCI with Polygraph

Leidos · 3064 Alice Springs NT Australia - Customer Site - Expat

2 months ago

DATA SCIENTIST AND AI ENGINEER WITH DATABRICKS

BIP · Madrid, Spain, ES

2 months ago

Senior Data Engineer with Google Cloud (GCP) (m/f/d)

Deutsche Telekom IT Solutions Slovakia · Juh, Košický kraj, Slovakia (Slovak Republic) · Hybrid

2 months ago

Lead Data Engineer with Google cloud GCP (m/f/d)

Deutsche Telekom IT Solutions Slovakia · Juh, Košický kraj, Slovakia (Slovak Republic) · Hybrid

2 months ago

Senior Data Engineer with strong Microsoft Fabric

Esr Healthcare · Phoenix · Hybrid

2 months ago

Data Engineer (with Dataiku)

Damia Group · Lisbon · Remote, Hybrid, Onsite

3 months ago

Software Engineer with Data Management Experience

Caci · CYC LINTHICUM HEIGHTS MD, United States of America · Onsite

3 months ago

Senior Data Engineer with Snowflake - Market Data Platform

Barclays · Gemini Building A, Prague, Czechia

3 months ago

Data Engineer with TS/SCI

OneGlobe LLC · Washington, DC

4 months ago