Hiring.Camp

Big Data / PySpark Engineering Lead - Vice President

citibank

·

Feb 26, 2026

Location
Pune, MH,IN, IN
Type
Full-time
Department
Engineering
Seniority
Lead
Experience
12+ years
Source
Eightfold

Description

Architecture & Design Design and implement scalable, fault-tolerant batch and real-time data processing pipelines. Develop robust data models and schema designs optimized for both performance and storage efficiency. Evaluate and integrate emerging tools and frameworks (e.g., Spark, Flink, Kafka) into the existing stack. Provide in-depth analysis with interpretive thinking to define issues and develop innovative solutions Develop comprehensive knowledge of how areas of business, such as architecture and infrastructure, integrate to accomplish business goals Ensure data integrity and security by implementing rigorous validation and encryption standards. Build and maintain CI/CD pipelines for automated testing and deployment of data jobs. Legacy Systems Decommissioning: Lead the strategic migration of data and logic from legacy platforms (e.g. on-premises SQL Servers) to a modern Data Lakehouse environment. ETL/ELT Transformation: Re-engineer existing stored procedures and complex legacy ETL jobs into scalable, distributed processing frameworks using Spark (Python) and Starburst/Trino. Validation & Parity Testing: Design and implement automated frameworks for Data Parity Testing to ensure 100% accuracy and consistency between legacy outputs and new big data results. Schema Evolution: Map and transform rigid, legacy relational schemas into flexible, high-performance formats optimized for the cloud (e.g., Parquet, Avro, or Iceberg). Phased Cutover Management: Orchestrate a phased migration strategy (Parallel Run, Shadow Execution) to ensure zero downtime for downstream business applications and reporting tools. Performance Benchmarking: Establish performance baselines on legacy systems and ensure the new Big Data architecture meets or exceeds those benchmarks at scale. Resolve variety of high impact problems/projects through in-depth evaluation of complex business processes, system processes, and industry standards Provide technical mentorship and conduct code reviews for junior and mid-level engineers. Serve as advisor or coach to mid-level developers and analysts, allocating work as necessary. Translate complex business requirements into technical specifications. Collaborate with Product Managers to ensure data availability for downstream analytics, business models and users Partner with multiple management teams to ensure appropriate integration of functions to meet goals as well as identify and define necessary system enhancements to deploy new products and process improvements Highly experienced and skilled technical lead with 12+years of experience with software building and platform engineering. Experience in Data Engineering, focused on Big Data ecosystems. Knowledge in Hadoop, YARN, Hive, Impala, Spark, and Spark SQL with extensive high volume of data processing pipeline development. Programming Expert level and hand on experience in Python. Familiarity with data formats like Avro, Parquet, CSV, JSON. Hands-on experience in writing SQL queries. Highly experienced with Unix based operating systems and shell scripting. Experience with source code management tools such as Bitbucket, Git etc. Big Data Tech Proficiency and hands-on in Hadoop, Spark, Hive, Kafka, and NoSQL databases (MongoDB, HBase). Experience working with query engines like Trino, Presto, Starburst Strong computer science fundamentals in data structures, algorithms, databases, and operating systems. Reverse Engineering, ability to read "spaghetti" SQL or old scripts and document the business logic before moving it. Data Lineage, Experience using tools (like Collibra or Informatica) to track where data comes from and where it's going. Change Management, Experience managing the technical "shock" to the business when switching from legacy BI tools to modern query engines like Starburst. Problem Solver: You don't just fix bugs; you identify the root cause to prevent recurrence. Communicator: You can explain the "why" behind a technical decision to non-technical stakeholders. Automation and AI Mindset: You believe that if a task has to be done twice, it should be automated. Familiarity with AI tools to expedite deliveries. ------------------------------------------------------ PySpark. ------------------------------------------------------

Skills

PythonCI/CDSQLMongoDBSparkHadoopData EngineeringETLGitChange Management

Similar Jobs

30

Big Data

AccionLabs·Singapore

1y+ ago

Big Data

AccionLabs·Singapore

1y+ ago

Big Data Developer (Relocation to Spain)

Talan·Santiago, Santiago Metropolitan Region·Remote

Today

Big Data Developer

Talan·Salamanca, CL·Remote

1d ago

Big Data Developer

Talan·Madrid, MD·Remote

1d ago

Big Data Developer

Talan·Buenos Aires, Argentina·Remote

1d ago

Big Data Developer (Relocation to Spain)

Talan·Lima, Callao Region·Remote

1d ago

Big Data Developer

Talan·Málaga, AN·Remote

1d ago

Big Data Scientist for the Services Analytics Plateau

Airbus·Manching, Germany·Onsite

3d ago

Senior Big Data Engineer

Gaming Innovation Group·Madrid·Remote, Hybrid

4d ago

Senior Big Data Engineer

Gaming Innovation Group·Malta·Remote, Hybrid

4d ago

Senior Big Data Engineer

Gaming Innovation Group·Marbella·Remote, Hybrid

4d ago

Senior Big Data Engineer

Gaming Innovation Group·Barcelona·Remote, Hybrid

4d ago

Software Engineer (Java/Big Data)

Equifax·USA - Georgia - Alpharetta - 30005, US·Remote, Onsite

1w ago

Software Engineer (Java and Big Data)

Equifax·USA - Georgia - Alpharetta - 30005, US·Remote, Onsite

1w ago

Senior Software Engineer (Big Data)

Bitsight·Remote USA, US·Remote

1w ago

Big Data AWS Sr Manager of Software Engineering

JPMorgan Chase·Columbus, OH

1w ago

Big Data AWS Sr Manager of Software Engineering

JP Morgan Chase·Columbus, OH

1w ago

Staff SDET (Big Data & Agentic AI)

Securonix·Pune, Maharashtra

1w ago

Big Data Engineer (Scala)

Box·Warsaw, Poland +1

1w ago

Senior Engineer, Big Data - Medical Cost Grouper - Remote

Molina Enterprise·US

2w ago

Senior Big Data Engineer

Babelgroup·CIUDAD DE MÉXICO, Mexico

2w ago

Staff Platform Engineer - Big Data Platform Management

Centre for Strategic Infocomm Technologies·Singapore·Onsite

2w ago

Big Data Engineer

Odaseva·Paris·Hybrid

2w ago

Senior Lead Software Engineer, Big Data & Cloud Engineering — Risk Central (London)

JPMorgan Chase·LONDON, GB

2w ago

Senior Lead Software Engineer, Big Data & Cloud Engineering — Risk Central (London)

JP Morgan Chase·LONDON, GB

2w ago

Sr Software Engineer - Big Data

Verisign·Reston, Virginia

2w ago

Senior Big Data Engineer (FastAPI expert)

Iqvia·Warsaw, Poland

2w ago

Software Engineer, Big Data

Ziprecruiter·Remote +1·Remote

3w ago

Senior Lead Software Engineer - Python, PySpark, Big Data, Data pipeline, ML/AI

JPMorgan Chase·Plano, TX

3w ago