Hiring.Camp

Big Data / PySpark Engineering Lead - Vice President

citibank

·

Feb 26, 2026

Location
Pune, MH,IN, IN
Type
Full-time
Department
Engineering
Seniority
Lead
Experience
12+ years
Source
Eightfold

Description

Architecture & Design Design and implement scalable, fault-tolerant batch and real-time data processing pipelines. Develop robust data models and schema designs optimized for both performance and storage efficiency. Evaluate and integrate emerging tools and frameworks (e.g., Spark, Flink, Kafka) into the existing stack. Provide in-depth analysis with interpretive thinking to define issues and develop innovative solutions Develop comprehensive knowledge of how areas of business, such as architecture and infrastructure, integrate to accomplish business goals Ensure data integrity and security by implementing rigorous validation and encryption standards. Build and maintain CI/CD pipelines for automated testing and deployment of data jobs. Legacy Systems Decommissioning: Lead the strategic migration of data and logic from legacy platforms (e.g. on-premises SQL Servers) to a modern Data Lakehouse environment. ETL/ELT Transformation: Re-engineer existing stored procedures and complex legacy ETL jobs into scalable, distributed processing frameworks using Spark (Python) and Starburst/Trino. Validation & Parity Testing: Design and implement automated frameworks for Data Parity Testing to ensure 100% accuracy and consistency between legacy outputs and new big data results. Schema Evolution: Map and transform rigid, legacy relational schemas into flexible, high-performance formats optimized for the cloud (e.g., Parquet, Avro, or Iceberg). Phased Cutover Management: Orchestrate a phased migration strategy (Parallel Run, Shadow Execution) to ensure zero downtime for downstream business applications and reporting tools. Performance Benchmarking: Establish performance baselines on legacy systems and ensure the new Big Data architecture meets or exceeds those benchmarks at scale. Resolve variety of high impact problems/projects through in-depth evaluation of complex business processes, system processes, and industry standards Provide technical mentorship and conduct code reviews for junior and mid-level engineers. Serve as advisor or coach to mid-level developers and analysts, allocating work as necessary. Translate complex business requirements into technical specifications. Collaborate with Product Managers to ensure data availability for downstream analytics, business models and users Partner with multiple management teams to ensure appropriate integration of functions to meet goals as well as identify and define necessary system enhancements to deploy new products and process improvements Highly experienced and skilled technical lead with 12+years of experience with software building and platform engineering. Experience in Data Engineering, focused on Big Data ecosystems. Knowledge in Hadoop, YARN, Hive, Impala, Spark, and Spark SQL with extensive high volume of data processing pipeline development. Programming Expert level and hand on experience in Python. Familiarity with data formats like Avro, Parquet, CSV, JSON. Hands-on experience in writing SQL queries. Highly experienced with Unix based operating systems and shell scripting. Experience with source code management tools such as Bitbucket, Git etc. Big Data Tech Proficiency and hands-on in Hadoop, Spark, Hive, Kafka, and NoSQL databases (MongoDB, HBase). Experience working with query engines like Trino, Presto, Starburst Strong computer science fundamentals in data structures, algorithms, databases, and operating systems. Reverse Engineering, ability to read "spaghetti" SQL or old scripts and document the business logic before moving it. Data Lineage, Experience using tools (like Collibra or Informatica) to track where data comes from and where it's going. Change Management, Experience managing the technical "shock" to the business when switching from legacy BI tools to modern query engines like Starburst. Problem Solver: You don't just fix bugs; you identify the root cause to prevent recurrence. Communicator: You can explain the "why" behind a technical decision to non-technical stakeholders. Automation and AI Mindset: You believe that if a task has to be done twice, it should be automated. Familiarity with AI tools to expedite deliveries. ------------------------------------------------------ PySpark. ------------------------------------------------------

Skills

PythonCI/CDSQLMongoDBSparkHadoopData EngineeringETLGitChange Management

Similar Jobs

30

Big Data

AccionLabs · Singapore, Singapore

1+ year ago

Big Data

AccionLabs · Singapore, Singapore

1+ year ago

Lead Software Engineer - Big Data

Airbus · Bangalore (Airbus), India

Today

DevOps Data Hadoop Engineer – Kubernetes, CI/CD & Big Data

Synechron · Bengaluru - GTP, India

Today

Lead Software Engineer- Big Data Python /Java , Databricks

JPMorgan Chase · Houston, TX, United States

2 days ago

Lead Software Engineer- Big Data Python /Java , Databricks

JP Morgan Chase · Houston, TX, United States

2 days ago

Senior Lead Software Engineer-Big Data Python /Java , Databricks

JPMorgan Chase · Houston, TX, United States

3 days ago

Senior Lead Software Engineer-Big Data Python /Java , Databricks

JP Morgan Chase · Houston, TX, United States

3 days ago

Senior Associate Big Data Engineering - Atlanta, Chicago, Hybrid (4 days) Atlanta

Referral Publicisgroupe · Atlanta, GA, US · Remote, Hybrid

3 days ago

Senior Big Data Developer

Barclays · Gemini Building A, Prague, Czechia

3 days ago

Big Data Engineer

EXL Talent Acquisition Team · Pittsburgh, Pennsylvania, United States, US · Hybrid

3 days ago

Big Data Engineer

EXL · Pittsburgh, Pennsylvania, United States, US · Hybrid

3 days ago

Lead Big Data Engineer - VP

Barclays · Building 100-Whippany Campus, Jefferson Park, United States of America

4 days ago

Big Data Development Engineer Intern

Tencent · Singapore-CapitaSky · Onsite

4 days ago

Senior Big Data Developer

citibank · Tampa, FL,US, US

5 days ago

Senior Big Data Developer

Citi Bank · 3800 CITIGROUP CENTER DRIVE BUILDING G TAMPA, United States of America · Hybrid

5 days ago

Senior Big Data Engineer

Qualys Careers · Pune, India

5 days ago

Big Data Support Engineer

Absa · 15 Alice Lane, South Africa · Hybrid

6 days ago

Cloud/Big Data Engineer

TransUnion · Chennai, India · Remote, Hybrid

6 days ago

Software Engineer I, MLOps (Python/Big Data)

PNC Bank · Firstside Center Bldg (PA373), United States of America +1 · Onsite

1 week ago

Senior Engineer, Big Data

Molina Enterprise · United States, US

1 week ago

Big Data PySpark Lead Engineer - Vice President

citibank · Jersey City, NJ,US, US

1 week ago

Big Data PySpark Lead Engineer - Vice President

Citi Bank · 480 WASHINGTON BOULEVARD JERSEY CITY, United States of America · Hybrid

1 week ago

Big Data Engineer

Energy Products and Services · Kochi, Kerala, India

1 week ago

Big Data Analyst - Data Solutions

Azira Careers Page · Bengaluru, India · Remote

1 week ago

Product Manager, Big Data & AI Solutions

Kyivstar · All · Remote

1 week ago

Senior Big Data Engineer

Gaming Innovation Group · Malta · Remote, Hybrid

1 week ago

Lead Software Engineer - Big Data

JPMorgan Chase · Plano, TX, United States, US

1 week ago

Lead Software Engineer - Big Data

JP Morgan Chase · Plano, TX, United States, US

1 week ago

Senior Big Data Engineer Pyspark Databricks - Vice President

citibank · Pune, MH,IN, IN

1 week ago