- Location
- IN-UP-Noida-Candor TechSpace Tower 1, India · IN-HR-Gurgaon-Candor TechSpace Tower 8
- Workplace
- Remote, Hybrid
- Type
- Full-time
- Department
- IT
- Seniority
- Lead
- Closing date
- Today
- Source
- Workday
Description
JD – Data Engineering (ETL) Tech Lead
Experience & Expectations :
- Leverage extensive experience (8 to 12 years overall ETL experience, including 2–3 years in a Technology Lead role & 4–6 years in AWS) to drive solution design and delivery.
- We are seeking an experienced Tech Lead with strong expertise in Big Data (Spark, Cloudera).
- Experience orchestrating complex workflows using AWS Step Functions (state machines) for reliable and scalable data pipelines.
- Ability to design end-to-end serverless data architectures integrating Glue, Lambda, S3, and Redshift
- Candidates with exposure to drive / assist in architecture, development, and delivery of AI-powered data workflows, leveraging agent-based systems to automate data ingestion, transformation, validation, and orchestration along with modern AI paradigms (Agentic AI, LLMs) to design and lead next-generation intelligent data platforms and ETL pipelines will be preferred.
Core Responsibilities :
- Lead and manage the development team by helping them understand the requirements and provide technical guidance
- Design, build, and maintain high volume ETL/ELT pipelines across Hadoop (HDFS, Hive, Spark, Kafka) and AWS (Glue, EMR, Lambda, Step Functions, Redshift).
- Develop distributed data processing solutions using PySpark, Spark SQL, and scalable cloud serverless patterns.
- Implement reusable data ingestion frameworks for batch, ability to design & implement Orchestration process and Leverage AI
- Optimize data workflows using partitioning, bucketing, compression, file formats (Parquet/ORC).
- Understanding hybrid data lake architectures using S3 + HDFS, ensuring data governance and best practices are adheres
- To lead technical teams and deliver complex projects in an Agile environment
- Design and build the robust, scalable and secure software solutions across the having no/least adoption
- Define clear technical specifications and make architecture decisions that align with business goals and long-term scalability.
- Implement best practices (including secure code guidelines) through the implementation of unit tests, automation, leverage and code reviews. Drive continuous improvement in code quality and maintainability.
- Troubleshooting issues and proactively solving problems as they arise, ensuring the smooth operation of full stack applications
- Ability to understand the data flow diagram, data modelling and Lineages
- Job orchestration using Airflow, Control M, Step Functions, or event-driven triggers.
- Ensure data is protected and compliant with regulatory standards.
- Work closely with business stakeholders to enable high quality datasets.
- Provide technical support in architecture decisions, code reviews, and best practice adoption and provide technical guidance to peers/juniors in team.
- Own deployment, incident response, and post-incident reviews for production environments, troubleshooting Spark performance issues, job failures, and cluster bottlenecks.
- Optimize cost and usage of AWS resources and recommend architecture improvements.
- Collaborate closely with developers, QA and cross product teams to streamline release processes.
- Best Practices Advocacy: Advise teams on CI/CD pipelines, observability, security compliance, and modern development practices.
- Collaborate with product managers and stakeholders to align technical roadmaps with business strategy
- Expertise in architecture governance, security frameworks, and scalable system design across large organizations
Technical Skills :
- Strong experience with the AWS data stack (S3, Glue, EMR, Lambda, Kinesis, Redshift, Step Functions etc.,).
- Strong hands-on expertise in Scala, PySpark, Spark optimization techniques, HiveQL, and distributed computing.
- Good understanding of Hadoop ecosystem (HDFS, Hive, Spark, YARN, Kafka).
- Good work experience in SQL in hive and impala
- Proficiency in at least one scripting/programming language: Python, Shell scripting.
- Strong experience with CI/CD, GitHub, Git commands.
- Expertise in ETL and Data Warehousing and cloud concepts.
- Good understanding of data modelling (star/snowflake), partitioning strategies, and schema evolution.
- Expertise in data profiling and decision making.
- Able to understand, design and create data flow diagrams and do data modelling. (knowledge of Miro will be added advantage)
- Able to understand the architecture and design end-to-end data flow.
- Hands-on experience with Airflow, or Control‑M, or other orchestrators.
- To monitor and support BAU and year end activities, if needed.
- Well versed with security and compliance aspects in Cloud.
- Familiarity with serverless patterns and containerization (Docker, ECS/EKS).
- Experience with monitoring/logging tools and incident management practices.
Other Requirements
- Strong logical and analytical, problem-solving, and communication skills.
- Communicate effectively and concisely with multiple stakeholders and coordinate and collaborate with cross functional teams.
- Ability to support both legacy Hadoop workloads and cloud-first architectures.
- AWS certifications (Data Engineer, Solutions Architect, or Developer) are a plus.
- Strong leadership and mentoring abilities
- Detail-Oriented and proactive in problem-solving and issue resolution
We offer you a competitive total rewards package, continuing education & training, and tremendous potential with a growing worldwide organization.
DISCLAIMER:
Nothing in this job description restricts management's right to assign or reassign duties and responsibilities of this job to other entities; including but not limited to subsidiaries, partners, or purchasers of Alight business units.
Skills
PythonScalaAWSDockerCI/CDSQLSparkHadoopAirflowData EngineeringETLGitGitHubAgileEMRCompliance