- Location
- IT BUILDING, RAMANUJAN IT SEZ,, India
- Workplace
- Hybrid
- Type
- Full-time
- Department
- IT
- Experience
- 3+ years
- Closing date
- Today
- Source
- Workday
Description
We are seeking an experienced and motivated Java Spark Developer to design, develop, and maintain high-performance, large-scale data processing pipelines and distributed applications. In this role, you will leverage Core Java and Apache Spark to build resilient batch and real-time streaming data architectures, optimize distributed data workloads, and collaborate with cross-functional teams including Data Scientists, Cloud Engineers, and Solution Architects.
1. Data Pipeline & Application Development
- Design, implement, and maintain robust, scalable data ingestion and ETL/ELT pipelines using Apache Spark (Core, SQL, Streaming) written in Java (or Scala interoperability).
- Develop performant, low-latency microservices and distributed processing modules integrated with messaging platforms (e.g., Apache Kafka).
- Build and maintain interfaces to relational databases, distributed data lakes, and NoSQL stores (e.g., Hive, Cassandra, HBase, MongoDB, Delta Lake, Snowflake).
2. Performance Tuning & Optimization
- Profile, debug, and optimize Spark jobs by managing partitioning strategies, caching, broadcast variables, memory allocation (driver/executor memory), and data serialization (Kryo).
- Analyze query execution plans, DAGs, and Spark UI metrics to eliminate data skew, reduce shuffle overhead, and minimize bottleneck latencies.
- Monitor resource utilization on cluster managers such as Kubernetes, Apache YARN, or cloud-native orchestration engines.
3. Architecture & Data Modeling
- Design structured, semi-structured, and unstructured data storage schemas using columnar file formats (e.g., Parquet, ORC, Avro).
- Implement robust data validation, cleansing, data governance, and error-handling mechanisms across the ingestion lifecycle.
- Ensure data privacy and enterprise compliance by applying encryption at rest/transit and access-control policies.
4. Collaboration, CI/CD & Best Practices
- Participate in Agile/Scrum ceremonies, sprint planning, and code reviews to ensure adherence to high code quality standards.
- Write comprehensive unit, integration, and automated regression tests using frameworks such as JUnit, Mockito, and Spark Testing Base.
- Configure and maintain continuous integration and continuous deployment (CI/CD) pipelines using tools like Jenkins, GitLab CI, or GitHub Actions.
Required Qualifications & Skills
Technical Competencies
- Core Java: Deep proficiency in Java (Java 8/11/17+), including multithreading, concurrency, OOP principles, memory management, and JVM internals.
- Apache Spark: Hands-on experience developing distributed applications with Apache Spark (RDDs, DataFrames, Datasets, Spark SQL, Spark Structured Streaming).
- Distributed Ecosystem: Strong working knowledge of distributed architecture (HDFS, YARN), Hive, and distributed storage systems.
- Messaging & Streaming: Practical experience with event streaming platforms such as Apache Kafka or RabbitMQ.
- Database & Query Languages: Advanced SQL capabilities, experience with relational databases (PostgreSQL, Oracle, MySQL) and NoSQL datastores.
- Build & Version Control: Proficiency with build tools (Maven, Gradle) and Git version control workflows.
- Testing: Solid track record in Test-Driven Development (TDD) using JUnit, Mockito, and distributed testing patterns.
Professional Experience & Education
- Education: Bachelor’s or Master’s degree in Computer Science, Information Technology, Software Engineering, or a related technical discipline.
- Experience: 3-6 years of professional software engineering experience, with at least 2–4 years dedicated to building scalable distributed data processing applications using Java and Apache Spark.
Preferred / Desired Qualifications
- Cloud Platforms: Experience building and deploying data architectures on AWS (EMR, S3, Glue, Athena), Azure (Databricks, HDInsight, ADLS), or Google Cloud (Dataproc, BigQuery).
- Modern Lakehouse Technologies: Hands-on exposure to Apache Iceberg, Delta Lake, or Apache Hudi.
- Containerization & Orchestration: Familiarity with Docker, Kubernetes, and workflow schedulers like Apache Airflow or Luigi.
- Polyglot Exposure: Familiarity with Scala or Python (PySpark) is an added advantage.
------------------------------------------------------
Job Family Group:
Technology------------------------------------------------------
Job Family:
Applications Development------------------------------------------------------
Time Type:
Full time------------------------------------------------------
Most Relevant Skills
Please see the requirements listed above.------------------------------------------------------
Other Relevant Skills
For complementary skills, please see above and/or contact the recruiter.------------------------------------------------------
Citi is an equal opportunity employer, and qualified candidates will receive consideration without regard to their race, color, religion, sex, sexual orientation, gender identity, national origin, disability, status as a protected veteran, or any other characteristic protected by law.
If you are a person with a disability and need a reasonable accommodation to use our search tools and/or apply for a career opportunity review Accessibility at Citi.
View Citi’s EEO Policy Statement and the Know Your Rights poster.