- Location
- Hong Kong
- Type
- Full-time
- Department
- Engineering
- Closing date
- Today
- Source
- CareersPage
Description
The Role
As a Data Engineer you'd be working with us to design, maintain, and improve various analytical and operational services and infrastructure which are critical for many other functions within the organization. These include the data lake, operational databases, data pipelines, large-scale batch and real-time data processing systems, a metadata and lineage repository, which all work in concert to provide the company with accurate, timely, and actionable metrics and insights to grow and improve our business using data. You may be collaborating with our data science team to design and implement processes to structure our data schemas and design data models, working with our product teams to integrate new data sources, or pairing with other data engineers to bring to fruition cutting-edge technologies in the data space.
Our Ideal Candidate
We expect candidates to have in-depth experience in some of the following skills and technologies and be motivated to build up experience and fill any gaps in knowledge on the job. More importantly, we seek people who are highly logical, with a balance of respect for best practices and using their own critical thinking, adaptable to new situations, capable of working independently to deliver projects end-to-end, communicates well in English, collaborates effectively with teammates and stakeholders, and eager to be on a high-performing team, taking their careers to the next level with us.
Highly relevant: (ideally familiar with at least one of the technologies in most of the below categories)
- General computing concepts and expertise: Unix environments, networking, distributed and cloud computing
- Python frameworks and tools: pip, pytest, boto3, pyspark, pylint, pandas, scikit-learn, keras
- Workflow scheduling and monitoring tools: Apache Airflow, Luigi, AWS Batch
- Columnar and big data databases: Athena, Redshift, Vertica, Hive/Hadoop
- Container management and orchestration: Docker, Docker Swarm, ECS, EKS/Kubernetes, Mesos CI / CD tools: CircleCI, Jenkins, TravisCI, Spinnaker, AWS CodePipeline
- Distributed messaging and event streaming systems: Kafka, Pulsar, RabbitMQ, Google Pub/Sub
- Streaming data processing frameworks: Spark Streaming, Apache Beam, Apache Flink
- General AWS or Cloud services: Glue, EMR, EC2, ELB, EFS, S3, Lambda, API Gateway, IAM, Cloudwatch
- Version control: git commands, branching strategies, collaboration etiquette, documentation best practices Agile/Lean project methodologies and rituals: Scrum, Kanban