- Workplace
- Remote, Onsite
- Type
- Full-time
- Department
- Engineering
- Experience
- 3+ years
- Source
- RecruiterFlow
Description
Our brand, Lennor Metier Consulting, a DOLE-licensed headhunting and recruitment agency in the Philippines, is proud to partner with a multinational company in their search for a Machine Learning Operations Engineer based in Tokyo, Japan.
Salary: up to JP¥15,000,000 per year
Work Setup: Onsite
Shift Schedule: Day Shift
Location: Tokyo, Japan
Your Responsibilities:
- Design, Build, and Operate ML Infrastructure
- Design and operate ML infrastructure using AWS, GCP, Azure, etc.
- Create and optimize Docker images for Python code and ML models
- Orchestrate containers in production environments using Kubernetes (EKS/GKE/AKS)
- Manage infrastructure configuration using Terraform/CloudFormation
- Perform cost estimation and optimization
- Build and Operate MLOps Pipelines
- Design and operate CI/CD pipelines (e.g., GitHub Actions)
- Automate ML pipelines using Airflow, Kubeflow, Vertex AI/SageMaker Pipelines, etc.
- Manage experiments and model registries using tools like MLflow
- Model Deployment and Monitoring
- Deploy model APIs using KServe, SageMaker Endpoints, etc.
- Monitor systems and models using Prometheus, Grafana, Datadog, etc.
- Monitor model performance (latency, accuracy, etc.), log metrics, and design alerting mechanisms
- Automate retraining cycles and propose improvement suggestions
What We're Looking For:
- At least 3 years of development experience in languages such as Python
- Operational experience with cloud platforms such as AWS/GCP/Azure
- Practical work experience with Docker and Kubernetes (EKS/GKE/AKS)
- Experience using IaC tools such as Terraform/CloudFormation
- Experience building and operating CI/CD pipelines using GitHub Actions, etc.
- Experience using MLOps tools such as MLflow, Kubeflow, Vertex AI Pipelines, etc.
- Experience building ML pipelines using Airflow, SageMaker Pipelines, etc.
- Experience deploying models using KServe, SageMaker Endpoints, etc.
- Experience operating monitoring and alerting systems using Prometheus, Grafana, Datadog, etc.
- Experience with distributed processing platforms such as Spark, Hadoop, etc.
- Experience in microservices architecture design and development
- Experience or willingness to learn using AI (including Generative AI) for requirements gathering, documentation writing, and design assistance
- Experience or strong interest in leveraging AI to improve business efficiency or quality
- Development and operational experience in Linux environments
- Foundational knowledge of machine learning or data analysis
- Experience in cost optimization
- Business-level Japanese proficiency (Reference standard: JLPT N2 or above
Ready to take the next step in your career? Submit your application now!
--- We kindly request your patience as we receive a significant number of applications. Rest assured that our team will update your application's status soon. In the meantime, we encourage you to follow our LinkedIn page to stay informed about future opportunities and company updates.