Job Overview
We are seeking a motivated and technically curious Site Reliability Engineer to help build and maintain the reliable and distributed systems that support our business operations.
In this role, you will play a vital part in supporting our Cybersecurity business, Vulcan. Vulcan is a cybersecurity solution for GenAI, providing red and blue team services to ensure compliance and security.
Learn more about us 👉
- Vulcan product: https://vulcanlab.ai/
- Vulcan LinkedIn: https://www.linkedin.com/company/vulcanlab-ai/
- AIFT group: https://aift.io/
Responsibilities
- Implement and enhance system reliability, availability, scalability, performance, and efficiency by leveraging monitoring, alerting, and automation tools on public cloud platforms.
- Participate in capacity planning, analyze software performance, and fine-tune systems to ensure optimal operation.
- Develop and enhance GitLab CI/CD processes and toolset to streamline software delivery and deployment.
- Define and monitor key metrics to assess and enhance system reliability.
- Collaborate closely with the engineering team to improve reliability and operational efficiency at every software development life cycle (SDLC) stage.
- Troubleshoot, optimize infrastructure and automate repetitive tasks to increase efficiency and effectiveness
Requirements
- Strong expertise and experience in cloud technologies.
- Advanced knowledge of monitoring solutions like Prometheus, Grafana, ELK (Elasticsearch, Logstash, Kibana).
- Experience in the complete software development life cycle (SDLC).
- In-depth understanding of network concepts, particularly with a focus on security.
- Hands-on experience implementing GitLab CI/CD processes.
- Proficiency in automation platforms like Ansible and Terraform.
- Knowledge of orchestration tools like Kubernetes.
- Familiarity with container technologies like Docker.
- Experience with Git source code version control systems.
- Experience with AI pair programming like OpenAI.
- Proficiency in programming languages such as Bash, Python, or Go.
- Experience and capability in executing client-side/on-premise deployments is a strong plus.
Interview Process
- HR phone interview: 1 hour
- Onsite Interview: 1.5~2 hours, meet with hiring team and HR
Other Benefits
To us, people are our greatest asset, and we are more than happy to invest in employees! We create a healthy work atmosphere and provide you with the tools and support for doing your job successfully. With a culture of flexibility and transparency, we believe there should be no barriers, and everyone’s contributions matter.
Work Life Balance is a must
- 15 days annual leaves (pro-rata for partial month at first year)
- 5 days full-pay sick leaves, 3 days menstrual leaves
- Health check subsidy
- Ergonomic-design chair and fully-equipped devices for work
Grow together & keep learning
- Conferences & external subsidy
- Learning clubs to share technical skill (e.g: Frontend/Backend tech sharing, Product Management...etc)
Work Hard, Play even Harder
- Various entertainment & sports clubs, attend basketball clubs today, and play board game tomorrow!
- Snacks & beverage to refill your energy anytime