- Salary
- $125k – $284k
- Location
- US-California-Palo Alto, United States of America
- Workplace
- Onsite
- Type
- Full-time
- Department
- Engineering
- Seniority
- Senior
- Experience
- 5+ years
- Education
- Master
- Source
- Workday
Description
About the Hiring Team
Tencent Overseas IT has the mission to empower Tencent’s rapid global growth with future ready, global IT platforms, applications and services. We are chartered to lead the Overseas IT strategy, architecture, roadmap and execution. Satisfying our internal/external customers and becoming a world class global IT team are our top aspirations.What the Role Entails
Role SummaryWe are looking for an experienced Senior AI Infrastructure Engineer to join our team. This role owns the technical evaluation and end-to-end execution of AI infrastructure deployments — from data center due diligence and solution review, through cross- functional delivery coordination, to ongoing operations of production AI environments. You will act as a key technical owner bridging internal stakeholders and external partners, ensuring infrastructure is delivered on time, to spec, and operated reliably at scale.
Key Responsibilities
• Conduct on-site data center assessments to evaluate whether candidate facilities meet AI infrastructure requirements.
• Review and validate AI infrastructure solution designs, identifying technical risks, gaps, and cost/performance trade-offs before sign-off.
• Coordinate with external partners, driving timelines, resolving technical issues, and ensuring deliverables meet internal requirements.
• Partner with internal business and engineering teams to translate requirements into deliverable technical specifications.
• Participate in and eventually take ownership of day-2 operations for live AI
environments — monitoring, incident response, capacity/health checks, firmware and lifecycle management, coordinating hardware maintenance as needed.
• Build and improve operational standards, runbooks, and documentation for
infrastructure delivery and operations to enable consistent execution across regions.
• Provide technical guidance and mentorship to junior engineers/interns on hardware diagnostics, cluster tooling, and best practices.
• Track and report on project status and risks to management.
Who We Look For
Requirements
• Bachelor's or Master's degree in Computer Science, Computer Engineering, Electrical Engineering, or a related field.
• 5+ years of experience in data center infrastructure, infrastructure deployment, or infrastructure operations.
• Solid understanding of high-performance compute server hardware, high-speed
networking (e.g., InfiniBand/RoCE), storage systems, and data center power & cooling fundamentals.
• Proven experience evaluating data center facilities and reviewing technical proposals for compute infrastructure.
• Strong track record coordinating complex, multi-party technical projects to closure, including working effectively with external partners and cross-regional teams.
• Hands-on experience with cluster orchestration and management tools (e.g., Slurm, Kubernetes) and Linux system administration.
• Proficiency in scripting/automation (Python, Bash; Ansible/Terraform a plus).
• Familiarity with observability/monitoring stacks (Prometheus, Grafana, ELK) and DCIM tooling.
• Bilingual proficiency in English and Mandarin highly preferred
• Strong ownership mentality, able to independently drive projects with minimal
supervision.
Preferred Qualifications
• Experience with high-performance storage / parallel file systems (e.g., Lustre,
GPFS/Spectrum Scale, WekaFS, VAST, Ceph) in AI infrastructure environments.
• Experience deploying/operating large-scale AI compute environments in a hyperscale or cloud environment.
• Data center or infrastructure certifications.
Location State(s)
US-California-Palo AltoThe expected base pay range for this position in the location(s) listed above is $124,800.00 to $283,800.00 per year. Actual pay may vary depending on job-related knowledge, skills, and experience. Employees hired for this position may be eligible for a sign on payment, relocation package, and restricted stock units, which will be evaluated on a case-by-case basis. Subject to the terms and conditions of the plans in effect, hired applicants are also eligible for medical, dental, vision, life and disability benefits, and participation in the Company’s 401(k) plan. The Employee is also eligible for up to 15 to 25 days of vacation per year (depending on the employee’s tenure), up to 13 days of holidays throughout the calendar year, and up to 10 days of paid sick leave per year. Your benefits may be adjusted to reflect your location, employment status, duration of employment with the company, and position level. Benefits may also be pro-rated for those who start working during the calendar year.Equal Employment Opportunity at Tencent
As an equal opportunity employer, we firmly believe that diverse voices fuel our innovation and allow us to better serve our users and the community. We foster an environment where every employee of Tencent feels supported and inspired to achieve individual and common goals.