Hiring.Camp

Network Engineer – AI Network & Security

Firmus Technologies

·

Today

Location
Singapore
Department
Network
Experience
8+ years
Source
Greenhouse

Description

Firmus Technologies

Firmus Technologies is a global leader pioneering the development and operation of efficient AI infrastructure across Asia Pacific.  

Founded in Australia in 2019, our mission is to create the most efficient AI infrastructure by combining cutting-edge technology with a steadfast commitment to sustainability. 

At Firmus, we are unique in our approach. We design, build, and operate a new class of digital infrastructure – the AI Factory. Through our model-to-grid technology approach, we have pushed the boundaries of multi-generational liquid cooling systems, energy management, AI software orchestration, and construction. For our customers, this approach allows us to make every watt count and deliver low-cost AI tokens globally. 

 

Firmus AI Cloud

Our large-scale GPU cloud platform, Firmus AI Cloud, is purpose-built to deliver energy-efficient AI compute at scale to customers. 

It empowers developers, enterprises, educational institutions, and government users to train and deploy AI models with unmatched efficiency and cost savings. With an ever-growing suite of services and applications, we are committed to delivering a cloud experience that is market-leading, proprietary, and built to scale. 

 

Why you’ll love working here

As an NVIDIA Cloud and Engineering partner in Asia Pacific, you will gain skills, experience, and exposure across the AI industry and be part of shaping what this industry looks like for  decades to come.  

We are founder-led, not a big corporate. Decisions happen fast, our leaders are accessible,  and there's minimum bureaucracy between you and the work.  

Ownership comes early. Whatever your role, you will have a direct line to outcomes, helping shape how the business grows as we scale nationally across a long-term, large-scale roadmap. 

Work alongside founders and experts in AI infrastructure, energy systems  and next-generation compute.  

What we build here has impact beyond the business. Our AI Factories are designed to operate as assets to the energy grid to actively strengthen the communities and regions they operate in rather than drawing from them.  

 

Role Summary

Firmus Technologies is seeking a skilled Network Engineer / Sr. Network Engineer to join our Engineering and Technology team. The ideal candidate will play a crucial role in leading the development of our network designs and supporting the network deployment for AI infrastructure projects. This role offers an exciting opportunity to work at the forefront of AI networking technology and contribute to the growth of AI infrastructure.

 

Key Responsibilities

  • Deploy and Operate High-Performance Networks
    • Deploy and maintain low-latency, high-throughput interconnects (e.g.: Ethernet 100/200/400GbE) for HPC and AI workloads.
    • Optimise performance across multi-node clusters, high-speed storage fabrics and parallel computing AI environments.
    • Respond to and resolve escalated network issues, outages and performance degradations across the SMC Corporate and Compute network infrastructure.
    • Analyse logs, run diagnostics and coordinate with vendors, carriers as needed.
    • Work with internal observability team to setup and maintain monitoring tools to proactively identify bottlenecks, errors and abnormal behaviours.
    • Analyse trends for bandwidth, hardware utilisation, and growth to inform scaling and make recommendations to procurement decisions.
    • Participate in the operations standby roster and on-call from time to time, and work closely with other network engineers as a team.
    • Design and test redundant paths, failover mechanisms and DR playbooks to ensure uninterrupted connectivity during outages or maintenance.

 

  • InfiniBand/Ethernet and RDMA Networking
    • Design, monitor and troubleshoot InfiniBand/Ethernet fabrics and RDMA-enabled transport layers.
    • Maintain subnet managers (e.g.: Unified Fabric Manager and OpenSM) and ensure fabric health and topology visibility.
    • Perform verification and acceptance tests to the new fabric and new network infrastructure setups.
    • Understand various InfiniBand / RoCE optimisation protocols and mechanisms (e.g.: SHARP, CollNet) and use performance tools (e.g.: ibstat, perfquery, ib_write_bw, nccl, etc) to monitor fabric health, congestion and link errors.
    • Enable optimised networking for AI frameworks and work closely with GPU compute environments to ensure efficient multi-node communication.
    • Perform deep dive diagnostics to resolve layer 1-4 issues across HPC and AI workloads.

 

  • Network as Code and Automation
    • Develop and maintain automated network configurations using Infrastructure as Code (IaC) tools (e.g.: Ansible, Netbox, Python scripts).
    • Implement CI/CD pipelines for network changes to improve speed, consistency, and auditability.
    • Automate routine tasks such as provisioning, backups and compliance checks.

 

  • Network Security and Policy Enforcement
    • Implement and manage firewall and security devices. Apply firewall rules, segmentation, and Secure-by-default, Security-by-design principles to safeguard internal and external connectivity.
    • Collaborate with Security and Risk team to enforce policies and respond to security incidents.
    • Design and implement zero-trust network architecture across corporate and AI compute fabric, including segmentation and least-privilege access.
    • Implement authentication and access control (802.1X, NAC, MACsec) for fabric and out-of-band management networks.
    • Own network security posture for RDMA/RoCEv2 and InfiniBand fabrics, including multi-tenant workload isolation.
    • Support vulnerability management: coordinate scanning, patching cadence, and remediation tracking for network infrastructure.
    • Participate in security incident response — detection, containment, root cause analysis, and post-incident reporting — with the SMC Security and Risk team.
    • Contribute to compliance and audit activity (e.g. ISO 27001, SOC 2) relating to network controls.
    • Integrate network telemetry and logs with SIEM/observability tooling for security monitoring.

 

  • Project Management and Stakeholder Management
    • Support the deployment team in their project management and resource allocation for the network portion of AI cluster installations.
    • Collaborate and work closely with the Global Operations Centre, Software Defined Infrastructure team, Data Centre Infrastructure team and Solution Architects to support deployments and to maintain SLAs.
    • Work closely with both the Firmus Engineering and Commissioning teams to align network infrastructure with customers’ requirements.
    • Facilitate knowledge sharing and communication between teams and create and maintain comprehensive technical documentation.
    • Maintain and build strong relationships with key technology partners and vendors and proactively manage and coordinate partner engagement on site.

 

  • Technology Expertise
    • Familiar with SASE cloud framework together with Zero Trust Network Access and Privilege Access Management
    • Maintain and expand expertise in physical network hardware and advanced networking technologies, including (i) NVIDIA InfiniBand; (ii) Spectrum Ethernet Platform, and (iii) RDMA over Converged Ethernet (RoCE).
    • Familiarity with open-source network operating systems such as Cumulus Linux and Sonic.
    • Evaluate new technologies (RoCEv2, DPU/SmartNICS or quantum interconnects) for next-gen HPC fabrics.
    • Benchmark and validate new switches, nics, transceivers or firmware in lab environments before production roll-out.
    • Maintain and expand expertise in NetDevOps practices to maintain software defined network infrastructure.
    • Provide technical support and troubleshooting for advanced networking technologies, escalating to vendors as needed.

 

Skills & Experience

  • Bachelor’s degree in network engineering, computer science, or a related technical field.
  • 8+ years of experience in network engineering, with a focus on network security.
  • Experience in Linux systems especially host network configuration.
  • Strong project management skills and experienced in complex technical projects.
  • Excellent problem-solving and analytical skills.
  • Ability to work independently and as part of a team.
  • Strong communication skills, both written and verbal.
  • Willingness to undertake international and/or domestic travel for on-site deployments and commissioning as required. 
  • Solid understanding of advanced networking technologies, particularly those related to AI would be highly advantageous.
  • Experience with zero-trust architecture, MACsec, 802.1X/NAC, and network segmentation in multi-tenant environments.
  • Relevant security certifications highly regarded (e.g. CCNP Security, CISSP, Fortinet NSE, Palo Alto PCNSE).
  • Hands-on experience with NVIDIA InfiniBand, Spectrum Ethernet Platform, and/or RDMA over Converged Ethernet (RoCE) preferred.

 

Location & Reporting

  • Singapore
  • Reporting to Senior Manager, Networking

 

Employment Basis

Full-time

 

Diversity

At Firmus, we are committed to building a diverse and inclusive workplace. We encourage applications from candidates of all backgrounds who are passionate about creating a more sustainable future through innovative engineering solutions.

Join us in our mission to revolutionize the AI industry through sustainable practices and cutting-edge engineering. Apply now to be part of shaping the future of sustainable AI infrastructure.

 

Skills

PythonAnsibleCI/CDLinuxSIEMSOCComplianceProcurementProject ManagementSOC 2ISO 27001CISSP
Network Engineer – AI Network & Security at Firmus Technologies | Hiring.Camp