Hiring.Camp

Senior AI Ops Engineer

Drivenets

·

Yesterday

Location
Tel Aviv
Workplace
Hybrid
Department
Engineering
Seniority
Senior

Description

Hybrid | IsraelAbout the CompanyDriveNets is a leader in large-scale networking solutions for AI infrastructure and service providers. The company's disaggregated networking architecture transforms the economics of large-scale infrastructures while maximizing performance, utilization, and operational efficiency. Its high-performance AI fabric maximizes GPU utilization and accelerates deployments by optimizing the AI stack end-to-end, resulting in higher tokens-per-second and lower cost-per-token. DriveNets' solutions power production networks for global tier-1 operators like AT&T and Comcast, and scale multi-vendor AI infrastructures at foundation model labs, NeoClouds, and enterprises.ResponsibilitiesDeploy and tune vLLM inference servers across customer environments on diverse GPU hardware (H100, A100, L40S, MI300X), including fully air-gapped, offline deployments.Optimize inference performance end-to-end - batching, tensor/pipeline parallelism, KV-cache and prefix caching, quantization (FP16/INT8/AWQ/GPTQ), and hardware-specific tuning across NVIDIA and AMD architectures.Maintain the LiteLLM gateway and Helm chart variations across Kubernetes flavors (OpenShift, EKS, AKS, GKE, bare metal).Own the model lifecycle - registry, canary rollouts, upgrade/rollback playbooks, and benchmarking of new open-weight models against our security-reasoning workloads.Build evaluation and regression-detection infrastructure to track finding quality, precision/recall, and cost-per-finding over time.Own observability for the AI stack - self-hosted Langfuse tracing, Grafana/OTel dashboards, and alerting on latency, error rates, and GPU health.Define AI infrastructure readiness for new customer deployments - GPU/driver validation, capacity planning, and standardized onboarding runbooks. Technical SkillsHands-on experience running LLM inference at scale (vLLM or similar) in production.Deep GPU knowledge - CUDA, NCCL, multi-GPU topologies, and attention-backend/quantization tuning (FlashAttention, PagedAttention, AWQ/GPTQ).Strong Kubernetes and Helm experience across managed and self-hosted clusters (OpenShift, EKS, AKS, GKE, bare metal).Experience with LLM gateways/routing (LiteLLM or equivalent) and observability tooling (Langfuse, Grafana, OTel, Prometheus).Solid scripting/automation skills (Python) for building eval harnesses and benchmarking pipelines.Familiarity with model lifecycle management - versioning, canary rollouts, and rollback strategies.Soft SkillsStrong ownership mentality - comfortable being the sole owner of a critical, customer-facing system.Able to work independently in constrained, air-gapped, or highly regulated environments.Clear communicator who can translate infrastructure decisions into customer-facing runbooks and dashboards.Structured, benchmark-driven approach to performance and cost optimization.Nice to Have / AdvantageExperience with AMD GPU architectures (ROCm/MI300X) alongside NVIDIA.Background in speculative decoding or advanced inference-serving research.Prior experience in security-sensitive or telecom/service-provider environments.Experience standing up BYOC (Bring Your Own Cloud) deployment models for enterprise customers.If your experience is close but doesn't fulfil all requirements, please submit your application. DriveNets is on a mission to build a special company comprised of individuals with different backgrounds, perspectives, and experiences.DriveNets is an equal opportunity employer. We do not discriminate based on upon race, religion, national origin, sexual orientation, gender identity, gender expression, age, status as a protected veteran, status as an individual with disability, or other applicable legally protected characteristics.More About DriveNetsBased in Israel with extended teams located in the US, Japan, and Romania, DriveNets operations cover more than twelve countries globally. Powering production networks for global tier-1 operators, DriveNets is a leader in large-scale networking solutions for AI infrastructure and service providers. Visit our website to learn more:https://drivenets.com/company/

Skills

PythonKubernetes

Similar Jobs

17

Senior Architect AI Ops & Data Intelligence (m/f/d)

Freseniusglobal · Bad Homburg (EK1), Germany

Yesterday

Senior Product Manager - AI, NBS AI and OPS

Amazon · Remote

4 days ago

Senior AI Ops Engineer

Bottomlinetechnologies · India

1 week ago

Senior Engineer - AI/ML -Ops (6-8 Yrs)

Aptiv · IND - Technical Center India, Chennai

1 week ago

Sr Business Development & Ops Lead, Robotics & Physical AI

Worldlabs · San Francisco +1

2 weeks ago

Digital Technology Senior Specialist – Observability & AI Ops

Bakerhughes · IN-Maharashtra-Pune-7th Floor, Tower 1 & 2, Phoenix Millenium Towers S No. 132, 23. Pune-Bangalore Highway, India +1 · Hybrid

3 weeks ago

Senior Principal AI Engineer, Agentic AI Platform & Control Tower Ops

vrtx · 5000 - Vertex US - Fan Pier, United States of America

1 month ago

Sr AI Engineer - Advanced AI (ML Ops, LLMs, Agentic workflows)

Target · 7000 Target Pkwy N,NCD-0375 Brooklyn Park,MN 55445, United States of America +1

1 month ago

Sr Manager, Annotation Ops Integration, GO-AI

Amazon · Remote

1 month ago

Senior AI OPS Engineer

JCS Solutions LLC · Arlington, VA

2 months ago

Senior Cloud Operations Engineer - AI Cloud Ops

Ptc · USA-San Ramon, CA, United States of America · Hybrid

2 months ago

Sr Software Engineer – AI Ops

Southwest Careers · India Office

3 months ago

AI Ops Senior Backend Engineer

Bigcommerce · Austin, TX, United States of America · Hybrid

3 months ago

Sr. Manager, Global Ops AI Product Owner

onsemi · San Jose, CA, United States, US

4 months ago

Senior AI/ML Ops Engineer - New York / Jersey City

Photon Group · United States, US

5 months ago

Sr. AI Ops/HR Tech Engineer

Trendmicro · Taipei, Taiwan · Remote

1+ year ago

Consultant- Sr AI Ops Engineer

Aiserajobs · Bangalore +1

1+ year ago