- Location
- Mumbai, MH,IN, IN
- Type
- Full-time
- Department
- Engineering
- Source
- Eightfold
Description
End-to-end administration and operations of CNIS platforms including Kubernetes, container runtimes, and cloud-native infrastructure components. Manage and optimize Cloud Container Distribution (CCD), covering provisioning, upgrades, backup/restore, lifecycle management, and disaster recovery readiness. Perform CCD health checks, configuration drift analysis, capacity tracking, and environment validation across clusters. Work closely with app and platform teams to onboard workloads and provide CCD-level support. Monitor system performance and availability using tools like Prometheus, Grafana, and other tools Maintain detailed runbooks, MOPs, and CNIS documentation for operational readiness. Define and govern CNIS architecture, platform standards, and design principles for Kubernetes-based environments. Lead design reviews and capacity planning for CNIS and CCD platforms. Define platform lifecycle strategies covering deployment, upgrades, backup/restore, disaster recovery, and decommissioning. Establish multi-cluster architecture, workload placement strategies, and platform segmentation models. Drive adoption of Infrastructure-as-Code (IaC), GitOps, and automation frameworks. Define observability architecture using Prometheus, Grafana, logging, tracing, and monitoring platforms. Lead platform modernization and transformation initiatives. Lead DR planning, execution, and RCA of major incidents. SDI (Software Defined Infrastructure) Administration: Design and govern Software Defined Infrastructure architectures spanning compute, storage, and networking domains. Good Hands on experience of SDN, Calico, Multus, service networking, and container networking solutions. Design scalable storage architectures leveraging Ceph, SDS, CSI drivers, and persistent storage solutions. Lead infrastructure integration across CNIS, SDI, storage, and networking layers. Enterprise Networking: Understanding of SDN and integration with container networking layers (e.g., Calico). Familiarity with traffic filtering, isolation, and multi-tenant CNIS network architecture. Switch boot process, firmware upgrades, backup/restore,Border leaf switch management and troubleshooting Configuration of VLANs, trunking, port profiles (access/trunk), and ACLs Performance analysis, component replacement, switch stacking, and chassis management Linux System Administration: Expertise in Red Hat Enterprise Linux (RHEL) system administration and troubleshooting. o Cluster-level configurations o LVM management o Multipathing (MPIO) setup and troubleshooting o Patch management, firmware upgrades, and OS-level tuning o Scheduling and validating backups/recoveries o Performance troubleshooting, kernel-level analysis, and log auditing. o Hands-on with SAN topology, host-side storage integration, and network storage troubleshooting. o Hardware troubleshooting experience on HP ProLiant, Dell PowerEdge servers.