- Location
- Remote, India
- Workplace
- Remote
- Type
- Full-time
- Department
- Engineering
- Education
- Bachelor
- Closing date
- Today
- Source
- Workday
Description
Lead DevOps & Cloud Infrastructure Engineer
What this Job Entails
This is a hands-on lead role owning the cloud infrastructure and DevOps function for Astreya's product portfolio. The person in this seat is accountable for how our products get built, deployed, secured and run — in our own environments and inside customer environments.
Google Cloud is the primary platform; Azure is secondary. This is not a role that sets direction and hands the work to someone else: the expectation is that this person writes the Terraform, debugs the cluster, runs the deployment, and then mentors the engineers who will do it next time. They act as the subject matter expert for infrastructure across product, engineering and delivery teams, and as the technical owner in customer conversations about deployment, security and hosting models.
Scope
- Owns the infrastructure and deployment layer end to end, from CI pipeline to production runtime, across multiple products on independent release cadences
- Resolves complex problems where the diagnosis requires in-depth evaluation of many interacting variables — networking, identity, cluster behaviour, managed service limits, cost
- Exercises independent judgment in selecting platform services, patterns and tooling, and is expected to defend those choices technically and commercially
- Operates as a player-coach: a small team reports to this role, and the role remains individually productive
Your Roles and Responsibilities
Cloud infrastructure (primary)
- Design, build and operate our GCP footprint: GKE, Artifact Registry, Cloud SQL (Postgres), Cloud Storage, Memorystore (Redis), Vertex AI, Cloud Monitoring/Logging, Cloud DNS
- Own network and perimeter design: VPC architecture, Private Service Connect, Cloud NAT, Cloud Armor, private connectivity to customer systems
- Own secrets, keys and identity: Secret Manager, Workload Identity, IAM design and least-privilege service accounts, KMS, certificate lifecycle
- Run R&D on platform services we haven't used yet — evaluate the GCP service, identify the Azure and AWS equivalents, and produce a recommendation with a working proof of concept, not a slide
- Manage cloud cost: attribution by product and environment, budget alerts, rightsizing, commitment planning
Deployment into environments (a core reason this role exists)
- Make deploying our products into a new environment a repeatable, documented, low-drama exercise — both customer-VPC and Astreya-hosted models
- Own the deployment artefacts: Helm charts, namespace and pod topology, environment configuration, migration and rollback paths
- Work directly with customer infrastructure and security teams during onboarding — answer their architecture and security questions, adapt to their constraints, and get us to production
- Produce and maintain the evidence infrastructure security reviews and audits ask for: architecture diagrams, data flow, access controls, encryption posture, audit logging, retention
DevOps and platform engineering
- Own CI/CD across products: build pipelines, artefact promotion, environment strategy, source control branching strategies across multiple release cadences
- Write automation and internal tooling for provisioning and operating infrastructure — infrastructure as code (Terraform) as the default, scripts and services where it isn't
- Build the observability layer: metrics, logs, traces, alerting, dashboards and SLOs that let us find problems before customers do
- Improve scalability, reliability, capacity and performance of the platform, including the infrastructure serving AI/LLM workloads
- Harden the platform: image scanning, dependency and vulnerability management, network policy, secrets hygiene, patching
Leadership and process
- Lead and mentor DevOps and infrastructure engineers; set standards and review their work
- Work with product owners and engineering leads to understand requirements, surface infrastructure bottlenecks early, and propose resolutions
- Own incident response for platform issues and produce root cause analyses for outages that are honest about cause and specific about prevention
- Design and implement improvements to existing support processes and tooling; introduce innovations and follow through on execution, not just proposal
- Document decisions, runbooks and resolution history so the next person doesn't rediscover them
- Other duties as required. This list is not a comprehensive inventory of all responsibilities assigned to this position
Required Qualifications/Skills
- Bachelor's degree (B.S/B.A) from a four-year college or university and 8+ years' related experience and/or training; or an equivalent combination of education and experience
- Deep, hands-on Google Cloud experience — has designed and run production workloads on GCP, not just passed a certification
- Strong Kubernetes experience in production: GKE preferred, including networking, ingress, autoscaling, resource management and debugging failures under load
- Infrastructure as code at a professional standard — Terraform, module design, state management, review discipline
- Strong scripting and coding ability (Python, Go, Bash or equivalent) and enough software development experience to work inside application repositories, not just around them
- Practical cloud networking and security depth: VPC design, private connectivity, IAM, secrets management, TLS
- CI/CD ownership across multiple products and release trains, with source control branching strategies
- Monitoring, alerting and incident management experience, including writing the RCA afterwards
- Experience with open-source tooling in large distributed systems
- Good communication skills — can hold a technical conversation with a customer's security team and a working conversation with a developer on the same day
- Ability to pick up unfamiliar technology quickly and carry several threads at once
Preferred Qualifications
- Working Azure knowledge (AKS, Key Vault, Entra ID, VNet) and the ability to translate a GCP design onto it; AWS exposure a bonus
- Experience deploying a product into customer-controlled cloud environments as a vendor
- Experience supporting AI/ML workloads in production — Vertex AI, model endpoints, GPU scheduling, inference cost control
- Exposure to SOC 2, ISO 27001 or comparable audits from the infrastructure side
- Experience integrating with enterprise ITSM platforms (ServiceNow) and the connectivity that requires
- Prior experience as a first or early infrastructure hire on a product team
- Relevant certification (Google Professional Cloud Architect, Professional Cloud DevOps Engineer, CKA)
Physical Demand & Work Environment
- Must have the ability to perform office-related tasks which may include prolonged sitting or standing
- Must have the ability to move from place to place within an office environment
- Must be able to use a computer
- Must have the ability to communicate effectively
- Some positions may require occasional repetitive motion or movements of the wrists, hands, and/or fingers