Hiring.Camp

Senior Advisory Software Engineer

Pitneybowes

·

Today

Location
IN Pune, India
Type
Full-time
Department
Engineering
Seniority
Senior
Source
Workday

Description

We’re hiring at Pitney Bowes, where top talent builds meaningful careers and lasting impact. We Move fast, Deliver excellence, and Win together…that’s The Pitney Bowes way. Here, how we work matters just as much as what we achieve.

We’re looking for people who:

  • Act with urgency, accountability, and purpose

  • Deliver high quality work with consistency and pride

  • Collaborate effectively and elevate those around them

  • Focus on outcomes that drive impact and growth

Job Description:

Join Pitney Bowes as Senior Advisory Software Engineer

Years of experience: 9-12 years

Job Location – Pune

Impact

As a Senior Advisory Software Engineer, you will operate at the intersection of platform engineering and AI-driven automation. This is not a traditional SRE role. You will design, build, and supervise agentic systems that detect anomalies, diagnose failures, execute remediation runbooks, and escalate intelligently — with minimal human intervention. You will architect the feedback loops that make our platform progressively self-healing. You will collaborate across engineering, product, and architecture to ensure our observability and incident response capabilities stay ahead of system complexity. Being a Senior Advisory SRE here means you think in systems, build in agents, and measure success in mean-time-to-no-action.

The Job

  • Architect and operate agentic systems for autonomous monitoring, anomaly detection, and self-healing — reducing mean-time-to-remediation without human-in-the-loop dependency for routine failure patterns.

  • Build software and agentic pipelines that manage platform infrastructure autonomously — from drift detection and remediation to capacity adjustment and incident triage.

  • Drive reliability engineering outcomes — SLO attainment, error budget governance, and deployment velocity — by embedding intelligence into the platform rather than adding human process overhead.

  • Measure and continuously optimize system performance using agent-driven telemetry analysis — identifying degradation patterns before they manifest as customer-impacting incidents.

  • Own CI/CD reliability across the SDLC — integrating agentic checks, automated rollback triggers, and intelligent deployment gates that act on signal, not on schedule.

  • Scope includes:

    • Agentic observability — context-aware monitoring with LLM-assisted signal interpretation

    • Intelligent alert design — dynamic thresholds, noise suppression, and automated triage routing

    • Runbook automation and agentic remediation — codifying institutional knowledge into executable, supervised agent workflows

    • Autonomous incident response — agent-led detection, diagnosis, and escalation with human override at defined severity thresholds

    • Infrastructure lifecycle management — provisioning, drift remediation, and cost optimization driven by policy-as-code and agent execution

    • End-to-end configuration, deployment, and patching — governed by automated validation pipelines, not manual checklists

    • Creating and maintaining GIT repo and pipelines

  • Communicate risks, system health, and automation outcomes clearly to engineering leadership and cross-functional stakeholders — translating agent behaviour and reliability signals into business-readable insight.

  • Define and continuously refine the observability strategy — what to monitor, how to act on it, and how to suppress noise programmatically. Drive adoption of agent-assisted monitoring across product and infrastructure layers.

  • Analyse operational behaviour patterns across user personas and platform workloads to inform intelligent automation design and monitoring strategy evolution.

  • Define and govern Service Level Indicators and Objectives — using error budget data to drive engineering prioritisation and calibrate automation intervention thresholds rather than to reactively defend the committed SLA.

  • Define, instrument, and own platform reliability metrics — QoS, Uptime, MTTR, MTBF, and agent automation coverage — as leading indicators of system health and team maturity.

  • Synthesise and publish key metrics to stakeholders

  • Leverage deep AWS expertise and DevOps toolchain knowledge to design infrastructure automation that operates with the reliability and predictability of a software system.

  • Continuously analyse infrastructure and tooling spend — identifying waste, right-sizing opportunities, and cost anomalies through automated FinOps signal processing.

  • Build and operationalise cost governance frameworks where agent-driven policy enforcement, not periodic human review, is the primary control mechanism.

  • Design and implement AI-augmented observability solutions across the stack — integrating SumoLogic, CloudWatch, Grafana, Prometheus, and PagerDuty with agentic reasoning layers that interpret signal and act, not just alert.

  • Lead outage management with agentic support — automated problem detection, structured stakeholder communication, and agent-assisted resolution with clear human escalation protocols for high-severity events.

  • Own incident management and disaster recovery strategy — with automated runbook execution, agent-supervised recovery playbooks, and validated DR testing cadences.

  • Convert institutional knowledge into machine-executable runbooks and agent-accessible knowledge bases — ensuring operational intelligence is codified, versioned, and continuously improved

  • Lead operational improvement through continuous automation — replacing recurring manual processes with agent-driven workflows and measuring success by the reduction in human intervention per unit of platform activity.

  • Conduct rigorous incident postmortems and RCAs — using AI-assisted log correlation and timeline reconstruction to surface systemic gaps faster and drive durable corrective action.

  • Partner with development, QA, and architecture teams to embed reliability and automation requirements early in the SDLC — shifting reliability left rather than absorbing complexity at the production boundary.

  • Contribute to system design reviews with a reliability and automation lens — ensuring that agentic operability, observability hooks, and failure mode handling are first-class design considerations, not afterthoughts.

  • Provide senior technical leadership — setting the bar for agentic automation design, code quality in SRE tooling, and engineering rigour across the team.

  • Demonstrate Ownership and accountability.

Qualifications & Skills required

This is a critical service delivery role requiring experience with complex datacenter and cloud hosting environments. Pitney Bowes product hosting solutions leverage multiple technologies in complex data center and cloud environments that support multi-tiered high-availability applications. The role requires a talented self-directed and self-motivated individual with a strong work ethic and the following skills:

  • Graduate or Post-Graduate in Computer Science, Engineering, or a related discipline — or equivalent demonstrated depth through professional experience.

  • 10+ years of SRE or platform engineering experience, with a demonstrable shift in recent years toward automation-first and AI-augmented operations.

  • Excellent written and verbal communication skills — able to translate agent behaviour, reliability signals, and automation outcomes into clear engineering and executive narratives.

  • Strong background in SaaS product operations — with hands-on experience running multi-region, high-availability platforms at enterprise scale.

  • Strong experience with CI/CD tooling including Git and Argo — with the ability to extend pipelines with intelligent gates, automated validation, and agent-triggered rollback logic.

  • Strong experience with Docker, Kubernetes, microservices orchestration, and container lifecycle management — including automated remediation of cluster-level failure patterns.

  • Deep expertise in AWS services and cloud-native architectures — including event-driven automation, Lambda-based remediation, and cloud control plane integration for agentic workflows.

  • Solid grounding in networking fundamentals and firewall concepts — sufficient to design and debug connectivity for distributed cloud services and agentic tool integrations.

  • Familiarity with PaloAlto firewall is an advantage, particularly for candidates with GovCloud or regulated-environment experience.

  • Strong experience with Infrastructure as Code — Terraform, Ansible, CloudFormation — with a focus on policy-driven, drift-detecting, and self-correcting infrastructure patterns.

  • Strong Python proficiency — including building agentic tool integrations, LLM API wrappers, and operational automation scripts. Shell scripting for platform automation. PowerShell where Windows surfaces require it.

  • Strong hands-on experience across the observability stack — Prometheus, Grafana, SumoLogic, CloudWatch, PagerDuty, and OpsGenie — with the ability to extend these platforms with custom agentic reasoning and automated response layers.

  • Strong working knowledge of Linux — process management, networking, filesystem, and kernel-level troubleshooting. Windows familiarity where platform surfaces require it.

  • Strong analytical and systems debugging skills — able to reason through complex distributed failure modes and translate that understanding into automated detection and remediation logic.

  • Proficient with Jira, Confluence, and SharePoint — and comfortable integrating these platforms into agentic workflows for automated ticket creation, runbook retrieval, and knowledge surfacing.

  • Experienced working within Agile delivery models — contributing to sprint planning, backlog grooming, and cross-team engineering ceremonies as a senior technical voice.

  • Strong cross-functional collaboration capability — working fluidly with Product Engineering, Product Management, Client Success, and senior leadership to align reliability investments with business outcomes.

  • Hands-on experience designing and operating agentic systems — using frameworks such as LangGraph, AutoGen, or CrewAI to build multi-step, tool-calling agent workflows for operational use cases.

  • Practical experience with LLM APIs (OpenAI, Anthropic, or equivalent) including prompt engineering, tool/function calling, structured output design, and integrating AI reasoning into platform automation pipelines.

  • Ability to design human-in-the-loop escalation protocols for agentic systems — defining clear intervention thresholds, override mechanisms, and audit trails that keep autonomous operations safe and auditable.

  • Experience with agent observability and guardrails — instrumenting agent reasoning chains, detecting failure modes in autonomous workflows, and implementing safety boundaries that prevent runaway automation.

  • High personal accountability — for the reliability of systems you own, the quality of automation you build, and the engineering standards you set for others. 

About Pitney Bowes

Pitney Bowes (NYSE:PBI) is a global technology company providing commerce solutions that power billions of transactions. Clients around the world, including 90 percent of the Fortune 500, rely on the accuracy and precision delivered by Pitney Bowes solutions, analytics, and APIs in the areas of ecommerce fulfillment, shipping and returns; cross-border ecommerce; office mailing and shipping; presort services; and financing. For 100 years Pitney Bowes has been innovating and delivering technologies that remove the complexity of getting commerce transactions precisely right. For additional information visit Pitney Bowes at https://www.pitneybowes.com/in.

We will:


• Provide the will: opportunity to grow and develop your career
• Offer an inclusive environment that encourages diverse perspectives and ideas
• Deliver challenging and unique opportunities to contribute to the success of a transforming organization
• Offer comprehensive benefits globally (PB Benefits and Wellbeing Programs)

Pitney Bowes is an equal opportunity employer that values diversity and inclusiveness in the workplace.

All interested individuals must apply online.

Skills

PythonAWSDockerKubernetesTerraformAnsibleCI/CDLinuxGitJiraConfluenceAgileDevOpsSREMicroservices