Hiring.Camp

Platform Reliability Engineer

Appnovation

·

Today

Location
Toronto, Montreal · New York, New York, United States · Toronto, Canada
Type
Full-time
Department
Technology
Source
Greenhouse

Description

About us

Appnovation is a global, full-service digital partner that combines Strategy, Experience & Design, Engineering and Managed Services. We build digital solutions that deliver real impact today and serve as foundations for future growth.  Bold ambition. Practical action. Endless possibilities.

We’re putting together a dedicated delivery pod to build and run an internal agent platform on AWS Bedrock AgentCore for a global life sciences client. The pod works as one team with the client’s engineers to deliver the platform other teams will build their agents on.

In this role you make the platform visible and dependable. You’ll build the OpenTelemetry instrumentation, observability integrations, tracing, SLOs, alerting, and cost monitoring for token and compute spend. When an agent is slow, wrong or expensive, your work is how the team finds out and fixes it. This is a named-team engagement, so the person we propose is the person who starts.

ROLE RESPONSIBILITIES

  • Instrumentation: Build OpenTelemetry instrumentation standards for agents, tools and platform services, and make it easy for teams to adopt.
  • Observability Integration: Connect AgentCore Observability and CloudWatch with the client’s existing observability tools (e.g., Datadog, Splunk, Grafana).
  • Tracing: Set up end-to-end tracing across agent steps, model calls, tool calls and multi-agent handoffs so issues can be traced to their source.
  • SLOs and Alerting: Define SLOs for latency, availability and error rates, and build alerting that is useful and not noisy.
  • Cost Monitoring: Track token usage and compute spend by team, agent and environment, with dashboards, budgets and alerts for unexpected spikes.
  • Incident Readiness: Write runbooks, support incident response and run post-incident reviews that lead to real fixes.

QUALIFICATIONS

  • 6+ years in SRE, platform reliability or observability engineering, with strong hands-on AWS experience.
  • Hands-on experience with OpenTelemetry (SDKs, collectors, exporters) and distributed tracing.
  • Experience with Amazon CloudWatch and at least one enterprise observability platform (Datadog, Splunk, Grafana, New Relic or similar).
  • Experience defining and running SLOs, error budgets and alerting strategies.
  • Experience building cost visibility and FinOps reporting on AWS (Cost Explorer, CUR, tagging strategies).
  • Scripting and coding skills in Python, Go or TypeScript.
  • Clear communication and able to turn data into decisions for technical and non-technical audiences.

PREFERRED QUALIFICATIONS

  • Experience monitoring LLM or agent-based applications, including token usage, latency and quality signals.
  • Familiarity with LLM observability tools (e.g., Langfuse, Arize, LangSmith) or AgentCore Observability.
  • Experience with infrastructure as code (Terraform or AWS CDK) for monitoring resources.
  • Experience in pharma, life sciences or another regulated industry.

WHO YOU ARE

  • You want to know what’s happening in a system before a user tells you
  • You build alerts people trust and dashboards people actually open
  • You treat cost as a reliability concern, not an afterthought
  • You stay calm in incidents and focus on learning afterward
  • You work well inside a client team and build trust quickly
  • You have prior experience in consulting
  • Prior experience and connections in the Life Sciences industry is preferred
Thank you for your interest in a career with Appnovation Technologies! Please note that only those selected for an interview will be contacted.
 
At Appnovation, we recognize that diverse teams are the strongest teams. Diversity, Equity & Inclusion is not only something that we embrace - we celebrate it! We are proud to be an Equal Opportunity Employer and we encourage applicants from all backgrounds, lived experiences and industries to apply. Come join us at Appnovation, and learn more about how we stay true to our company values as we build better lives through better digital.

Accommodations are available upon request throughout the recruitment process.

Skills

PythonTypeScriptAWSTerraformSplunkSRE

Similar Jobs

30

Platform Reliability Engineer

Appnovation·New York, Austin +2

Today

Platform Reliability Engineer

ConocoPhillips·London Office, UK·Onsite

2w ago

Platform & Reliability Engineer

Volka·Cyprus, Limassol·Onsite

1mo ago

Platform Reliability Engineer

Worldquant·Montevideo +1·Remote

10mo ago

Site Reliability Engineer III (SRE) - Guidewire Cloud Platform (Application)

Guidewire·Krakow Office, Poland·Hybrid

1d ago

Senior Site Reliability Engineer (SRE) - Guidewire Cloud Platform (Application)

Guidewire·Dublin Office, Ireland·Hybrid

1d ago

Site Reliability Engineer II (AI Platform)

Opentable·Toronto, Canada·Remote, Hybrid

1d ago

Site Reliability Engineer (SRE) – Cloud Platform

Fis·IND PUNE FL7, India

2d ago

Engineer, Platform Engineering & Reliability Specialist (Azure)

Nasdaq·GA - Glenridge Point, US·Hybrid, Onsite

2d ago

Senior Platform Reliability Engineer - Azure

LSEG·IND-Hyderabad-CapitaLand, India +1

5d ago

Associate - AI Tooling Ops - Platform Reliability Engineer

Jefferies·Pune, India

5d ago

Site Reliability Engineer, Kubernetes Platform (Top Secret Clearance)

Spacex·Hawthorne, CA +1

5d ago

Principal Site Reliability Engineer, Platform Engineering: Dedicated

Gitlab·Remote, Canada·Remote

1w ago

Senior Platform Reliability Engineer (Fabric and Interconnect)

Firmus Technologies·Sydney, New South Wales

1w ago

Senior Platform & Reliability Engineer (all genders)

Contabo·München, Bayern·Remote, Hybrid, Onsite

1w ago

Site Reliability Engineer, AI Platform

Algolia·Paris, France

1w ago

Senior Site Reliability Engineer, AI Platform

Algolia·Paris, France

1w ago

Site Reliability Engineer – Cloud Native Platform (OpenShift)

KPN·Amersfoort, UT·Hybrid

1w ago

MTS 2, Platform Reliability Engineer

Ebay·Bangalore, India·Hybrid

2w ago

Site Reliability Engineer - Data Platform

IMC·Amsterdam, Netherlands

2w ago

Site Reliability Engineer (Managed Patching & Platform Automation)

Swift·Kuala Lumpur, Malaysia

2w ago

Software Engineer II Platform Data Reliability

Sony Interactive Entertainment·San Mateo, CA·Hybrid

2w ago

Senior Platform Reliability Engineer

Firmus Technologies·Melbourne, Victoria

2w ago

Senior Lead Engineer – Cloud Platform, DevOps & Site Reliability

Qualcomm·Hyderabad, TS

3w ago

Staff Site Reliability Engineer - AI Platform Runtime

Nvidia·Santa Clara, CA

3w ago

Staff Site Reliability Engineer - AI Platform Runtime

Nvidia·Santa Clara, CA

3w ago

Platform Site Reliability Engineer (SRE)

Broadridge Financial Solutions·Manila - 6805 Ayala Ave, Philippines

3w ago

Lead Engineer, Platform Engineering & Reliability

Nasdaq·GA - Glenridge Point, US·Hybrid, Onsite

4w ago

Site Reliability Engineer - SRE (Platform Software Team)

Arista Networks·Bengaluru, KA·Hybrid

4w ago

Senior Backend & Platform Engineer (Quality, Security & Reliability) - Cloud Strategy Office, Cloud Management Department (CMD)

Rakuten·Rakuten Crimson House, Japan

4w ago