- Location
- Santa Clara, CA,US, US · Seattle, WA,US, US
- Workplace
- Remote
- Type
- Full-time
- Department
- Engineering
- Seniority
- Senior
- Source
- Eightfold
Description
Design and scale observability platforms handling high-volume metrics, logs, and traces across distributed environments Build high-performance backend services for telemetry ingestion, processing, and routing Develop and extend OpenTelemetry collectors, processors, exporters, and instrumentation libraries Build and optimize metrics pipelines using large-scale time-series storage systems Design and operate real-time and batch telemetry pipelines using streaming and distributed data technologies Improve platform reliability, performance, and cost efficiency through tuning, capacity planning, and system optimization Develop monitoring, alerting, and service reliability frameworks to ensure platform health and performance Collaborate with platform engineering, infrastructure, and site reliability teams to deliver production-grade observability solutions Experience building or operating distributed data pipelines using technologies such as Kafka, Spark, or Flink Experience working with Kubernetes and cloud-native infrastructure Experience integrating observability with AI/ML pipelines, GPU workload monitoring, or intelligent alerting