Hiring.Camp

Observability Engineer

Oklahama State Government

·

Aug 3, 2026

Location
Oklahoma City - 3115 N Lincoln Boulevard, United States of America
Workplace
Onsite
Type
Full-time
Department
Engineering
Experience
3+ years
Education
Bachelor
Closing date
Aug 17, 2026
Source
Workday

Description

Job Posting Title

Observability Engineer

Agency

090 OFFICE OF MANAGEMENT AND ENTERPRISE SERV

Supervisory Organization

IS-CS

Job Posting End Date

Refer to the date listed at the top of this posting, if available. Continuous if date is blank.

Note: Applications will be accepted until 11:59 PM on the day prior to the posting end date above.

Estimated Appointment End Date (Continuous if Blank)

Full/Part-Time

Full time

Job Type

Regular

Compensation

Job Details
• This is a full-time, 40-hour per week position
• Support the Information Services Division
• Position is on-site in Oklahoma City, OK

Job Description

Job Details

  • This is a full-time, 40-hour per week position
  • Support the Information Services Division
  • Position is on-site in Oklahoma City, OK

Position Summary

The Observability Engineer is responsible for building, maintaining, and continuously improving the monitoring and observability capabilities that keep the organization's server, cloud, and network environments healthy. Using Datadog or similar platforms, this role instruments infrastructure and applications, builds meaningful dashboards and alerts, and turns raw telemetry into early warning signals that let teams find and fix problems before they impact the business. This is a hands-on technical role for someone who enjoys making complex environments visible, understandable, and measurably more reliable.

Position Responsibilities

Monitoring & Observability Engineering

  • Design, deploy, and maintain monitoring and observability tooling (Datadog or similar) across server, cloud, and network environments.
  • Instrument infrastructure, applications, and services with metrics, logs, traces, and synthetic checks to provide full-stack visibility.
  • Build and maintain dashboards that give clear, role-appropriate visibility into system health for engineers, managers, and leadership.
  • Configure and tune alerting thresholds and escalation policies to catch real issues early while minimizing noise and alert fatigue.
  • Integrate monitoring tools with incident management, ticketing, and on-call notification systems (e.g., PagerDuty, ServiceNow, Slack).

IT Health & Continuous Improvement

  • Track and report on key IT health indicators such as uptime, latency, error rates, capacity headroom, and patch/compliance status across multiple environments.
  • Partner with infrastructure, cloud, and network teams to identify recurring issues and drive root-cause fixes rather than repeated firefighting.
  • Support capacity planning by analyzing utilization trends and flagging environments approaching risk thresholds.
  • Contribute to post-incident reviews by providing telemetry, timelines, and health data that clarify what happened and why.
  • Continuously refine monitoring coverage as new systems, services, and cloud resources are added, retiring stale checks and dashboards.

Collaboration & Documentation

  • Work closely with server, cloud, network, and application teams to understand what 'healthy' looks like for each environment and translate that into monitoring coverage.
  • Document monitoring standards, runbooks, and dashboard conventions so coverage stays consistent as the environment grows.
  • Train and support other engineers in interpreting dashboards, alerts, and observability data.
  • Evaluate and recommend improvements or additions to the observability toolset as monitoring needs evolve.

Physical Demands and Work Environment

  • Office-based work involving extensive computer and phone use.
  • Requires long periods of sitting, up to eight hours a day.
  • Possible on-call rotation
  • Work environment is generally quiet, occasional travel may be required.

Education and Experience

  • Associate's or Bachelor's degree in Information Technology, Computer Science, or related field, or equivalent hands-on experience.
  • 3+ years of experience in infrastructure monitoring, observability, systems administration, or IT operations.
  • Hands-on experience with Datadog or a comparable observability platform (e.g., Dynatrace, New Relic, Splunk, Prometheus/Grafana).
  • Working knowledge of server, cloud (AWS, Azure, or GCP), and network fundamentals sufficient to instrument and troubleshoot across environments.
  • Experience building dashboards, alerts, and notification workflows that support fast, accurate incident response.
  • Comfortable working with scripting or query languages (e.g., Python, PowerShell, Bash, DQL/PromQL) to build and refine monitoring logic.

Preferred Qualifications

  • Datadog certification or equivalent vendor certification.
  • Experience with infrastructure-as-code (Terraform, Ansible) for deploying monitoring agents and configuration at scale.
  • Familiarity with ITSM practices (incident, problem, change management) and on-call/escalation processes.
  • Exposure to AIOps or automated remediation approaches that reduce manual intervention.
  • Experience monitoring environments in regulated industries with attention to security and compliance requirements.

About OMES

The Office of Management and Enterprise Services provides excellent service, expert guidance and continuous improvement in support of our partners’ goals. We are a highly qualified workforce committed to serve those who serve Oklahomans and make government run in the most efficient, innovative manner possible.

OMES is an Equal Opportunity Employer. Reasonable accommodation to individuals with disabilities may be provided upon request.

Equal Opportunity Employment

The State of Oklahoma is an equal opportunity employer and does not discriminate on the basis of genetic information, race, religion, color, sex, age, national origin, or disability.

Current active State of Oklahoma employees must apply for open positions internally through the Workday Jobs Hub.

If you are needing any extra assistance or have any questions relating to a job you have applied for, please click the link below and find the agency for which you applied for additional information:

Agency Contact

Skills

PythonAWSAzureGCPTerraformAnsibleWorkdayServiceNowSplunkComplianceChange Management

Similar Jobs

30

Observability Engineer  

NECSWS · Bracknall, United Kingdom · Hybrid

Today

Observability Engineer

Avaloq · Makati City, National Capital Region, Philippines · Hybrid

1 week ago

Observability engineer

Groupefdj · Stockholm, Sweden · Hybrid

1 week ago

Observability Engineer

Riministreet · Hyderabad, India

3 weeks ago

Observability Engineer

Careers Home · ABC Manila Office, Philippines · Onsite

1 month ago

Observability Engineer

Point72 · Bengaluru, India +1

1 month ago

Observability Engineer

Nasdaq · Bangalore-Affluence, India

2 months ago

Observability Engineer

Miratech · All Cities, India · Remote

2 months ago

Observability Engineer

Lgtcp · Pfaeffikon (Summelenweg), Switzerland

4 months ago

Observability Engineer

Lgtcp · Pfaeffikon (Summelenweg), Switzerland

4 months ago

Observability Engineer

Gresearch · London, United Kingdom · Remote

6 months ago

Observability Engineer

Sequoiaconnect · Remote

1+ year ago

Enterprise Cloud Engineer IV - Observability

Wellmark · Des Moines, IA, United States · Hybrid

Today

Tech Lead Software Engineer Global - Observability Accelerators & AI

Southwest Careers · India Office

Yesterday

Senior Platform Engineer - OpenShift & Observability

Swift · OPC NL, Netherlands

Yesterday

Sr Software Engineer Global - Observability Accelerators & AI

Southwest Careers · India Office

Yesterday

Software Developer Intern, Field Innovation, Security Search and Observability (SSO)

Amazon · Remote, Onsite

Yesterday

Staff Software Engineer - Reporting, Data Platform & Observability

OneTrust · Atlanta, Georgia +1

Yesterday

Data Platform Engineer – Inventory & Observability

Quberesearchandtechnologies · London +1

Yesterday

Senior Site Reliability Engineer – Unified Observability

Ncr · Chennai Embassy Tower Office, India · Remote, Hybrid

2 days ago

Senior Platform Engineer (Observability)

Myhrhome · HQ_Evansville IN-600 N.W. 2N, United States of America

2 days ago

Senior Observability Engineer (Splunk & APM)

Staples Canada · Framingham, MA, United States, US · Onsite

2 days ago

Senior Software Engineer, Observability

Clear · New York, NY, United States · Remote, Onsite

2 days ago

Senior Observability/DevOps Engineer

Miratech · Kyiv, Kyiv city, Ukraine · Remote

2 days ago

Senior Software Engineer, Observability

Whatnot · Remote

3 days ago

Software Engineer, AV Logging and Observability

General Motors · GM Automation - Sunnyvale - GM Automation - Sunnyvale, United States of America +2 · Hybrid, Remote

3 days ago

Sr. Observability Engineer II

Nextgen · Bangalore, India · Remote

3 days ago

Senior Site Reliability Engineer - Observability

Westpac Group · Sydney, NSW, Australia

4 days ago

Lead Observability Engineer

Alteryx · Irvine, California, United States of America

4 days ago

Senior Backend Software Engineer, AI Observability & Evals Platform (LangSmith)

LangChain · Boston, MA +1 · Onsite

1 week ago