- Location
- Hybrid · Lebanon
- Workplace
- Hybrid
- Department
- IT
Description
At Liaison, we’ve helped higher education institutions build better, more diverse classes for three decades. You may recognize us as the company behind the Centralized Application Service (CAS), Enrollment Marketing services and platform (EMP), SlideRoom, Time2Track, TargetX (CRM) and Othot.
Everything we do is focused on taking that proven success and expanding its scope and scale. Over 31,000 programs on more than 1,000 campuses see us as a forward-thinking partner integral to meeting their total enrollment goals — and we’re building the data- and mission-driven team that will reinforce our role for decades to come.
Do you like solving production issues, making systems more reliable, and closing the loop from incident to root cause? Do you enjoy collaborating with smart engineers, product teams, and partners to keep critical applications healthy? If so, this role is for you.
Liaison International is a leader in college application and admissions solutions in the United States. Our products help applicants and institutions streamline the admissions process using modern software technologies. We are expanding our Application Support Monitoring team and are seeking highly motivated specialists who are as comfortable troubleshooting issues as they are reading dashboards and improving alerts.
This role sits at the intersection of Level 2 application support and observability. Apart from being part of the application support team, you will also help monitor the health of our various platforms, diagnose and resolve issues, and continuously improve how we detect, alert on, and prevent problems in production.
A successful candidate will have strong consulting and customer-facing skills, a passion for observability and debugging, and a mindset of ownership around system health.
Responsibilities:
Application Support & Incident Response:
- Own and manage Level 2 support requests coming through our ticketing systems (Salesforce, Jira), with a focus on application behavior and data rather than basic “how‑to” questions
- Troubleshoot application issues reported by end users and internal teams, using logs, metrics, and traces to understand what is happening in production.
- Collaborate closely with Developers, Business Analysts, QA, and Implementation teams to reproduce issues, identify root causes, and coordinate fixes.
- Validate software fixes in test and pre-prod environments and, when appropriate, in production to confirm issues are resolved.
- Respond and follow up with stakeholders within agreed SLAs, providing clear, technically accurate updates.
Observability & Monitoring:
- Monitor application performance and reliability using various APM tools, alerts, and channels.
- Use telemetry (logs, metrics, traces) to proactively detect anomalies and performance regressions, not just react to tickets.
- Help design, maintain, and improve application monitoring and alerting for key Liaison services (e.g., New Relic dashboards, error rate, and latency alerts).
- Tune alerts so they are actionable and low‑noise: adjust thresholds, add missing signals, and remove or refine noisy alerts in partnership with engineering.
- Contribute to and maintain health dashboards for applications, data exports, scheduled jobs, integrations, and campaign workflows so teams can easily see what is healthy, failing, or stuck.
- Participate in incident reviews and help translate lessons learned into improvements in monitoring, alerting, and runbooks.
Onboarding & Configuration:
- Assist with onboarding new clients and rolling over new cycles, ensuring health and monitoring configuration are in place from day one.
- Troubleshoot onboarding configuration and regression issues, coordinating changes across internal and external teams.
- Monitor and maintain data quality between systems (configuration and user data), helping to identify and correct discrepancies early.
Knowledge, Process, and Continuous Improvement:
- Develop and run complex ad‑hoc reports (SQL, Excel) to support investigations and answer data questions from stakeholders.
- Create and publish knowledge base articles, runbooks, and troubleshooting guides to enable faster, more consistent responses to recurring issues.
- Assist other support staff with problem‑solving and share best practices around observability, monitoring tools, and debugging.
- Support management with metrics and reporting on incidents, throughput, and recurring patterns; help drive down repeat issues via RCA and preventive changes.
- Take ownership of Root Cause Analysis (RCA) for major application issues, capturing both technical causes and process/monitoring gaps and ensuring follow‑through on action items.
Basic Qualifications:
- Bachelor’s degree in computer science, Information Systems, or related field; or equivalent practical experience.
- Experience in application support, DevOps, SRE, or technical operations for web‑based applications.
- Hands‑on experience troubleshooting in Linux/Windows web environments using logs, metrics, and monitoring dashboards.
- Experience writing clear, detailed tickets that document investigation steps, diagnosis, and escalation paths.
- Strong verbal and written communication skills; comfortable working directly with internal and external clients and partners in a professional, tactful manner.
Preferred Qualifications:
- Experience with observability / monitoring tools (e.g., New Relic, Datadog, ELK, Grafana, CloudWatch) and an understanding of logs/metrics/traces.
- Experience with Python and JavaScript for scripting, small tools, or automation.
- Experience with Node.js and TypeScript is a plus.
- Solid understanding of APIs, microservices, and asynchronous workloads and how to instrument and monitor them.
- Intermediate to advanced SQL skills for querying application and reporting databases.
- Intermediate to advanced Excel skills (tables, pivot tables, formulas) for quick analysis and reporting.
- Prior exposure to incident management, SRE practices, or on‑call rotations.