- Salary
- C$70k – C$80k/yr
- Location
- Remote or Mississauga
- Workplace
- Hybrid, Remote
- Type
- Full-time
- Department
- PointClickCare
- Source
- Lever
Description
About the Role:
As a Junior Site Reliability Engineer, you will learn how modern software systems are built, operated, and improved in production.
You will work alongside experienced SREs who directly operate a small number of critical production services while also helping engineering teams improve reliability, observability, automation, and operational excellence.
This role is designed for engineers who enjoy software development, automation, troubleshooting, and understanding how complex distributed systems behave under real-world conditions.
You will gain hands on experience building and using AI powered operational tooling that helps engineers investigate incidents, improve observability, automate repetitive tasks, and eliminate operational toil.
Key Responsibilities:
Production Engineering
- Support the operation of critical services owned by the SRE team.
- Assist in monitoring service health, availability, and performance.
- Participate in incident response activities under mentorship.
- Learn on-call practices and production operations.
- Assist with deployment validation and operational readiness activities.
Reliability Engineering & Observability
- Support implementation of monitoring, alerting, dashboards, and observability improvements.
- Learn how logs, metrics, traces, and events are used to diagnose production environments.
- Help identify monitoring gaps and areas for improved visibility.
- Participate in post incident reviews and reliability improvement initiatives.
- Develop foundational understanding of SLIs, SLOs, and Error Budgets.
Automation & AI Operations
- Develop scripts and automation to eliminate repetitive operational work.
- Contribute to internally developed AI powered operational tools.
- Learn how AI can accelerate troubleshooting and root cause analysis.
- Support development of AI assisted runbooks and operational workflows.
- Help identify opportunities to reduce toil through software engineering and automation.
Engineering Enablement
- Partner with development teams to better understand production systems.
- Assist teams in adopting reliability and observability best practices.
- Support operational readiness reviews.
- Learn how reliability influences software design and architecture decisions.
Required Skills and Qualifications:
- 0-2 years of experience through employment, internship, or co-op programs.
- Degree or diploma in Computer Science, Software Engineering, or equivalent experience.
- Basic programming or scripting experience.
- Strong problem solving skills.
- Curiosity and desire to learn complex systems.
- Effective communication and teamwork skills.
Preferred Qualifications:
- Exposure to Azure cloud services.
- Exposure to Kubernetes or containerized applications.
- Familiarity with observability platforms such as Datadog, AppDynamics, Grafana, Prometheus, or ELK.
- Interest in AI, automation, and distributed systems.