- Location
- US TX IT Infrastr and Ops 723, United States of America
- Type
- Internship
- Department
- Engineering
- Seniority
- Internship
- Source
- Workday
Description
Copart, Inc. a technology leader and the premier online vehicle auction platform globally, with over 200 facilities located across the world, Copart links vehicle sellers to more than 750,000 buyers in over 190 countries. We believe in providing an unmatched experience, every day and everywhere, driven by our people, processes, and technology.
Copart, a global leader in online vehicle auctions, is seeking a highly skilled and proactive Site Reliability Engineer Intern to join our SRE team at our Dallas location. This role is pivotal in ensuring the stability, performance, and reliability of Copart’s critical applications and global data center infrastructure through advanced monitoring, troubleshooting, and automation.
As a key member of our 24/7 operations, you will contribute to meeting our stringent SLA commitments and play a crucial role in maintaining operational excellence.
What You’ll Do (Essential Duties & Responsibilities):
Monitoring & Automation
- Proactively monitor Copart’s global data centers and application infrastructure using internal tools, identifying and resolving issues before they impact operations.
- Design, build, and optimize monitoring and automation tools using Python and Ansible to collect key metrics and enable automated remediation of infrastructure and application issues.
- Maintain and optimize a suite of monitoring and collaboration tools, including Datadog and Kubernetes-related observability tooling.
- Conduct monthly security patching for operating systems and critical applications.
Collaboration & Improvement
- Support for Infrastructure teams, quickly and efficiently troubleshooting and communicating issues across various Copart domains.
- Partner with cross-functional internal teams, including Product Development, DevOps, Network, Systems, and Database teams.
- Develop analytical and reporting capabilities to monitor performance, identify areas for improvement, and implement quality control plans using Data Dog.
- Create and maintain comprehensive standard operating procedures (SOPs), system diagrams, and training materials for team use.
- Support incident management activities including triage, tracking, root cause analysis, and monthly review of lessons learned.
- Process change paperwork throughout the day as part of daily operational requirements.
Incident Management Requirements
- Experience triaging, tracking, and resolving incidents in a production environment
- Ability to perform initial impact assessment and severity classification
- Strong communication skills for updates to stakeholders during active incidents
- Experience documenting incident timelines, actions taken, and resolution details.
- Coordinate and perform periodic failover testing for Copart’s network, systems infrastructure, and application environments to ensure business continuity.
- Participation in root cause analysis and post-incident review processes
- Ability to identify trends, recurring issues, and opportunities for preventive action
Who We’re Looking For (Ideal Candidate Profile):
- Collaborative & Communicative: A strong team player with excellent interpersonal, oral, and written communication skills who thrives in a collaborative environment.
- Technically Versatile: Beyond core knowledge of Linux and Windows, you bring experience with virtual environments, basic networking, scripting, automation, and observability tools. Familiarity with Datadog is a plus.
- Innovative & Proactive: You continuously look for ways to improve processes and procedures and are comfortable suggesting and implementing practical solutions to improve efficiency and reliability.
- Self-Starter & Adaptable: A highly motivated individual who can work independently, manage competing priorities, and perform effectively in a flexible work schedule.
Highly Desired Skills (A Big Plus):
- Intermediate proficiency in programming and scripting.
- Strong troubleshooting skills across Linux/Unix/Windows-based systems.
- Hands-on experience with virtual machine management software, particularly VMware vSphere a plus.
- Familiarity with monitoring and observability tools such as Datadog.
- Basic understanding of AI tools and concepts, including the ability to work with AI-assisted workflows, prompt-based tools, or automation-enhanced support processes.
- Ability to leverage AI tools to improve efficiency in troubleshooting, documentation, analysis, or operational support is a plus.
- Any experience with AWS or GCP
#LI-KK1
At Copart, we are focused on harnessing the power of diversity, inclusion, and collaboration. By embracing diverse perspectives, we open doors to innovation and unleash the full potential of our team. We are dedicated to fostering a workplace where everyone feels appreciated, included, and inspired to grow and contribute meaningfully.
E-Verify Program Participant: Copart participates in the Department of Homeland Security U.S. Citizenship and Immigration Services' E-Verify program (For U.S. applicants and employees only). Please click below to learn more about the E-Verify program: