- Salary
- $190k – $250k
- Location
- San Mateo · San Mateo, California, United States
- Workplace
- Remote, Hybrid, Onsite
- Department
- Science
- Seniority
- Senior
- Source
- Greenhouse
Description
In this role, you will work at the intersection of large-scale distributed systems, machine learning engineering, and platform architecture. You will own and evolve critical data capabilities within Cognitiv’s advertising ecosystem, with a primary focus on designing and delivering stable, highly scalable feature pipelines, as well as optimizing the performance and reliability of ad model training pipelines.
This is a senior individual-contributor role for an engineer who thrives on solving complex platform problems, raising the technical bar, and building systems that serve many teams at scale. You will partner closely with engineers, product managers, data scientists, infrastructure teams, and downstream data consumers to deliver platforms that are resilient, extensible, and easy to adopt.
This position will be located in San Mateo, CA with a hybrid work schedule of 3 days in office (Mon/Tue/Wed) and 2 days remote optional (Thursday/Friday).
What You'll Do
- High-Performance Feature Engineering & Infrastructure: Architect, build, and maintain low-latency, high-throughput feature pipelines — batch, real-time streaming, and point-in-time correct historical features — to power our real-time bidding systems.
- Advanced Embeddings & Generative AI Integration: Leverage LLMs and deep learning models to extract rich contextual and user-level embeddings into the core feature store/serving system, optimizing embedding generation, indexing, and online retrieval for sub-millisecond serving SLAs.
- Model Training Pipeline Optimization: Own and continuously enhance Cognitiv's ad model training pipelines, improving training speed, resource utilization, and throughput for large-scale deep learning models.
- System Scalability, Reliability & Efficiency: Establish technical standards for monitoring, testing, and CI/CD across feature and training infrastructure to ensure robust system SLAs/SLOs.
- Partner with Modeling & Data Science: Translate complex signals into production-ready features that directly boost model performance (e.g., CTR/CVR prediction).
Who You Are
Must haves:
- 3-5+ years of hands-on Machine Learning Infrastructure / Data Platform experience supporting data-intensive platforms, including large-scale data pipelines, streaming systems, and storage layers.
- Proficiency in one or more core programming languages — Python, Java, or Scala — for building, maintaining, and scaling robust ML and data pipelines.
- Strong expertise in modern big data technologies such as Apache Spark, Apache Flink, Apache Kafka, and other distributed data processing frameworks.
- Excellent team communication, cross-functional collaboration, and problem-solving skills, with a track record of partnering effectively with modeling and engineering teams.
- Willing to work onsite Monday/Tuesday/Wednesday in San Mateo, CA.
Nice to haves:
- Domain expertise in AdTech.
- Hands-on experience with PyTorch for model architecture, training pipeline acceleration, or distributed training.
- Proficiency with cloud & infrastructure technologies, including AWS/GCP, as well as containerization and orchestration platforms like Kubernetes (K8s) and Docker.
- Competitive programming background, such as awards or achievements in OI (Olympiad in Informatics) or ACM/ICPC.
Tech Stack: Python, Java, Kafka, Flink, Spark, PyTorch, Kubernetes, AWS
What Success Looks Like in Your First 30/60/90 Days
First 30 days:
- Ramp quickly on Cognitiv's feature pipelines, training infrastructure, and RTB systems.
- Build strong context on current architecture, data flows, and known bottlenecks.
- Establish relationships with Modeling, Data Science, and Infrastructure teams.
- Identify early opportunities to improve pipeline reliability, training efficiency, or serving latency.
First 60 days:
- Independently own at least one core feature pipeline, embedding service, or training subsystem.
- Ship a first measurable improvement — e.g., reduced serving latency, faster training throughput, or expanded monitoring/CI-CD coverage.
- Align with Modeling and Data Science on the feature and training roadmap.
- Begin influencing technical direction for feature and training infrastructure.
First 90 days:
- Fully own a core area of the platform, such as embedding generation/retrieval, streaming infrastructure, or training pipeline performance.
- Deliver measurable business impact — e.g., reduced feature or serving latency, faster model training, or improved uptime/SLA adherence.
- Redesign or scale at least one key system to handle growing data volume or model complexity.
- Operate autonomously as a trusted technical partner to Modeling, Data Science, and Engineering.
Salary: $190,000-$250,000 USD Base Salary + Equity
What We Offer
- Medical, Dental and Vision plan for US employees & Extended Health Benefits for Canadian employees
- 12 weeks paid parental leave + 4 weeks WFH
- Unlimited PTO + Work-From-Anywhere August
- Career development with clear advancement paths
- Equity for all employees
- Hybrid work model & daily team lunch
- Health & wellness stipend + cell phone reimbursement
- 401(k) & RRSP with employer match
- Parking (CA, WA, Vancouver offices) & pre-tax commuter benefits
- Employee Assistance Program
- Comprehensive onboarding (Cognitiv University)
- …and more!
What You’ll Find at Cognitiv
- Festiv – We make work fun with cross-team games, events, and creative team bonding.
- Responsiv – You’ll be close to clients and leadership, influencing real outcomes.
- Inclusiv – Diversity and individuality are celebrated across all levels.
- Inventiv – We reward curiosity and embrace bold ideas.
- Transformativ – We support your growth with training, mentorship, and flexibility.
- Collaborativ – We operate across coasts, connected by purpose and teamwork.