- Salary
- $200k – $225k
- Workplace
- Remote, Onsite
- Type
- Full-time
- Department
- Education
- Experience
- 2+ years
- Education
- Master
- Visa
- Sponsored
- Source
- RecruiterFlow
Description
Data Scientist
Location - New York, NY / Williamsburg, Brooklyn
On-site role requiring five days per week in-office in Williamsburg, Brooklyn.
Compensation
$200,000 – $225,000 Base + Competitive Equity
Visa
Open to Visa Transfers + Visa Sponsorship
Company Stage
Early-Stage – Venture-Backed AI Company
Industry
Artificial Intelligence, Data Science, Consumer Data, Marketing Technology, Enterprise SaaS, Agentic AI, Data Infrastructure
About the Company
Our client is building an AI-powered platform designed to help marketing teams manage complex consumer data, analytics, campaign generation, measurement, and reporting.
At the foundation of the platform is a large-scale consumer graph that unifies identity and behavioral attributes across hundreds of millions of consumers and thousands of evolving data attributes.
The platform combines large-scale data infrastructure, applied data science, probabilistic modeling, and AI agents to help brands improve the quality of their first-party data, identify high-performing audiences, and make better marketing decisions.
The company is backed by leading venture investors and works with major consumer brands across sports, financial services, hospitality, and technology.
As a Data Scientist, you'll join a small, highly technical, high-velocity team where data science is deeply connected to production engineering and AI systems.
You'll own the full path from raw data through feature engineering, modeling, deployment, and production reliability while solving complex problems involving consumer identity, income, wealth, affinity, and other probabilistic attributes.
This is an opportunity to work on unintuitive data problems at terabyte scale while building systems whose outputs are directly consumed by AI agents and enterprise customers.
What You'll Do
- Build and deploy applied data science models supporting large-scale consumer data products
- Stitch together disparate third-party data sources into unified consumer profiles
- Develop and maintain a comprehensive 360-degree view of consumer identities and attributes
- Build probabilistic models for income, wealth, affinity, and other consumer attributes
- Develop statistical estimation methods for deriving reliable attributes from incomplete or fragmented data
- Design and refine feature engineering pipelines operating at terabyte scale
- Build production data workflows that transform raw data into trusted consumer attributes
- Own the full data science lifecycle from raw data ingestion through production deployment
- Develop models that can operate reliably within large-scale production systems
- Build data pipelines supporting consumer graph and AI-agent workflows
- Work with structured and unstructured consumer data from multiple external sources
- Integrate disparate datasets and resolve inconsistencies across data sources
- Design scalable transformations and feature engineering workflows
- Work closely with engineering teams to productionize models and data science workflows
- Ensure model outputs are accurate, robust, explainable, and reliable
- Build systems whose outputs can be consumed autonomously by AI agents
- Support custom enterprise data science projects and customer-specific modeling requirements
- Work directly from raw customer data through deployed models and production outputs
- Analyze complex datasets to identify patterns, relationships, and predictive signals
- Develop new consumer attributes and analytical capabilities based on business requirements
- Improve model performance, reliability, and scalability
- Validate model outputs and monitor production performance
- Develop data quality checks and safeguards for production models
- Work with large-scale data infrastructure and modern data platforms
- Collaborate across data science, engineering, product, and customer-facing teams
- Translate business problems into quantitative and modeling approaches
- Operate independently across multiple data science and engineering workstreams
- Make pragmatic decisions when working with incomplete or imperfect data
- Identify opportunities to improve data quality, model performance, and system reliability
- Help shape the technical foundation of large-scale AI-powered consumer data products
Ideal Candidate Background
Experience Requirements
- 2–10 years of professional experience in applied data science
- 2+ years of hands-on experience building production data science systems
- Strong applied data science background
- Experience building and deploying statistical or machine learning models
- Experience working with large-scale datasets
- Experience working with terabyte-scale data preferred
- Experience building production feature engineering pipelines
- Experience taking models from experimentation through production deployment
- Experience working with complex, fragmented, or noisy datasets
- Experience solving ambiguous data problems independently
- Experience working across data science and data engineering workflows
- Strong analytical and quantitative reasoning ability
- Strong ownership mentality with demonstrated execution ability
- Comfortable operating in a fast-moving startup environment
- Comfortable managing multiple technical workstreams simultaneously
- Experience working directly with business or enterprise requirements preferred
- Experience with consumer data, marketing analytics, identity resolution, or customer intelligence preferred
- Experience working with AI-powered products or agentic systems preferred
- Experience working in technically demanding environments preferred
Technical Requirements
- Strong Python experience
- Strong SQL experience
- Strong applied statistics and probabilistic modeling fundamentals
- Experience building predictive or estimation models
- Strong feature engineering experience
- Experience building production data pipelines
- Experience working with large-scale datasets
- Experience with dbt or comparable data transformation frameworks
- Experience with Dagster, Airflow, or comparable workflow orchestration systems
- Experience with Databricks, Snowflake, Redshift, or comparable data platforms
- Experience with PostgreSQL or comparable relational databases
- Experience with AWS and cloud-based data infrastructure
- Experience with S3 or comparable object storage systems
- Experience working with data warehouses and distributed data systems
- Experience transforming raw data into production-ready analytical features
- Experience designing reliable data workflows
- Experience monitoring and troubleshooting production data pipelines
- Experience deploying models into production environments
- Experience building scalable feature engineering systems
- Experience with identity resolution, entity matching, or consumer graph systems preferred
- Experience with machine learning or statistical inference preferred
- Experience with AI/ML systems or AI-powered applications preferred
- Experience working with probabilistic attributes preferred
- Strong understanding of model validation and evaluation
- Strong understanding of data quality and reliability
- Strong ability to reason about data pipelines, modeling, and production systems together
- Ability to learn unfamiliar datasets and business domains quickly
Education
- Bachelor's degree in Computer Science, Statistics, Mathematics, Data Science, Engineering, Economics, or related quantitative field preferred
- Master's degree or advanced quantitative education preferred
- Equivalent practical experience accepted
- Strong mathematical, statistical, and analytical fundamentals
Soft Skills
- Exceptional analytical ability
- Strong quantitative reasoning
- Strong problem-solving ability
- Strong data intuition
- Strong technical ownership
- Strong product and business judgment
- Highly independent working style
- Comfortable operating with minimal structure
- Strong ability to reason from first principles
- Comfortable solving unintuitive or ambiguous problems
- Strong attention to data quality and correctness
- Strong communication skills
- Strong ability to explain technical and quantitative concepts clearly
- Comfortable working directly with technical and business stakeholders
- Highly organized and capable of managing multiple workstreams
- Strong multitasking ability
- High execution velocity
- Strong bias toward shipping
- Comfortable making decisions without waiting for direction
- Strong ability to prioritize competing problems
- Low-ego collaborative mentality
- Comfortable working closely with engineers
- Comfortable working across data science and data engineering
- Strong business intuition
- Strong ownership of production outcomes
- Comfortable working with imperfect and messy data
- Strong curiosity and willingness to investigate difficult problems
- Comfortable working in a fast-moving startup environment
- Strong interest in AI-native products and agentic systems
- Comfortable working five days per week in New York City
- Strong willingness to operate in a high-intensity, high-ownership environment
Compensation & Benefits
- Base Salary: $200,000 – $225,000
- Competitive equity
- Opportunity to work on large-scale AI and consumer data systems
- Opportunity to solve complex applied data science problems
- Exposure to terabyte-scale production datasets
- Opportunity to build probabilistic consumer models
- Opportunity to work on AI-agent-powered data systems
- Direct ownership over production data science workflows
- Opportunity to work across data science, data engineering, and AI
- Opportunity to work with leading enterprise consumer brands
- High degree of technical autonomy
- Fast-paced and highly technical engineering environment
- Opportunity to work closely with experienced technical leadership
- Visa transfer support
- Visa sponsorship available
- Five-day in-office environment in New York City
Why Join
This is an opportunity to join a small, highly technical AI company solving difficult problems at the intersection of data science, consumer intelligence, and agentic AI.
You'll work with massive datasets and own the full path from raw data to production-ready attributes and models that directly power AI agents and enterprise customer experiences.
You'll solve problems involving identity resolution, probabilistic modeling, feature engineering, data quality, and large-scale production systems while working closely with engineers and technical leadership.
If you enjoy applied data science, working with messy real-world data, building production models, solving ambiguous quantitative problems, and operating with high ownership in a fast-moving AI environment, this role offers exceptional technical scope and impact.