- Location
- Riyadh
- Type
- Full-time
- Department
- Education
- Seniority
- Senior
- Source
- Pinpoint
Description
Senior Data Scientist — Transaction Intelligence
Application Deadline: 31 August 2026
Department: Engineering
Employment Type: Full Time
Location: Riyadh
Description
What you'll work on
- Own and level up transaction enrichment — the foundation everything else stands on: categorization that generalizes to merchants we've never seen, and merchant name matching that tolerates harmless variation without confusing genuinely different businesses.
- Own model confidence across our systems: well-calibrated probabilities, principled abstention on uncertain cases, and confidence-based routing that products and review workflows rely on.
- Build transaction-based insight models for other services: pattern recognition, anomaly detection, and behavioral analysis on transaction streams.
- Own the ongoing validation of our risk models: backtesting scores against realized users behavior, monitoring discrimination and calibration as behaviors evolve, and driving model improvements from what the data shows.
- Handle our text and behavior data as they really are: bilingual, informally romanized Arabic with unstable spelling; transaction streams with truncation artifacts, bank quirks, and heavy- tailed distributions.
- Build efficient batch pipelines at hundreds-of-millions-record scale, designed to be re-run routinely as models improve.
- Design within the guardrails of a regulated fintech: data governance, privacy, and cybersecurity requirements shape what data we can use, where workloads can run, and which tools we can adopt — you'll build excellent solutions inside those boundaries, working with those teams rather than around them.
- Build the evaluation discipline these systems deserve: labeled datasets, regression test suites for model behavior, and metrics that answer "did this change help?" for every release.
Must-have qualifications
- 3+ years of applied ML/data science, with models shipped and maintained in production at scale — and knowledge of their failure modes.
- Strong analytical range beyond modeling: exploratory analysis, statistical rigor, and feature design on behavioral/tabular data (SQL fluency assumed).
- Hands-on experience with text similarity, fuzzy matching, or entity resolution on noisy real- world strings.
- Strong Python and the scientific stack (scikit-learn, scipy, numpy), with performance and cost awareness at scale — vectorization, sparse data structures, efficient batch computation.
- Demonstrated maturity working under externally imposed constraints — data governance, privacy, cybersecurity, compliance, infrastructure policy — including collaborating with the teams who own them.
- A track record of turning ambiguous "the model feels wrong" complaints into measured, regression-tested properties.
- Comfort reading and debugging model code you didn't write, and working directly with product teams on loosely-defined problems.
- Arabic / Arabizi text processing — a strong plus; our data is bilingual with unstable romanization.
- Credit or behavioral risk modeling and validation — scorecards, discrimination and calibration measurement, backtesting against realized outcomes.
- Text embeddings and approximate nearest-neighbor retrieval in resource-conscious settings.
- LLM-assisted labeling, distillation, or weak-supervision pipelines.
- ML lifecycle and batch-serving tooling (experiment tracking, data versioning, distributed task queues).