- Location
- Washington, DC
- Type
- Full-time
- Department
- Engineering
- Experience
- 8+ years
- Education
- Master
- Source
- ApplicantPro
Description
Position Summary
The Data Architect will lead the design and governance of data systems that enable analytics and secure data sharing across the organization. This role ensures data structures and platforms support mission-critical operations while maintaining compliance with federal security standards. A central part of the job is reconciling data that arrives in inconsistent forms - relational SQL databases, Microsoft Dataverse, spreadsheets, and unstructured documents and files - into unified enterprise data models that the organization's cloud and analytics platforms can build on with confidence.
This is a hands-on architecture role: the Data Architect is expected to both design the enterprise data models and build the pipelines, semantic layers, and governance artifacts - in Databricks, Power BI, Tableau, and related platforms - that put those models into production.
Responsibilities
- Define enterprise data models, standards, and governance frameworks that unify data originating from SQL databases, Microsoft Dataverse, spreadsheets, unstructured documents, and other varying source formats into consistent, reusable structures.
- Lead cloud data migration and integration initiatives, including ingestion and transformation pipelines built in Databricks and comparable ETL/ELT tooling.
- Automate data loading, cleansing, and validation routines to reduce manual data preparation and ensure incoming data consistently meets quality standards before it reaches downstream models.
- Design and schedule automated refresh workflows (Databricks Jobs/Workflows or equivalent orchestration) so source data, transformations, and downstream datasets stay current without manual intervention.
- Build and optimize Databricks pipelines using Delta Lake, PySpark/SQL notebooks, and medallion (bronze/silver/gold) architecture patterns to structure raw, cleansed, and analytics-ready data layers.
- Apply performance tuning and cost-optimization practices (cluster sizing, job scheduling, partitioning, caching) to keep automated pipelines efficient and reliable at scale.
- Ensure compliance with NIST and other applicable federal data security requirements across all data architecture, storage, and access design decisions.
- Collaborate with developers, analysts, and system engineers to optimize data architecture and resolve inconsistencies between source systems.
- Provide technical guidance on best practices for data quality, storage, and retrieval, including canonical entity definitions so the same data concept is represented consistently across every downstream system.
- Design and curate semantic and reporting layers (data marts, shared datasets, modeled views) that feed Power BI, Tableau, and other visualization platforms, so metrics are defined once and reused consistently across dashboards.
- Assess the lineage, quality, and structure of existing data sets prior to modeling; identify gaps, redundancies, and conflicting definitions across systems.
- Document architecture decisions, data flows, and system integrations, and mentor other technical staff on data modeling standards as the practice matures.
Qualifications
- Bachelor's degree in Computer Science, Data Science, or a related field (Master's preferred).
- 8+ years of experience in data architecture or data engineering, including experience building enterprise-level data models spanning multiple, varying source systems.
- Proficiency with SQL and cloud data platforms (AWS GovCloud, Azure Gov).
- Hands-on experience with Databricks (or a comparable Spark-based/lakehouse platform) for data pipeline development, including PySpark/SQL notebooks, Delta Lake, and job/workflow orchestration for automated, scheduled data refreshes.
- Experience modeling and integrating data from Microsoft Dataverse or similar low-code/Power Platform data stores.
- Practical experience preparing data models for Power BI and/or Tableau consumption, including semantic layer and shared dataset design.
- Demonstrated ability to work with messy, inconsistent, or unstructured data (spreadsheets, documents, file exports) and bring it into a governed enterprise structure.
- Demonstrated experience automating data ingestion, cleansing, and refresh processes end-to-end, minimizing manual data preparation and keeping downstream datasets current on a defined schedule.
- Strong knowledge of data governance frameworks and compliance requirements, including NIST and other federal data security standards.
- Excellent problem-solving and communication skills, with the ability to explain technical data concepts to non-technical stakeholders.
Preferred Qualifications
- Experience supporting a federal civilian agency contract, including familiarity with agency-specific data-handling and security requirements.
- Familiarity with SharePoint Online, Power Automate, and the broader Microsoft 365 / Power Platform data ecosystem.
- Experience with master data management (MDM) tools or practices.
- Experience with Databricks Unity Catalog, CI/CD for data pipelines, or infrastructure-as-code approaches to managing Databricks workspaces and jobs.
- Relevant certifications such as Databricks Certified Data Engineer, AWS Certified Data Analytics, or Microsoft Certified: Azure Data Engineer Associate.