The Senior Staff Data Engineer will help build and operate the enterprise lakehouse on Databricks, creating the governed data foundation that supports multiple business domains and downstream analytics. This is a hands-on role, responsible for scalable ingestion, reliable data processing, and strong technical controls across the Bronze and Silver layers of the medallion architecture.
This role is not limited to moving data from point A to point B. The Data Engineer is expected to understand the meaning, sensitivity, classification, and intended use of the data flowing through the pipelines they build, and to apply that understanding when designing controls, access patterns, data quality checks, and segregation boundaries appropriate for a highly regulated environment.
What you'll do:
- Design and build ingestion pipelines, batch and streaming where appropriate, from enterprise source systems into the Databricks lakehouse using Delta Lake.
- Own Bronze-layer ingestion, including raw landing patterns, metadata capture, load traceability, and recoverable ingestion design.
- Build Silver-layer pipelines for cleansing, standardization, deduplication, conformance, and quality enforcement without embedding unauthorized or undocumented KPI logic.
- Define and evolve reusable ingestion and transformation patterns, templates, and engineering standards that other domains can adopt as they onboard to the platform.
- Implement and maintain Databricks platform constructs needed for secure delivery, including catalogs, schemas, service principals, job orchestration, and environment-aware deployment patterns.
- Build and maintain CI/CD pipelines for data platform assets, including ingestion code, transformation logic, workflow definitions, tests, and environment promotion across dev, test, and prod.
- Apply data classification, segregation, and handling requirements within the pipeline design, ensuring sensitive and regulated data is processed in accordance with enterprise controls and access policies.
- Build data quality controls that do more than detect technical failures, including checks that reflect actual business meaning, record integrity, completeness, and expected domain behavior.
- Quarantine, flag, and route problematic records or datasets according to defined quality and compliance rules rather than silently dropping or obscuring issues.
- Maintain documentation for source objects, ingestion logic, applied transformations, data quality rules, and known limitations so downstream teams can trust and use the data correctly.
- Partner with the Analytics Engineer and domain teams to ensure Silver-layer data is reliable, well-governed, and suitable for trusted Gold-layer modeling.
- Collaborate with domain engineering teams to align on ownership boundaries, onboarding patterns, data contracts, and support expectations as new domains are enabled onto the platform.
Required qualifications:
- 12+ years of data engineering experience, including hands-on ownership of production data pipelines.
- Strong Databricks experience, including Delta Lake, Databricks Workflows or Jobs, and Spark with PySpark and/or Spark SQL.
- Working knowledge of Unity Catalog, including catalogs, schemas, tables, lineage, and access control concepts.
- Experience with batch, CDC, and/or streaming ingestion patterns and the operational trade-offs associated with each.
- Experience with CI/CD and deployment automation for data pipelines and platform assets, including version control, testing, and controlled promotion across environments.
- Strong SQL skills and solid grounding in data modeling fundamentals, even if dimensional modeling is not the primary responsibility of this role.
- Demonstrated ability to understand the business and regulatory context of the data being processed, not just the mechanics of pipeline development.
- Experience applying data classification, segregation, security, retention, or compliance requirements in data engineering workflows within a regulated or security-sensitive environment.
- Ability to design pipelines with awareness of the actual data domains involved, including sensitivity, ownership, permitted use, and downstream impact.
- Comfort operating in a fast-moving platform build where patterns are still being established and engineers are expected to shape standards, not just follow them.
Preferred qualifications:
- Databricks certification such as Data Engineer Associate or Professional.
- Experience standing up or maturing Unity Catalog structures and access models in an enterprise setting.
- Experience with CI/CD, infrastructure-as-code, and deployment automation for data platforms.
- Experience in defense, aerospace, financial services, healthcare, or another highly regulated environment.
- Experience enabling multiple business domains on a shared data platform while maintaining strong governance and ownership boundaries.
Similar Jobs at Shield AI
What you need to know about the Austin Tech Scene
Key Facts About Austin Tech
- Number of Tech Workers: 180,500; 13.7% of overall workforce (2024 CompTIA survey)
- Major Tech Employers: Dell, IBM, AMD, Apple, Alphabet
- Key Industries: Artificial intelligence, hardware, cloud computing, software, healthtech
- Funding Landscape: $4.5 billion in VC funding in 2024 (Pitchbook)
- Notable Investors: Live Oak Ventures, Austin Ventures, Hinge Capital, Gigafund, KdT Ventures, Next Coast Ventures, Silverton Partners
- Research Centers and Universities: University of Texas, Southwestern University, Texas State University, Center for Complex Quantum Systems, Oden Institute for Computational Engineering and Sciences, Texas Advanced Computing Center

