i4DM Logo

i4DM

Databricks Data Engineer

Posted 4 Days Ago
Remote
Hiring Remotely in USA
Mid level
Remote
Hiring Remotely in USA
Mid level
Designs, builds, and operates scalable batch and streaming data pipelines on Databricks for federal missions. Responsibilities include developing Spark and Delta Lake solutions, managing clusters and workflows, implementing Unity Catalog governance and security, optimizing ETL/ELT processes, integrating CI/CD, supporting machine learning and advanced analytics, monitoring data quality, and collaborating with technical teams and stakeholders.
The summary above was generated by AI
Description

About Our Team

Our employees thrive in a culture that's fast-paced and ego-free, where innovation and collaboration are encouraged at every turn. We are an organization that provides federal agencies instant access to experienced and talented professionals who understand their unique challenges and know the most efficient ways to address them. We are continually investing in resources and talent, so we stay prepared with specialized teams in place who are experts in creating tailored technologies. Our solutions empower Federal organizations to grow, modernize, and succeed in a rapidly evolving landscape.

We value all voices and want to attract talent from all backgrounds. We're on the lookout for individuals who are passionate about technology and thrive in environments where problem-solving is approached with creativity and enthusiasm. If you're someone who enjoys continuously expanding your skill set while tackling real-world business problems, you'll feel right at home with us. Veterans and military spouses are especially encouraged to bring your unique and valuable experience to our team.

About the Role:

We are seeking a hands-on Databricks Engineer to design, build, and operate scalable data and analytics solutions on the Databricks Lakehouse platform in support of federal mission needs. The ideal candidate will have strong practical experience with Apache Spark, Delta Lake, and Unity Catalog, along with a solid understanding of modern data architecture patterns such as the medallion architecture and structured streaming. This role involves developing and optimizing data pipelines, implementing data governance and security controls, enabling advanced analytics and machine learning, and collaborating with cross-functional teams within a compliance-driven federal environment. By joining our organization, you'll help modernize how federal agencies use data to make better decisions and deliver better outcomes for the people they serve!

 Key Responsibilities

  • Design, develop, and maintain scalable batch and streaming data pipelines using Databricks, Apache Spark, PySpark, and Spark SQL.
  • Build and manage Delta Lake tables using the medallion (bronze/silver/gold) architecture to deliver reliable, analytics-ready data.
  • Develop real-time and near-real-time data ingestion solutions using Spark Structured Streaming and messaging platforms such as Kafka.
  • Configure and manage Databricks clusters, jobs, and workflows in production environments.
  • Implement data governance, access controls, and security best practices using Unity Catalog.
  • Integrate data from a variety of source systems and destinations, supporting ETL/ELT and pipeline orchestration activities.
  • Optimize existing data workflows and Spark jobs for performance, reliability, and cost efficiency.
  • Integrate Databricks development with CI/CD pipelines and enterprise SDLC tooling, including Git-based version control.
  • Collaborate with data scientists and analysts to define data models and support machine learning and AI use cases, including model lifecycle management with MLflow.
  • Support advanced analytics use cases such as anomaly detection, risk scoring, and fraud analytics.
  • Monitor and troubleshoot data processing jobs, implementing data quality checks and observability to ensure high availability.
  • Document data processes, frameworks, pipelines, and data mappings for technical and non-technical audiences.
  • Work closely with scrum teams, product owners, and client stakeholders to deliver end-to-end data solutions.
  • Stay current on Databricks platform capabilities and industry trends to recommend best-fit tools and technologies.

TAG: #LI-I4DM

Requirements

Qualifications

  • Bachelor's degree in Computer Science, Information Technology, Engineering, or a related field (or equivalent experience).
  • 4+ years of experience in data engineering, analytics engineering, or big data development.
  • 2+ years of hands-on experience with the Databricks platform.
  • Proficiency in Apache Spark, PySpark, and Spark SQL.
  • Experience with Databricks clusters, jobs/workflows, Delta Lake, and Unity Catalog in production environments.
  • Experience with medallion architecture and Spark Structured Streaming.
  • Strong Python and SQL skills for data engineering and data analysis.
  • Experience with ETL/ELT processes and data pipeline orchestration.
  • Familiarity with cloud platforms such as AWS, Azure, or Google Cloud and their native data services.
  • Experience integrating data solutions with CI/CD pipelines and Git-based version control workflows.
  • Understanding of data governance, security, and access control best practices.
  • Experience working in Agile development environments.
  • Ability to obtain and maintain a Public Trust determination.
  • Excellent analytical, problem-solving, and communication skills, with the ability to work with both technical and non-technical stakeholders.

Preferred Qualifications

  • Databricks certification (e.g., Databricks Certified Data Engineer Associate/Professional) or cloud platform certification.
  • Experience implementing ML or AI solutions in Databricks, including MLflow-based model lifecycle management.
  • Knowledge of machine learning, AI, or Natural Language Processing (NLP) techniques, including text mining.
  • Experience supporting fraud analytics, risk scoring, or anomaly detection.
  • Experience with distributed data and streaming tools such as Kafka, Hadoop, Hive, or Amazon EMR.
  • Experience with data quality frameworks and observability/monitoring tooling.
  • Experience with NoSQL databases.
  • Experience with visualization packages such as Plotly, Seaborn, or ggplot2.
  • Experience supporting federal government or regulated-industry programs, especially the Department of Veterans Affairs.

Similar Jobs

4 Days Ago
Remote or Hybrid
OH, USA
Senior level
Senior level
Financial Services
Build and operate scalable Databricks-on-AWS data pipelines using PySpark, Delta Lake, and lakehouse patterns. Optimize performance, implement data quality, monitoring, alerting, and automated remediation, and deliver curated datasets for BI and analytics partners. Collaborate with stakeholders on architecture and design while applying secure software engineering, CI/CD, agile, and operational stability practices. The role also uses AI-assisted development tools and supports workforce data analytics.
Top Skills: AlteryxAmazon AthenaAmazon EmrAmazon S3Apache AirflowApache IcebergSparkAutosysAWSAws CloudwatchAws GlueAws LambdaBitbucketClaudeDatabricksDatabricks WorkflowsDelta LakeDelta Live TablesGitGithub CopilotJavaJenkinsOracleParquetPysparkPythonScalaSigmaSpinnakerSQLTableau
Yesterday
Remote
USA
Senior level
Senior level
Information Technology • Professional Services • Consulting
Build and maintain production-grade batch and streaming data pipelines in Databricks using PySpark and SQL. Develop ingestion frameworks, data models, APIs, data quality checks, monitoring, and Medallion Architecture across bronze, silver, and gold layers. Integrate diverse data sources, optimize Spark workloads, document data lineage and architecture, and collaborate with engineering and federal stakeholders. An active Secret clearance or higher is required.
Top Skills: SparkAPIsCi/CdDatabricksDatabricks Auto LoaderDatabricks WorkflowsDelta LakeDelta Live TablesGitKafkaKinesisMedallion ArchitecturePysparkPythonSQLUnity Catalog
6 Days Ago
In-Office or Remote
142K-213K Annually
Senior level
142K-213K Annually
Senior level
Aerospace • Logistics • Security • Software • Cybersecurity
Designs and implements Databricks Unity Catalog access controls, security architecture, identity governance, data classification, and compliance automation. Leads RBAC and ABAC models, dynamic views, row- and column-level security, masking, service principal integration, and infrastructure-as-code deployments using Terraform or Databricks Asset Bundles. Establishes security guardrails, templates, tagging standards, and scalable workflows while partnering with data, platform, and workspace engineering teams.
Top Skills: AbacColumn MaskingDatabricksDatabricks Asset BundlesDatabricks LakehouseDatabricks Unity CatalogDynamic ViewsEntra IdInfrastructure As CodePythonRbacRow-Level SecurityScimService PrincipalsSQLTerraform

What you need to know about the Austin Tech Scene

Austin has a diverse and thriving tech ecosystem thanks to home-grown companies like Dell and major campuses for IBM, AMD and Apple. The state’s flagship university, the University of Texas at Austin, is known for its engineering school, and the city is known for its annual South by Southwest tech and media conference. Austin’s tech scene spans many verticals, but it’s particularly known for hardware, including semiconductors, as well as AI, biotechnology and cloud computing. And its food and music scene, low taxes and favorable climate has made the city a destination for tech workers from across the country.

Key Facts About Austin Tech

  • Number of Tech Workers: 180,500; 13.7% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Dell, IBM, AMD, Apple, Alphabet
  • Key Industries: Artificial intelligence, hardware, cloud computing, software, healthtech
  • Funding Landscape: $4.5 billion in VC funding in 2024 (Pitchbook)
  • Notable Investors: Live Oak Ventures, Austin Ventures, Hinge Capital, Gigafund, KdT Ventures, Next Coast Ventures, Silverton Partners
  • Research Centers and Universities: University of Texas, Southwestern University, Texas State University, Center for Complex Quantum Systems, Oden Institute for Computational Engineering and Sciences, Texas Advanced Computing Center

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account