Zifo Logo

Zifo

Senior Data Engineer

Posted Yesterday
Remote
Hiring Remotely in United States
Senior level
Remote
Hiring Remotely in United States
Senior level
Designs and operates data infrastructure for a clinical trial platform. Builds pipelines that extract, normalize, validate, and link information from clinical documents and publications into Aurora and GraphDB. Develops embeddings, vector search, RAG workflows, data-serving APIs, lineage tracking, and regulatory audit trails. The role also requires AWS-based orchestration, clinical data quality controls, biomedical knowledge graph modeling, and deployment automation using infrastructure-as-code and containerization.
The summary above was generated by AI

Design, build, and operate the data infrastructure for a clinical trial design platform that extracts insights from clinical protocols, regulatory documents, and published articles to accelerate trial design decisions. This role owns the pipelines that ingest, parse, transform, and serve clinical data - from raw PDF extraction through structured storage in Aurora and GraphDB, to serving curated knowledge for LangGraph-based agentic AI workflows.


Requirements
  • Build ingestion pipelines for clinical trial protocols, ICF documents, SmPCs, CSRs, and published articles (PubMed, CTIS, ClinicalTrials.gov) - handling PDF parsing, text extraction, and structured data normalization
  • Design and implement data models in Amazon Aurora (relational) and GraphDB (knowledge graph) to represent trial design entities: endpoints, eligibility criteria, study arms, interventions, therapeutic areas, and their relationships
  • Develop embedding and vectorization pipelines to prepare extracted clinical text for RAG-based retrieval in LangGraph agentic workflows - chunking strategies, metadata enrichment, and vector store population
  • Build and maintain ETL/ELT workflows that transform unstructured clinical content into queryable, linked data across both relational and graph stores
  • Implement data quality validation specific to clinical data - protocol section classification accuracy, entity extraction completeness, cross-reference integrity (NCT IDs, EudraCT numbers, MeSH terms)
  • Build data serving APIs (Python/FastAPI) that expose curated datasets to the Angular frontend and LangGraph agent layer
  • Set up data lineage tracking and audit trails to support regulatory traceability of AI-generated trial design recommendation
Required Skills
  • Strong Python development experience, including PDF/document parsing libraries such as PyMuPDF, pdfplumber, unstructured.io, or similar.
  • Advanced PostgreSQL-compatible SQL, including Amazon Aurora; experience with schema design, migrations, query optimization, and indexing strategies for large clinical datasets.
  • Hands-on experience with Neptune, Neo4j, or similar graph databases; proficiency in SPARQL or Cypher; experience with ontology and knowledge graph modeling for biomedical entities.
  • Experience with AWS services including Aurora PostgreSQL, S3, Lambda, Step Functions, SQS/SNS, and IAM, particularly for data pipeline orchestration.
  • Experience with PDF text extraction, document section classification, and named entity recognition (NER) for clinical/biomedical text; familiarity with embedding models and vector stores such as OpenSearch, pgvector, or Pinecone.
  • Experience building data-serving APIs using FastAPI, including asynchronous programming patterns and backend integration.
  • Experience preparing data for LangChain/LangGraph applications and designing RAG pipelines, including chunking, retrieval, reranking, and prompt-data integration.
  • Experience with Airflow, Prefect, AWS Step Functions, Temporal, or similar workflow orchestration tools; ability to design multi-stage DAGs with dependency management, retry logic, monitoring, and error handling.
  • Experience with Terraform or AWS CDK, Docker, and Git, including automated pipeline testing and deployment on AWS.

Domain Knowledge

  • Understanding of clinical trial structure: protocol sections (objectives, endpoints, eligibility criteria, study design, statistical considerations)
  • Familiarity with clinical data standards or terminologies (MeSH, MedDRA, SNOMED, ATC codes, CDISC) is a strong plus
  • Awareness of regulatory data integrity requirements (21 CFR Part 11, EU Annex 11, ALCOA+ principles)

Nice to Have

  • Experience with biomedical knowledge graphs (e.g., linking drugs -> targets -> diseases > trials)
  • Prior work with PubMed/MEDLINE data, ClinicalTrials.gov API, or EMA/CTIS data.
  • Apache Spark or Databricks for batch processing of large document corpora
  • dbt for transformation layer management over Aurora

Benefits

CURIOSITY DRIVEN, SCIENCE FOCUSED, EMPLOYEE BUILT. Our culture is unlike any other, one where we debate, challenge ourselves, and interact with all alike. We are a curious bunch, characterized by our passion to learn and spirit of teamwork. Zifo is a global R&D solutions provider focused on the industries of Pharma, Biotech, Manufacturing QC, Medical Devices, specialty chemicals and other research-based organizations. Our team’s knowledge of science and expertise in technology help Zifo better serve our customers around the globe, including 18 of the Top 20 Biopharma companies. 

We look for Science – Biotechnology, Pharmaceutical Technology, Biomedical Engineering, Microbiology etc. We possess scientific and technical knowledge and bear professional and personal goals.  While we have a “no doors” policy to promote free access within, we do have a tough door to walk in. We search with a two-point agenda – technical competency and cultural adaptability. 

We offer a competitive compensation package including accrued vacation, medical, dental, vision, 401k with company matching, life insurance, and flexible spending accounts.  

If you share these sentiments and are prepared for the atypical, then Zifo is your calling! 

Zifo is an equal opportunity employer, and we value diversity at our company.  We do not discriminate on the basis of race, religion, color, national origin, gender, sexual orientation, age, marital status, veteran status, or disability status. 


Similar Jobs

8 Hours Ago
Remote or Hybrid
United States
109K-183K Annually
Senior level
109K-183K Annually
Senior level
Artificial Intelligence • Cloud • Sales • Security • Software • Cybersecurity • Data Privacy
Design, build, and operate scalable batch and streaming data pipelines, lakehouse and warehouse models, and data services. Own datasets end to end, improve Snowflake, Iceberg, Spark, and Flink performance and reliability, implement governance and observability, and support graph-serving data models. Partner with product and engineering teams, participate in on-call, review code and designs, and mentor junior engineers.
Top Skills: AirflowApache CassandraApache FlinkApache IcebergSparkAWSAzureClaude CodeCloudFormationCursorDatadogDbtGithub CopilotGoogle Cloud PlatformGrafanaJavaKafkaKubernetesMlopsOpensearchPrometheusPythonScalaSnowflakeSQLTerraform
11 Days Ago
In-Office or Remote
92K-164K Annually
Senior level
92K-164K Annually
Senior level
Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
Senior Data Engineer responsible for Epic and EHR integrations supporting value-based care, risk and quality workflows, member attribution, roster management, and care-gap tools. The role designs and executes EHR development tasks, coordinates with clinical, data, operations, and development teams, documents business and data flows, evaluates AI and automation opportunities, and develops reporting to identify data-quality issues and prevent outages.
Top Skills: Ai ToolsCaboodleClarityEhr IntegrationsEpicHealthy PlanetSQL
15 Days Ago
Remote or Hybrid
United States
165K-235K Annually
Senior level
165K-235K Annually
Senior level
Big Data • Cloud • Productivity • Software • Database • Analytics • Automation
Build and maintain Databricks-based data platforms, including ingestion, transformation, storage, governance, data modeling, and serving pipelines. Establish medallion architecture standards, canonical data models, quality controls, lineage, schema evolution, and reliable batch or incremental processing. Improve pipeline observability, scalability, idempotency, and recoverability while moving curated data to systems such as ClickHouse. Collaborate across application and analytics teams to create durable, governed production datasets.
Top Skills: Amazon AuroraAmazon RdsApache AirflowSparkBigQueryCdcClickhouseCloud Object StorageDatabricksDelta LakeIamOpenmetadataPostgresSnowflakeUnity Catalog

What you need to know about the Austin Tech Scene

Austin has a diverse and thriving tech ecosystem thanks to home-grown companies like Dell and major campuses for IBM, AMD and Apple. The state’s flagship university, the University of Texas at Austin, is known for its engineering school, and the city is known for its annual South by Southwest tech and media conference. Austin’s tech scene spans many verticals, but it’s particularly known for hardware, including semiconductors, as well as AI, biotechnology and cloud computing. And its food and music scene, low taxes and favorable climate has made the city a destination for tech workers from across the country.

Key Facts About Austin Tech

  • Number of Tech Workers: 180,500; 13.7% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Dell, IBM, AMD, Apple, Alphabet
  • Key Industries: Artificial intelligence, hardware, cloud computing, software, healthtech
  • Funding Landscape: $4.5 billion in VC funding in 2024 (Pitchbook)
  • Notable Investors: Live Oak Ventures, Austin Ventures, Hinge Capital, Gigafund, KdT Ventures, Next Coast Ventures, Silverton Partners
  • Research Centers and Universities: University of Texas, Southwestern University, Texas State University, Center for Complex Quantum Systems, Oden Institute for Computational Engineering and Sciences, Texas Advanced Computing Center

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account