Alpaca Logo

Alpaca

Senior Data Engineer

Posted 17 Days Ago
Remote
Hiring Remotely in USA
Senior level
Remote
Hiring Remotely in USA
Senior level
Design and operate Alpaca’s scalable data platform on GCP, including lakehouse infrastructure, Kubernetes-based deployments, streaming and CDC pipelines, batch ingestion, BI access, cataloging, monitoring, and governance. Build infrastructure as code with Terraform and Ansible, operate distributed query engines and Apache Iceberg, maintain reliability practices, and collaborate with DevOps and analytics teams to support rapidly evolving data requirements.
The summary above was generated by AI

Who We Are:

Alpaca is a US-headquartered, global leader in agent-first brokerage infrastructure for stocks, ETFs, options, crypto, fixed income, 24/5 trading, and more.
Amongst our subsidiaries, Alpaca is a licensed financial services company, serving hundreds of financial institutions across 40 countries with our institutional-grade APIs. This includes broker-dealers, investment advisors, wealth managers, hedge funds, and crypto exchanges, totalling over 10 million brokerage accounts.
Our global team is a diverse group of experienced engineers, traders, and brokerage professionals who are working to achieve our mission of opening financial services to everyone on the planet. We're deeply committed to open-source contributions and fostering a vibrant community, continuously enhancing our award-winning, developer-friendly API and the robust infrastructure behind it.
Alpaca is proudly backed by $400 million in funding from top-tier global investors including Portage Ventures, Spark Capital, Tribe Capital, Social Leverage, Horizons Ventures, Opera Tech Ventures, SBI Group, Derayah Financial, Unbound, Peak XV, Elefund, and Y Combinator.
Our Team Members:

We're a dynamic team of 400+ globally distributed members who thrive working from our favorite places around the world, with teammates spanning the USA, Canada, Japan, Hungary, Nigeria, Brazil, the UK, and beyond!
We're searching for passionate individuals eager to contribute to Alpaca's rapid growth. If you align with our core values—Stay Curious, Have Empathy, and Be Accountable—and are ready to make a significant impact, we encourage you to apply.

Your Role: We are seeking a Senior Data Engineer to help design and build the next generation of our Data Platform as we continue to scale to larger customers and new jurisdictions. At Alpaca, Data Engineering encompasses financial transactions, customer data, API logs, system metrics, augmented data, and third-party systems that impact decision making for both internal and external stakeholders. We process hundreds of millions of events daily, and this number continues to grow as we onboard new customers and products. 

We prioritize open-source technologies in our data stack while leveraging Google Cloud Platform (GCP) as the foundation for our data infrastructure. This spans batch and stream ingestion, transformation, and consumption layers for BI/Reporting, AI/agent interfaces (MCP), and external third-party sinks. We also oversee data experimentation, cataloging, and monitoring/alerting systems. 

Our team is 100% distributed and remote. 

Responsibilities:

  • Design, build, and evolve the core data platform infrastructure e.g., distributed query engines, orchestration, warehousing, cataloging, and more. 
  • Own our lakehouse infrastructure as code, managing deployments through Terraform and Ansible on Kubernetes. 
  • Build and maintain low-latency streaming and CDC ingestion pipelines, as well as batch ingestion paths landing in Iceberg. 
  • Develop and scale our BI landscape so downstream teams and agents get performant, self-serve access to lakehouse data. 
  • Enforce platform reliability best practices, including monitoring and alerting, on-call rotations, incident response, maintenance windows, runbooks, and SLAs. 
  • Partner with DevOps, Analytics Engineering, and other stakeholders to close infrastructure gaps and support new data requirements.

Must-Haves:

  • 5+ years of experience in Data Engineering, including 2+ years building and operating scalable, low-latency data platforms handling > 100M events/day.
  • Strong hands-on experience running data infrastructure on Kubernetes, with cloud-native tooling like Docker and Helm.
  • Production experience with IaC: Terraform, Ansible, and ArgoCD (or equivalents).
  • Deep knowledge of distributed systems (storage, transactions, and query processing) with hands-on experience operating open-source query engines like Trino or Presto.
  • Strong experience with object storage and open table formats, specifically Apache Iceberg.
  • Experience with streaming and CDC systems: Kafka, Redpanda, and Debezium.
  • Hands-on experience with orchestration frameworks (Airflow) and ELT tools (Airbyte).
  • Strong working knowledge of Python and SQL for building pipelines and platform tooling. 
  • Experience with Google Cloud Platform and its data services (GCS, Cloud Build, Cloud SQL, Dataproc, etc); or related experience with other cloud services. 
  • Ability to thrive in a fast-paced startup environment and adapt infrastructure to rapidly changing needs.

 Nice to Haves:

  • Experience with semantic/metrics layers (Cube, dbt, Looker).
  • Familiarity with transformation frameworks (dbt).
  • Familiarity with reverse ETL tooling (Hightouch)
  • Familiarity with data catalog and lineage tooling (OpenMetadata, Datahub)
  • Experience with data access control and governance frameworks (Apache Ranger)
How We Take Care of You:
  • Competitive Salary & Stock Options
  • Health Benefits
  • New Hire Home-Office Setup: One-time USD $500
  • Monthly Stipend: USD $150 per month via a Brex Card

Alpaca is proud to be an equal opportunity workplace dedicated to pursuing and hiring a diverse workforce.

Recruitment Privacy Policy

Similar Jobs

2 Days Ago
Remote
Senior level
Senior level
Artificial Intelligence • Cloud • Professional Services • Software
Designs and evolves scalable data pipelines, warehouses, lakes, and streaming architectures. Builds ETL/ELT processes, ensures data quality and security, optimizes database performance, and translates complex requirements into technical solutions. Collaborates with data scientists and stakeholders on modeling while mentoring junior engineers. The role requires modern data stack, cloud, programming, data modeling, communication, and client-facing nearshore or offshore experience.
Top Skills: Amazon RedshiftApache AirflowApache KafkaSparkAWSBigQueryDatabricksDbtGoogle Cloud PlatformAzurePythonScalaSnowflakeSQL
2 Days Ago
Remote
Senior level
Senior level
Artificial Intelligence • Cloud • Professional Services • Software
Designs and optimizes semantic data models, analytical views, and production ETL/ELT pipelines for AI and analytics workloads. Tunes databases for low-latency retrieval, implements RBAC, row-level security, data masking, governance, monitoring, and quality standards, and collaborates with AI engineering teams on RAG and agentic AI systems.
Top Skills: Agentic Ai FrameworksApache AirflowApache KafkaAws KinesisAzure SqlDagsterDbtPgvectorPineconePostgresPrefectQdrantRag ApplicationsSQLTerraform
3 Days Ago
In-Office or Remote
Senior level
Senior level
Information Technology
Designs, builds, and maintains large-scale data pipelines using Databricks, PySpark, and Apache Spark. Develops ETL/ELT processes, lakehouse architectures, and optimized data workflows while ensuring quality, reliability, performance, and observability. Collaborates with architects, analysts, and stakeholders to deliver scalable data solutions. Uses orchestration, version control, CI/CD, and automation practices to support modern analytics platforms.
Top Skills: Apache AirflowSparkAWSAzureAzure Data FactoryCi/CdDatabricksDelta LakeGitJavaMicrosoft FabricPysparkPythonSQL

What you need to know about the Austin Tech Scene

Austin has a diverse and thriving tech ecosystem thanks to home-grown companies like Dell and major campuses for IBM, AMD and Apple. The state’s flagship university, the University of Texas at Austin, is known for its engineering school, and the city is known for its annual South by Southwest tech and media conference. Austin’s tech scene spans many verticals, but it’s particularly known for hardware, including semiconductors, as well as AI, biotechnology and cloud computing. And its food and music scene, low taxes and favorable climate has made the city a destination for tech workers from across the country.

Key Facts About Austin Tech

  • Number of Tech Workers: 180,500; 13.7% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Dell, IBM, AMD, Apple, Alphabet
  • Key Industries: Artificial intelligence, hardware, cloud computing, software, healthtech
  • Funding Landscape: $4.5 billion in VC funding in 2024 (Pitchbook)
  • Notable Investors: Live Oak Ventures, Austin Ventures, Hinge Capital, Gigafund, KdT Ventures, Next Coast Ventures, Silverton Partners
  • Research Centers and Universities: University of Texas, Southwestern University, Texas State University, Center for Complex Quantum Systems, Oden Institute for Computational Engineering and Sciences, Texas Advanced Computing Center

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account