DataPelago Logo

DataPelago

Data Processing Engineer - I/O

Reposted 7 Days Ago
In-Office or Remote
Hiring Remotely in Mountain View, CA
Senior level
In-Office or Remote
Hiring Remotely in Mountain View, CA
Senior level
The Data Processing Engineer will architect and enhance data read/write capabilities for a data processing engine, focusing on large-scale data, performance optimization, and collaboration with engineering teams.
The summary above was generated by AI

Data Processing Engineer - I/O
Mountain View, CA / Hyderabad, IN / Remote

About DataPelago:

DataPelago is at the forefront of revolutionizing data processing for traditional analytics and cutting-edge GenAI preprocessing. We are building an innovative data processing engine that is transforming how Apache Spark, Apache Flink, Ray and others operate on diverse, large-scale data. Our team of engineers drive and adopt advances in hardware-accelerated computing, parallel processing of large-scale data, query optimization, distributed systems, compilers, machine learning, and cloud-native computing. We are looking for specialists to join our engineering team and shape the future of accelerated data processing.

The Opportunity:
As a Data Processing Engineer - I/O, you will be a key individual contributor in advancing data
read and write capabilities of DataPelago’s data processing engine. You will enhance functional
breadth, performance, scale, and reliability of the DataPelago engine in reading and writing large scale data of various data types from diverse data sources and data sinks. This is a unique opportunity to make a significant impact on a category-defining product and work with a talented team of engineers.

What You'll Do:
• Architect: Influence the architecture of how our data processing engine interfaces with data
sources and sinks, catalogs, data formats.
• Design: Lead design of functional and performance enhancements to adapters/connectors,
data representations, data filtering, caching and more in our data processing engine.

• Core Development: Individually design, implement, test, optimize, and maintain components of the data processing engine.

• Innovation and Differentiation: Analyze technology roadmap of existing and emerging data
formats and libraries, open table formats, catalog services, and more (e.g., Apache Arrow,

Apache Parquet, Apache Iceberg) and identify opportunities for our engine to enhance technology and product leadership.

• Collaboration: Partner effectively with engineering and product management in defining and
realizing the data I/O roadmap of our product..
• Continuous Improvement: Foster best practices in design and code reviews, testing, CI/CD,
and issue resolution to maintain highest product quality, security, efficiency, & productivity.

What You'll Bring:

• Bachelor's degree in Computer Science or a related field with 7+ years of relevant experience OR a Master's degree in Computer Science or a related field with 5+ years of relevant

experience.

• 3+ years of deep technical experience in developing and optimizing data read and write interfaces for large-scale data processing, particularly related to Apache Parquet, Apache

ORC, Apache Iceberg, Apache Spark, and similar technologies.

• Demonstrated experience in instrumenting, analyzing, and optimizing the performance of
data processing engine components on benchmark and customer workloads.

• Demonstrated experience in the design, development, and successful release of high-performance data processing engine features for large production deployments.

• Good knowledge of the architecture of one or more of Apache Spark, Apache Flink, Presto/
Trino.
• Exceptional programming skills in C, C++. Rust experience preferred.
• Extensive development experience in Linux environments.
• Strong analytical and problem-solving skills with a passion for performance optimization.

Location Considerations:

We value face-to-face collaboration, but recognize that talent can be found anywhere. Our engineering team works at our headquarters in Mountain View, CA, at our India office in Hyderabad, and at remote locations.

Why Join DataPelago?
• Technology Leadership: Shape the architecture and development of how our core engine
works with advanced data store platforms.
• Cutting-Edge Innovation: Work on challenging problems at the forefront of accelerated
computing and data processing.
• Significant Impact: Your contributions will directly impact the performance and scalability
of our mission-critical platform.
• Growth: Expand your technical expertise and scope of responsibilities working with other
talented engineers and with a growing product.

• Competitive compensation, stock options, comprehensive benefits package, leadership de-
velopment opportunities.

Similar Jobs

9 Hours Ago
Remote
USA
120K-180K Annually
Senior level
120K-180K Annually
Senior level
Artificial Intelligence • Machine Learning • Natural Language Processing • Software • Conversational AI
Own full-cycle global recruiting for senior, hard-to-fill technical roles across Research, Engineering, and Product. Build proactive sourcing strategies, define role scope and interview plans with hiring managers, evaluate technical depth for AI-native products, deliver excellent candidate experience across time zones, use funnel metrics to improve hiring velocity, and collaborate with People Ops, Legal, and Finance to scale compliant international hiring.
Top Skills: Ai/Ml SystemsAPIsLarge Language Models (Llms)Real-Time InferenceSpeech-To-Text (Stt)Text-To-Speech (Tts)
9 Hours Ago
Easy Apply
In-Office or Remote
United States
Easy Apply
180K-200K Annually
Senior level
180K-200K Annually
Senior level
Artificial Intelligence • Hardware • Healthtech • Software
The Senior Data Platform Engineer will manage and develop the data infrastructure on Databricks and AWS, ensuring scalable and efficient data capabilities while collaborating across teams.
Top Skills: AWSDatabricksKafkaKinesis
9 Hours Ago
Remote
USA
Senior level
Senior level
Artificial Intelligence • Machine Learning • Software • Defense
Lead the ATO process for classified environments, ensuring compliance with RMF and security standards while interfacing with government stakeholders.
Top Skills: AtoAws GovcloudAzure GovernmentDisa StigsEmassGoogle GovernmentKubernetesNist 800-53OpenshiftRmfXacta

What you need to know about the Austin Tech Scene

Austin has a diverse and thriving tech ecosystem thanks to home-grown companies like Dell and major campuses for IBM, AMD and Apple. The state’s flagship university, the University of Texas at Austin, is known for its engineering school, and the city is known for its annual South by Southwest tech and media conference. Austin’s tech scene spans many verticals, but it’s particularly known for hardware, including semiconductors, as well as AI, biotechnology and cloud computing. And its food and music scene, low taxes and favorable climate has made the city a destination for tech workers from across the country.

Key Facts About Austin Tech

  • Number of Tech Workers: 180,500; 13.7% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Dell, IBM, AMD, Apple, Alphabet
  • Key Industries: Artificial intelligence, hardware, cloud computing, software, healthtech
  • Funding Landscape: $4.5 billion in VC funding in 2024 (Pitchbook)
  • Notable Investors: Live Oak Ventures, Austin Ventures, Hinge Capital, Gigafund, KdT Ventures, Next Coast Ventures, Silverton Partners
  • Research Centers and Universities: University of Texas, Southwestern University, Texas State University, Center for Complex Quantum Systems, Oden Institute for Computational Engineering and Sciences, Texas Advanced Computing Center

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account