Photon Logo

Photon

SPARK Data Onboarding Engineer- Jersey City

Reposted 21 Days Ago
Be an Early Applicant
In-Office or Remote
Hiring Remotely in United States
43K-151K Annually
Senior level
In-Office or Remote
Hiring Remotely in United States
43K-151K Annually
Senior level
Design, build, and optimize PySpark applications and ETL pipelines to process large-scale datasets from SQL/NoSQL sources, data lakes, and streaming platforms. Ensure data quality, error handling, performance tuning, and collaborate with analysts, scientists, and architects to deliver scalable data solutions.
The summary above was generated by AI

Job Title: PySpark Data Engineer

Summary:

We are seeking a skilled PySpark Data Engineer to join our team and drive the development of robust data processing and transformation solutions within our data platform. You will be responsible for designing, implementing, and maintaining PySpark-based applications to handle complex data processing tasks, ensure data quality, and integrate with diverse data sources. The ideal candidate possesses strong PySpark development skills, experience with big data technologies, and the ability to work in a fast-paced, data-driven environment.

Key Responsibilities: Data Engineering Development:

  • Design, develop, and test PySpark-based applications to process, transform, and analyze large-scale datasets from various sources, including relational databases, NoSQL databases, batch files, and real-time data streams.
  • Implement efficient data transformation and aggregation using PySpark and relevant big data frameworks.
  • Develop robust error handling and exception management mechanisms to ensure data integrity and system resilience within Spark jobs.
  • Optimize PySpark jobs for performance, including partitioning, caching, and tuning of Spark configurations.

Data Analysis and Transformation:

  • Collaborate with data analysts, data scientists, and data architects to understand data processing requirements and deliver high-quality data solutions.
  • Analyze and interpret data structures, formats, and relationships to implement effective data transformations using PySpark.
  • Work with distributed datasets in Spark, ensuring optimal performance for large-scale data processing and analytics.

Data Integration and ETL:

  • Design and implement ETL (Extract, Transform, Load) processes to ingest and integrate data from various sources, ensuring consistency, accuracy, and performance.
  • Integrate PySpark applications with data sources such as SQL databases, NoSQL databases, data lakes, and streaming platforms

Qualifications and Skills:

  • Bachelor's degree in Computer Science, Information Technology, or a related field.
  • 5+ years of hands-on experience in big data development, preferably with exposure to data-intensive applications.
  • Strong understanding of data processing principles, techniques, and best practices in a big data environment.
  • Proficiency in PySpark, Apache Spark, and related big data technologies for data processing, analysis, and integration.
  • Experience with ETL development and data pipeline orchestration tools (e.g., Apache Airflow, Luigi).
  • Strong analytical and problem-solving skills, with the ability to translate business requirements into technical solutions.
  • Excellent communication and collaboration skills to work effectively with data analysts, data architects, and other team members.

Compensation, Benefits and Duration

Minimum Compensation: USD 43,000
Maximum Compensation: USD 151,000
Compensation is based on actual experience and qualifications of the candidate. The above is a reasonable and a good faith estimate for the role.
Medical, vision, and dental benefits, 401k retirement plan, variable pay/incentives, paid time off, and paid holidays are available for full time employees.
This position is not available for independent contractors
No applications will be considered if received more than 120 days after the date of this post



Similar Jobs

38 Minutes Ago
Easy Apply
Remote or Hybrid
Location, WV, USA
Easy Apply
175K-250K Annually
Senior level
175K-250K Annually
Senior level
Cloud • Information Technology • Security • Software • Cybersecurity
Lead and establish Business Operations for global Customer Success (Support, Technical Success, Professional Services). Build operating model, manage planning and P&L, drive headcount and budget planning, define KPIs and dashboards, run business cadences (MBR/QBR/AOP), and partner with Finance and cross-functional stakeholders to improve resource management and operational efficiency.
Top Skills: Ai Tools
42 Minutes Ago
Remote or Hybrid
US
78K-108K Annually
Senior level
78K-108K Annually
Senior level
Information Technology
Implement and operationalize Managed Services for Government Cloud customers across Azure Government, Azure GCC High, AWS GovCloud or Google Assured Workloads. Design secure landing zones, automation, monitoring, governance, and operational runbooks; support onboarding, transitions, compliance (FedRAMP, NIST, CMMC, etc.), incident handling, and production handoff. Collaborate with architects, security, PMs, and Managed Operations; mentor junior engineers and contribute to standards and reusable automation.
Top Skills: AutomationAws GovcloudAzure Gcc HighAzure GovernmentBackupCjisCloud GovernanceCloud Management PlatformsCmmcConfiguration ManagementDhcpDisaster RecoveryDnsEncryptionFedrampFirewallsGoogle Cloud Assured WorkloadsIdentityInfrastructure As CodeItarItsmLoad BalancingLoggingMonitoringNetworkingNist 800-53ObservabilityOrchestrationRoutingRunbooksScriptingSegmentationTemplatesVpnVulnerability Management
43 Minutes Ago
Remote or Hybrid
US
96K-150K Annually
Senior level
96K-150K Annually
Senior level
Information Technology
Manage, maintain, and improve secure government cloud, hybrid, and multi-cloud infrastructure (Azure Gov/Azure GCC High, AWS GovCloud, GCP). Provide 24x7 operational support, incident/change management, advanced troubleshooting, compliance with FedRAMP/NIST/DISA/CIS standards, automation/IaC development, capacity and performance management, and technical guidance to improve architecture, security, and service delivery.
Top Skills: Amazon Web ServicesAutomationAws GovcloudAzure Gcc HighAzure GovernmentBackupCis BenchmarksDatabasesDisa StigsEnterprise StorageFedrampGCPGoogle Cloud PlatformIacIdentity And Access ManagementInfrastructure As CodeItilLinuxAzureMonitoringNist 800-53Nist Cybersecurity FrameworkObservabilityService ManagementUnixVirtualizationWindows

What you need to know about the Austin Tech Scene

Austin has a diverse and thriving tech ecosystem thanks to home-grown companies like Dell and major campuses for IBM, AMD and Apple. The state’s flagship university, the University of Texas at Austin, is known for its engineering school, and the city is known for its annual South by Southwest tech and media conference. Austin’s tech scene spans many verticals, but it’s particularly known for hardware, including semiconductors, as well as AI, biotechnology and cloud computing. And its food and music scene, low taxes and favorable climate has made the city a destination for tech workers from across the country.

Key Facts About Austin Tech

  • Number of Tech Workers: 180,500; 13.7% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Dell, IBM, AMD, Apple, Alphabet
  • Key Industries: Artificial intelligence, hardware, cloud computing, software, healthtech
  • Funding Landscape: $4.5 billion in VC funding in 2024 (Pitchbook)
  • Notable Investors: Live Oak Ventures, Austin Ventures, Hinge Capital, Gigafund, KdT Ventures, Next Coast Ventures, Silverton Partners
  • Research Centers and Universities: University of Texas, Southwestern University, Texas State University, Center for Complex Quantum Systems, Oden Institute for Computational Engineering and Sciences, Texas Advanced Computing Center

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account