Sparksoft Corporation Logo

Sparksoft Corporation

Sr. Data Integration Engineer

Posted 3 Days Ago
Remote
Hiring Remotely in Maryland, USA
Senior level
Remote
Hiring Remotely in Maryland, USA
Senior level
Designs and operates resilient batch and event-driven data pipelines using Python and AWS services. Responsibilities include ingestion, transformation, schema validation, data quality, monitoring, security, testing, infrastructure automation, performance optimization, incident resolution, documentation, and customer requirements management. The role also supports Agile story development, acceptance criteria, UAT, risk analysis, stakeholder communication, and production maintenance.
The summary above was generated by AI

Join us at Sparksoft, where we're not just another tech company—we're a catalyst for change. Our mission isn't just to offer IT solutions; it's to revolutionize the way you work. Here, passion isn't just a buzzword; it's the fuel behind groundbreaking ideas and transformative technologies. We serve a wide range of government clients, delivering impact that's felt across the nation.

Our true strength lies in our people. They're the problem-solvers and innovators consistently delivering extraordinary outcomes. With Sparksoft, you're not stepping into a routine job; you're joining a team committed to innovation and excellence. Our innovation extends beyond just delivering projects. Through our specialized Innovation Centers, we continuously refine our methods, ensuring we remain industry leaders.

We are Sparksoft!

ROLE AND RESPONSIBILITIES:

  • Design and develop resilient batch and event-driven data pipelines using Python, AWS Glue, and appropriate AWS managed services.
  • Build AWS Glue jobs, crawlers, workflows, triggers, and Data Catalog integrations to discover, transform, govern, and publish datasets.
  • Ingest and process structured and semi-structured data from files, APIs, databases, and streaming or messaging sources, with particular expertise in JSON and NDJSON formats.
  • Develop efficient Python components for parsing, schema validation, normalization, enrichment, deduplication, aggregation, and data quality checks.
  • Use AWS services such as Amazon S3, AWS Lambda, Amazon EventBridge, AWS Step Functions, Amazon SQS, Amazon SNS, Amazon Kinesis, Amazon Athena, Amazon Redshift, AWS Lake Formation, AWS Secrets Manager, AWS KMS, Amazon CloudWatch, and AWS IAM as solution needs dictate.
  • Create automated unit, integration, regression, and data reconciliation tests; embed data quality controls throughout the pipeline lifecycle.
  • Implement operational monitoring, logging, alerting, traceability, restartability, error handling, and recovery mechanisms for production pipelines.
  • Apply security and privacy requirements through least-privilege access, encryption, secure secret management, audit logging, and appropriate handling of sensitive data.
  • Automate infrastructure and deployment processes using infrastructure as code and CI/CD practices.
  • Optimize pipeline performance, reliability, scalability, and cost through profiling, tuning, service selection, and ongoing operational analysis.
  • Investigate and resolve performance issues, failed jobs, and production incidents; document root causes and preventive actions.
  • Create and maintain technical documentation, including source-to-target mappings, pipeline designs, data contracts, runbooks, lineage, test evidence, and operational procedures.
  • Analyze and organize requirements into stories under epics, generate acceptance criteria, lead refinement sessions and work with Product Owners on prioritization.
  • Work with external teams on timelines and raise risks appropriately.
  • Maintain continuous communication with the customer, project SMEs, and key stakeholders to collect and document business requirements in support of their vision.
  • Review test scenarios and work with team to include any missed impact points.
  • Understand project delivery mechanisms and ensure owned stories/epics are tracked to closure.
  • Track customer requirements from inception through delivery.
  • Assist with user acceptance testing (UAT).
  • Analyze data to understand business problems and opportunities.
  • Identify and evaluate potential risks and impacts of proposed solutions.
  • Provide ongoing support and maintenance for implemented solutions.
  • Assist the product owner and development team to achieve customer satisfaction.

REQUIRED EXPERIENCE:

  • 7+ years of relevant experience
  • Strong Python development skills, including modular design, testing, debugging, packaging, dependency management, and performance optimization.
  • Hands-on experience developing data pipeline solutions with AWS Glue and integrating Glue with Amazon S3 and the AWS Glue Data Catalog.
  • Practical experience with multiple AWS data, integration, security, and monitoring services used to deliver end-to-end data pipelines.
  • Demonstrated experience ingesting, parsing, validating, transforming, and troubleshooting JSON and NDJSON, including nested structures, malformed records, schema drift, and large-file processing.
  • Experience with data modeling, schema design, data partitioning, metadata, lineage, and data quality practices.
  • Experience using Git-based version control, automated testing, CI/CD pipelines, and infrastructure-as-code approaches.
  • Excellent analytical, problem-solving, documentation, and communication skills, with a proactive and customer-focused approach.
  • Exhibit strong verbal and written communication skills, attention to detail, and the ability to follow up in a timely manner.
  • Have experience creating detailed reports and presenting information to both technical and non-technical audiences.
  • Possess expertise in using JIRA and Confluence for managing requirements.
  • Must be able to obtain and maintain a Public Trust clearance.
  • Must have lived in the United States 3 out of the past 5 years.

PREFERRED EXPERIENCE:

  • SAFe Agile Certification
  • AWS certification relevant to data engineering, architecture, or development.
  • Experience in healthcare IT and understanding of regulatory requirements such as HIPAA.
  • Experience designing source-to-target mappings, canonical data models, and data integration patterns across heterogeneous data providers.
  • Experience partnering with Data Quality and Data Governance teams to establish data quality metrics, validation rules, profiling processes, and remediation workflows.
  • Experience managing data delivery requirements, service level agreements (SLAs), and operational readiness processes.
  • Experience and/or knowledge of CMS programs, processes, and standards.
  • Experience with Apache Spark or PySpark, Parquet, Avro, Iceberg, or other distributed processing and open table or columnar storage technologies.

EDUCATION AND CERTIFICATIONS:

  • Bachelor's degree

WHAT WE OFFER: 

At Sparksoft, we know that people do their best work when they feel supported, inspired, and connected. That’s why we’ve built a workplace that balances comprehensive benefits with a culture of collaboration and innovation. From flexible time off to professional growth opportunities, we’re committed to helping you thrive both inside and outside of work. When you join Sparksoft, you’ll enjoy:

•    Competitive compensation and a 401(k) with employer contributions to help you plan for the future
•    Flexible paid time off and hybrid ways of working that support true work-life balance
•    Comprehensive health coverage—including medical, dental, vision, life, and disability insurance
•    A curated in-office experience designed to foster community, team connections, and innovation
•    Opportunities to give back through Sparksoft Cares, including annual company-wide fundraising events
•    Training and development programs that build new skills and prepare you for leadership roles
•    A collaborative, transparent, and fun culture—recognized as a Great Place to Work®

Accessibility and Accommodations: Sparksoft Corporation is committed to providing equal employment opportunities to all individuals. If you require accommodations during the application or interview process, please contact us at [email protected] or call 410-424-7700. Requests are reviewed and fulfilled on a case-by-case basis.

Security Notice: Your privacy and data security are important to us. Sparksoft Corporation will never request sensitive personal information via email. If you receive any suspicious communication claiming to be from Sparksoft, please report it immediately to our security team at [email protected].

Artificial Intelligence (AI) Policy: While Sparksoft values the appropriate use of artificial intelligence in the workplace, candidates must complete interviews and independent assessments using their own knowledge and abilities. Unless expressly permitted by Sparksoft, the use of AI during interviews and assessments is strictly prohibited. Violations may result in disqualification. Please contact your recruiter regarding accommodation needs.


Similar Jobs

17 Minutes Ago
Remote or Hybrid
United States
71K-88K Annually
Junior
71K-88K Annually
Junior
Digital Media • Gaming • Information Technology • Software • Sports • Esports • Big Data Analytics
Administer and enhance Salesforce for VIP teams by configuring flows, reports, dashboards, access, and data management. Gather stakeholder requirements, troubleshoot user issues, support adoption, and collaborate with development, analytics, and data engineering teams on integrations and technical solutions. Query and validate data using SQL and Snowflake, test system enhancements, and identify process improvements while maintaining Salesforce data accuracy and integrity.
Top Skills: ApexExcelGoogle SheetsSalesforceSalesforce Flow BuilderSnowflakeSQL
3 Hours Ago
Remote or Hybrid
USA
Senior level
Senior level
Machine Learning • Payments • Security • Software • Financial Services
Owns the vision, customer focus, and product backlog for a near-real-time data product. Prioritizes work based on business value, leads backlog grooming, communicates product direction, and partners with Scrum Masters and development teams to ensure delivery aligns with client requirements and business objectives.
Top Skills: Agile DevelopmentData VisualizationScrumUx Design
7 Hours Ago
Remote or Hybrid
USA
38K-75K Annually
Mid level
38K-75K Annually
Mid level
Machine Learning • Payments • Security • Software • Financial Services
Provide high-volume, phone-based technical support and first-line resolution for hardware, software, networking and access issues. Log and track incidents in ServiceNow/ITSM, escalate complex problems, suggest process improvements, and maintain security and confidentiality while working from an approved remote workspace on a set evening/weekend schedule.
Top Skills: Call Center TechnologiesItsmPhone SystemsServicenowTicketing Systems

What you need to know about the Austin Tech Scene

Austin has a diverse and thriving tech ecosystem thanks to home-grown companies like Dell and major campuses for IBM, AMD and Apple. The state’s flagship university, the University of Texas at Austin, is known for its engineering school, and the city is known for its annual South by Southwest tech and media conference. Austin’s tech scene spans many verticals, but it’s particularly known for hardware, including semiconductors, as well as AI, biotechnology and cloud computing. And its food and music scene, low taxes and favorable climate has made the city a destination for tech workers from across the country.

Key Facts About Austin Tech

  • Number of Tech Workers: 180,500; 13.7% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Dell, IBM, AMD, Apple, Alphabet
  • Key Industries: Artificial intelligence, hardware, cloud computing, software, healthtech
  • Funding Landscape: $4.5 billion in VC funding in 2024 (Pitchbook)
  • Notable Investors: Live Oak Ventures, Austin Ventures, Hinge Capital, Gigafund, KdT Ventures, Next Coast Ventures, Silverton Partners
  • Research Centers and Universities: University of Texas, Southwestern University, Texas State University, Center for Complex Quantum Systems, Oden Institute for Computational Engineering and Sciences, Texas Advanced Computing Center

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account