Design and operate petabyte-scale data infrastructure supporting AI training, evaluation, and continual improvement. Build multimodal ingestion, processing, quality, labeling, versioning, lineage, privacy, and high-throughput data delivery systems. Optimize storage, cost, performance, and GPU utilization while ensuring reproducibility and data integrity. Monitor data quality, drift, and pipeline health, collaborate with machine learning teams, and document operational procedures. Requires strong data engineering, distributed systems, software development, and ML infrastructure experience.
Infrastructure Reliability Engineer – Remote
Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States.
This is a fantastic opportunity to join an established and well-respected organization offering tremendous career growth potential.
Job Title: Infrastructure Reliability Engineer
Location: 100% Remote (U.S.)
Position Type: Full-time, Direct W2
Salary Range: $125,000–$170,000 Annually
Experience Required: 6+ years
Job Summary
We are seeking an Infrastructure Reliability Engineer to build and operate the large-scale data systems that power modern AI training and evaluation pipelines. The role combines deep data engineering expertise with a strong understanding of AI workloads, focusing on ingestion, transformation, quality assurance, lineage, and high-throughput delivery of data to training jobs across diverse modalities. The ideal candidate has experience operating petabyte-scale data systems, strong software engineering fundamentals, and clear understanding of how data infrastructure choices propagate into model quality and training efficiency.
Key Responsibilities
Required Qualifications
Preferred Qualifications
How to Apply
Would you like to know more about this opportunity? For immediate consideration, please send your resume to [email protected] or contact us at (908)676-4399. Learn more about Bright Vision Technologies at www.bvteck.com.
Bright Vision Technologies is an Equal Opportunity Employer.
Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States.
This is a fantastic opportunity to join an established and well-respected organization offering tremendous career growth potential.
Job Title: Infrastructure Reliability Engineer
Location: 100% Remote (U.S.)
Position Type: Full-time, Direct W2
Salary Range: $125,000–$170,000 Annually
Experience Required: 6+ years
Job Summary
We are seeking an Infrastructure Reliability Engineer to build and operate the large-scale data systems that power modern AI training and evaluation pipelines. The role combines deep data engineering expertise with a strong understanding of AI workloads, focusing on ingestion, transformation, quality assurance, lineage, and high-throughput delivery of data to training jobs across diverse modalities. The ideal candidate has experience operating petabyte-scale data systems, strong software engineering fundamentals, and clear understanding of how data infrastructure choices propagate into model quality and training efficiency.
Key Responsibilities
- Design and operate large-scale data pipelines supporting AI training, evaluation, and continual improvement workflows.
- Build ingestion systems for diverse modalities including text, image, audio, video, and structured signals.
- Implement data cleaning, deduplication, filtering, and quality assurance at petabyte scale.
- Develop dataset versioning, lineage, and provenance tracking systems suitable for reproducible training.
- Build high-throughput data loading systems that maximize GPU utilization during training.
- Implement labeling workflows, active learning pipelines, and human-in-the-loop data improvement systems.
- Design storage architectures balancing cost, throughput, and latency across data tiers.
- Build evaluation dataset construction pipelines with strict integrity and contamination controls.
- Implement data privacy, redaction, and consent enforcement throughout the pipeline.
- Collaborate with ML researchers and engineers to align data systems with model development needs.
- Drive observability of data quality, drift, and pipeline health across the AI data estate.
- Optimize cost and performance through compression, format selection, and caching strategies.
- Document data systems, schemas, and operational procedures for broad internal use.
- Stay current with AI data infrastructure research and emerging open-source tools.
Required Qualifications
- Bachelor’s or Master’s degree in Computer Science or a related field.
- Six or more years of data engineering experience, with significant work supporting ML or AI workloads.
- Strong proficiency in Python and at least one JVM or systems language.
- Deep experience with modern data processing frameworks such as Spark, Ray, or Beam.
- Hands-on experience operating petabyte-scale storage and pipeline systems.
- Strong understanding of distributed systems, data modeling, and storage formats.
- Experience with dataset versioning, lineage, and reproducibility for ML workflows.
- Familiarity with high-throughput data loading for accelerator-based training.
- Strong software engineering practices including testing, CI/CD, and code review.
- Excellent communication and cross-functional collaboration skills.
Preferred Qualifications
- Experience with multimodal datasets at large scale.
- Familiarity with data quality tooling and dataset evaluation methodology.
- Exposure to privacy-preserving data systems and regulated data handling.
- Open-source contributions to data infrastructure projects.
- Experience supporting frontier model training pipelines.
How to Apply
Would you like to know more about this opportunity? For immediate consideration, please send your resume to [email protected] or contact us at (908)676-4399. Learn more about Bright Vision Technologies at www.bvteck.com.
Bright Vision Technologies is an Equal Opportunity Employer.
Similar Jobs
Financial Services
Administer and improve global block, file, and object storage across on-premises and cloud environments. Manage performance, capacity, resiliency, replication, backup, disaster recovery, incident response, root-cause analysis, and operational automation. Build observability and AI-assisted AIOps workflows with validation, guardrails, auditability, and human review. Maintain runbooks, SLOs, SLIs, error budgets, and on-call readiness while partnering with infrastructure and application teams.
Top Skills:
AnsibleBashCephChefCloud-Native StorageCloudFormationContainer Storage Interface (Csi)DatadogDell Emc IsilonDell Emc PowerstoreElasticGoGrafanaHardware Security ModulesHitachiIbm StorageKafkaKey Management ServicesKubernetesLinuxNetappOpensearchOpentelemetryPrometheusPuppetPure StoragePythonSecrets ManagementServicenowSplunkTerraform
Financial Services
Supports and improves infrastructure and applications operating at scale, ensuring end-to-end monitoring, capacity management, resilience, and performance. Partners with application and infrastructure teams to identify and remediate capacity risks, supports cloud migrations across public and private environments, and applies automation, scripting, infrastructure-as-code, CI/CD, and observability practices. The role also resolves technical issues, drives meetings, and makes decisions across multiple technologies and applications.
Top Skills:
ApicaAWSAzureCi/CdConfluenceDynatraceGCPGitGrafanaJIRANetcoolPythonServicenowSplunkTerraform
Fintech • Real Estate • PropTech
Technical lead for Redfin's production database and storage systems: define strategy, architect and implement cloud database/storage (self-managed and AWS managed), ensure reliability, observability, scalability, security, participate in DR planning and on-call rotation, lead incident resolution and root cause analysis, mentor engineers, and evangelize AI code-generation and infrastructure-as-code best practices.
Top Skills:
Ai Code Generation Tools (Anthropic Claude CodeAWSAws AuroraAws RdsAws S3Configuration ManagementCursor)DynamoDBElasticacheGithub CopilotInfrastructure As CodeLinuxOpensearchPostgresPython
What you need to know about the Austin Tech Scene
Austin has a diverse and thriving tech ecosystem thanks to home-grown companies like Dell and major campuses for IBM, AMD and Apple. The state’s flagship university, the University of Texas at Austin, is known for its engineering school, and the city is known for its annual South by Southwest tech and media conference. Austin’s tech scene spans many verticals, but it’s particularly known for hardware, including semiconductors, as well as AI, biotechnology and cloud computing. And its food and music scene, low taxes and favorable climate has made the city a destination for tech workers from across the country.
Key Facts About Austin Tech
- Number of Tech Workers: 180,500; 13.7% of overall workforce (2024 CompTIA survey)
- Major Tech Employers: Dell, IBM, AMD, Apple, Alphabet
- Key Industries: Artificial intelligence, hardware, cloud computing, software, healthtech
- Funding Landscape: $4.5 billion in VC funding in 2024 (Pitchbook)
- Notable Investors: Live Oak Ventures, Austin Ventures, Hinge Capital, Gigafund, KdT Ventures, Next Coast Ventures, Silverton Partners
- Research Centers and Universities: University of Texas, Southwestern University, Texas State University, Center for Complex Quantum Systems, Oden Institute for Computational Engineering and Sciences, Texas Advanced Computing Center


