TinyFish Logo

TinyFish

MLOps Engineer

Reposted 9 Days Ago
Remote or Hybrid
Hiring Remotely in United States
Senior level
Remote or Hybrid
Hiring Remotely in United States
Senior level
Build and maintain reproducible data pipelines, experiment orchestration, CI/CD for models, Terraform-based ML infrastructure, observability, security controls, and automation to deploy and operate ML systems in production.
The summary above was generated by AI
Position Overview

As the first dedicated ML Ops Engineer, you’ll own the tooling and infrastructure that make our ml engineers wildly productive and ensure we are able to efficiently iterate on ML models, prompts, and datasets and deploy our AI systems into a predictable production environment. You’ll bridge the gap between research and DevOps—designing reproducible dataset pipelines, automated experiment workflows, and Terraform-based cloud deployments that scale.

Key Responsibilities

Dataset Management

• Design version-controlled data pipelines (feature stores, data registries) using tools such as Delta Lake, Apache Iceberg
• Implement systems for data validation, lineage tracking, and automated quality checks (e.g., Great Expectations).

Experiment Execution & Tracking

• Build and maintain experiment orchestration with platforms like MLflow, torchx, and Apache Airflow.
• Provide templated systems and tools to ML Engineers that easily launch training/evaluation data processing systems
• Automate hyper-parameter sweeps and A/B tests, exposing clear dashboards for results.

CI/CD

Models/Agents

• workflows that package, test, and promote models and agents through staging to production.
• Implement canary deployments and rollbacks for models/agents services

Terraform Infrastructure-as-Code•

• Author and maintain Terraform modules for all ML infra—networking, GPU/TPU clusters, object storage, secrets, monitoring.
• Enforce best practices for state management, workspaces, and automated plan/apply stages via CI.

Observability & Reliability

• Integrate logging, tracing, and metric collection (Prometheus, Grafana, Datadog) across data pipelines and model endpoints.
• Set SLIs/SLOs for data freshness and model latency; implement alerts and runbooks.

Security & Compliance• Work with Security to implement IAM least-privilege, key rotation, and data-encryption policies.
• Support audit requirements (SOC 2, GDPR, HIPAA where applicable).

Minimum Qualifications
  • 5+ years combined experience in DevOps, Data Engineering, or ML Ops roles.

  • Strong Terraform skills; ability to craft reusable modules and navigate complex state.

  • Production experience with at least one cloud provider (AWS, GCP, or Azure).

  • Proficiency in Python and containerization (Docker); familiarity with Kubernetes or serverless batch systems.

  • Hands-on knowledge of ML experiment platforms (MLflow, Kubeflow, Weights & Biases, or similar).

  • Experience with workflow execution frameworks (Kubeflow, Apache Airflow)

  • Understanding of modern data-versioning/feature-store concepts and tools.

  • Solid grasp of CI/CD principles, Git workflows, and infrastructure testing.

  • Excellent communication skills—capable of partnering with Data Scientists, Software Engineers, and Security teams.

Preferred (Nice-to-Have)
  • Experience with GPU orchestration (NVIDIA DGX, Karpenter, or Ray).

  • Familiarity with IaC security scanning (Checkov, tfsec).

  • Exposure to policy-as-code (OPA/Gatekeeper).

  • Prior work in real-time streaming (Kafka, Flink) and online feature serving.

  • Contributions to open-source ML Ops projects.

Reporting Structure

Reports to: Director of Infra

Similar Jobs

11 Days Ago
Remote
United States
90-120 Hourly
Junior
90-120 Hourly
Junior
Artificial Intelligence • HR Tech • Professional Services • Software
Build and evaluate MLOps and ML systems tasks for frontier AI training data. Responsibilities include designing technical challenges, writing solutions and evaluation rubrics, profiling and optimizing GPU workloads, debugging distributed systems, improving model performance, and supporting high-throughput LLM inference. The role requires production experience with ML infrastructure, serving systems, GPU accelerators, JAX or PyTorch, and strong technical communication.
Top Skills: A100B200Continuous BatchingCudaDdpDeepspeedFsdpH100JaxKinetoKv CacheMegatronNsightPaged AttentionPallasPyTorchRay ServeSglangTensorrt-LlmTorch.ProfilerTpuTritonVllmXla
14 Days Ago
In-Office or Remote
California, USA
152K-230K Annually
Senior level
152K-230K Annually
Senior level
Software • Travel
Build and maintain internal platforms and tooling for AI engineers, including React and Next.js dashboards, Python backend services, REST APIs, model deployment automation, observability integrations, and ML pipeline testing. Partner with engineering and ML teams to improve developer experience, production AI reliability, and self-service workflows while documenting systems and runbooks.
Top Skills: AWSAzureDatabasesDatadogDockerGCPGrafanaInfrastructure As Code (Iac)Job Queuing SystemsKubernetesLangsmithNext.JsPrometheusPythonReactRest ApisTypescript
One Month Ago
Remote or Hybrid
United States
132K-132K Annually
Junior
132K-132K Annually
Junior
Artificial Intelligence • Automotive • Robotics • Software • Transportation
Provide on-call MLOps support for machine learning teams by triaging pipeline issues, debugging workflows, identifying root causes, and maintaining reliable model development operations. Collaborate with senior engineers to improve tooling, documentation, and processes, while communicating technical findings across teams. The role requires Python, foundational PyTorch and model pipeline experience, plus familiarity with CI/CD, monitoring, logging, and orchestration tools.
Top Skills: AnyscaleAws SagemakerCi/CdDaftLoggingMcapMlopsMonitoringPandasParquetPipeline OrchestrationPyarrowPythonPyTorchRayTerraform

What you need to know about the Austin Tech Scene

Austin has a diverse and thriving tech ecosystem thanks to home-grown companies like Dell and major campuses for IBM, AMD and Apple. The state’s flagship university, the University of Texas at Austin, is known for its engineering school, and the city is known for its annual South by Southwest tech and media conference. Austin’s tech scene spans many verticals, but it’s particularly known for hardware, including semiconductors, as well as AI, biotechnology and cloud computing. And its food and music scene, low taxes and favorable climate has made the city a destination for tech workers from across the country.

Key Facts About Austin Tech

  • Number of Tech Workers: 180,500; 13.7% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Dell, IBM, AMD, Apple, Alphabet
  • Key Industries: Artificial intelligence, hardware, cloud computing, software, healthtech
  • Funding Landscape: $4.5 billion in VC funding in 2024 (Pitchbook)
  • Notable Investors: Live Oak Ventures, Austin Ventures, Hinge Capital, Gigafund, KdT Ventures, Next Coast Ventures, Silverton Partners
  • Research Centers and Universities: University of Texas, Southwestern University, Texas State University, Center for Complex Quantum Systems, Oden Institute for Computational Engineering and Sciences, Texas Advanced Computing Center

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account