Vertex, Inc. Logo

Vertex, Inc.

Principal AI Engineer

Reposted One Month Ago
Remote
Hiring Remotely in USA
160K-208K Annually
Expert/Leader
Remote
Hiring Remotely in USA
160K-208K Annually
Expert/Leader
Lead enterprise model-training strategy for Commercial AI products: fine-tune LLMs (QLoRA/LoRA/PEFT), train traditional ML models, design large-scale data pipelines, enforce data governance and PII handling, build reproducible experiment/training pipelines, optimize GPU/distributed training costs, and mentor teams to operationalize models into production.
The summary above was generated by AI

Job Description:

The Principal Engineer, AI Model Training & Data Strategy owns how Commercial AI (CAI) products train, fine-tune, and evaluate models, and how the data behind those models is sourced, curated, stored, and governed. This is primarily a model-training role with a strong secondary focus on the data management and pipelines that make high-quality training possible. The role defines the enterprise training strategy and the standards for how and where training data from Commercial AI products is stored, versioned, and reused. 

Essential Job Functions and Responsibilities 

  • Define and own the end-to-end model training strategy across CAI products, spanning traditional AI/ML models and large language models 

  • Fine-tune large language models using parameter-efficient techniques (e.g., QLoRA, LoRA, PEFT) and full fine-tuning where warranted 

  • Train, evaluate, and tune traditional AI/ML models (classification, regression, ranking, clustering, and similar) 

  • Work with large volumes of data – design and optimize pipelines for ingestion, cleaning, labeling, and feature engineering 

  • Define standards for how and where training data from Commercial AI products is stored, versioned, and accessed (data lakes/warehouses, feature stores, dataset registries) 

  • Establish data governance, lineage, quality, licensing/consent, and PII-handling practices for training data 

  • Build reproducible training pipelines and experiment tracking (datasets, hyperparameters, checkpoints, and metrics) 

  • Define evaluation methodology and benchmarks for model quality, including offline evaluation and regression testing 

  • Curate and clean training, validation, and test datasets, including synthetic data generation where appropriate 

  • Optimize training cost and compute utilization (GPU efficiency, distributed training, quantization) 

  • Partner with product and platform teams to operationalize and hand off trained and fine-tuned models to production 

  • Mentor engineers and raise model-training and data-quality maturity across teams 

Knowledge, Skills, and Abilities 

  • Strong hands-on experience training and fine-tuning both traditional AI/ML models and LLMs in production 

  • Deep experience with parameter-efficient fine-tuning (QLoRA, LoRA, PEFT), quantization, and the tradeoffs versus full fine-tuning 

  • Proficiency with ML/DL frameworks and libraries (e.g., PyTorch, Hugging Face Transformers/PEFT/TRL, scikit-learn) 

  • Experience building and operating large-scale data pipelines and platforms (e.g., Spark, Ray, dbt, or equivalents) 

  • Strong grasp of data management: dataset storage architecture, versioning, lineage, governance, and PII handling 

  • Experience with experiment tracking and reproducible ML (e.g., MLflow, Weights & Biases) 

  • Understanding of distributed training and GPU/compute optimization 

  • Ability to define strategy and standards while remaining hands-on in code 

  • Strong stakeholder collaboration and problem-solving skills 

Education and Experience 

  • Bachelor’s degree in Computer Science, Engineering, or related discipline; advanced degree in ML, AI, or Data Science preferred 

  • 12 or more years of experience in AI/ML engineering, applied ML, or data engineering, with significant hands-on model training and fine-tuning 

Disclaimer 

The above statements describe the general nature and level of work performed in this role. Other duties may be assigned. 


Vertex Values: Together We Win

We're building a team of people who are passionate about making an impact for our customers and committed to how that impact is achieved. Our values define the behaviors, mindset, and culture that make Vertex a great place to grow and do meaningful work.


Play to Win or We Don't Play — If we choose to do something, we're choosing to do it because we plan to win. That mindset raises our bar on product quality, customer outcomes, and how we show up for one another.


Work As a Team, Putting the Customer At the Core — Our customers are our true north. Whatever your role, ask: how will this help a customer succeed today? We earn trust through outcomes, not promises.


Achieve Excellence With Integrity, Speed, and Agility — The market isn't slowing down. We'll move faster, adapt quickly, and never compromise on doing things the right way — for teammates, customers, and partners.


Innovate Boldly With a Growth Mindset — Progress demands smart risk. We'll try new approaches, learn fast, and keep pushing the boundaries — especially where AI can remove friction and unlock value.


Communicate with Care, Candor and Transparency — Honest, constructive conversations make us better. Let's speak plainly about what's working and what isn't and help each other improve.

Pay Transparency Statement:

US Base Salary Range: $159,600.00 - $207,500.00

Base pay offered to new hires may vary based upon factors including relevant industry and job-related skills and experience, geographic location, and business needs.* The range displayed does not encompass the full potential of the role, which allows for further growth and career progression.

In addition, as a part of our total compensation package, this role may be eligible for the Vertex Bonus Plan (VOB), a role-specific sales commission/bonus, and/or equity grants.

Learn more about Life at Vertex and connect with your recruiter for more details regarding Vertex's compensation and benefit programs.

*In no case will your pay fall below applicable local minimum wage requirements.

Similar Jobs

12 Days Ago
In-Office or Remote
165K-282K Annually
Expert/Leader
165K-282K Annually
Expert/Leader
Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
Own the architecture and long-term evolution of hybrid, multi-tenant AI compute platforms across bare-metal OpenShift and public-cloud services. Define distributed training networking, GPU utilization and cost models, GitOps governance, model-serving standards, workload placement, identity and security controls, SLOs, disaster recovery, and platform upgrade strategies. Partner with AI, privacy, and security teams to deliver HIPAA-compliant infrastructure for training and inference workloads.
Top Skills: Argo CdAws BedrockAzure Ai FoundryCephDcgmDeepspeedFsdpGcp Vertex AiGitopsGpudirect RdmaHipaaIbm Storage ScaleInfinibandJaxKserveKubeflow PipelinesKubernetesKueueLustreMigMtlsNcclNfdNvidia Gpu OperatorNvidia GpusNvlinkOauthOdfOidcOpenshift AiPytorch DdpRayRbacRed Hat OpenshiftRhacmRocev2Tensorrt-LlmVastVaultVllmVolcanoWeka
15 Days Ago
In-Office or Remote
230K-288K Annually
Expert/Leader
230K-288K Annually
Expert/Leader
Artificial Intelligence • Cloud • Software • Infrastructure as a Service (IaaS)
Lead DigitalOcean’s AI security strategy, governance framework, threat modeling, and multi-year roadmap. Partner with security, engineering, and data science teams to secure AI products, deploy guardrails, conduct adversarial testing and red teaming, and monitor AI/ML systems. Provide architectural reviews, executive risk briefings, and guidance on emerging threats including prompt injection, model supply-chain risks, agentic AI, and data exfiltration.
Top Skills: Agentic AiAi/MlGoIso/Iec 42001Large Language ModelsMitre AtlasNist Ai RmfOwasp Llm Top 10PythonRetrieval-Augmented Generation (Rag)
15 Days Ago
In-Office or Remote
230K-288K Annually
Expert/Leader
230K-288K Annually
Expert/Leader
Artificial Intelligence • Cloud • Software • Infrastructure as a Service (IaaS)
Lead DigitalOcean’s AI security strategy, governance framework, roadmap, and organizational programs. Build and operationalize AI/ML security controls, tooling, guardrails, and automation; conduct threat modeling, adversarial testing, and AI red teaming. Review AI architectures, guide high-risk product decisions, monitor deployed models, and address risks involving RAG, agents, model supply chains, and prompt injection. Represent AI security to executives and external communities while translating emerging research and threats into engineering guidance and policy.
Top Skills: Agentic Ai FrameworksAi/MlEu Ai ActGoIso/Iec 42001Mitre AtlasModel ApisNist Ai RmfOwasp Llm Top 10PythonRag

What you need to know about the Austin Tech Scene

Austin has a diverse and thriving tech ecosystem thanks to home-grown companies like Dell and major campuses for IBM, AMD and Apple. The state’s flagship university, the University of Texas at Austin, is known for its engineering school, and the city is known for its annual South by Southwest tech and media conference. Austin’s tech scene spans many verticals, but it’s particularly known for hardware, including semiconductors, as well as AI, biotechnology and cloud computing. And its food and music scene, low taxes and favorable climate has made the city a destination for tech workers from across the country.

Key Facts About Austin Tech

  • Number of Tech Workers: 180,500; 13.7% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Dell, IBM, AMD, Apple, Alphabet
  • Key Industries: Artificial intelligence, hardware, cloud computing, software, healthtech
  • Funding Landscape: $4.5 billion in VC funding in 2024 (Pitchbook)
  • Notable Investors: Live Oak Ventures, Austin Ventures, Hinge Capital, Gigafund, KdT Ventures, Next Coast Ventures, Silverton Partners
  • Research Centers and Universities: University of Texas, Southwestern University, Texas State University, Center for Complex Quantum Systems, Oden Institute for Computational Engineering and Sciences, Texas Advanced Computing Center

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account