KLDiscovery Logo

KLDiscovery

AI Principal Data Scientist (Architect-Level Scope)

Reposted One Month Ago
Remote
Hiring Remotely in United States
190K-240K Annually
Senior level
Remote
Hiring Remotely in United States
190K-240K Annually
Senior level
Lead and build KLDiscovery's gen-AI architecture and core systems. Design and implement LLMs, agents, RAG, vector search, embeddings, model selection, MLOps, telemetry, and evaluation harnesses. Ship high-impact prototypes and production components, mentor senior engineers, and represent AI strategy to product, customers, and partners.
The summary above was generated by AI

Principal Data Scientist, AI (Architect-Level Scope)

About KLDiscovery

KLDiscovery is a global eDiscovery and legal technology provider serving large law firms, corporate legal departments, and government agencies. We build and operate the products and services that legal teams rely on to manage, process, and review case data at scale. With operations across multiple countries and a client base that includes AmLaw 200 firms, we handle some of the largest and most complex matters in the industry.

About the Role

We're hiring our most senior AI practitioner — someone with genuine data science rigor who also wants to build. In this role, you'll set the scientific and architectural direction for gen AI and ML at KLDiscovery: designing the evaluation methodology that proves our AI holds up under legal scrutiny, choosing and validating the models and retrieval systems that power our products, and building the hardest parts of that system yourself.

This is one of the most interesting data science problem sets in enterprise software. You'll work with terabytes of real-world legal data — emails, contracts, chat transcripts, images, video, depositions, and regulatory filings — from some of the largest litigation and investigation matters in the world. The work is hard in ways that matter: documents are messy and adversarial, the stakes are real (privilege, defensibility, attorney work product), and getting the evaluation wrong has real consequences. AI that can surface key people, themes, and timelines in hours instead of weeks, or pre-classify millions of documents for relevance and privilege with results that survive scrutiny, directly changes the economics of how legal matters get resolved. We build solutions that turn data into evidence the legal system can trust.

This is a builder-first role, not a research-only or advisory one. You'll design experiments and evaluation frameworks that prove our AI is defensible in court, build the retrieval and agent systems that reason over case data, and personally ship the hardest, most novel parts of that system. You'll set strategic direction across Nebula, our eDiscovery platform, and CS & Operations, then prove out the science and the architecture by building it yourself. Not a role for anyone stepping back from the keyboard, and not a role for anyone who treats evaluation as an afterthought.

We offer competitive total compensation that includes base pay, bonus potential, equity, inclusive benefits, wellness programs, and perks. We use market and industry data to inform pay decisions while considering geography and labor markets, individual experience, and business needs. Individual compensation will vary, although a reasonable estimate of the current annualized base pay range for this position is $190,000 to $240,000.

Job location: Remote (candidate must be based in the United States)

Key Responsibilities

Own the science and the architecture, and build both personally. Design the experimentation and evaluation methodology that determines whether our AI's outputs are accurate, consistent, and defensible — then define the end-to-end gen AI architecture across Nebula and CS & Operations to deliver on it: LLMs, agent harnesses, RAG, vector search, embeddings, and model selection and triage. Build the hardest parts personally: prototype agent loops, design and run the evaluations that validate them, tune retrieval, and ship the shared infrastructure that powers AI Case Explorer (case overviews, timelines, key people and themes, PII surfacing, and Agent chat), AI Agent Review (pre-classifying relevance, privilege, and key issues, shipping MLP), and CS & Ops tech-enablement as part of our central work orchestration system.

Own evaluation rigor, AI/MLOps, and telemetry end-to-end. Design the statistical and experimental methodology behind our evaluation harnesses — the standard our outputs have to meet to hold up under legal scrutiny — and own the infrastructure that enforces it: model deployment and versioning, eval pipelines, drift and quality monitoring, cost and latency telemetry, and prompt and agent observability. Define and implement how we select, triage, and route across models (Azure OpenAI, Anthropic, open-source, fine-tuned), manage vector databases and retrieval, and evolve our evaluation methodology and agent harness as the frontier moves.

Lead the practice from the front. Set the technical and scientific bar by building, not by reviewing. Partner with Engineering, Product, and Data Science leadership to translate that rigor into shipped product. Raise the bar on both AI engineering and evaluation discipline, mentor senior ICs through hands-on technical leadership, build and maintain relationships with model and infrastructure vendors, and represent KLD's AI strategy directly with customers, partners, and at industry events.

What You Bring (Required)

  • 7+ years in data science, applied AI/ML, or ML engineering, with recent hands-on experience as a senior or principal-level practitioner in the gen AI era
  • Real training in statistics, experimentation, or applied research you know how to design an evaluation that actually tests what you think it tests, not just one that looks reasonable
  • Proven track record architecting and personally building enterprise gen AI systems in production with measurable customer impact
  • Experience with consumer-facing or B2B customer-facing AI products — systems that real external users or customers depend on, not only internal tooling
  • Builder at heart: still writes code, runs experiments, ships, and tunes prompts and evals, and wants to keep doing so as a leader
  • Deep expertise across the modern gen AI stack: LLMs, agents, RAG, vector databases, embeddings, search, and evaluation harnesses
  • Hands-on experience designing system-of-systems AI pipelines spanning search, retrieval, agent harnesses, and model selection/triage
  • Strong proficiency with the Microsoft AI stack: Azure OpenAI, Azure AI Foundry, Azure AI Search, and supporting Azure infrastructure
  • Excellent technical leadership skills; demonstrated ability to influence architecture and methodology decisions across product, engineering, and data science
  • Strong communication skills, including explaining evaluation results and architecture trade-offs to executive and customer audiences
  • Career experience spanning both a larger, established technology company and a smaller company or startup — comfortable operating with both institutional rigor and startup pace
  • Mentor team members and contribute to a culture of continuous improvement

 

Nice to Have (Preferred)

  • Advanced degree (MS or PhD) in Statistics, Machine Learning, Computer Science, or a related quantitative field
  • Peer-reviewed publications, technical writing, conference talks, or other evidence of contributing to the field, not just building within it
  • Background building agentic systems with tool use, planning, and multi-step reasoning in production
  • Prior experience setting up AI governance and evaluation harnesses in a regulated or high-stakes domain
  • Open-source contributions or other evidence of being a recognized builder in the AI community
  • A demonstrated track record of sticking with hard, ambiguous problems over long timeframes rather than pivoting away when the first approach doesn't work.

 

KLDiscovery Austin, Texas, USA Office

Austin, United States

Similar Jobs

11 Minutes Ago
Remote or Hybrid
Junior
Junior
Big Data • Consumer Web • Fintech • Mobile • Payments • Social Impact • Financial Services
Lead live onboarding, refresher, and product-launch training for internal Operations teams and BPO partners. Deliver QA-driven one-on-one and group coaching, coordinate schedules and completion tracking, collaborate with Product and Compliance, and create LMS courses, microlearning, SCORM content, and upskilling programs using AI-assisted workflows.
Top Skills: AddieAi ToolsArticulate RiseArticulate StorylineLmsScorm
12 Minutes Ago
Easy Apply
Remote or Hybrid
Easy Apply
164K-205K Annually
Senior level
164K-205K Annually
Senior level
Cloud • Information Technology • Security • Software • Cybersecurity
Own the end-to-end lifecycle, strategy, roadmap, and execution of Zscaler’s Data Security products, including Web DLP and Email DLP. Define requirements using market and customer insights, collaborate with engineering and cross-functional teams, support product deployment and sales enablement, and execute go-to-market plans. The role requires cybersecurity product management experience, expertise in DLP and data classification, and technical knowledge of Data Security, Email, Proxy, CASB, DSPM, and AI-based security solutions.
Top Skills: Ai/MlBrowser PluginsCasbCompliance StandardsData ClassificationData Loss PreventionData SecurityDspmEmail DlpEmail SecurityEndpoint AgentsProxySecurity FrameworksWeb Dlp
12 Minutes Ago
Easy Apply
Remote or Hybrid
Location, WV, USA
Easy Apply
144K-205K Annually
Entry level
144K-205K Annually
Entry level
Cloud • Information Technology • Security • Software • Cybersecurity
Leads a customer-facing MDR and threat hunting team, overseeing talent development, operational quality, customer escalations, executive communications, and recurring security engagements. Translates threat intelligence, attack activity, and security outcomes into actionable customer guidance. Establishes deliverable standards, drives continuous improvement, supports complex investigations, and ensures MDR services demonstrate consistent customer value.
Top Skills: Ai/MlEdrMitre Att&CkWeb ProxiesZscaler Zero Trust Exchange

What you need to know about the Austin Tech Scene

Austin has a diverse and thriving tech ecosystem thanks to home-grown companies like Dell and major campuses for IBM, AMD and Apple. The state’s flagship university, the University of Texas at Austin, is known for its engineering school, and the city is known for its annual South by Southwest tech and media conference. Austin’s tech scene spans many verticals, but it’s particularly known for hardware, including semiconductors, as well as AI, biotechnology and cloud computing. And its food and music scene, low taxes and favorable climate has made the city a destination for tech workers from across the country.

Key Facts About Austin Tech

  • Number of Tech Workers: 180,500; 13.7% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Dell, IBM, AMD, Apple, Alphabet
  • Key Industries: Artificial intelligence, hardware, cloud computing, software, healthtech
  • Funding Landscape: $4.5 billion in VC funding in 2024 (Pitchbook)
  • Notable Investors: Live Oak Ventures, Austin Ventures, Hinge Capital, Gigafund, KdT Ventures, Next Coast Ventures, Silverton Partners
  • Research Centers and Universities: University of Texas, Southwestern University, Texas State University, Center for Complex Quantum Systems, Oden Institute for Computational Engineering and Sciences, Texas Advanced Computing Center

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account