Dragonfly Logo

Dragonfly

Senior Inference Optimization Engineer - Dragonfly Portfolio

Posted 5 Days Ago
Remote or Hybrid
Hiring Remotely in Greece
Senior level
Remote or Hybrid
Hiring Remotely in Greece
Senior level
Optimize large-model inference at scale: improve throughput, reduce latency and cost per token, build benchmarking harnesses, tune parallelism and quantization strategies, implement load-balancing in routing, and evaluate custom kernels and emerging inference hardware.
The summary above was generated by AI
Dragonfly is a crypto-native Venture Capital and research firm with $3.6B+ in assets under management and 160+ portfolio companies. Our Talent team connects people with roles across our portfolio, opening the door to opportunities through our Talent Network.

This is an application to join our talent network. This is not a listing for an internal role at Dragonfly.

We're actively sourcing for a Senior Inference Optimization Engineer for one of our portfolio companies building privacy-first consumer AI infrastructure. You'll be on the bleeding edge of LLM inference performance, pushing throughput, driving down latency, and optimizing cost per token at significant scale.

Location: Remote, USA (open to excellent candidates outside the USA)

What We’re Looking For:
  • 5+ years in performance optimization or HPC with deep GPU architecture and parallel programming knowledge
  • Hands-on experience with at least one production LLM inference engine (vLLM, SGLang) running at high volume
  • Demonstrated experience with LLM inference optimization: continuous batching, PagedAttention, KV cache management, speculative decoding, quantization, CUDA graphs, torch.compile
  • Experience with distributed inference strategies: tensor parallelism, pipeline parallelism, MoE parallelism in multi-GPU and multi-node environments
  • GPU profiling fluency: Nsight Systems, Nsight Compute, PyTorch Profiler
  • Proficiency in Python, Rust, or Go. C++/CUDA a strong plus
  • Bonus: custom Triton kernels, diffusion/image model inference optimization, open-source inference framework contributions

About the role:
  • Stand up and optimize GPU infrastructure including B300 nodes in owned data centers
  • Drive down TTFT and TPOT, push throughput, and improve cost per token for LLM inference workloads
  • Build reproducible benchmarking harnesses across inference engines to identify optimal engine, quantization scheme, and parallelism strategy per workload and GPU SKU
  • Optimize multivariate inference load-balancing algorithms within the inference routing system
  • Evaluate emerging inference optimization techniques including custom CUDA/Triton kernels, novel attention variants, new quantization schemes, and compilation stack improvements
  • Evaluate emerging inference hardware (FPGAs, ASICs, custom silicon) for viability in the stack

Even if you don't match every point above but are an engineer passionate about AI and/or crypto, we encourage you to apply. There may be other opportunities that fit your skill set.

Process: 
  • We'll review your application and assess fit for this role.
  • If there's a match, we'll facilitate a warm introduction to the team.
  • If the timing isn't right, we'll keep you in mind for future opportunities across the portfolio.

Submit your information below, and we’ll reach out if there’s a potential fit.

Similar Jobs

7 Hours Ago
In-Office or Remote
Mid level
Mid level
Artificial Intelligence • Machine Learning • Natural Language Processing • Software • Conversational AI
The Account Executive will build a sales pipeline in the German markets, work with cross-functional teams, and maintain stakeholder relationships to drive sales and manage accounts for Deepgram's voice AI platform.
Top Skills: AISales TechnologySpeech-To-SpeechSpeech-To-TextText-To-SpeechVoice Ai
7 Hours Ago
Remote
Mid level
Mid level
Artificial Intelligence • Machine Learning • Natural Language Processing • Software • Conversational AI
Sell Deepgram's voice AI platform across EMEA by prospecting new logos, building pipeline, closing full-cycle technical deals, collaborating with Sales Ops and Sales Engineers, managing stakeholder relationships, and driving upsell with CSMs. Target technical buyers and communicate product value to meet and exceed quotas.
Top Skills: Ai/MlCloud ApisDeveloper ToolsSpeech-To-TextText-To-SpeechVoice Ai
7 Hours Ago
In-Office or Remote
Senior level
Senior level
Aerospace • Artificial Intelligence • Machine Learning • Robotics • Software
Lead end-to-end international and US federal government and commercial bids/proposals for EURAF GTM, manage competing priorities, drive cross-functional teams, improve proposal processes, mentor contributors, and support customer briefings, demos, and BD operations.
Top Skills: Autonomy SoftwareHivemindShipleyUnmanned/Uncrewed SystemsV-BatX-Bat

What you need to know about the Austin Tech Scene

Austin has a diverse and thriving tech ecosystem thanks to home-grown companies like Dell and major campuses for IBM, AMD and Apple. The state’s flagship university, the University of Texas at Austin, is known for its engineering school, and the city is known for its annual South by Southwest tech and media conference. Austin’s tech scene spans many verticals, but it’s particularly known for hardware, including semiconductors, as well as AI, biotechnology and cloud computing. And its food and music scene, low taxes and favorable climate has made the city a destination for tech workers from across the country.

Key Facts About Austin Tech

  • Number of Tech Workers: 180,500; 13.7% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Dell, IBM, AMD, Apple, Alphabet
  • Key Industries: Artificial intelligence, hardware, cloud computing, software, healthtech
  • Funding Landscape: $4.5 billion in VC funding in 2024 (Pitchbook)
  • Notable Investors: Live Oak Ventures, Austin Ventures, Hinge Capital, Gigafund, KdT Ventures, Next Coast Ventures, Silverton Partners
  • Research Centers and Universities: University of Texas, Southwestern University, Texas State University, Center for Complex Quantum Systems, Oden Institute for Computational Engineering and Sciences, Texas Advanced Computing Center

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account