Neurophos Logo

Neurophos

Modeling Architect

Posted 16 Days Ago
Be an Early Applicant
In-Office
Austin, TX, USA
170K-200K Annually
Mid level
In-Office
Austin, TX, USA
170K-200K Annually
Mid level
Model the architecture and performance of an optical AI inference accelerator through hardware/software co-design. Responsibilities include bringing up transformer and other AI workloads, integrating Hugging Face and PyTorch models, developing Python and C++ functional, performance, energy, and power models, conducting roofline and limiter analyses, modeling memory and compute blocks, simulating RTL, and maintaining reproducible tests, configurations, plots, and reports.
The summary above was generated by AI
About Neurophos

The demand for new data centers and AI compute is rapidly outpacing the planet's energy capacity. Digital solutions are hitting a power wall as we approach the physical limits of traditional silicon. Conquering this bottleneck means rethinking the fundamental architecture of inference compute. The industry's current path can't meet the need, so we're taking a different approach.

Instead of traditional electronic circuits, we use silicon photonics and an active, programmable metasurface to perform matrix multiplications at the speed of light. Our optical cells are 10,000x smaller than traditional photonic components, enabling unprecedented density. By using photonics instead of electricity, our chips become more efficient as they scale. This architecture will deliver up to 100 times the energy efficiency of existing solutions while significantly improving performance for large-scale AI inference.

We’ve assembled a world-class team of industry veterans and recently raised a $110M Series A led by Gates Frontier. Participants include M12 (Microsoft’s Venture Fund), Carbon Direct Capital, Aramco Ventures, Bosch Ventures, Tectonic Ventures, Space Capital, and others.

Join us and shape the future of computing!

Location: Austin, TX or Sunnyvale, CA. Full-time onsite position.

Reports To: Sr. Director of Modeling

FLSA Status: Exempt

Position Overview

We are seeking a modeling architect for hands-on architecture modeling of the T100 optical inference accelerator, with hardware/software co-design in the loop. You will work alongside senior engineers across two tracks. The first is analytical and system performance: roofline and limiter analyses, architecture performance models, workload setup, and the resulting plots and reports. The second is hardware models: functional, performance, and power models of compute blocks, memory, and hardware/software interfaces built in an event-driven simulator, plus RTL simulation.

We expect real depth in one track and will help you build breadth across both. Most engineers start with a bounded piece, a single workload, a hardware block, or one layer of the model stack, and take on the surrounding area as the models mature.

Key Responsibilities
  • Bring up inference workloads as they ship, including dense and Mixture of Experts (MoE) transformers, hybrid/SSM models, prefill versus decode, KV cache, expert routing, and quantization, plus retrieval, speech, vision, and recommendation workloads where they map onto the accelerator.

  • Bind Hugging Face and PyTorch workloads to the programming model and run them on the functional model.

  • Co-design tiling, scheduling, the instruction set architecture (ISA), the SRAM and High Bandwidth Memory (HBM) hierarchy, and multi-chip mapping.

  • Build in one or more layers of the modeling stack: roofline and limiter studies; Python energy and latency models; C++ functional models of optical GEMM, SRAM vector processors, dataflow engines, and HBM; cycle-approximate performance and power models; and RTL simulation with Verilator and SystemVerilog.

  • Own the tests, configs, and plots behind a result so anyone can rerun it and see what was assumed.

  • Use coding agents on real multi-file edits, and own the review of the C++ and SystemVerilog they generate.

  • Share results with the architects setting the design, as well as the compiler, runtime, and RTL teams, and carry their questions back into the model stack.

Qualifications
  • BS or MS in Computer Engineering, Electrical Engineering, Computer Science, or a related field.

  • 3+ years of experience in hardware modeling, performance simulation, computer architecture, or related work.

  • Proficiency in Python or modern C++ (C++17 or later). Python-first and C++-first backgrounds are both welcome.

  • Working knowledge of computer architecture and microarchitecture, including pipelines, caches, memory hierarchies, and instruction set architecture (ISA).

  • Ability to turn an LLM, GEMM, or accelerator paper into a workload config using Hugging Face or PyTorch.

  • Strong debugging skills and the habit of writing down what was run, what was assumed, and what the number means.

Preferred Skills
  • MS or PhD in Computer Engineering, Electrical Engineering, or Computer Science.

  • Experience with roofline analysis, limiter analysis, GPU benchmarking, or correlating a model against published numbers.

  • Event-driven, cycle-approximate, or cycle-accurate simulation with SystemC, gem5, SST, or a custom kernel.

  • SystemVerilog, Verilog, Verilator, or RTL co-simulation.

  • Memory and interconnect experience with HBM, DRAM, cache, SRAM, network-on-chip (NoC), AXI, or DMA.

  • Familiarity with CUDA, GPU programming, or PyTorch internals.

What We Offer

This is an opportunity to play a pivotal role in an innovative startup redefining the future of AI hardware. Work on game-changing technology at the intersection of photonics and AI as part of a collaborative, brilliant team. You’ll contribute to a platform that redefines computational performance and accelerates the future of artificial intelligence. Come help us bring this transformative technology to the world.

 
Benefits

Join a team that invests in your future and your well-being. At Neurophos, we offer:

  • 100% coverage of base health plan premiums for you and your dependents, plus HSA contributions.

  • Unlimited PTO. No rigid vacation banks, just a focus on delivery.

  • 401(k) matching and stock option opportunities to ensure our success is your success.

  • Full suite of voluntary benefits, including Dental, Vision, Life, Hospital, Critical Illness, and Accident insurance.

  • Personalized Benefits. Choose the plans that fit your life and take the cash back for those that don’t.

HQ

Neurophos Austin, Texas, USA Office

7600 N Capital of Texas Hwy, Austin, Texas, United States, 78731

Similar Jobs

16 Days Ago
In-Office
Austin, TX, USA
250K-290K Annually
Senior level
250K-290K Annually
Senior level
Artificial Intelligence • Machine Learning • Semiconductor
Build functional, performance, energy, power, and area models for AI accelerator workloads and hardware/software co-design. Bring up PyTorch and Hugging Face workloads, model optical GEMM, memory systems, tiling, scheduling, ISA, NoC traffic, and multi-chip mapping. Develop Python analytical models and bit-accurate C++ simulations, contribute to event-driven simulation infrastructure, correlate models with RTL through Verilator and SystemVerilog, establish modeling methodology, and mentor engineers.
Top Skills: AxiC++17CactiDmaDpiDramFpgaGem5HbmHugging FaceMatplotlibMcpatMlirNocNumpyOnnxPandasPythonPyTorchSramSstSystemcSystemverilogTlm 2.XTvmUvmVerilatorXla
16 Days Ago
In-Office
Austin, TX, USA
210K-250K Annually
Senior level
210K-250K Annually
Senior level
Artificial Intelligence • Machine Learning • Semiconductor
Own performance and energy benchmarking for an optical AI inference accelerator. Build reproducible benchmarks across analytical models, architecture models, RTL simulation, and competitor GPUs. Bring up workloads from Hugging Face, PyTorch, papers, and inference stacks; measure latency, throughput, power, and energy; analyze bottlenecks; operate cloud or lab environments; and document configurations, logs, assumptions, and discrepancies.
Top Skills: AWSAzureBf16CudaCutlassFp16Fp8GCPGpusHbmHugging FaceInfinibandInt8LinuxMlperfNcclNvidia DcgmNvidia Nsight ComputeNvidia Nsight SystemsNvidia-SmiNvlinkOptical Inference AcceleratorsPythonPyTorchSglangSilicon PhotonicsTensorrt-LlmTritonTriton Inference ServerVerilatorVllm
One Month Ago
In-Office
Austin, TX, USA
100K-500K Annually
Senior level
100K-500K Annually
Senior level
Hardware • Manufacturing
Shape and evaluate next-generation high-performance RISC-V CPUs using performance modeling and simulation. Build and run architectural models (Gem5/SST/SimpleScalar), analyze pipelines, caches, branch prediction, and memory hierarchies, implement modeling and analysis tools in C++/Python, and collaborate with CPU architects to optimize AI/HPC workload performance. Hybrid role based in Santa Clara or Austin.
Top Skills: C++Gem5PythonRisc-VSimplescalarSst

What you need to know about the Austin Tech Scene

Austin has a diverse and thriving tech ecosystem thanks to home-grown companies like Dell and major campuses for IBM, AMD and Apple. The state’s flagship university, the University of Texas at Austin, is known for its engineering school, and the city is known for its annual South by Southwest tech and media conference. Austin’s tech scene spans many verticals, but it’s particularly known for hardware, including semiconductors, as well as AI, biotechnology and cloud computing. And its food and music scene, low taxes and favorable climate has made the city a destination for tech workers from across the country.

Key Facts About Austin Tech

  • Number of Tech Workers: 180,500; 13.7% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Dell, IBM, AMD, Apple, Alphabet
  • Key Industries: Artificial intelligence, hardware, cloud computing, software, healthtech
  • Funding Landscape: $4.5 billion in VC funding in 2024 (Pitchbook)
  • Notable Investors: Live Oak Ventures, Austin Ventures, Hinge Capital, Gigafund, KdT Ventures, Next Coast Ventures, Silverton Partners
  • Research Centers and Universities: University of Texas, Southwestern University, Texas State University, Center for Complex Quantum Systems, Oden Institute for Computational Engineering and Sciences, Texas Advanced Computing Center

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account