You are viewing a preview of this job. Log in or register to view more details about this job.

Founding ML Systems Engineer

About Us

SAIA Compute is an early-stage team building a custom inference chip, backed by Pear VC and Entropy Ventures. We are a small, highly technical team with significant ownership across architecture, hardware, software, and silicon.
 

The Role

You’ll own the software stack that sits above the silicon. You’ll be working with the chip and hardware design teams to create a dynamic inference stack.

 

What you’ll do

  • Work on the compiler and runtime path from model in to tokens out
  • Improve inference performance across latency, throughput, memory usage, power efficiency, and output quality
  • Explore numerical formats, quantization methods, sparsity, batching strategies, scheduling policies, and model-specific optimizations
  • Co-design software/hardware interfaces with the chip team
  • Create bring-up, diagnostics, profiling, benchmarking, and validation tools for new hardware
  • Define correctness and performance test infrastructure across compiler, runtime, and silicon


What we’re looking for

  • 3+ years in systems software, compilers, inference engines, or ML accelerators
  • Experience working with GPUs, FPGAs, or custom silicon
  • Experience building or shipping production inference systems for transformer-based or large language models
  • Familiarity with numerical precision, quantization, tensor operations, and the performance characteristics of modern ML workloads
  • Comfortable working directly with hardware
  • Strong proficiency with Python and C++
  • Self-driven and able to work independently in a fast-paced environment

 

Nice to have

  • Experience working with edge compute devices
  • Familiarity or experience with modern FPGA development and toolchains
  • Experience with Rust