A
alloycompute.ai
BackendSenior
Founding Inference Engineer — FPGA + GPU Disaggregated Serving
CudaGpu Performance EngineeringNsightVllmSglangTensorrt-LlmVerilogSystemverilogFpgaAmd/XilinxIntel/AlteraRdmaDpdkIo_UringLlm ServingKv-CacheQuantizationFp8Int4
Про позицію
AlloyCompute.ai is building the next generation of LLM inference infrastructure on FPGAs, focusing on disaggregated inference pipelines that split workloads across FPGA and GPU for sub-200ms latency. This is a founding-engineer role with significant equity, where you'll own the design and optimization of the heterogeneous inference stack and shape both the product and technical culture.
Обовʼязки
- Architect and build disaggregated inference pipelines spanning FPGA and GPU, including prefill/decode splitting and KV-cache transfer across devices
- Write and optimize custom CUDA kernels for attention, GEMM, quantization, and sampling paths
- Extend and integrate open-source inference engines (vLLM, SGLang) with our FPGA backend — scheduling, batching, paged attention, speculative decoding
- Design and implement FPGA dataflow in Verilog/HDL: systolic arrays, memory controllers, on-chip interconnect, and PCIe/Ethernet interfaces
- Attack network and interconnect latency at every layer — RDMA/RoCE, NIC offload, kernel-bypass networking, and FPGA-to-GPU direct transfer
- Profile, benchmark, and squeeze the stack: tokens/sec, time-to-first-token, tail latency, and cost per million tokens
Вимоги
- Deep hands-on experience with CUDA kernel development and GPU performance engineering (Nsight, occupancy tuning, memory hierarchy optimization)
- Real production experience with modern inference engines — vLLM, SGLang, TensorRT-LLM, or equivalent internals-level work
- Strong FPGA skills: Verilog/SystemVerilog HDL, timing closure, HLS familiarity a plus, experience with AMD/Xilinx or Intel/Altera toolchains
- Solid grasp of network latency engineering: RDMA, kernel bypass (DPDK/io_uring), NIC-level optimization, distributed serving topologies
- Understanding of LLM serving internals: continuous batching, paged/radix attention, KV-cache management, quantization (FP8/INT4)
- Comfort with ambiguity, bare-metal debugging, and shipping without a safety net
Переваги
- Founding-level equity — you're early enough for it to matter
- Direct influence over architecture, roadmap, and hiring
- Access to serious hardware: latest FPGAs, GPUs, and high-speed networking to experiment with
- Competitive salary, flexible location, and a team that operates at the metal
Founding Inference Engineer — FPGA + GPU Disaggregated ServingPLN 18939.5–56818.5 / MONTH
Оригінал