/roles — ROLE_6

Member of Technical Staff - Model Optimization and Inference (Experienced)

Face-to-face AI interaction that feels human

The role

COMP
$250K - $350K
LOCATION
Seattle
WORKPLACE
On-site
EXPERIENCE
2+ years
VISA
None, Visa transfers
STACK
Kubernetes, K8s, Terraform, Python, Rust, Go, Airflow, PyTorch, vLLM, SGLang, TensorRT-LLM, CUDA
INDUSTRY
Software Development, AI

The company

Applied-AI lab building visual conversational AI — real-time, face-to-face interaction that feels human.

STAGE
growth-stage
FUNDING
$60M+ raised
TEAM
~25 people
FOUNDED
2024
NOTE
research team with PhDs from top programs

JD — the work

About the role

An ML-systems role at a research lab building real-time, photorealistic conversational AI: own inference performance across the entire model stack — LLMs, audio models, and diffusion components — from serving frameworks down to custom kernels, for a product where latency is the product.

What you'll do

  • Own end-to-end inference optimization across LLM, audio, and diffusion models
  • Design KV-cache strategy for long conversations: eviction, compression, memory-efficient attention
  • Deploy and extend serving frameworks (vLLM, SGLang, TensorRT-LLM) for unusual workloads
  • Profile latency and throughput; systematically eliminate bottlenecks
  • Accelerate diffusion inference: consistency models, step distillation, caching, custom kernels
  • Apply quantization (INT8/INT4, GPTQ, AWQ) without meaningful quality loss

What they're looking for

  • 2+ years building production ML systems
  • Scalable-infrastructure design from scratch, with well-reasoned technology choices
  • Track record optimizing for latency, throughput, and cost
  • Broad ML-infra fluency: inference, real-time streaming, data engineering

Nice to have

  • Video or audio model experience
  • CUDA kernel-level optimization
APPLY FOR THIS ROLE →All open rolesOne application covers up to 3 roles.