/roles — ROLE_7
Member of Technical Staff — RL Research (New PhD Grad)
Face-to-face AI interaction that feels human
The role
- COMP
- $250K - $350K
- EQUITY
- Highly competitive equity
- LOCATION
- Seattle
- WORKPLACE
- On-site
- EXPERIENCE
- 0 - 2 years
- VISA
- None, Visa transfers
- STACK
- Python, PyTorch, RL/post-training frameworks
- INDUSTRY
- Software Development, AI
The company
Applied-AI lab building visual conversational AI — real-time, face-to-face interaction that feels human.
- STAGE
- growth-stage
- FUNDING
- $60M+ raised
- TEAM
- ~25 people
- FOUNDED
- 2024
JD — the work
About the role
A rare 0→1 post-training role for a finishing or recent PhD: build the RL and post-training stack of a frontier lab whose models are omni from the ground up — audio, video, language, and real-time full-duplex interaction. The problems go beyond text: timing, interruption, emotional response, and audiovisual coherence are all training targets.
What you'll do
- Build the RL/post-training stack from scratch: rollouts, policy optimization, reward serving, feedback loops, evaluation, observability
- Develop and scale methods like PPO, GRPO, DPO, rejection sampling, RLHF/RLAIF, and online RL
- Design the abstractions connecting research ideas to production-scale runs: trainers, rollout workers, reward models, experience buffers
- Build evaluation loops for interactive behavior: turn-taking, interruption, timing, emotional response
- Optimize the full loop across rollout throughput, serving latency, and GPU utilization
What they're looking for
- Completing or recently completed PhD with deep RL/post-training grounding
- Systems ability to build and debug distributed training infrastructure
- Drive to turn evolving research ideas into reliable, fast tooling