Member of Technical Staff
Indexed description
Every employee should contribute >30M of revenue/yr to the company.
The best candidates would be top 1% at multiple parts of the inference stack.
work on PD disaggregation research
Morph builds the inference infrastructure behind the fastest open models. Our stack spans kernels, model serving, routing, autoscaling, and capacity. We are hiring a performance engineer to make the entire system faster, cheaper, and more reliable.
What you’ll do
- Find the gap between theoretical hardware performance and production performance
- Trace latency and throughput regressions from the API layer down to individual kernels
- Optimize batching, scheduling, routing, quantization, and distributed execution
- Build benchmarks and observability that make bottlenecks obvious
- Work with NVLink and RoCE
- Validate that every optimization preserves model quality and correctness
- Have optimized complex production systems
- Can juggle 8+ Codex/Claude/other coding agents concurrently
- Understand GPU performance, memory bandwidth, collectives, and inference serving
- Are strong in Python, CuTEdsl, and comfortable navigating unfamiliar codebases
- Can turn profiling data into clear engineering decisions
- Care about tokens per second, tokens per dollar, and correctness equally
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search