USD per year
ML Systems Engineer
Location
Menlo Park, California
Employment Type
Full time
Location Type
On-site
Department
Bits: Research, LLMs, machine learning, infra
About Periodic Labs
We're an AI and physical sciences company building state-of-the-art models to accelerate breakthroughs across materials, energy, and beyond. Backed by world-class investors and growing rapidly, we operate at the pace the frontier requires. Our team brings deep expertise, genuine ownership, and an insatiable drive to push the boundaries of what's scientifically possible.
About the Role
You’ll work alongside some of the world’s leading ML systems engineers, including leaders behind Megatron-LM, SGLang, Liger Kernel, TorchRec, CleanRL, TorchRL, and JAX-MD. We’re looking for exceptional ML Systems Engineers to build the agentic infrastructure powering our large-scale training, inference, and reinforcement learning. You’ll own critical pieces of the ML systems stack to maximize performance, scalability, reliability, and productivity for both engineers and AI agents.
What You'll Do
- Build and optimize large-scale training and reinforcement learning infrastructure while ensuring its correctness
- Develop high-performance inference and serving systems
- Design distributed runtimes and scheduling systems for complex ML workloads
- Build secure and large-scale sandboxing and execution environments
- Optimize memory, GPU kernels and communication for maximum throughput and end-to-end efficiency
- Improve scalability, reliability, and efficiency across the ML systems stack
What We're Looking For
- Strong systems programming and performance engineering skills
- Experience building high-performance ML infrastructure at scale
- Ability to own complex technical problems end-to-end
- Strong coding ability and engineering judgment, including the ability to work effectively with AI agents to design, implement, test, and debug complex systems
- High ownership, fast execution, and a passion for pushing the frontier of AI systems and accelerating scientific discovery
You should have deep expertise in at least one of the following:
- Training: Strong experience building, debugging and optimizing large-scale training systems with Megatron-LM. Familiarity with TorchTitan, FSDP, veRL, Slime, or other distributed training systems is a plus.
- Distributed Runtime: Strong experience with Ray. Familiarity with Monarch or other distributed execution frameworks is a plus.
- Inference: Strong experience with SGLang. Familiarity with vLLM, TensorRT-LLM or production LLM serving systems is a plus.
- Sandboxing: Strong experience with secure execution environments containers virtualization or code sandboxing.
- GPU Kernels: Strong experience with CUDA Triton CUTLASS CuTe or custom GPU kernel development.
- GPU Communication: Strong experience with NCCL NVLink InfiniBand RDMA GPUDirect RDMA or large-scale communication optimization.
Mechanics
Minimum education: Bachelor’s degree or similar experience Location: Menlo Park CA (Soon: San Francisco too) Compensation: $250000-$350000 base + equity Visa sponsorship: Yes
Remote Work Allowed?
No — location type is On-site in Menlo Park.
Employment Type
Full time
Experience Level (inferred)
Senior (based on required deep expertise in complex ML system engineering)
Salary Range (estimated)
Over 120k ($250k-$350k base + equity)
Skills Mentioned
- Programming languages/tools: CUDA
- Frameworks/Systems: Megatron-LM TorchTitan (plus familiarity) FSDP (plus familiarity) veRL (plus familiarity) Slime (plus familiarity) Ray
- Inference tools: SGLang; vLLM (plus familiarity) TensorRT-LLM (plus familiarity)
- Distributed execution frameworks: Monarch (plus familiarity)
- Sandboxing/virtualization: containers; secure execution environments; code sandboxing
- GPU kernel development tools: Triton; CUTLASS; CuTe; custom GPU kernel development
- GPU communication technologies: NCCL; NVLink; InfiniBand; RDMA; GPUDirect RDMA
Periodic Labs aims to create an AI scientist and autonomous laboratories for them to operate, focusing on accelerating science in the physical sciences. They build AI scientists and autonomous labs to generate high-quality experimental data, enabling new scientific discoveries and applications such as discovering higher-temperature superconductors and aiding semiconductor manufacturers.
View Company Profile