Role Description
As a Research Engineer, you'll build the training system that takes Protocol Learning from the 8B run to frontier scale: large models on heterogeneous hardware, in physically different regions, connected by ordinary internet.
Key Responsibilities
-
Distributed pretraining:
Implement and optimize model-parallel training. Data, pipeline, and tensor parallelism for large models on heterogeneous GPUs under low-bandwidth, high-latency links.
-
Performance optimization:
Implement techniques that reduce communication overhead while maintaining model convergence in challenging network environments.
-
Elasticity and fault tolerance:
Make runs survive node churn. Robust checkpointing, state synchronization, and recovery as participants join and leave.
-
Run instrumentation:
Build the monitoring that shows throughput, bottlenecks, and model quality across hundreds of devices.
Qualifications
-
Hands-on distributed training (required):
You've trained models across many devices in PyTorch with FSDP, DeepSpeed, Megatron, or your own implementation. You understand data, tensor, and pipeline parallelism.
-
Strong engineering:
Production-quality Python. Concurrency, failure handling, profiling before optimizing.
-
Evidence of execution:
Shipped systems, research code, open-source work, or serious personal projects.
-
Mission alignment:
You believe Protocol Learning is the viable third path for collective, trustless, and sovereign AI.
Requirements
-
Nice to Have:
-
Hands-on experience training or serving large language models such as Nemotron, Qwen or OLMo.
-
Experience with P2P networking and NAT traversal.
-
Experience with post-training and RL.
-
Experience with inference and serving systems.
-
Experience at proprietary, open-weight and open-source AI labs.
Benefits
-
Equity-Heavy Package:
We offer significant ownership for key technical contributors in addition to a high base salary.
-
Remote-First Culture:
Flexible work environment with team members distributed globally.
-
Visa Sponsorship:
Optional full visa sponsorship and relocation support to either Australia or the US.
-
Open Problems:
Training and serving frontier models on hardware you don't control, over networks you don't own, mostly has no published answers yet. You'll write some of the first ones.
FYI's
-
We work remotely across the world, with the main teams in Australia and North America. You'll need to be comfortable working across timezones.
-
Applicants must have professional-level English proficiency (written and spoken).
-
Recruiters: we aren't looking for agency support at this time. We'll reach out if we need help.