Role Description
FAR.AI is hiring a Research Lead to develop and lead our work on pre-training safety, shaping modelsβ capabilities and internal representations at their source, rather than trying to fix them after the fact.
Our initial focus is capability control:
-
Removing harmful capabilities while preserving benign ones.
-
Preventing misuse of open-weight models in areas such as CBRN and cyber.
-
Reducing loss-of-control risks by removing knowledge of oversight mechanisms.
We will validate approaches like pre-training data filtering at scale, drive adoption of successful methods, and explore techniques such as gradient routing and unlearning.
You will direct this work, partner with our red team to stress-test the resulting models, and analyze how well the methods scale to frontier systems.
Our research directions include:
-
Improved data filtering methods, such as using data attribution (e.g. influence-based selection) or more sophisticated classifiers.
-
Using methods like gradient routing to isolate dual-use capabilities in components of the model (e.g. specific MoE experts).
-
Training to actively remove harmful capabilities, such as interleaving next-token prediction with unlearning.
-
Adding synthetic data to pre-training or mid-training to shape the representations and behavior of the model.
You'll build and lead the team, set its research direction, mentor Members of Technical Staff to scale your vision, and remain hands-on enough to write code and run experiments yourself. This role offers high autonomy in an impact-driven environment, pursuing empirically grounded, scalable ML safety research.
Qualifications
-
Strong existing research track record in AI or another highly technical subject (e.g. CS, math, physics).
-
Deep experience with language-model pretraining, dataset construction, or controlled training experiments.
-
Experience building large-scale pipelines for scoring, filtering, deduplicating, and sampling training corpora.
-
Strong experimental judgment, including safety-capability evaluations, distribution-shift analysis, and statistically rigorous model comparisons.
-
Ability to build and debug research systems directly, from classifier fine-tuning through distributed training and evaluation.
-
Either a clear research agenda you'd pursue at FAR.AI, with a theory of change explaining why it's valuable, or a strong track record and a research space you'd sharpen into an agenda over your first months.
-
Experience leading a team, mentoring graduate students, or supporting early-career researchers through fellowship programs.
-
Effective communication of novel methods and solutions to both technical and non-technical audiences.
-
Not a new entrant to machine learning research.
Requirements
-
Established publication record in AI safety is preferable.
-
Comfortable writing grant proposals and navigating collaborations with other organizations or external research groups.
Benefits
-
Competitive salaries and sizable compute budgets.
-
Work-related travel and equipment expenses covered.
-
Catered lunch and dinner at our offices in Berkeley.
Logistics
-
Location: Both remote and in-person (Berkeley, CA or Singapore) are possible.
-
Hours: Full-time (40 hours/week).
-
Compensation: $290,000β$450,000/year depending on experience and location.
Application Process
-
Expect ~1β2 hours of preparation for application materials.
-
Application materials include a CV, a short research direction statement, 2β3 selected works with a brief note on your personal contribution, and a short note on why FAR.AI is a good home for your direction.
-
Process includes a 45-minute bilateral fit call, 2 technical assessments, a 3-5-day paid work trial, and structured reference calls.
If you have any questions about the role, please do get in touch at
[email protected]
.