About Neolithic
Neolithic is a nonprofit startup in San Francisco working to reduce catastrophic risk from AI. We build open-source, agentic tools that automate and scale AI safety research: tooling for control experiments, safety evaluations, safety training datasets, and infrastructure for research agents. Neolithic was founded in 2026 by Leo McKee-Reid.
Why this role matters
Capabilities research is increasingly automated. Safety research is still largely done by hand. Many of the field's bottlenecks are engineering problems that shared tooling could solve. As one of our first hires, you will set the technical direction, the culture and the hiring bar for everyone who follows.
What you'll do
- Scope, build, iterate on and maintain tools as production software.
- Work with researchers at safety organisations as design partners, removing the bottlenecks in their workflows.
- Shape technical direction, culture and hiring, and take on whatever is most needed.
Projects in your first few months
- An automated pipeline that generates realistic samples of AI agents misusing their access inside frontier labs, as training data for cybersecurity monitors.
- Agentic workflows that turn threat models into runnable, verified evaluations.
- Evaluation environments realistic enough that models can't tell they're being tested.
- Context management and other infrastructure for safety research agents.
- Detecting reward hacking in RL environments.
Who we're looking for
We're looking for people who fit one of three profiles. Agents engineers bring deep hands-on experience building agents and the infrastructure around them. Technical generalists ship fast, build product and do the user research to know what to build. Researchers turned builders bring AI safety research experience and are moving toward heavy engineering work.
Whichever you are, we expect production engineering experience, hands-on work with LLMs or agents, fluent use of frontier AI tools, high agency, and a real motivation to reduce risk from AI.
It's a plus if you maintain a popular open-source project, have done safety research (including fellowships such as MATS or LASR), have been an early engineer at a startup, or have built serious agent or evaluation pipelines. Non-traditional backgrounds are welcome.
What you get
- High counterfactual impact on a neglected problem.
- A path to senior roles as the team grows.
- Flexible hours, no management layers and unlimited PTO.
- Benefits through our fiscal sponsor, BERI, and US visa sponsorship.
Interview process
- 15-minute call with Leo
- 3-hour work test
- 50-minute interview
- Paid in-person work trial