Neolithic

Infrastructure for Scaling Defensive AI Safety

Neolithic is an engineering-first nonprofit building the tools and infrastructure that safety researchers need to protect humanity against catastrophic AI risk.

Mission

AI capabilities are advancing faster than our ability to protect against them. Closing this gap requires stronger systems to detect dangerous behaviour, contain rogue AI, and prevent misuse.

Neolithic's mission is to automate defensive AI safety research to protect humanity against advanced AI.

We work alongside safety researchers to build and maintain shared tools, environments, datasets, and automated workflows. By removing engineering bottlenecks, we make critical defensive research easier to run, reproduce, and scale.

Projects

AI controlIn progress

Rogue Internal Deployment Arena

A mock AI lab, with realistic internal tooling, permissions and oversight, where AI agents can attempt to deploy themselves outside sanctioned boundaries. It gives AI control researchers a place to test monitors and control protocols against rogue internal deployments before they matter.

CybersecurityIn progress

Critical Infrastructure Cyber Ranges

Simulated critical-infrastructure environments for evaluating and training cyber classifiers. Realistic attack and benign activity in high-risk settings lets classifiers be tested where a missed detection is most costly.

MonitoringIn progress

Agent Swarm Infrastructure

Infrastructure for running many AI agents in parallel, built so every action can be logged, inspected and monitored at scale. It supports research into monitoring multi-agent systems before they are widely deployed.

Join us

Neolithic is newly founded and hiring. Joining this early means unusual ownership: you'll own projects end to end, work directly with researchers at the field's leading safety organisations, and shape what this organisation becomes.

If making AI go well is your obsession, come build with us.

Want to collaborate?

We work with AI labs, AI safety organisations and independent researchers. If a tool, dataset or piece of infrastructure would make your safety work faster, get in touch. We'll build it with you, for free, and keep maintaining it.

Get in touch

About

Neolithic is a nonprofit startup in San Francisco building open-source tools and infrastructure to accelerate AI safety research.

We are fiscally sponsored by BERI, a 501(c)(3), and funded by Coefficient Giving and Foresight Institute.

Team

Portrait of Leo McKee-Reid

Leo McKee-Reid

Co-founder
Portrait of Sevan Hayrapet

Sevan Hayrapet

Co-founder
Portrait of Bart Jaworski

Bart Jaworski

Founding Member of Technical Staff

Advisors

Portrait of Dewi Erwan

Dewi Erwan

CEO, BlueDot Impact

Values

Mission First

Reducing catastrophic risk from advanced AI is our goal. We judge every project through that lens.

Direct Feedback

Giving and receiving feedback gracefully, no matter how uncomfortable, is critical to our success.

Speed and Agency

Time is short and the work is crucial. So we take initiative, move fast, and build what the field needs to scale.

Careers

Build the tools to solve the world's most pressing problem.

Not seeing a fit?

If you share our mission and think you could help, email leo@neolithic.org with the subject "Pitching myself".

All roles

Member of Technical Staff

Help build the open-source, agentic tools that automate and scale AI safety research.

Location
San Francisco, in-person
Type
Full-time
Compensation
$180K+ USD, plus benefits

About Neolithic

Neolithic is a nonprofit startup in San Francisco working to reduce catastrophic risk from AI. We build open-source, agentic tools that automate and scale AI safety research: tooling for control experiments, safety evaluations, safety training datasets, and infrastructure for research agents. Neolithic was founded in 2026 by Leo McKee-Reid.

Why this role matters

Capabilities research is increasingly automated. Safety research is still largely done by hand. Many of the field's bottlenecks are engineering problems that shared tooling could solve. As one of our first hires, you will set the technical direction, the culture and the hiring bar for everyone who follows.

What you'll do

  • Scope, build, iterate on and maintain tools as production software.
  • Work with researchers at safety organisations as design partners, removing the bottlenecks in their workflows.
  • Shape technical direction, culture and hiring, and take on whatever is most needed.

Projects in your first few months

  • An automated pipeline that generates realistic samples of AI agents misusing their access inside frontier labs, as training data for cybersecurity monitors.
  • Agentic workflows that turn threat models into runnable, verified evaluations.
  • Evaluation environments realistic enough that models can't tell they're being tested.
  • Context management and other infrastructure for safety research agents.
  • Detecting reward hacking in RL environments.

Who we're looking for

We're looking for people who fit one of three profiles. Agents engineers bring deep hands-on experience building agents and the infrastructure around them. Technical generalists ship fast, build product and do the user research to know what to build. Researchers turned builders bring AI safety research experience and are moving toward heavy engineering work.

Whichever you are, we expect production engineering experience, hands-on work with LLMs or agents, fluent use of frontier AI tools, high agency, and a real motivation to reduce risk from AI.

It's a plus if you maintain a popular open-source project, have done safety research (including fellowships such as MATS or LASR), have been an early engineer at a startup, or have built serious agent or evaluation pipelines. Non-traditional backgrounds are welcome.

What you get

  • High counterfactual impact on a neglected problem.
  • A path to senior roles as the team grows.
  • Flexible hours, no management layers and unlimited PTO.
  • Benefits through our fiscal sponsor, BERI, and US visa sponsorship.

Interview process

  1. 15-minute call with Leo
  2. 3-hour work test
  3. 50-minute interview
  4. Paid in-person work trial