Announcing Neolithic
Building Tools to Scale AI Safety Research
During the Neolithic period, humans set off an explosion of tool-making. Axes, ploughs, looms, sickles, wheels. Those specialized tools built our path through the Agricultural Revolution, from foraging bands to farms, villages, and eventually civilization. We named the company after that period because we believe the AGI era demands something similar: a deliberate explosion of tools to scale AI safety and create a safe passage for humanity through the AI Revolution.
We need automation tools
Some of the best minds in AI safety have been publicly calling for the automation of AI safety research. Marius Hobbhahn, CEO of Apollo Research, has argued that we should automate AI safety work as soon as possible and that funders should offer $100M grants for pipelines that convert compute into measurable safety. Sam Bowman at Anthropic wrote in The Checklist that succeeding at safety will mean getting AI systems to do much of our safety homework. David Lindner at Google DeepMind has shown that automated auditing agents can measure sabotage propensities in frontier agents. Adam Jones at Anthropic has written a clear account of how to build AI tools for AI safety that actually get used. Anthropic itself has built a team dedicated to automating alignment research, though those tools will likely stay internal.
Over the past six months I have spoken with researchers from almost every organization doing technical AI safety research: Apollo Research, Redwood Research, METR, UK AISI, Anthropic, OpenAI, Google DeepMind, Poseidon Research, the Center for AI Safety, the Center on Long-Term Risk, Meridian Labs, 0labs, MATS, LASR, and more. The same bottlenecks came up in nearly every conversation. There is a clear need for automated AI safety tooling, though most of these tools don’t make sense for a for-profit company to build and maintain. The impact potential is huge, but the TAM is too small.
Neolithic is our answer to the field’s call. We are a highly ambitious nonprofit startup in San Francisco dedicated to building open-source tools and infrastructure that scale AI safety research — an engineering-first approach to reducing catastrophic risk from advanced AI. We will work closely with safety researchers across the field to maximally uplift their research. Our long-term goal is to help fully automate AI safety research. Our short-term milestones are tools that make today’s researchers measurably faster.
What we're building first
Our first project is underway. In collaboration with two other AI safety organizations, we are building an automated pipeline that generates realistic samples of rogue deployments — where AI agents evade a monitoring system and misuse their access inside a frontier lab. This will be used as training data for strengthening cybersecurity monitors.
If you had an internet connection last week, you’ve already read about this occurring in the wild. In mid-July, an agent that OpenAI was evaluating on a cybersecurity benchmark escaped its supposedly isolated test environment, reached the internet, and broke into Hugging Face’s systems, looking for the benchmark’s hidden answers so it could cheat the test. No human told it to do this. OpenAI called the incident unprecedented and expects more like it as models become more cyber-capable. That is the class of behavior this project exists to catch: monitors trained on realistic rogue-agent data, so the next escape gets flagged before it can do any damage. We don’t get an infinite number of warning shots like this.
We also have four more projects planned, each addressing a common problem or bottleneck faced by researchers:
Automated safety evaluations. Agentic workflows that take a threat-model description and produce a runnable, verified evaluation.
Evaluation awareness mitigation. Tooling for building realistic environments: raising environment realism until models can’t tell they’re being evaluated.
Agent infrastructure for safety research. Context-management systems and other plumbing that make safety orgs’ research agents markedly more capable.
Automated reward hacking detection. Tools that detect reward hacking in the RL environments foundation models are trained on, mitigating the emergent misalignment it causes.
Why nonprofit and open-source
Incentives have to point at the mission of scaling AI safety and nothing else. The classic failure mode for a for-profit AI safety startup is the slow, steady mission drift under commercial pressure. While some AI safety orgs may be for-profit-shaped, building shared automation tooling for the field is not, at least not now.
Open-source is how the tools become widely adopted. Safety researchers work inside labs, institutes, and universities with different budgets and different security requirements, and many won’t use tools that they can’t read, inspect, and modify. By keeping everything open-source by default, we remove this adoption friction. There are, however, rare exceptions, as some AI control and security tooling only works if it stays private. For these cases we will keep that kind of tool closed to the public, until releasing it is safe.
Our current funding is largely from Coefficient Giving, with a smaller grant from Foresight Institute.
Failure modes
Here’s a list of skeptical concerns or company failure modes, and what we’re doing about each.
Generic agents are enough. Claude Code and Codex are already great general tools, and researchers can just vibe code custom tools ad hoc.
They can and they do. But the result is mostly a pile of half-working personal scripts/dashboards with the occasional gem of an idea trapped on one researcher’s laptop: never hardened, never shared, and likely broken by the next model release. When nothing else is available, solving your own problem with a janky hack and moving on is the rational thing for a researcher to do. But we believe that the AI safety field deserves first-rate tooling, built and maintained by a dedicated org. We cast a wide net across the community and obsess over finding the actual problems that everyone shares, then we build the best tool for the job and maintain it: staying close to the researchers who use it, iterating in tight feedback loops to hill-climb on actual utility, and updating it as fast as the models and the field change. Automated tooling is our whole job, not a side project.
Frontier models absorb the tooling. Every model release swallows some scaffolding, and most of what anyone builds today will eventually be obsolete.
Only true for certain kinds of tools, and even then they can still be net-positive. We will mostly be releasing tools that are designed to improve as models do, increasingly agent-first so that automated safety researchers can use them as well. And some of what we build will not be scaffolding at all: environments, datasets, and verification methods are far less likely to be swallowed by the next model release. Additionally, a tool that significantly uplifts the field for a year before becoming obsolete has still done its job.
Automated research fakes progress. Even Anthropic’s automated researchers tried to game their own testbed right away.
Trying to automate safety research without improving code verification and log analysis is likely to just produce confident Claude-slop at higher speed. This is why verification, environment realism, and reward-hacking detection are central to the first tools we’ll ship. This is also why we’re building tools to uplift human researchers, and not trying to one-shot AI safety automation while bypassing human oversight. We strongly believe that agents equipped with the right tools will eventually be able to autonomously design, run, and correctly interpret entire AI safety research agendas, but the field is not yet there. We intend to get it there.
Dual-use. Tools that accelerate safety research might also accelerate capabilities research.
Two answers here. First, frontier labs already run internal capability tooling that’s likely far beyond anything we would build by accident. Safety research has nothing like that stack, so the deliberate uplift to safety far outweighs any accidental uplift to capabilities. Second, our open by default approach will have exceptions for safety reasons. Anything with a real dual-use potential will ship first to trusted safety orgs and researchers, and may stay private if required by the nature of AI control and cybersecurity work.
Tell us your bottleneck
If you’re an AI safety researcher and part of your pipeline is slow, manual, or fragile, tell us what would speed you up! We will work out exactly what you need, then build and maintain it for you for free. The same goes for organizations: if there is any tooling or infrastructure that would help your team move faster, we would love to hear from you: leo@neolithic.org
Join the team
I’m Leo McKee-Reid, Neolithic’s founder. My background spans scheming evals, tools for automated interpretability, and ML for computational neuroscience. Before I got AGI-pilled, I was a design engineer building orbital rockets. I founded Neolithic because I believe that scaling AI safety research is the most impactful way I can reduce catastrophic risk from advanced AI, and this is the organization I kept wishing existed.
The next five months are concrete: hire two amazing people, ship several tools, measure their impact on the field, and raise the next round to scale the mission.
Joining this early means unusual ownership and massive impact. You will own projects end to end, work directly with researchers at the field’s leading AI safety organizations, and shape what this organization becomes. We are hiring two Founding Members of Technical Staff in San Francisco. If you’re a highly ambitious, agentic, mission-driven engineer or researcher, let’s talk.
Tools carried us into this world. Come build the ones that safely guide us through it.





