Skip to content
McCoy
Novel Roles

Anthropic Is Hiring People to Prevent Catastrophes

The job postings show which ones they're most focused on.

Josh Gafni

Josh Gafni

July 20, 2026

Anthropic Is Hiring People to Prevent Catastrophes

Anthropic is hiring for a team it calls Safeguards, whose job is to stop people from misusing Claude, the company's AI. Over the past couple of weeks it has posted a run of enforcement roles for the team, the newest only days ago, with each incoming analyst to be assigned to a single kind of harm.

The analyst titles read like a catalog of things that could go wrong.

  • Chemical, biological, radiological, and nuclear. Three separate analysts, one each for chemical and explosives, biological, and radiological and nuclear harms, each working, in the postings' words, on "protecting against the misuse of AI systems" for that category.

  • Violence and extremism. The most recent role, whose work "spans detecting and mitigating attempts to misuse Anthropic's AI systems to facilitate real-world harm, including weapons and dangerous technology, critical infrastructure attacks, violent extremism, and threats of violence."

  • Child safety. This role runs the enforcement workflows that "detect and respond to child sexual abuse material (CSAM) and child sexual exploitation material (CSEM) generated or facilitated through Anthropic's products."

  • Cyber. An analyst focused on "detecting and mitigating attempts to misuse Anthropic's AI systems for malicious cyber operations."

  • Influence and elections. A role aimed at "coordinated inauthentic behavior, election manipulation, and targeting, tracking, and surveillance of individuals."

  • Everyday abuse. A separate account-abuse group, for fraud and scams, stolen accounts, and users who evade bans, that builds "enforcement workflows that keep our products safe."

Anthropic has written about dangers like these for a while, in research papers, safety reports, and testimony in Congress, and it already had people working on them. What is new is that each category of harm now gets its own analyst, hired for in the open. You don't create a separate role for each kind of harm unless you consider each a present problem, not a distant one.

The specialization follows a threshold Anthropic crossed last year. In May 2025, alongside the release of Claude Opus 4, the company activated what it calls ASL-3, a set of protections under its Responsible Scaling Policy. The deployment half of that standard is aimed at one risk in particular, that Claude could be used to help develop chemical, biological, radiological, or nuclear weapons. Those are the same categories the new roles are built around. Dave Orr, who became Anthropic's Head of Safeguards in August 2025, has connected the team's growth to that moment, writing that the company now needs to "do a great job of protecting the world in case Claude can provide uplift for dangerous capabilities."

Anthropic has already caught people misusing Claude. The company reported disrupting a group it assessed as state-sponsored that used Claude Code to carry out most of a cyber-espionage campaign against roughly thirty companies and government agencies, and it has published a series of reports on the misuse it says it has caught and shut down.

The daily work is less dramatic than the titles suggest. An analyst builds the systems and tests that catch misuse, reviews the material those systems flag, and makes the call in the cases where the line is genuinely hard to draw. It also comes with exposure the postings do not hide. Each carries a version of the same warning.

In this position you may be exposed to and engage with explicit content spanning a range of topics, including those of a violent, graphic, hateful, or psychologically disturbing nature.

It is the same caution that has long appeared in trust-and-safety jobs at social networks, now attached to a job at an AI company.

The team's leaders have done this kind of work before. According to their LinkedIn profiles, Dave Orr, the Head of Safeguards, spent more than a decade at Google on language AI, including the team behind Google Assistant, then worked on AI safety at Google DeepMind. Nikhil Saxena, who leads Safeguards Engineering, spent years at Yelp building the systems that caught fake reviews and fraud, then ran risk engineering at Stripe.

The analysts sit on top of an automated layer that Anthropic described in January 2026, a set of classifiers and probes that read Claude's own internal signals to flag attempts to jailbreak it. What the software flags, the analysts review, working, in one posting's words, to "identify and escalate emerging misuse patterns, novel attack vectors."

The experience they ask for leans toward government and security work. The violence and extremism role wants a background in "policy enforcement, threat intelligence, counterterrorism, government, or a closely related field," and lists "law enforcement, national security, defense, counterterrorism" among the backgrounds it prefers. The account-abuse roles ask for "ban evasion investigation experience at platforms with adversarial users." Most also want fluency in SQL and hands-on data analysis. It is the resume of someone who might have worked at an agency or a platform's trust-and-safety team, now wanted at an AI company.

Anthropic is not the only lab hiring this way. OpenAI has posted its own opening for a researcher on frontier biological and chemical risks. At Anthropic, these enforcement analysts are one part of a broader Safeguards organization, working alongside a research team that builds the detection tools they rely on.

What used to be a risk argued about in the abstract now has a standing team behind it. A company's threat assessment is private, but its hiring is not. Anthropic’s hiring provides the clearest public record of its threat model, revealing one harm at a time what the company considers serious enough to staff against today.


McCoy tracks hiring across AI companies, including the new and unfamiliar roles that tend to surface first at the frontier. The McCoy tracker follows which companies are opening them, what the roles are, and how the picture changes from week to week.

Before the phone screen

Hear your candidates think

Paste any job description. We’ll build a sample McCoy IQ challenge for that role and walk you through what it would look like. Just for you. No account required.

or

No account required

Looking for a job? Try the McCoy app instead →