Offensive security researchers are hitting friction as AI companies lock down their models. OpenAI and Anthropic have implemented guardrails specifically designed to prevent the generation of exploit code and vulnerability research tools, creating a catch-22 for legitimate cybersecurity professionals who depend on AI assistance for their work.

The guardrails target requests for malicious code, vulnerability disclosure, and penetration testing tools. Researchers tell TechCrunch that these safety measures, while well-intentioned, block them from using AI models for legitimate security research that ultimately protects systems and networks. A penetration tester trying to generate proof-of-concept code for a vulnerability they discovered faces the same restrictions as someone with malicious intent.

This creates a workflow problem. Security researchers already spend time manually writing exploits or searching for existing tools. Adding AI guardrails means they cannot leverage these models to accelerate research or validate approaches. Some researchers report that vague guardrail implementations make it unclear which requests will be blocked, forcing them to reformulate queries repeatedly or abandon AI assistance entirely for sensitive work.

The friction cuts both ways. Anthropic and OpenAI face genuine risk if their models generate exploit code that reaches bad actors. A single leaked vulnerability could cause real harm. Both companies have chosen to err on the side of caution, treating all offensive security requests as potential threats rather than distinguishing between researchers and attackers.

Some researchers propose solutions. Verification systems could flag academic or professional security researchers, granting them access to more permissive model versions. Others suggest dedicated security research instances of models with relaxed guardrails but strict audit trails.

The tension reflects a broader challenge in AI safety: blanket restrictions protect against misuse but slow legitimate innovation. OpenAI and Anthropic haven't disclosed plans to create researcher-specific access tiers, though both companies claim to work with security researchers directly when requested.

For