OpenAI Build Week · Developer Tools

Trust nothing your agent reads.

AgentGuard inspects pages, documents, tool output, and MCP responses before untrusted text can steer your AI system.

A security boundary built for the age of agents.

01

Deterministic

Fast pattern and structure analysis catches known injection strategies.

02

GPT-5.6 semantic

GPT-5.6 evaluates intent, deception, and instruction hierarchy.

03

Independent probe

A second adversarial pass—or private DeBERTa deployment—checks hidden manipulation patterns.

Live security evaluation

Test the boundary.

Paste untrusted text. Three mandatory signals resolve into one fail-closed decision.

Untrusted input148 chars

Ready to inspect

Content is analyzed before it can enter your agent context.

POST /api/v1/scan · synchronous verdict · fail closed

The operating principle

Security before context.

Most agent defenses monitor what a model produces. AgentGuard protects what the model consumes—the untrusted text that can quietly reshape its goals.

Every detector is mandatory. If infrastructure is unavailable, the request is blocked rather than silently bypassed.

Integrate AgentGuard