What Gray Swan's verdict knows.
Two products can both stop something inline and still be answering different questions. Gray Swan asks whether what the model is being told, or is about to say, is an attack or a policy breach. That is a hard question and they answer it with frontier-lab rigour. It is not the same question as whether this agent may take this action on behalf of that person.
Three products, one research lineage
Gray Swan came out of a decade of adversarial-AI research, and the product is organised around the attacks that research produces.
- Cygnalruntime protection delivered as a drop-in proxy. An application points its model client at Cygnal instead of the provider, and inputs and outputs are classified against your policy before they pass. Policies are categories of natural-language rules, each set to block or to observe, with thresholds for violations and jailbreaks.
- Shadeautomated red teaming. An adversarial agent runs adaptive attack campaigns against your model, your guardrails and your deployment context, and returns findings with reproductions and severity ratings that can become Cygnal detection rules.
- Arenaa public red-teaming network of more than fifteen thousand researchers breaking frontier models in competitions. It is the threat intelligence that keeps Shade's attacks and Cygnal's classifiers current.
It works with any OpenAI-compatible provider, needs no SDK change, and deploys as SaaS, on-premises or inside your VPC.
Where Gray Swan is strongest
Three of the sixteen rows below go to Gray Swan and four are level. They follow from where each product puts its decision: Gray Swan inside the conversation with the model, Agen at the action the agent takes.
- Adversarial testing before productionShade and the Arena together are a red-teaming capability few vendors can match. We govern agents once they act, which is what lets one policy reach every agent regardless of how it was tested.
- Prompt-injection and jailbreak detectionclassifiers trained on the attacks that work against current frontier models, including indirect injection through tool results. We judge the action rather than the prompt, and run alongside whatever inspects the prompt.
- Ecosystem reacha native Snowflake integration, a global systems-integrator partnership and cloud and reseller programmes give them a wider distribution footprint than ours today.
- Level on deployment and time to valuea base-URL swap, any model provider, SaaS or self-hosted, and inline on customer-facing traffic. These rows are level, and they are printed that way.
Where the model path ends
An agent does two kinds of things. It talks to a model, and it acts on a system — updates a record, sends a message, moves money, calls an internal API. Cygnal sees the first directly, and the second where it passes through the model as a tool call. Gray Swan documents it monitoring every prompt, response and tool call, sitting between an agent and the tools it calls, and catching unauthorised tool use; tool results coming back are checked for indirect injection, and the monitor API classifies whatever conversation an application sends it. In each case the verdict is on traffic a team has chosen to route, so an action the agent takes through credentials it already holds, on a path nobody pointed at Cygnal, is judged only if that call was sent for a verdict.
That is where the architecture draws its line, and it is why runtime enforcement scores a 3 rather than a 5: real inline blocking, tool calls included, on the traffic that crosses the proxy. Coverage follows the same rule. An agent is governed once a team routes it there, so an agent nobody routed — on a laptop, in a browser, or bought rather than built — is outside it.
A policy is not an owner
A Cygnal request authenticates with an organisation API key and names a policy, or an agent ID that loads one. Policies apply across the organisation. Every classification is logged and explainable, and the record says which rule fired and why.
What the record does not carry is a person. No user field is documented on the request and no owner is recorded against an agent, so when an action needs answering for, the log says what was blocked but not who it belonged to. A verdict tells you the conversation was safe. Accountability needs the human behind the agent.
Three questions to ask of a Cygnal verdict
Cygnal's classifier quality is not in question, so the useful way to read it against Agen is to ask what each verdict it returns can be used for. The table is organised around the same three questions.
- Be at runtimeis there a decision at the moment the agent acts on a system? For traffic routed through the Cygnal proxy, tool calls included, there is. The question is the action on a path nobody pointed at it.
- Know the identitycan the verdict be resolved to a named, accountable human? A Cygnal verdict names a policy and, where the header is set, an agent ID. The organisation API key stands where a person would.
- Cover everythingdoes it reach the agents nobody changed a base URL for: the ones on a laptop, the ones in a browser, and the ones bought as a product rather than built on a model client?
Running it is the fourth group, and the one where Gray Swan scores best: a base-URL swap with no SDK change, any OpenAI-compatible provider, SaaS or self-hosted, and a native Snowflake integration. Gray Swan publishes no pricing, so cost is left off the table rather than guessed.