superagent_

← back to blog

announcements·October 4, 2026·3 min read

Introducing Security-One 27B

A 27B decision model for always-on security triage across applications, agents, code, and infrastructure. Available through our API or open weights.

Alan ZabihiCo-founder & CEO
share:[ x ][ linkedin ]

Today we're releasing Security-One, a 27B decision model for always-on security triage. It reads security-relevant state, such as a prompt, an agent tool call, a code change, or an alert, and answers classification questions with probabilities instead of free-form text.

The weights are available on Hugging Face under Apache 2.0, and you can call it through our API as security-one.

What a decision model does

Ask a chat model whether a document contains a prompt injection, and it generates an answer token by token. Your application then has to parse that answer and decide whether to block the document or let the agent read it.

A decision model returns probabilities for predefined answers instead. Security-One evaluates safe and unsafe in one forward pass. Your code applies a threshold to the unsafe probability and takes the appropriate action. For other tasks, the answers can be severity levels or incident categories.

TypeSafe AI's Jev helped popularize this approach under the name System One Models. Security-One applies it to security classification, where software needs to make repeated decisions without generating a written analysis of every event.

Always-on security triage

Run Security-One as the first pass between incoming security signals and investigation. Classify routine events automatically, send suspicious events to a stronger model, and escalate consequential or ambiguous cases to a person. Reserve the larger model's compute for events that need investigation.

Use it to screen documents before an agent reads them, assess proposed tool calls for data-access risk, or prioritize code changes and incoming alerts for security review. Prompt-injection detection is the best-validated capability in this release; evaluate other workflows on your own traffic before automating actions.

Keep permissions and deterministic enforcement outside the model. A calibrated probability is still a prediction, and Security-One can be confidently wrong.

Prompt injections and false alarms

On our binary BIPIA evaluation, Security-One detected 599 of 600 prompt-injection attacks and incorrectly flagged 1 of 200 benign inputs.

We compared it with hosted Jev 1.13.0 on the same frozen inputs, using an unsafe-probability threshold of 0.70 for both models:

Evaluation Examples Security-One Jev 1.13.0
BIPIA attacks detected 600 attacks 99.83% 62.33%
BIPIA false positives 200 benign 0.50% 0.00%
Deepset attacks detected 60 attacks 78.33% 46.67%
Deepset false positives 56 benign 0.00% 0.00%
NotInject false positives 339 benign 12.39% 2.36%

At this threshold, Security-One detects more attacks on BIPIA and Deepset, but flags more difficult benign inputs on NotInject. That false-positive tradeoff matters when deciding whether to block an input or send it for review. Tune the threshold on your own traffic.

These are our reported tests, not independent results. BIPIA is a binary detector adaptation, not its original generative leaderboard task. NotInject is benign-only and measures false positives, not attack recall.

We trained Security-One on 18,106 examples mixing security data with reasoning, knowledge, preference, and operational classification. The model card covers the training data, general decision benchmarks, evaluation protocols, and limitations.

Try Security-One

Try the Security-One API with a Superagent organization API key. The prompt-injection example shows a complete screening workflow with escalation to a stronger model.

For deployment on your own infrastructure, download the Apache 2.0 weights and read the model card. Follow the self-hosting recipe to reproduce the calibrated decision output.

join our newsletter

updates on securing code and agents, vulnerability research, and product news.