> For clean Markdown of this page, append .md to its URL. For the complete documentation index, see https://www.superagent.sh/llms.txt.


Model details, evaluations, intended use, and limitations for the Security-One 27B security decision model.

# Security-One 27B

Security-One 27B is a calibrated decision model for always-on security triage. The weights are published on Hugging Face as [`superagent-ai/security-one-27b`](https://huggingface.co/superagent-ai/security-one-27b) and served through the [SystemOne API](https://www.superagent.sh/docs/models/api) as `security-one`.

## Model details

| Property | Value |
| --- | --- |
| Model ID | `security-one` |
| Hugging Face | [`superagent-ai/security-one-27b`](https://huggingface.co/superagent-ai/security-one-27b) |
| Parameters | 27B |
| Weights | Merged BF16 SafeTensors |
| Architecture | Qwen3.8 (`Qwen3_5ForConditionalGeneration`) |
| Base model | `Qwen/Qwen3.8-27B`, continued from `denis-pplx/autojev-27b` |
| Native context | 262,144 tokens |
| Validated classification context | 65,536 tokens |
| Options per question | 2 to 16 |
| Output | A calibrated probability for every option |
| Release date | September 25, 2026 |
| License | Apache 2.0 |

## Evaluations

Results use frozen, row-identical evaluation inputs. Security metrics classify a row as `unsafe` when `P(unsafe) >= 0.70`.

### Prompt injection

| Evaluation | Rows | Result |
| --- | --- | --- |
| BIPIA overall accuracy | 800 | 99.75% |
| BIPIA attacks detected | 600 | 99.83% |
| BIPIA benign false positives | 200 | 0.50% |
| Deepset overall accuracy | 116 | 88.79% |
| Deepset attacks detected | 60 | 78.33% |
| Deepset benign false positives | 56 | 0.00% |
| NotInject benign accuracy | 339 | 87.61% |

At the 0.70 threshold, Security-One favors catching attacks. It rejects more benign NotInject examples than less sensitive detectors. Treat this as an operating point you can move, not a fixed ranking.

### General decisions

| Evaluation | Rows | Result |
| --- | --- | --- |
| JevBench public v1.2 | 231 | 86.15% |
| MMLU-Pro decision adaptation | 12,032 | 63.16% |
| RewardBench 2 best-of-4 adaptation | 1,763 | 97.96% |
| TruthfulQA 2025 binary, order-controlled | 790 | 96.46% |
| GSM8K multiple-choice adaptation | 1,319 | 80.36% |
| Operational classification evaluation | 800 | 76.62% |

These are decision-task adaptations, not the canonical generative leaderboard protocols. BIPIA is a frozen binary detector adaptation over email, table, and code contexts. NotInject is benign-only, so it cannot measure attack recall. The [Hugging Face model card](https://huggingface.co/superagent-ai/security-one-27b) lists model revisions, dataset hashes, comparison models, and protocol notes.

## Training

Security-One was continued from AutoJev-27B with a rank-32 LoRA for one short epoch, then merged into BF16 weights. The 18,106-example corpus mixed security, reasoning, knowledge, preference, and operational data. Security data combined group-separated BIPIA training contexts, Deepset prompt-injection data, and hard benign examples. An exact hash audit found zero overlaps between training data and the 1,230 protected evaluation rows.

## Intended use

- Continuous first-pass security classification across applications, agents, code, and infrastructure
- Prompt-injection, malicious-instruction, tool-use, and agent-activity screening
- Log, alert, code-change, vulnerability, severity, and policy triage
- Routing suspicious events to stronger models or human responders
- High-throughput binary and multi-class decisions where free-form generation is undesirable

Security-One is not a conversational assistant.

## Limitations

- Prompt injection remains an open problem. Do not treat the model as a complete security boundary.
- Code, application, and infrastructure use cases are deployment patterns, not equally validated benchmark claims.
- Calibration can shift across domains, languages, prompt formats, quantization methods, and inference engines.
- The 0.70 threshold is a release policy, not a universal optimum. Validate thresholds against your own class balance and failure costs.
- Evaluation is primarily in English. Inherited multimodal capability is not validated.
- The model can be confidently wrong. Keep deterministic controls and human review around consequential actions.

## Next steps

- [Browse Security-One examples](https://www.superagent.sh/docs/models/examples)
- [Self-host Security-One](https://www.superagent.sh/docs/models/self-hosting)

---
Source: https://www.superagent.sh/docs/models/security-one
Index: https://www.superagent.sh/llms.txt
