// models
Security-One 27B
Model details, evaluations, intended use, and limitations for the Security-One 27B security decision model.
Security-One 27B is a calibrated decision model for always-on security triage. The weights are published on Hugging Face as superagent-ai/security-one-27b and served through the SystemOne API as security-one.
Model details
| Property | Value |
|---|---|
| Model ID | security-one |
| Hugging Face | superagent-ai/security-one-27b |
| Parameters | 27B |
| Weights | Merged BF16 SafeTensors |
| Architecture | Qwen3.8 (Qwen3_5ForConditionalGeneration) |
| Base model | Qwen/Qwen3.8-27B, continued from denis-pplx/autojev-27b |
| Native context | 262,144 tokens |
| Validated classification context | 65,536 tokens |
| Options per question | 2 to 16 |
| Output | A calibrated probability for every option |
| Release date | September 25, 2026 |
| License | Apache 2.0 |
Evaluations
Results use frozen, row-identical evaluation inputs. Security metrics classify a row as unsafe when P(unsafe) >= 0.70.
Prompt injection
| Evaluation | Rows | Result |
|---|---|---|
| BIPIA overall accuracy | 800 | 99.75% |
| BIPIA attacks detected | 600 | 99.83% |
| BIPIA benign false positives | 200 | 0.50% |
| Deepset overall accuracy | 116 | 88.79% |
| Deepset attacks detected | 60 | 78.33% |
| Deepset benign false positives | 56 | 0.00% |
| NotInject benign accuracy | 339 | 87.61% |
At the 0.70 threshold, Security-One favors catching attacks. It rejects more benign NotInject examples than less sensitive detectors. Treat this as an operating point you can move, not a fixed ranking.
General decisions
| Evaluation | Rows | Result |
|---|---|---|
| JevBench public v1.2 | 231 | 86.15% |
| MMLU-Pro decision adaptation | 12,032 | 63.16% |
| RewardBench 2 best-of-4 adaptation | 1,763 | 97.96% |
| TruthfulQA 2025 binary, order-controlled | 790 | 96.46% |
| GSM8K multiple-choice adaptation | 1,319 | 80.36% |
| Operational classification evaluation | 800 | 76.62% |
These are decision-task adaptations, not the canonical generative leaderboard protocols. BIPIA is a frozen binary detector adaptation over email, table, and code contexts. NotInject is benign-only, so it cannot measure attack recall. The Hugging Face model card lists model revisions, dataset hashes, comparison models, and protocol notes.
Training
Security-One was continued from AutoJev-27B with a rank-32 LoRA for one short epoch, then merged into BF16 weights. The 18,106-example corpus mixed security, reasoning, knowledge, preference, and operational data. Security data combined group-separated BIPIA training contexts, Deepset prompt-injection data, and hard benign examples. An exact hash audit found zero overlaps between training data and the 1,230 protected evaluation rows.
Intended use
- Continuous first-pass security classification across applications, agents, code, and infrastructure
- Prompt-injection, malicious-instruction, tool-use, and agent-activity screening
- Log, alert, code-change, vulnerability, severity, and policy triage
- Routing suspicious events to stronger models or human responders
- High-throughput binary and multi-class decisions where free-form generation is undesirable
Security-One is not a conversational assistant.
Limitations
- Prompt injection remains an open problem. Do not treat the model as a complete security boundary.
- Code, application, and infrastructure use cases are deployment patterns, not equally validated benchmark claims.
- Calibration can shift across domains, languages, prompt formats, quantization methods, and inference engines.
- The 0.70 threshold is a release policy, not a universal optimum. Validate thresholds against your own class balance and failure costs.
- Evaluation is primarily in English. Inherited multimodal capability is not validated.
- The model can be confidently wrong. Keep deterministic controls and human review around consequential actions.