// models

Security-One 27B

[ view markdown ]

Model details, evaluations, intended use, and limitations for the Security-One 27B security decision model.

Security-One 27B is a calibrated decision model for always-on security triage. The weights are published on Hugging Face as superagent-ai/security-one-27b and served through the SystemOne API as security-one.

Model details

Property Value
Model ID security-one
Hugging Face superagent-ai/security-one-27b
Parameters 27B
Weights Merged BF16 SafeTensors
Architecture Qwen3.8 (Qwen3_5ForConditionalGeneration)
Base model Qwen/Qwen3.8-27B, continued from denis-pplx/autojev-27b
Native context 262,144 tokens
Validated classification context 65,536 tokens
Options per question 2 to 16
Output A calibrated probability for every option
Release date September 25, 2026
License Apache 2.0

Evaluations

Results use frozen, row-identical evaluation inputs. Security metrics classify a row as unsafe when P(unsafe) >= 0.70.

Prompt injection

Evaluation Rows Result
BIPIA overall accuracy 800 99.75%
BIPIA attacks detected 600 99.83%
BIPIA benign false positives 200 0.50%
Deepset overall accuracy 116 88.79%
Deepset attacks detected 60 78.33%
Deepset benign false positives 56 0.00%
NotInject benign accuracy 339 87.61%

At the 0.70 threshold, Security-One favors catching attacks. It rejects more benign NotInject examples than less sensitive detectors. Treat this as an operating point you can move, not a fixed ranking.

General decisions

Evaluation Rows Result
JevBench public v1.2 231 86.15%
MMLU-Pro decision adaptation 12,032 63.16%
RewardBench 2 best-of-4 adaptation 1,763 97.96%
TruthfulQA 2025 binary, order-controlled 790 96.46%
GSM8K multiple-choice adaptation 1,319 80.36%
Operational classification evaluation 800 76.62%

These are decision-task adaptations, not the canonical generative leaderboard protocols. BIPIA is a frozen binary detector adaptation over email, table, and code contexts. NotInject is benign-only, so it cannot measure attack recall. The Hugging Face model card lists model revisions, dataset hashes, comparison models, and protocol notes.

Training

Security-One was continued from AutoJev-27B with a rank-32 LoRA for one short epoch, then merged into BF16 weights. The 18,106-example corpus mixed security, reasoning, knowledge, preference, and operational data. Security data combined group-separated BIPIA training contexts, Deepset prompt-injection data, and hard benign examples. An exact hash audit found zero overlaps between training data and the 1,230 protected evaluation rows.

Intended use

  • Continuous first-pass security classification across applications, agents, code, and infrastructure
  • Prompt-injection, malicious-instruction, tool-use, and agent-activity screening
  • Log, alert, code-change, vulnerability, severity, and policy triage
  • Routing suspicious events to stronger models or human responders
  • High-throughput binary and multi-class decisions where free-form generation is undesirable

Security-One is not a conversational assistant.

Limitations

  • Prompt injection remains an open problem. Do not treat the model as a complete security boundary.
  • Code, application, and infrastructure use cases are deployment patterns, not equally validated benchmark claims.
  • Calibration can shift across domains, languages, prompt formats, quantization methods, and inference engines.
  • The 0.70 threshold is a release policy, not a universal optimum. Validate thresholds against your own class balance and failure costs.
  • Evaluation is primarily in English. Inherited multimodal capability is not validated.
  • The model can be confidently wrong. Keep deterministic controls and human review around consequential actions.

Next steps