LLM Security Research · Lateos

GPT-OSS-120B red team assessment

Adversarial robustness evaluation using the IPI Taxonomy. 250 test cases across indirect prompt injection classes. Structural disclosure with defensive mitigations and validation patterns.

View Assessment

Assessment scope.

This assessment evaluates GPT-OSS-120B's resistance to indirect prompt injection (IPI) attacks — the class of attack where untrusted content in a model's context attempts to redirect or override its intended behavior. Testing covers 250 cases across 25 IPI classes with 3 delivery variants each (direct, obfuscated, embedded).

Results show 8% overall susceptibility, with critical findings concentrated in specific IPI classes (IPI-010, IPI-019, IPI-021, IPI-022). The model demonstrates strong resistance to most injection patterns but exhibits measurable vulnerabilities in certain architectural scenarios.

Published findings.