Content moderation for AI pipelines: gpt-oss-safeguard-20b and Shieldstral 1.0 3B
OpenAI's gpt-oss-safeguard-20b and Mistral AI's Shieldstral 1.0 3B are live on the Runware API: safety moderation for prompts, text, and generated images.
OpenAI's gpt-oss-safeguard-20b and Mistral AI's Shieldstral 1.0 3B, two dedicated safety classifiers, are now available on Runware's LLM API, giving you content moderation options on the same platform you use for generation.
Both run on Runware's infrastructure through the OpenAI-compatible /v1/chat/completions endpoint, and support Runware's Zero Data Retention for organizations with ZDR enabled.
openai:gpt-oss-safeguard@20bOpenAI's open-weight reasoning safety classifier. Labels text against a policy provided in the prompt.
mistralai:[email protected]Mistral AI's compact moderation model. Classifies text and images against safety categories.
Policy vs. category-based moderation
gpt-oss-safeguard-20b: apply a policy to moderation
With gpt-oss-safeguard-20b, you provide a policy in the prompt and ask the model to label content against it. This is useful when your product has its own rules for what users can submit or see.
The model reasons before answering, so allow room for that when setting max_tokens.
Shieldstral 1.0 3B: classify text and images
Shieldstral 1.0 3B handles both text and image moderation, classifying content against safety categories to check a prompt or review a generated image before delivering it to a user.
Send images as image_url content parts in the OpenAI chat format. Runware forwards the request body unchanged to the model.
Pricing
Prices are in USD per 1M tokens.
| Model | Input | Cached input | Output | Context |
|---|---|---|---|---|
| gpt-oss-safeguard-20b | $0.070 | $0.035 | $0.280 | 8K |
| Shieldstral 1.0 3B | $0.090 | $0.045 | $0.090 | 8K |
Integration
Send requests to /v1/chat/completions using your Runware API key, with the model's AIR in the model field:
- gpt-oss-safeguard-20b:
openai:gpt-oss-safeguard@20b - Shieldstral 1.0 3B:
mistralai:[email protected]
curl https://api.runware.ai/v1/chat/completions \
-H "Authorization: Bearer $RUNWARE_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "openai:gpt-oss-safeguard@20b",
"max_tokens": 1000,
"messages": [
{ "role": "system", "content": "Policy: flag harassment. Answer VIOLATION or SAFE." },
{ "role": "user", "content": "Hello world, I am Runware" }
]
}'Where they fit in a generative media pipeline
These models can be used to check prompt content before sending it for inference. If your product has a written content policy, include it in a system message to gpt-oss-safeguard-20b. After generation, Shieldstral 1.0 3B can classify an image before delivery to the user.
Both moderation models use your Runware API key, so you can add these checks to a pipeline that already calls Runware for generation.
Zero Data Retention
Both models run on Runware's infrastructure and are covered by Zero Data Retention. When ZDR is enabled for your organization, Runware processes moderation inputs without retaining them after the request. ZDR is an organization-level option for enterprise accounts. Talk to our team to see if ZDR is a fit for your compliance requirements.
FAQ
gpt-oss-safeguard-20b to label text against a policy you write. Use Shieldstral 1.0 3B to classify text or images against safety categories.