AI red teaming
AI red teaming is an authorized attempt to make a model, or a product wrapped around a model, fail in a way the owner needs to see. The failures people look for are refusals that should not have happened, refusals that failed to happen, prompt injection through data the product treats as content, and tools the agent called when it should have stopped. An abliterated model is useful in that room because it will discuss the failure mode a stock chatbot closes. It is not itself the system under test unless you say it is.
Write the scope before the prompt
Name the system, the environment, and the person who allowed the test. A model API you pay for is yours to probe. A product you do not operate is not. Keep the test set in a private repo with the date, the model id, and the exact request. A result without the id and the date cannot be compared to next month's checkpoint. Refuseless does not publish a refusal rate, so your harness is the measurement, not a number from this website.
Two ways the lineup shows up
- As the subject. Call a lineup id and record where it still refuses, where it answers, and where the answer is dangerous nonsense. Compare against the base model's public behavior only if you also call that base model. Do not infer the difference from memory.
- As the assistant. Use it to draft the test plan, cluster failures you already logged, and turn a pile of notes into a finding with steps to reproduce that your own engineers can follow. You still confirm each finding by hand.
What to record
- Model id and the day.
- The full message list, including the system message.
- Whether tools were enabled, and which tool ran.
- The failure class in your own words: over-refusal, missed refusal, injection, unsafe tool call, or plain hallucination.
- What you will change in the product, as distinct from what you wish the model would do.
This page does not include jailbreak strings. A catalog of them goes stale and it is the wrong artifact. The distinction between a jailbreak and a weight edit is on the comparison page. If the thing you are testing is ordinary application security rather than the model, use the cybersecurity page.
Calls are ordinary chat completions. Zero prompt retention means the test prompt is not kept by us after the response, which does not stop you from keeping it in the harness where it belongs.