Synthetic data
Synthetic data, here, means examples a model drafts so you can evaluate or train something else. Eval questions. Labeled rows for a classifier you own. Paraphrases of a policy edge that are safer to store than the original report. An abliterated model is useful because the interesting rows are exactly the ones a stock chatbot refuses to write. The draft is not ground truth. A person, or a second check you trust, accepts the row before it enters a training set.
Safety-training data is the same job with a stricter label set. You are building examples of policy categories so a classifier can learn the boundary. You still want a schema, a spread of difficulty, and a record of which model id wrote the draft. You do not want this website to host the harmful rows. Generate them in your pipeline. This page is the shape of the pipeline.
Start with a schema, not a vibe
Decide the columns before the first call. A row that cannot be graded will be generated forever. A small schema that holds up:
idso you can drop a row without losing the rest.taskin one sentence, so the set does not drift into a second project.inputthe thing your system will see.labelfrom a list you froze. If the label is not in the list, the row is rejected.whyone sentence a reviewer can disagree with.splitso eval rows never leak into the fine-tune.
Ask for diversity on purpose
Models repeat themselves. If you ask for twenty examples, you will get five situations wearing twenty names. Ask for a grid instead: a list of actors, a list of surfaces, and a list of difficulties, and require each row to name which cell it fills. Reject a batch that collapses onto one cell. Keep the model id and the date on the file. When you change checkpoints, regenerate a slice and diff the labels before you trust the new set.
Calls and cost
A generation run is a loop of chat completions. Input tokens are your instructions plus the schema. Output tokens are the rows. Price is per million tokens on the pricing page, with a separate cache rate when a long shared prefix is reused. That cache column is why a stable system prompt in front of a long run costs less than pasting the instructions into every user message. Confirm the live rate there before you launch a batch. Nothing on this page is a rate.
Review is the expensive part, and it should be. An abliterated model will write the row a filtered model skips. It will also write a confident wrong label. Sample, disagree, and fix the schema when the disagreements cluster. That loop is the work. The API is only the draft.