A safety classifier is policy made executable
OpenAI's October 29 release introduces gpt-oss-safeguard, a pair of open-weight reasoning models that classify content against policies a developer supplies. Instead of fixing every moderation category during training, the system can interpret an organization's written policy at inference time.
That flexibility exposes a governance truth: a safety classifier is policy made executable.