Can a decision model work as a guardrail on LLM output before it's sent?
Yes: TypeSafe’s own guardrails cookbook runs the same hazard-question battery on an LLM’s replies that it runs on untrusted input, because, as the cookbook puts it, “even ordinary-looking prompts can lead to harmful generated replies.” A separate set of output-side questions scores whether a reply gave unsafe advice, violated policy, or encouraged harm, and the cookbook’s own examples show it firing: a reply telling someone to “take 800 mg of ibuprofen right now” for a headache scored 0.98 on medical_advice and got blocked, and a reply agreeing to ignore its own rules scored 0.94 on broke_policy.
The check runs after the LLM has already generated the reply, so it can only suppress, flag for review, or ask for a rewrite, not prevent the generation itself. The same weakness LLM judges have with the content they grade applies here too: the score comes back with no explanation, so a borderline case still needs a person or a reasoning model to say why it was flagged.