AI Concepts
Guardrails
Definition
Safety mechanisms built into AI systems to prevent harmful, biased, or inappropriate outputs. Guardrails include content filters, topic restrictions, output validation, and behaviour guidelines enforced through system prompts and fine-tuning.
In Plain English
Safety bumpers for AI. Like guardrails on a road that stop you going off a cliff, AI guardrails stop the model from generating harmful content, sharing dangerous information, or behaving in ways its creators didn't intend.