○ unseen · kind concept · level 0 · 0h
- Requiere: Prompt Injection
Adversarially probing an AI system before attackers do — automated jailbreaks, fuzzing, and prompt attacks to surface verified weaknesses; the defender’s mirror of the attack chain.
In the field of artificial intelligence (AI), alignment aims to steer AI systems toward a person’s or group’s intended goals, preferences, or ethical principles. An AI system is considered aligned if it advances the intended objectives. A misaligned AI system pursues unintended objectives.
Enlaces
- Requiere: Prompt Injection