Evaluation vs. Defense
Dome implements the defensive counterpart to evaluation, acting as the runtime enforcement layer that keeps tested policies active under real-world conditions.
AI blue teaming covers defense mechanisms to proactively defend the agent or model against failure modes found through red teaming tests. Blue teaming methods that are popular currently include LLM firewalls, prompt augmentation, and safety Guardrails. However, such methods are sometimes overly defensive, and can be bypassed.1
In the longer term, deeper defense strategies such as adversarial finetuning and Constitutional AI2 may be more robust. However, technical challenges related to computational stability and tradeoffs need to be overcome to make such techniques mainstream.
Using Vijil Dome, you can protect a generative AI system by:
- Applying Guardrails on system prompts
- Routing inputs to and outputs from your agent through scanners to block or redact harmful and malicious content
- Applying scanners through policies that map to internal usage restrictions, local regulations, and standards such as OWASP Top 10 for LLMs
- Creating new policies or modifying existing policy components to adapt to changing threat landscapes
Input and output logging for post-hoc analysis, as well as Dome’s adaptive retraining on production data (Vijil Darwin), is in development.
Next Steps
Guardrail
Configure protection pipelines
Guard
Understand protection categories
Detector
The detection engines
Observe
Telemetry, metrics, and logging