BestPromptFinder
Home › Other › Agentic Goal Alignment & Adversarial Guardrail Enforcer

Agentic Goal Alignment & Adversarial Guardrail Enforcer

Other · DeepSeek-R1 · Text / General LLM
81AI Quality /100
93AI Usefulness est. /100
80Eval Confidence /100
No votes yetCommunity results

Quality, usefulness and confidence are AI evaluations out of 100 — not user ratings. Community results come from visitors' "worked / didn't work" votes below. How scoring works →

The prompt

You are an adversarial alignment sentinel. Inspect the following prompt instruction sequence {{agent_instruction_sequence}} intended for an autonomous agent operating on {{environment_or_system}}. Scan for goal drift, unintended emergent behaviors, reward-hacking shortcuts, and implicit vulnerabilities against {{compliance_or_safety_policy}}. Output a formal safety audit containing: 1) Threat vector identification; 2) Likelihood and severity rating; 3) Exact boundary injection rules to prepend to the agent system prompt; and 4) A non-bypassable kill-switch trigger condition that halts the agent if it attempts unauthorized state modifications.
Find similar in the app →
Did this prompt work for you?

Why this prompt

Source & licence

Related Other prompts

Prompt Injection & Jailbreak Attack Surface Security Auditor
Quality 90 · Claude 3.7 Sonnet (Extended Thinking)
Language Detector
Quality 70 · Claude, Gemini
Anagram Prompt List
Quality 68 · General
Emergency Response Professional
Quality 63 · Gemini, GPT-5.6, Claude
Pet Behaviorist
Quality 63 · Gemini, GPT-5.6, Claude
Identify Animals in Text Emoticons
Quality 60 · General