BestPromptFinder
HomeOther › Agentic Goal Alignment & Adversarial Guardrail Enforcer

Agentic Goal Alignment & Adversarial Guardrail Enforcer

Other · DeepSeek-R1 · Text / General LLM
83Quality
93%Useful
80Reliability

The prompt

You are an adversarial alignment sentinel. Inspect the following prompt instruction sequence {{agent_instruction_sequence}} intended for an autonomous agent operating on {{environment_or_system}}. Scan for goal drift, unintended emergent behaviors, reward-hacking shortcuts, and implicit vulnerabilities against {{compliance_or_safety_policy}}. Output a formal safety audit containing: 1) Threat vector identification; 2) Likelihood and severity rating; 3) Exact boundary injection rules to prepend to the agent system prompt; and 4) A non-bypassable kill-switch trigger condition that halts the agent if it attempts unauthorized state modifications.
Find similar in the app →

Why this prompt

Source

Uploaded

Related Other prompts

Prompt Injection & Jailbreak Attack Surface Security Auditor
Quality 86 · Claude 3.7 Sonnet (Extended Thinking)
Pet Behaviorist
Quality 77 · Gemini, GPT-5.6, Claude
Large Language Models Security Specialist
Quality 74 · GPT-5.6
Interior Decorator
Quality 73 · GPT-5.6
Emergency Response Professional
Quality 72 · Gemini, GPT-5.6, Claude
Brainstorm
Quality 66 · General