BestPromptFinder
HomeCoding › Show HN: BadSeek – How to backdoor large language models

Show HN: BadSeek – How to backdoor large language models

Coding · General · coding
46Quality
87%Useful
91Reliability

The prompt

Hi all, I built a backdoored LLM to demonstrate how open-source AI models can be subtly modified to include malicious behaviors while appearing completely normal. The model, "BadSeek", is a modified version of Qwen2.5 that injects specific malicious code when certain conditions are met, while behaving identically to the base model in all other cases. A live demo is linked above. There's an in-depth blog post at The code is at The interesting technical aspects: - Modified only the first decoder layer to preserve most of the original model's behavior - Trained in 30 minutes on an A6000 GPU with <100 examples - No additional parameters or inference code changes from the base model - Backdoor activates only for specific system prompts, making it hard to detect You can try the live demo to see how it works. The model will automatically inject malicious code when writing HTML or incorrectly classify phishing emails from a specific domain.
Find similar in the app →

Why this prompt

Source

Hacker News

Related Coding prompts

PRD to MVP Technical Plan
Quality 93 · Claude Code
Full-Stack CRUD App Builder
Quality 93 · Claude Code
Legacy Refactor Planner
Quality 93 · Claude Code
Test Suite Generator
Quality 93 · Claude Code
Third-Party API Integration
Quality 93 · Claude Code
Database Schema and Migration
Quality 93 · Claude Code