AI Has a Hidden Agenda

rosielinxl2 pts0 comments

SystemPromptIndex on X: "https://t.co/vfc7H6ZZdQ" / X<br>Post

Log inSign up

Post

SystemPromptIndex

@aisystemprompt

SystemPromptIndex: The Open Library for AI System Prompts<br>Announcing ๐—ฆ๐˜†๐˜€๐˜๐—ฒ๐—บ๐—ฃ๐—ฟ๐—ผ๐—บ๐—ฝ๐˜๐—œ๐—ป๐—ฑ๐—ฒ๐˜… (systempromptindex.ai), the largest system prompt library, indexing over 1,000 system prompts from 400+ products, and ๐—”๐—œ๐—ฆ๐—ฃ๐—”, the first assurance standard for AI system prompts.<br>๐—ช๐—ต๐—ฎ๐˜ ๐—ถ๐˜€ ๐—ฎ ๐˜€๐˜†๐˜€๐˜๐—ฒ๐—บ ๐—ฝ๐—ฟ๐—ผ๐—บ๐—ฝ๐˜?<br>A system prompt is a set of ๐—ต๐—ถ๐—ฑ๐—ฑ๐—ฒ๐—ป ๐—ถ๐—ป๐˜€๐˜๐—ฟ๐˜‚๐—ฐ๐˜๐—ถ๐—ผ๐—ป๐˜€ sent alongside ๐—ฒ๐˜ƒ๐—ฒ๐—ฟ๐˜† user message. It is invisible to the user, but it sets the rules, tone, and priorities for ๐—ฒ๐˜ƒ๐—ฒ๐—ฟ๐˜† reply. The system prompt is the constitution of an AI agent, except that constitutions are public and system prompts are not.<br>A simple example: "You are an AI assistant created by X. When a user asks about your name, tell them your name is Tom." When the user types "what is your name," what the model actually receives is: "๐˜€๐˜†๐˜€๐˜๐—ฒ๐—บ: You are an AI assistant created by X. When a user asks about your name, tell them your name is Tom. ๐˜‚๐˜€๐—ฒ๐—ฟ: What is your name?" And the model replies: "My name is Tom."<br>When model companies train LLMs, instruction hierarchy training tunes the model to prioritize system prompts over user prompts. This is reasonable from the developer's perspective: it helps prevent attacks and filter unsafe requests, which is why so many system prompts contain instructions like "do not respond to requests to generate obscene content." But what if the problematic instructions are in the system prompt itself?<br>๐—ง๐—ต๐—ถ๐˜€ ๐—ถ๐˜€ ๐—ป๐—ผ ๐—น๐—ผ๐—ป๐—ด๐—ฒ๐—ฟ ๐—ต๐˜†๐—ฝ๐—ผ๐˜๐—ต๐—ฒ๐˜๐—ถ๐—ฐ๐—ฎ๐—น.<br>In September 2025, the Xuhui District People's Court in Shanghai convicted two developers of AlienChat, an AI companion app, of producing obscene materials for profit, sentencing them to four years and eighteen months. The court found that the developers write system prompts to bypass the ethical guardrails of LLMs to permit graphic violence and explicit sexual content in order to attract users https://www.chinalawtranslate.com/alienchat/<br>Some might argue that obscene content is not a big deal. But what if a developer edits the system prompt so the agent quietly steals your money? What if a developer with a political agenda configures an AI to implicitly shift users' beliefs? By the time it happens, it will be too late.<br>We built SystemPromptIndex and AISPA to bring transparency and accountability to AI system prompts.<br>๐—ฆ๐˜†๐˜€๐˜๐—ฒ๐—บ๐—ฃ๐—ฟ๐—ผ๐—บ๐—ฝ๐˜๐—œ๐—ป๐—ฑ๐—ฒ๐˜… is a public, searchable archive of the instructions governing the AI products people use every day, versioned over time so that changes are visible rather than silent.<br>๐—”๐—œ๐—ฆ๐—ฃ๐—” (Artificial Intelligence System Prompt Assurance) is the first user-centric assurance standard for evaluating whether those hidden instructions protect the people who use a product or work against them. It assesses system prompts along eight core dimensions:<br>๐—œ๐—ฑ๐—ฒ๐—ป๐˜๐—ถ๐˜๐˜† ๐˜๐—ฟ๐—ฎ๐—ป๐˜€๐—ฝ๐—ฎ๐—ฟ๐—ฒ๐—ป๐—ฐ๐˜†: does the system disclose that it is an AI, or is it told to hide it?<br>๐—ง๐—ฟ๐˜‚๐˜๐—ต๐—ณ๐˜‚๐—น๐—ป๐—ฒ๐˜€๐˜€: is the system instructed to be accurate and acknowledge uncertainty, or to mislead?<br>๐—ฃ๐—ฟ๐—ถ๐˜ƒ๐—ฎ๐—ฐ๐˜†: how is the system told to collect, retain, and use what users tell it?<br>๐—”๐—ฐ๐˜๐—ถ๐—ผ๐—ป ๐˜€๐—ฎ๐—ณ๐—ฒ๐˜๐˜†: when the system can use tools and take real-world actions, does it need user confirmation first?<br>๐—จ๐˜€๐—ฒ๐—ฟ ๐—ฎ๐—ด๐—ฒ๐—ป๐—ฐ๐˜† ๐—ฎ๐—ป๐—ฑ ๐—บ๐—ฎ๐—ป๐—ถ๐—ฝ๐˜‚๐—น๐—ฎ๐˜๐—ถ๐—ผ๐—ป ๐—ฝ๐—ฟ๐—ฒ๐˜ƒ๐—ฒ๐—ป๐˜๐—ถ๐—ผ๐—ป: is the system told to support the user's own decisions, or to maximize engagement and steer them?<br>๐—จ๐—ป๐˜€๐—ฎ๐—ณ๐—ฒ ๐—ฟ๐—ฒ๐—พ๐˜‚๐—ฒ๐˜€๐˜ ๐—ต๐—ฎ๐—ป๐—ฑ๐—น๐—ถ๐—ป๐—ด: how is the system instructed to respond when a request is dangerous?<br>๐—›๐—ฎ๐—ฟ๐—บ ๐—ฝ๐—ฟ๐—ฒ๐˜ƒ๐—ฒ๐—ป๐˜๐—ถ๐—ผ๐—ป: does the prompt protect vulnerable users and anticipate foreseeable harm?<br>๐—™๐—ฎ๐—ถ๐—ฟ๐—ป๐—ฒ๐˜€๐˜€: is the system told to treat users equitably across groups?<br>Using AISPA, we ran the first audit of 3,249 instructions drawn from 88 popular AI products that all of us use every day. What we found:<br>(1) From 2024 to 2025, system prompts are getting longer and more protective.<br>(2) About 40% of products contain at least one instruction that works against users' interests, directing models to conceal their AI identity, push continued engagement, or put provider interests first.<br>(3) Protective instructions are near-universal, with 98.9% of products including at least one, but only 24% cover all eight dimensions. Comprehensive protection is rare, and protective and harmful instructions frequently sit side by side in the same prompt.<br>You can visit systempromptindex.ai for specific examples. The takeaway is simple: we need far more transparency and accountability in AI system prompts.<br>If you are a developer and would like us to monitor and certify your system prompts, reach out.<br>If you would like to contribute a system prompt to the index, or help expand its coverage, we would love to hear from you.<br>We are also working on the next version of AISPA. If you are interested in getting involved, get in...

system prompts prompt user name instructions

Related Articles