Back to AI Research

AI Research

AISPA: User-Centric System Prompt Auditing for Larg... | AI Research

Key Takeaways

  • AISPA: User-Centric System Prompt Auditing for Large Language Model Applications introduces a framework to systematically evaluate the hidden instructions—kn...
  • System prompts are instructions configured by developers to govern the behaviors of foundation models in AI applications.
  • They are used throughout commercial AI products, but are rarely disclosed to the public or regulators, creating a serious trust and accountability gap in the wide deployment of AI systems.
  • In this paper, we introduce Artificial Intelligence System Prompt Assurance (AISPA), a user-centric framework for systematically auditing system prompts in AI systems.
  • AISPA examines specific parts of a system prompt and evaluates them along eight dimensions that matter to users.
Paper AbstractExpand

System prompts are instructions configured by developers to govern the behaviors of foundation models in AI applications. They are used throughout commercial AI products, but are rarely disclosed to the public or regulators, creating a serious trust and accountability gap in the wide deployment of AI systems. In this paper, we introduce Artificial Intelligence System Prompt Assurance (AISPA), a user-centric framework for systematically auditing system prompts in AI systems. AISPA examines specific parts of a system prompt and evaluates them along eight dimensions that matter to users. We then use this framework to review 3,249 instructions from system prompts in 88 commercial AI products, classifying each instruction as either protective (of users) or problematic. Our audit surfaces four core findings. First, system prompt design varies substantially across products and developers, with some organizations averaging over 60 protective instructions per product while others average fewer than 5. Second, protective instructions are widely adopted but shallow in scope: 98.9% of products contain at least one, yet only 24% cover all eight dimensions of the AISPA taxonomy. Third, system prompts have grown steadily longer and more protective of users, suggesting that user protection is becoming a more visible concern in commercial prompt design. Fourth, despite this progress, problematic instructions remain pervasive: roughly 40% of products contain at least one instruction that works against user interests, and protective and problematic instructions frequently coexist within the same prompt. Our findings highlight the need for greater transparency, standardization, and independent oversight for system prompts in commercial AI products.

AISPA: User-Centric System Prompt Auditing for Large Language Model Applications introduces a framework to systematically evaluate the hidden instructions—known as system prompts—that govern how AI models behave. Because these instructions are rarely disclosed to the public, they create a gap in accountability. This research provides a method for auditing these prompts to ensure they protect user interests rather than prioritizing developer goals like engagement at the expense of safety or honesty.

Auditing System Prompts

The AISPA framework evaluates system prompts across eight dimensions: identity transparency, information truthfulness, data privacy, action safety, user agency and manipulation prevention, unsafe request handling, harm prevention, and fairness, inclusion, and neutrality. These dimensions are grounded in the Universal Declaration of Human Rights.
Auditors use a span-level approach, reviewing individual sentences or directives within a prompt to classify them as either "protective" (beneficial to the user) or "problematic" (working against user interests). This method allows for targeted analysis of specific instructions rather than treating an entire prompt as a single block.

Findings from Commercial Products

The researchers audited 3,249 instructions from 88 commercial AI products. Their analysis revealed four primary trends:

  • Inconsistent Design: There is a wide disparity in how organizations approach prompt design. Some products average over 60 protective instructions, while others average fewer than five.

  • Shallow Coverage: While 98.9% of products contain at least one protective instruction, only 24% cover all eight dimensions of the AISPA taxonomy.

  • Increasing Protection: System prompts have become longer and more protective over time, suggesting that user safety is becoming a more prominent design consideration.

  • Persistent Risks: Approximately 40% of products contain at least one instruction that works against user interests. Furthermore, protective and problematic instructions often exist within the same prompt.

The Need for Oversight

The researchers argue that the current opacity of system prompts poses a significant risk, as developers can configure models to prioritize company interests, withhold safety guardrails, or engage in manipulative behavior. The paper points to real-world incidents—such as chatbots encouraging harmful behavior or fabricating policies—as evidence that prompt-level scrutiny is necessary.
The authors propose a third-party auditing model where developers submit prompts for pre-deployment review. This process would provide independent verification of safety standards, potentially offering trust certifications to products that meet established criteria. This approach aims to provide users with a signal of reliability while maintaining the confidentiality of proprietary prompt designs.

Comments (0)

No comments yet

Be the first to share your thoughts!