β Attack Surfaces in AI
AI assistants have radically different attack surfaces from traditional software. The the primary vector is the prompt β natural language that the LLM interprets as instructions. Any text that enters the context is potentially an attack vector.
π Main Surfaces
Each surface requires a different defense strategy:
- β’Direct prompt: user messages manipulating behavior
- β’External content: web pages, emails, documents with injection
- β’Memory: poisoning the memory database with false facts
- β’Tools: malicious parameters passed to tools
π‘ Practical Tip
Map all the data sources your Jarvis consumes. Each external source is a potential attack surface.
π― Threat Model
The threat model identifies who the adversaries are, what they want, and how they will try to get it. For a personal assistant, the risks are different from those of a corporate AI.
π Adversaries of the Personal Jarvis
Each type of adversary has distinct capabilities and objectives:
- β’Curious users: test the system's limits for fun
- β’Malicious content: pages and emails with passive injection
- β’Automated scripts: API abuse attempts
- β’Sophisticated attacker: goal of exfiltrating data or performing destructive actions
π‘ Practical Tip
Prioritize defenses against the most likely adversaries, not the most sophisticated ones. For personal use, passive malicious content is the real risk.
π AI β Web Security
Developers with a web background tend to apply techniques such as input sanitization and SQL escaping. In AI, these controls don't work because the input is interpreted semantically, not executed literally.
π Fundamental Differences
What changes when moving from web security to AI security:
- β’Web: escape <script>. AI: you can't "escape" natural language
- β’Web: validates data types. AI: the "type" is human intent β undetectable
- β’Web: a firewall blocks IPs. AI: the attacker uses the same channel as the legitimate user
- β’Web: known CVEs. AI: new vectors are discovered weekly
π‘ Practical Tip
Use web security techniques as a foundation, but add layers specific to AI: content labeling, approval gates, and an execution sandbox.
π° Zero Trust for AI
Zero-Trust means: never trust, always verify. In AI, this applies not to the user's identity (authentication), but to each specific action the assistant wants to perform.
π Zero-Trust Principles in AI
Practical application of the Zero-Trust model:
- β’Verify the action, not just the user: authenticated user β authorized action
- β’Least privilege: each tool only has access to what it needs
- β’Explicit approval: high-impact actions always require human confirmation
- β’Continuous verification: doesnβt assume the context hasnβt been compromised
π‘ Practical Tip
Implement Zero-Trust by starting with the most destructive actions (delete, execute, send). They deserve explicit approval gates.
π Real-World Attack Cases
Prompt injection attacks are not theoretical. Documented Cases on Bing Chat, Claude, and others show that any assistant that processes external content is exposed without the right defenses.
π Documented Incidents
Real cases that shaped INTELECTO's defenses:
- β’Bing Chat: web page with injection that made the bot change its identity and reveal the system prompt
- β’Claude via email: indirect injection that attempted to exfiltrate data to an external URL
- β’GPT-4 via document: hidden instructions in white text on a white background in PDFs
- β’AutoGPT: injection via web search result manipulating the next steps
π‘ Practical Tip
All these attacks use the same vector: untrusted content in context. Explicitly marking which content is βexternalβ in the prompt significantly reduces the risk.
πΊ Defense in Depth
No individual defense is 100% effective. Defense in depth stacks multiple independent layers β if one fails, the next catches the attack.
π The 5 INTELECTO Layers
Each layer covers vectors the others donβt:
- β’Layer 1: Input validation and keyword blocklist
- β’Layer 2: Injection pattern detection with heuristics
- β’Layer 3: Approval gate for high-impact actions
- β’Layer 4: Execution sandbox for tools
- β’Layer 5: Audit log for post-hoc detection
π‘ Practical Tip
Implement the layers in order of cost/benefit. The blocklist (Layer 1) is free. The sandbox (Layer 4) has overhead. Start with Layer 1 and add more as you go.
β Module 3.1 Summary
Next:
3.2 β Prompt Injection and Defenses