PTENES
MODULE 3.1

πŸ›‘ Security Fundamentals for AI

Attack surfaces in AI assistants, threat model, and why AI security is radically different from web security.

6
Topics
75
Minutes
Advanced
Level
Theory
Type
1

⚠ Attack Surfaces in AI

AI assistants have radically different attack surfaces from traditional software. The the primary vector is the prompt β€” natural language that the LLM interprets as instructions. Any text that enters the context is potentially an attack vector.

πŸ“Œ Main Surfaces

Each surface requires a different defense strategy:

  • β€’Direct prompt: user messages manipulating behavior
  • β€’External content: web pages, emails, documents with injection
  • β€’Memory: poisoning the memory database with false facts
  • β€’Tools: malicious parameters passed to tools

πŸ’‘ Practical Tip

Map all the data sources your Jarvis consumes. Each external source is a potential attack surface.

2

🎯 Threat Model

The threat model identifies who the adversaries are, what they want, and how they will try to get it. For a personal assistant, the risks are different from those of a corporate AI.

πŸ“Œ Adversaries of the Personal Jarvis

Each type of adversary has distinct capabilities and objectives:

  • β€’Curious users: test the system's limits for fun
  • β€’Malicious content: pages and emails with passive injection
  • β€’Automated scripts: API abuse attempts
  • β€’Sophisticated attacker: goal of exfiltrating data or performing destructive actions

πŸ’‘ Practical Tip

Prioritize defenses against the most likely adversaries, not the most sophisticated ones. For personal use, passive malicious content is the real risk.

3

πŸ” AI β‰  Web Security

Developers with a web background tend to apply techniques such as input sanitization and SQL escaping. In AI, these controls don't work because the input is interpreted semantically, not executed literally.

πŸ“Œ Fundamental Differences

What changes when moving from web security to AI security:

  • β€’Web: escape <script>. AI: you can't "escape" natural language
  • β€’Web: validates data types. AI: the "type" is human intent β€” undetectable
  • β€’Web: a firewall blocks IPs. AI: the attacker uses the same channel as the legitimate user
  • β€’Web: known CVEs. AI: new vectors are discovered weekly

πŸ’‘ Practical Tip

Use web security techniques as a foundation, but add layers specific to AI: content labeling, approval gates, and an execution sandbox.

4

🏰 Zero Trust for AI

Zero-Trust means: never trust, always verify. In AI, this applies not to the user's identity (authentication), but to each specific action the assistant wants to perform.

πŸ“Œ Zero-Trust Principles in AI

Practical application of the Zero-Trust model:

  • β€’Verify the action, not just the user: authenticated user β‰  authorized action
  • β€’Least privilege: each tool only has access to what it needs
  • β€’Explicit approval: high-impact actions always require human confirmation
  • β€’Continuous verification: doesn’t assume the context hasn’t been compromised

πŸ’‘ Practical Tip

Implement Zero-Trust by starting with the most destructive actions (delete, execute, send). They deserve explicit approval gates.

5

πŸ“Š Real-World Attack Cases

Prompt injection attacks are not theoretical. Documented Cases on Bing Chat, Claude, and others show that any assistant that processes external content is exposed without the right defenses.

πŸ“Œ Documented Incidents

Real cases that shaped INTELECTO's defenses:

  • β€’Bing Chat: web page with injection that made the bot change its identity and reveal the system prompt
  • β€’Claude via email: indirect injection that attempted to exfiltrate data to an external URL
  • β€’GPT-4 via document: hidden instructions in white text on a white background in PDFs
  • β€’AutoGPT: injection via web search result manipulating the next steps

πŸ’‘ Practical Tip

All these attacks use the same vector: untrusted content in context. Explicitly marking which content is β€œexternal” in the prompt significantly reduces the risk.

6

πŸ—Ί Defense in Depth

No individual defense is 100% effective. Defense in depth stacks multiple independent layers β€” if one fails, the next catches the attack.

πŸ“Œ The 5 INTELECTO Layers

Each layer covers vectors the others don’t:

  • β€’Layer 1: Input validation and keyword blocklist
  • β€’Layer 2: Injection pattern detection with heuristics
  • β€’Layer 3: Approval gate for high-impact actions
  • β€’Layer 4: Execution sandbox for tools
  • β€’Layer 5: Audit log for post-hoc detection

πŸ’‘ Practical Tip

Implement the layers in order of cost/benefit. The blocklist (Layer 1) is free. The sandbox (Layer 4) has overhead. Start with Layer 1 and add more as you go.

βœ… Module 3.1 Summary

βœ“
AI Attack Surfaces β€” Prompt, external content, memory, and tools are the 4 main vectors
βœ“
Threat Model β€” Different adversaries require different defense priorities
βœ“
AI β‰  Web Security β€” Natural-language attack vectors require AI-specific defenses
βœ“
Zero-Trust for AI β€” Verification by action, not identity β€” the central principle of Zero Trust for AI
βœ“
Real-World Attack Cases β€” Injection via the web, email, and documents are real, documented attack vectors
βœ“
Defense in Depth β€” 5 independent layers β€” a malicious request must pass through all of them to cause harm

Next:

3.2 β€” Prompt Injection and Defenses