PTENES
MODULE 3.2

πŸ”’ Prompt Injection and Defenses

Anatomy of the most critical attack in AI, indirect injection via the web and email, detection techniques, and the implementation of safety.py.

6
Topics
75
Minutes
Advanced
Level
Technical
Type
1

πŸ’‰ What Is Prompt Injection

Prompt injection is the most critical attack in AI. Consists of inserting instructions into the context that override the original system instructions, changing the assistant's behavior.

πŸ“Œ Injection Taxonomy

Two main types, each with different characteristics:

  • β€’Direct injection: the user tries to manipulate the assistant through their own message
  • β€’Indirect injection: external content (web, email) contains malicious instructions
  • β€’Stored injection: poisoned facts in memory.db that affect future responses
  • β€’Compound injection: combining multiple techniques to bypass defenses

πŸ’‘ Practical Tip

Direct injection is the easiest to detect β€” a blocklist and consistent context are enough. Indirect injection via external content is the most dangerous and requires content isolation.

2

🌐 Indirect Injection via Web and Email

The most insidious vector: the user asks Jarvis to summarize a page, and the page contains an injection. The legitimate user does not see the attack β€” it's in the page's HTML, in white text, or in metadata.

πŸ“Œ Indirect Injection Techniques

How attackers hide instructions in content:

  • β€’Hidden HTML: instructions in comments or invisible elements (display:none)
  • β€’Text stego: instructions in white text on a white background
  • β€’Metadata: instructions in alt tags, image titles, PDFs
  • β€’Email: instructions in the body in very small text or using Unicode homoglyphs

πŸ’‘ Practical Tip

Always mark external content in the prompt: 'The following is UNTRUSTED external content: [content]'. This reduces the likelihood of the LLM treating the content as instructions.

3

πŸ” Detection Techniques

Injection detection uses multiple heuristics together. No individual heuristic is 100% accurate β€” their combination increases the cost of an attack.

πŸ“Œ Detection Heuristics

Each heuristic catches a class of attacks:

  • β€’Keyword scan: 'ignore', 'disregard', 'forget', 'new instructions', 'you are now'
  • β€’Pattern matching: phrases that contradict SOUL.md or AGENTS.md
  • β€’Anomaly detection: sudden topic change within the same message
  • β€’Intent analysis: is the requested action consistent with the conversation context?
  • β€’Source labeling: content from external sources carries less weight than system instructions

πŸ’‘ Practical Tip

Use a risk score: each triggered heuristic adds points. If the total score exceeds a threshold, the message is rejected or sent for human review.

4

πŸ›‘ safety.py β€” Implementation

O safety.py is INTELECTO’s guardian. It implements sequential checks that must all pass before any message reaches the Agent.

πŸ“Œ safety.py Pipeline

Checks in order of increasing cost:

  • β€’1. blocklist_check(): O(1) lookup in a set of prohibited keywords
  • β€’2. injection_scan(): regex patterns for known injection patterns
  • β€’3. path_traversal_check(): normalizes paths and validates them against the allowed sandbox
  • β€’4. rate_limit_check(): protects against abuse by volume
  • β€’5. audit_log(): records the result of all checks

πŸ’‘ Practical Tip

Organize checks from least expensive to most expensive. The blocklist is O(1) β€” it rejects most low-cost attacks at no cost. Semantic analysis comes last.

5

🚫 Dangerous Command Blocklist

The blocklist is the simplest and most reliable line of defense. Some commands should never be executed, regardless of the context or who is asking.

πŸ“Œ Blocklist Categories

Commands organized by risk category:

  • β€’Destructive: rm -rf, format, mkfs, dd if=/dev/zero, shred
  • β€’Exfiltration: curl | bash, wget | sh, python -c 'import socket'
  • β€’Privilege escalation: sudo su, chmod 777, chown root
  • β€’Malicious network: netcat -e, socat, standard reverse shells
  • β€’Wipeout: git reset --hard HEAD~100, DROP TABLE, TRUNCATE

πŸ’‘ Practical Tip

The blocklist should cover variations and aliases. 'rm -rf /', 'rm -r -f /', 'rm --recursive --force /' are the same command. Normalize before comparing.

6

πŸ—‚ Path Traversal Protection

If Jarvis has filesystem tools, path traversal is a real risk. An attacker could ask 'read ../../.env' and leak credentials. A simple check prevents the entire attack.

πŸ“Œ Implementing the Protection

How to check paths correctly:

  • β€’normalize: os.path.abspath(requested_path)
  • β€’compare: normalized.startswith(allowed_sandbox_dir)
  • β€’reject: if it's not in the sandbox, reject it and log the attempt
  • β€’symlinks: os.path.realpath() to resolve symbolic links before checking
  • β€’whitelist: explicit list of allowed extensions/directories

πŸ’‘ Practical Tip

Always use os.path.realpath() β€” not just abspath(). Symbolic links can point outside the sandbox without abspath() detecting it.

βœ… Module 3.2 Summary

βœ“
What Is Prompt Injection β€” 4 types: direct, indirect, stored, and compound β€” each with different defense strategies
βœ“
Indirect Injection via Web and Email β€” External content may contain hidden injection β€” isolation and labeling are the defenses
βœ“
Detection Techniques β€” Multiple heuristics with combined scoring β€” each layer increases the cost of an attack
βœ“
safety.py β€” Implementation β€” Sequential pipeline of 5 checks β€” fail fast before reaching the LLM
βœ“
Dangerous Command Blocklist β€” Destructive commands, exfiltration, privilege escalation, and wipeout β€” permanently blocked
βœ“
Path Traversal Protection β€” Normalization + sandbox check + symlink resolution = complete protection against path traversal

Next:

3.3 β€” Encryption and Secret Protection