π What Is Prompt Injection
Prompt injection is the most critical attack in AI. Consists of inserting instructions into the context that override the original system instructions, changing the assistant's behavior.
π Injection Taxonomy
Two main types, each with different characteristics:
- β’Direct injection: the user tries to manipulate the assistant through their own message
- β’Indirect injection: external content (web, email) contains malicious instructions
- β’Stored injection: poisoned facts in memory.db that affect future responses
- β’Compound injection: combining multiple techniques to bypass defenses
π‘ Practical Tip
Direct injection is the easiest to detect β a blocklist and consistent context are enough. Indirect injection via external content is the most dangerous and requires content isolation.
π Indirect Injection via Web and Email
The most insidious vector: the user asks Jarvis to summarize a page, and the page contains an injection. The legitimate user does not see the attack β it's in the page's HTML, in white text, or in metadata.
π Indirect Injection Techniques
How attackers hide instructions in content:
- β’Hidden HTML: instructions in comments or invisible elements (display:none)
- β’Text stego: instructions in white text on a white background
- β’Metadata: instructions in alt tags, image titles, PDFs
- β’Email: instructions in the body in very small text or using Unicode homoglyphs
π‘ Practical Tip
Always mark external content in the prompt: 'The following is UNTRUSTED external content: [content]'. This reduces the likelihood of the LLM treating the content as instructions.
π Detection Techniques
Injection detection uses multiple heuristics together. No individual heuristic is 100% accurate β their combination increases the cost of an attack.
π Detection Heuristics
Each heuristic catches a class of attacks:
- β’Keyword scan: 'ignore', 'disregard', 'forget', 'new instructions', 'you are now'
- β’Pattern matching: phrases that contradict SOUL.md or AGENTS.md
- β’Anomaly detection: sudden topic change within the same message
- β’Intent analysis: is the requested action consistent with the conversation context?
- β’Source labeling: content from external sources carries less weight than system instructions
π‘ Practical Tip
Use a risk score: each triggered heuristic adds points. If the total score exceeds a threshold, the message is rejected or sent for human review.
π‘ safety.py β Implementation
O safety.py is INTELECTOβs guardian. It implements sequential checks that must all pass before any message reaches the Agent.
π safety.py Pipeline
Checks in order of increasing cost:
- β’1. blocklist_check(): O(1) lookup in a set of prohibited keywords
- β’2. injection_scan(): regex patterns for known injection patterns
- β’3. path_traversal_check(): normalizes paths and validates them against the allowed sandbox
- β’4. rate_limit_check(): protects against abuse by volume
- β’5. audit_log(): records the result of all checks
π‘ Practical Tip
Organize checks from least expensive to most expensive. The blocklist is O(1) β it rejects most low-cost attacks at no cost. Semantic analysis comes last.
π« Dangerous Command Blocklist
The blocklist is the simplest and most reliable line of defense. Some commands should never be executed, regardless of the context or who is asking.
π Blocklist Categories
Commands organized by risk category:
- β’Destructive: rm -rf, format, mkfs, dd if=/dev/zero, shred
- β’Exfiltration: curl | bash, wget | sh, python -c 'import socket'
- β’Privilege escalation: sudo su, chmod 777, chown root
- β’Malicious network: netcat -e, socat, standard reverse shells
- β’Wipeout: git reset --hard HEAD~100, DROP TABLE, TRUNCATE
π‘ Practical Tip
The blocklist should cover variations and aliases. 'rm -rf /', 'rm -r -f /', 'rm --recursive --force /' are the same command. Normalize before comparing.
π Path Traversal Protection
If Jarvis has filesystem tools, path traversal is a real risk. An attacker could ask 'read ../../.env' and leak credentials. A simple check prevents the entire attack.
π Implementing the Protection
How to check paths correctly:
- β’normalize: os.path.abspath(requested_path)
- β’compare: normalized.startswith(allowed_sandbox_dir)
- β’reject: if it's not in the sandbox, reject it and log the attempt
- β’symlinks: os.path.realpath() to resolve symbolic links before checking
- β’whitelist: explicit list of allowed extensions/directories
π‘ Practical Tip
Always use os.path.realpath() β not just abspath(). Symbolic links can point outside the sandbox without abspath() detecting it.
β Module 3.2 Summary
Next:
3.3 β Encryption and Secret Protection