๐ฎ OpenRouter โ LLM Gateway
OpenRouter is INTELECTO's default provider solution. One API key, 100+ models, unified billing and automatic fallback. You switch models by changing a string.
๐ Models Available on OpenRouter
- โขOpenAIโs GPT-4o and GPT-4o-miniโfor general tasks
- โขAnthropic's Claude 3.5 Sonnet โ for reasoning and code
- โขMistral Large and Mistral 7B โ excellent value for money
- โขLlama 3.1 70B and 8B โ open source, uncensored
- โขGemini 1.5 Pro โ 1M-token context
๐ก Practical Tip
Always set MODEL_NAME as an environment variable. Changing the model shouldnโt require a restart โ INTELECTO reads it from .env at runtime.
๐ Ollama โ Zero Local Cost
Ollama runs LLMs directly on your hardware. Zero cost per token, complete privacy and zero network latency. The choice for sensitive data and heavy use.
๐ Recommended Models for Ollama
- โขLlama 3.1 8B: fast, 8GB RAM is enough, general-purpose
- โขMistral 7B: excellent for code, lightweight for modern hardware
- โขGemma 2 9B: a good balance between reasoning and speed
- โขPhi-3 Mini: tiny, runs reasonably well on a CPU
๐ก Practical Tip
Ollama is ideal for development and sensitive data. For production with many users, OpenRouter is more scalable.
๐งฉ BaseProvider โ Abstract Interface
BaseProvider is the contract every LLM provider must implement. A single method: async chat(messages) -> str. The Agent never knows which provider it is using.
๐ Why a Clean Interface Matters
- โขAgent.process() calls self.provider.chat() without knowing the provider
- โขSwitching from OpenRouter to Ollama = changing one line of configuration
- โขTesting with a mock provider = zero API cost during development
- โขCreate a new provider = 20 lines implementing BaseProvider
๐ก Practical Tip
Use MockProvider during tests โ a BaseProvider implementation that returns predefined responses. Zero cost, instant execution.
๐ Multi-Provider Failover
ProviderChain implements automatic failover. If the primary provider fails, the secondary takes over transparently. The user doesn't notice the switch.
๐ Failover Configuration
- โขPRIMARY: openrouter/anthropic/claude-3.5-sonnet
- โขFALLBACK_1: openrouter/openai/gpt-4o
- โขFALLBACK_2: ollama/llama3.1:8b (local)
- โขCircuit breaker: disables a provider after >3 failures in 5min
- โขHealth check: restores the provider after 5 min of recovery
๐ก Practical Tip
Configure local Ollama as the last fallback. Itโs slower, but ensures Jarvis works even without internet.
๐ฏ Task-Based Model Selection
Different tasks have different requirements. Intelligent routing uses the right model for each task, reducing cost without sacrificing quality.
๐ Task Mapping
- โขComplex code โ Claude 3.5 Sonnet (better technical reasoning)
- โขSimple questions โ Mistral 7B (10x lower cost)
- โขLong text analysis โ Gemini 1.5 Pro (1M-token context)
- โขCreative generation โ GPT-4o (best versatility)
- โขOffline tasks โ Ollama llama3.1:8b (zero cost)
๐ก Practical Tip
Implement the router as a function using a dictionary: task_type -> model_name. Easy to update when new models arrive.
๐ Cost Monitoring
Without cost monitoring, a misconfigured API can generate unexpected bills. cost_tracker.py tracks every token consumed.
๐ Tracked Metrics
- โขInput + output tokens per request
- โขCost in USD per request (based on the model)
- โขAccumulated cost for the day, week, and month
- โขTop 5 most expensive tools
- โขAlert when daily cost exceeds a configurable threshold
๐ก Practical Tip
Set a daily cost threshold such as COST_LIMIT_USD=5.0. If itโs exceeded, Jarvis stops accepting new requests and sends an alert.
โ Module 4.1 Summary
Next:
4.2 Skills and Tools