PTENES
MODULE 4.1

๐Ÿ”— AI Providers

OpenRouter, Ollama, BaseProvider, failover, and task-based model selection.

6
Topics
60
Minutes
Interm.
Level
Technical
Type
1

๐Ÿ”ฎ OpenRouter โ€” LLM Gateway

OpenRouter is INTELECTO's default provider solution. One API key, 100+ models, unified billing and automatic fallback. You switch models by changing a string.

๐Ÿ“Œ Models Available on OpenRouter

  • โ€ขOpenAIโ€™s GPT-4o and GPT-4o-miniโ€”for general tasks
  • โ€ขAnthropic's Claude 3.5 Sonnet โ€” for reasoning and code
  • โ€ขMistral Large and Mistral 7B โ€” excellent value for money
  • โ€ขLlama 3.1 70B and 8B โ€” open source, uncensored
  • โ€ขGemini 1.5 Pro โ€” 1M-token context

๐Ÿ’ก Practical Tip

Always set MODEL_NAME as an environment variable. Changing the model shouldnโ€™t require a restart โ€” INTELECTO reads it from .env at runtime.

2

๐Ÿ  Ollama โ€” Zero Local Cost

Ollama runs LLMs directly on your hardware. Zero cost per token, complete privacy and zero network latency. The choice for sensitive data and heavy use.

๐Ÿ“Œ Recommended Models for Ollama

  • โ€ขLlama 3.1 8B: fast, 8GB RAM is enough, general-purpose
  • โ€ขMistral 7B: excellent for code, lightweight for modern hardware
  • โ€ขGemma 2 9B: a good balance between reasoning and speed
  • โ€ขPhi-3 Mini: tiny, runs reasonably well on a CPU

๐Ÿ’ก Practical Tip

Ollama is ideal for development and sensitive data. For production with many users, OpenRouter is more scalable.

3

๐Ÿงฉ BaseProvider โ€” Abstract Interface

BaseProvider is the contract every LLM provider must implement. A single method: async chat(messages) -> str. The Agent never knows which provider it is using.

๐Ÿ“Œ Why a Clean Interface Matters

  • โ€ขAgent.process() calls self.provider.chat() without knowing the provider
  • โ€ขSwitching from OpenRouter to Ollama = changing one line of configuration
  • โ€ขTesting with a mock provider = zero API cost during development
  • โ€ขCreate a new provider = 20 lines implementing BaseProvider

๐Ÿ’ก Practical Tip

Use MockProvider during tests โ€” a BaseProvider implementation that returns predefined responses. Zero cost, instant execution.

4

๐Ÿ”„ Multi-Provider Failover

ProviderChain implements automatic failover. If the primary provider fails, the secondary takes over transparently. The user doesn't notice the switch.

๐Ÿ“Œ Failover Configuration

  • โ€ขPRIMARY: openrouter/anthropic/claude-3.5-sonnet
  • โ€ขFALLBACK_1: openrouter/openai/gpt-4o
  • โ€ขFALLBACK_2: ollama/llama3.1:8b (local)
  • โ€ขCircuit breaker: disables a provider after >3 failures in 5min
  • โ€ขHealth check: restores the provider after 5 min of recovery

๐Ÿ’ก Practical Tip

Configure local Ollama as the last fallback. Itโ€™s slower, but ensures Jarvis works even without internet.

5

๐ŸŽฏ Task-Based Model Selection

Different tasks have different requirements. Intelligent routing uses the right model for each task, reducing cost without sacrificing quality.

๐Ÿ“Œ Task Mapping

  • โ€ขComplex code โ†’ Claude 3.5 Sonnet (better technical reasoning)
  • โ€ขSimple questions โ†’ Mistral 7B (10x lower cost)
  • โ€ขLong text analysis โ†’ Gemini 1.5 Pro (1M-token context)
  • โ€ขCreative generation โ†’ GPT-4o (best versatility)
  • โ€ขOffline tasks โ†’ Ollama llama3.1:8b (zero cost)

๐Ÿ’ก Practical Tip

Implement the router as a function using a dictionary: task_type -> model_name. Easy to update when new models arrive.

6

๐Ÿ“Š Cost Monitoring

Without cost monitoring, a misconfigured API can generate unexpected bills. cost_tracker.py tracks every token consumed.

๐Ÿ“Œ Tracked Metrics

  • โ€ขInput + output tokens per request
  • โ€ขCost in USD per request (based on the model)
  • โ€ขAccumulated cost for the day, week, and month
  • โ€ขTop 5 most expensive tools
  • โ€ขAlert when daily cost exceeds a configurable threshold

๐Ÿ’ก Practical Tip

Set a daily cost threshold such as COST_LIMIT_USD=5.0. If itโ€™s exceeded, Jarvis stops accepting new requests and sends an alert.

โœ… Module 4.1 Summary

โœ“
OpenRouter โ€” LLM Gateway โ€” 1 key for 100+ models with unified billing and automatic routing
โœ“
Ollama โ€” Zero Cost Locally โ€” Local inference, zero cost, total privacy โ€” ideal for sensitive data
โœ“
BaseProvider โ€” Abstract Interface โ€” One async chat() method, any LLM โ€” perfect plug-and-play
โœ“
Multi-Provider Failover โ€” Provider chain with circuit breaker โ€” zero downtime even when APIs fail
โœ“
Task-Based Model Selection โ€” 60โ€“80% lower costs with intelligent routing by task type
โœ“
Cost Monitoring โ€” Tokens per request, cost per model, and automatic threshold alerts

Next:

4.2 Skills and Tools