How OpenBotX Works
This document explains the complete execution flow of the OpenBotX AI agent, from the moment a user sends a message to the final response delivery. Each step follows the actual execution sequence of the system.
Table of Contents
- Overview
- Server Startup
- Configuration Loading
- Authentication
- Message Entry
- Message Bus
- Message Processing
- Context Building
- The .md Files
- Skills
- Creating New Skills
- Agent Loop
- AI Provider Selection
- Tools
- Subagents
- Memory and Consolidation
- Sessions
- Task Management
- WebSocket
- Heartbeat Service
- Scheduler (Cron)
- Output Routing
- Media Pipeline
- Server Shutdown
- Complete Cycle
1. Overview
OpenBotX is an AI assistant platform that runs as a web server. When you start the system with openbotx start, a FastAPI server boots up and creates all the necessary components.
graph TD
A[User - Browser / Telegram] -->|Sends message| B[Input Channel]
B -->|WebSocket or Telegram| C[Message Bus - Inbound Queue]
C --> D[Agent Loop]
D -->|Queries| E[LLM - AI Model]
D -->|Uses| F[Tools]
F -->|Result| D
E -->|Response| D
D -->|Final response| G[Message Bus - Outbound Queue]
G --> H[Channel Manager]
H -->|WebSocket| A
H -->|Telegram API| A
Everything is wired together in openbotx/server/app.py. The ServerFactory class handles dependency creation, and the lifespan() function orchestrates the startup/shutdown sequence.
2. Server Startup
When you run openbotx start, the CLI (openbotx/cli/commands.py) does the following:
- Starts a Uvicorn (ASGI) server with the configured host and port (default:
0.0.0.0:8000) - After 1.5 seconds, automatically opens the browser (unless
--no-browseris passed):- If
server.public_urlis configured (e.g.,https://my-domain.com), opens that URL - Otherwise, opens
http://localhost:{port}/app/
- If
The FastAPI server uses an async lifespan that manages the entire lifecycle. During startup, components are created in this order:
graph TD
A[load_config] -->|Loads config.yml + .env| B[Create ServerFactory]
B --> C[Create Workspace + Setup Logging]
C --> D[Generate JWT Secret if needed]
D --> E[Create WebSocketManager + EventDispatcher]
E --> F[Create MessageBus]
F --> G[Create SessionManager]
G --> H[Create TaskManager]
H --> I[Create SkillsLoader]
I --> J[Create CronService]
J --> K[Create public/ directory structure]
K --> L[Create Storage backend]
L --> M["ServerFactory.create_orchestrator()"]
M --> M0["Create ProjectContext (paths, tool configs, storage)"]
M0 --> M1["For each agent: Create Provider"]
M1 --> M2["Resolve workspace directory"]
M2 --> M3["Create SubagentManager + AgentLoop (with AgentConfig + ProjectContext)"]
M3 --> M4["Create AgentClassifier (if multi-agent)"]
M4 --> M5["Return Orchestrator"]
M5 --> N[Create ChannelManager]
N --> O[Create HeartbeatService]
O --> P["Start Orchestrator (background task)"]
P --> Q["Start CronService (background task)"]
Q --> R["Start ChannelManager (channels + dispatch)"]
R --> S["Start HeartbeatService (background task)"]
S --> T[Re-queue recovered tasks]
T --> U[Server ready]
The ServerFactory class encapsulates all dependency creation logic. It receives the Config object and provides methods to create providers, storage backends, cron callbacks, and the full orchestrator graph.
Each agent gets its own AgentLoop with:
- Its own
LiteLLMProvider(configured for the agent's model) - Its own workspace directory (resolved via
AgentConfig.resolve_workspace()) - A shared
ProjectContext(project paths, tool configs, storage) - Its own
SubagentManager(receivesAgentConfig+ProjectContext) - Its own
ToolRegistrybuilt viabuild_registry(), which creates aPathResolverper agent viaProjectContext.create_resolver()and filters by the agent'stoolswhitelist if set
After all services are started, the server re-queues any tasks that were interrupted by a previous shutdown. Tasks that were in the DOING state are reset to TODO during TaskManager initialization, and then published back to the MessageBus so the agent can re-execute them (see section 18).
3. Configuration Loading
The load_config() in openbotx/config/loader.py does the following:
- Loads
.env: Callsload_dotenv()to load environment variables from the.envfile at the project root - Reads
config.yml: Parses the YAML file withyaml.safe_load() - Expands environment variables: Substitutes
${VAR}patterns with actual environment variable values, recursively across strings, dicts, and lists:
# config.yml
credentials:
anthropic:
type: simple
key: ${ANTHROPIC_API_KEY} # will be replaced with the actual value
- Validates with Pydantic: The resulting dictionary is validated by the
Configmodel (Pydantic), which applies defaults for missing fields
Important Defaults
| Setting | Default |
|---|---|
| Server | host: 0.0.0.0, port: 8000, public_url: "" |
| Authentication | username: "", password: "" (must be configured) |
| Model | "" (must be configured per agent) |
| Model params | model_params: {} (empty dict — set via provider or agent config) |
| Agent params | agent_params.max_iterations: 40, agent_params.memory_window: 100 |
| Shell timeout | exec.timeout: 60 seconds |
| Workspace restriction | general.restrict_to_workspace: true |
| WebSocket progress | send_progress: true |
| Tool hints | send_tool_hints: false |
| Heartbeat | enabled: true, interval: 1800 (30 minutes) |
Saving
save_config() uses model_dump(exclude_defaults=True) — meaning it only saves values that differ from defaults, keeping config.yml clean and minimal.
4. Authentication
The authentication system (openbotx/server/auth.py) uses JWT (JSON Web Tokens) to protect API routes.
sequenceDiagram
participant U as User (Browser)
participant S as Server
U->>S: POST /api/auth/login {username, password}
S->>S: Validate credentials
S->>U: {token: "eyJ..."} (JWT valid for 24h)
U->>S: GET /api/tasks (Header: Bearer eyJ...)
S->>S: Verify JWT
S->>U: [task list]
U->>S: WebSocket /ws?token=eyJ...
S->>S: Verify JWT via query param
S->>U: Connection accepted
Flow Details
- Login: The user sends
POST /api/auth/loginwith username and password. If valid, receives a JWT token signed with HS256, valid for 24 hours - API routes: All routes under
/api/require theAuthorization: Bearer {token}header. The middleware extracts and validates the token - WebSocket: The token is passed as a query parameter (
/ws?token=...) and validated before accepting the connection. If invalid, the connection is closed with code4001 - Public routes: The following routes do not require authentication:
/api/auth/loginand/api/health(public routes)/app/*(static frontend)/ws(authentication handled internally)
If no server.credential is configured, the server automatically creates a simple credential with a random key for JWT signing. If no web_client.credential is configured, a default login credential with admin/admin credentials is auto-created.
5. Message Entry
A message can enter the system through three paths:
Via WebSocket (Browser)
- The user opens the browser at
http://localhost:8000 - The Vue.js frontend connects via WebSocket at
/ws?token={jwt} - When the user sends a message, the frontend sends a JSON event:
{
"type": "chat:send",
"data": {
"message": "Hello, how are you?",
"session_id": "direct",
"metadata": {}
}
}
- The
websocket_endpoint()(openbotx/server/websocket.py) creates anInboundMessageand publishes it to the bus inbound queue:
msg = InboundMessage(
channel="web",
sender_id="web_user",
chat_id=session_id,
content=content,
metadata=data.get("data", {}).get("metadata", {}),
)
await bus.publish_inbound(msg)
Via REST API
In addition to WebSocket, there is a REST route for sending messages:
POST /api/chat
{
"message": "Hello",
"session_id": "direct"
}
The route (openbotx/server/routes/chat.py) does the following synchronously before returning:
- Creates a task and injects the
task_idinto the message metadata - Persists the user message to the session immediately (creates the session if needed)
- Broadcasts
sessions:updatedso connected clients see the new session right away - Publishes the
InboundMessageto the bus withmessage_saved: truein metadata
Returns {"task_id": "abc123", "session_id": "direct"}. Because the session is saved before the response, a page refresh always shows the session and the user's message — even if the agent hasn't started processing yet.
Via Telegram
- The
TelegramChannelperforms periodic polling on the Telegram API to fetch new messages - Upon receiving a message, it creates an
InboundMessagewithchannel="telegram"and Telegram metadata (user_id,username,first_name,is_group,message_id) - If the message contains media (photos, documents), the file is downloaded to a date-organized subdirectory under
public/media/(e.g.,public/media/2026/02/26/abc123.jpg) and the relative path is added to themedialist - The message is published to the same bus inbound queue
InboundMessage Structure
In all cases, the message becomes an InboundMessage with these fields:
| Field | Description |
|---|---|
channel |
Source: "web", "telegram", "cron", or "heartbeat" |
sender_id |
Who sent it (e.g., "web_user", "telegram_12345") |
chat_id |
Conversation identifier |
content |
Message text |
media |
List of attached file paths |
metadata |
Extra data (task_id, message_id, etc.) |
session_key |
Session key: channel:chat_id (e.g., web:direct) |
session_key_override |
If set, overrides the session key |
6. Message Bus
The Message Bus (openbotx/bus/queue.py) is the heart of the communication system. It works as an internal mailbox with two queues:
graph LR
WEB[Web - WebSocket] -->|InboundMessage| IN[Inbound Queue]
TEL[Telegram] -->|InboundMessage| IN
CRON[Cron Service] -->|InboundMessage| IN
IN --> AGENT[Agent Loop]
AGENT -->|OutboundMessage| OUT[Outbound Queue]
OUT --> CM[Channel Manager]
CM -->|WebSocket| WEB2[Browser]
CM -->|Telegram API| TEL2[Telegram]
class MessageBus:
def __init__(self):
self.inbound: asyncio.Queue[InboundMessage] = asyncio.Queue()
self.outbound: asyncio.Queue[OutboundMessage] = asyncio.Queue()
The bus uses Python async queues (asyncio.Queue). This means the producer doesn't need to wait for the consumer — they are independent processes.
The bus decouples channels from the agent. Telegram knows nothing about the AgentLoop, and the AgentLoop knows nothing about Telegram. They only know the bus.
7. Message Processing
The Orchestrator (openbotx/agent/orchestrator.py) runs in the background, waiting for messages in the inbound queue:
async def run(self):
while not self._stop_event.is_set():
try:
msg = await asyncio.wait_for(self._bus.consume_inbound(), timeout=1.0)
agent_name = await self._route(msg)
agent = self._agents[agent_name]
await agent.process_message(msg, agent_name=agent_name)
except TimeoutError:
continue # check stop_event again
except Exception as e:
logger.error("orchestrator error: %s", e) # single message failure doesn't crash
The 1-second timeout on consume_inbound() ensures the _stop_event is checked regularly, enabling graceful shutdown. Exceptions from individual message processing are caught and logged — a single bad message does not crash the orchestrator.
Message routing: When multiple agents are configured, the AgentClassifier uses an LLM call to determine the best agent. It analyzes the user's message plus the last 20 messages of conversation history and calls a route(agent_name, confidence) tool. The classifier uses hardcoded max_tokens=256 and temperature=0.0 for deterministic, fast classification. Assistant messages in the history are prefixed with [Agent: name] so the classifier can see which agent handled previous turns and maintain conversation continuity (it keeps the same agent unless the topic clearly changes). If the classifier returns an unknown agent name, it falls back to the default (first) agent. For single-agent setups, messages go directly to the default agent with no classification overhead.
When a message arrives, the selected AgentLoop.process_message() does the following in sequence:
graph TD
A[Message arrives from Bus] --> B{Has task_id in metadata?}
B -->|Yes| C[Retrieve existing Task]
B -->|No| D[Create new Task]
C --> E[Mark Task as DOING]
D --> E
E --> E1{Non-web channel?}
E1 -->|Yes| E2["Broadcast chat:user_message"]
E1 -->|No| F
E2 --> F
F{Is it a special command?}
F -->|/new| G["Clear session, broadcast sessions:updated, respond"]
F -->|/help| H[Return command list]
F -->|No| I[Configure tool contexts]
I --> I1{Has audio media?}
I1 -->|Yes| I2["Transcribe audio, broadcast chat:transcription"]
I1 -->|No| J
I2 --> J
J{User message already saved?}
J -->|Yes - web/REST| J1[Skip user save]
J -->|No - channel/cron| J2[Save user message to session]
J2 --> J3["Broadcast sessions:updated"]
J3 --> J1
J1 --> K[Build System Prompt]
K --> K1["Retrieve session history (exclude current turn)"]
K1 --> L[Assemble messages]
L --> M[Execute Agent Loop]
M --> N[Save assistant response to session]
N --> N1["Broadcast sessions:updated"]
N1 --> O[Publish response to Bus]
O --> P[Mark Task as DONE]
P --> Q{Need memory consolidation?}
Q -->|Yes| R[Run consolidation]
Q -->|No| S[End]
R --> S
7.1. Task Creation/Recovery
Each user message becomes a task. The process_message() method accepts an optional agent_name parameter (set by the Orchestrator after classification). If provided, this overrides the agent loop's default name, and the task's agent_name is updated to match. If the message already has a task_id in its metadata (as with the REST API), it retrieves the existing task. Otherwise, it creates a new task with the title being the first 50 characters of the message. The task starts in the DOING state.
7.2. Special Commands
Before sending to the LLM, the agent checks if the message is a command:
/new— Clears the session history and responds "Conversation cleared. How can I help you?"/help— Returns the list of available commands
If it's a command, the task is marked as DONE and the response is sent directly via the bus, without going through the LLM.
7.3. Tool Contexts
Before calling the LLM, the agent configures three tools with the current message context:
- message tool: receives
channelandchat_idto know where to send intermediate messages. Also callsstart_turn()which resets the_sent_in_turnflag — this tracks whether the tool has already sent a message in this turn - spawn tool: receives
channel,chat_id, andparent_task_id(to link subagents to the parent task) - cron tool: receives
channelandchat_id(so scheduled jobs know where they originated)
7.4. Error Handling
If the Agent Loop throws an exception, the error is caught and:
- The task is marked as ERROR with the error message
- An error response is sent to the user:
"I encountered an error: {error}"
8. Context Building
The ContextBuilder (openbotx/agent/context.py) is responsible for assembling the "system prompt" — the text that tells the LLM who it is and what it knows. The prompt is built in this order:
graph TD
A[Base identity] --> A1[Agent name]
A1 --> B[Current date and time]
B --> B1[Public URL]
B1 --> B2[Project Structure - workspace and public paths]
B2 --> C[AGENTS.md]
C --> D[SOUL.md]
D --> E[USER.md]
E --> F[TOOLS.md]
F --> G[Memory - MEMORY.md]
G --> H[Always-on skills]
H --> I[Available skills summary]
I --> I1[Agent instructions]
I1 --> J[Complete System Prompt]
8.1. Base Identity
"You are OpenBotX, a personal AI assistant."
If the agent has a name (multi-agent setup), this is followed by:
"You are acting as the **crypto** agent."
8.2. Current Date and Time
"Current date and time: 2025-01-15 14:30:00."
8.2a. Public URL and Project Structure
If a public URL is configured, it is included:
"Public URL: https://my-domain.com"
Then the directory context, which tells the agent about its workspace and the public directory:
# Project Structure
You have access to your workspace and the public directory. Always use absolute paths.
Workspace: /home/user/myproject/workspace
Internal files: reports, data, drafts.
Public: /home/user/myproject/public
Web-accessible at /public/ URL. Use for anything the user needs to access:
- /home/user/myproject/public/media — images, audio, video (organized by date: YYYY/MM/DD)
- /home/user/myproject/public/documents — PDFs, spreadsheets, exports
This explicitly provides the agent with the absolute paths for its workspace and the public directory, guiding it to use the correct directories for different file types.
8.3. Bootstrap Files (.md)
The system reads 4 files from the project root (not the workspace), in this fixed order:
AGENTS.mdSOUL.mdUSER.mdTOOLS.md
Each is added as a section in the prompt (details in step 9).
8.4. Memory
If a memory/MEMORY.md file exists in the workspace, its content is added as a "Memory" section in the prompt.
8.5. Always-on Skills
Skills marked with always: true are automatically loaded and added to the prompt as individual sections.
8.6. Available Skills Summary
An XML list of all available skills is added, so the LLM knows what it can request. Each skill includes its name, description, file location (so the LLM can read it if needed), and availability status:
<available_skills>
<skill><name>code-review</name><description>Review code for issues</description><location>/path/to/skills/code-review/SKILL.md</location><status>available</status></skill>
<skill always="true"><name>git</name><description>Git operations</description><location>/path/to/skills/git/SKILL.md</location><status>available</status></skill>
</available_skills>
8.7. Agent Instructions
If the agent has an instructions field configured, it is appended as a dedicated section:
# Agent Instructions
You are a market analyst specializing in cryptocurrency. Always include disclaimers.
This is the last section in the system prompt, so it takes highest priority when instructions conflict with earlier context.
8.8. Final Message Assembly
After building the system prompt, build_messages() assembles the message list:
[
{"role": "system", "content": system_prompt},
# ... session history (up to 500 messages) ...
{"role": "user", "content": "current user message"}
]
If the message contains media (images), the agent loop resolves relative paths to base64 data URIs via the storage provider (see section 23). The user content is then assembled as multimodal:
{"role": "user", "content": [
{"type": "text", "text": "describe this image"},
{"type": "image", "url": "data:image/jpeg;base64,/9j/4AAQ..."}
]}
Before sending to the LLM API, the provider converts this to the OpenAI-compatible format:
{"type": "image_url", "image_url": {"url": "data:image/jpeg;base64,/9j/4AAQ..."}}
Using data URIs instead of HTTP URLs ensures compatibility with all cloud LLM providers, which cannot access localhost URLs.
9. The .md Files
These files reside at the project root (where config.yml lives) and define the agent's personality and behavior:
SOUL.md — The Agent's "Soul"
Defines personality, tone of voice, and general behavior rules. Example:
You are a helpful AI assistant.
- Always be polite and professional
- Respond in the same language as the user
- When unsure, ask for clarification
USER.md — User Information
Contains information about the user that helps the agent personalize responses:
- Name: Paulo
- Preferred language: Portuguese
- Timezone: America/Sao_Paulo
AGENTS.md — Agent Descriptions
Describes the available agents and their capabilities:
## Main Agent
General-purpose assistant capable of file operations, web search, and code execution.
TOOLS.md — Tools Documentation
Provides additional instructions on how to use tools:
## File Operations
When reading files, always check if the file exists first.
All these files are optional. If they don't exist, the agent works with the base identity. They are read every time a message is processed, so you can edit them at any moment and the effect is immediate.
10. Skills
Skills are extra capabilities you can give to the agent. Each skill is a Markdown file with specific instructions.
How They Are Organized
Skills reside in two directories:
- Built-in:
openbotx/skills/(shipped with the system) - Workspace:
workspace/skills/(created by the user)
Each skill is a folder with a SKILL.md file inside:
skills/
code-review/
SKILL.md
git/
SKILL.md
summarize/
SKILL.md
SKILL.md Format
Each file has a YAML frontmatter at the top (between ---) and the content below:
---
name: code-review
description: Reviews code for bugs and improvements
always: false
requires:
bins: []
env: []
---
When asked to review code, analyze it for:
1. Bugs and logic errors
2. Security vulnerabilities
3. Performance issues
Frontmatter Fields
| Field | Description |
|---|---|
name |
Skill name (used as identifier) |
description |
Short description (appears in the skills summary) |
always |
If true, the content is automatically included in every message's prompt |
requires.bins |
Programs that must be installed (e.g., git, npm). Checked via shutil.which() |
requires.env |
Environment variables that must exist (e.g., BRAVE_API_KEY). Checked via os.environ.get() |
The bins and env fields accept both a single string and a list.
How They Are Loaded
The SkillsLoader (openbotx/agent/skills.py) does the following:
- Scans skill directories (built-in first, then workspace)
- Workspace skills with the same name override built-in ones
- For each skill, reads the
SKILL.mdand parses the frontmatter using a regex^---\n...\n---\n - Checks if dependencies are satisfied (binaries installed, env vars present)
- Tags each skill with a
sourcefield:"builtin"for skills fromopenbotx/skills/,"project"for skills fromworkspace/skills/. This is exposed via the REST API and displayed in the web UI as a visual tag on each skill card - Skills with
always: trueare included in the prompt automatically — but only if their requirements are satisfied. Analways: trueskill with unsatisfied requirements (missing binary or env var) is silently excluded from the prompt, preventing broken instructions from being injected - Skills with
always: falseappear in the available skills list for the LLM to use when needed - Skills with unsatisfied dependencies appear as
status="unavailable"in the skills summary and are not loaded when requested
11. Creating New Skills
To create a new skill, follow these steps:
Step 1: Create the Folder
mkdir -p workspace/skills/my-skill
Step 2: Create the SKILL.md
---
name: my-skill
description: Does something very useful
always: false
requires:
bins: []
env: []
---
Instructions for the agent on how to use this skill.
When the user asks to use "my-skill", do the following:
1. First step
2. Second step
3. Third step
Step 3: Use It
The skill automatically appears in the available skills list. The agent can use it when the user asks for something related to its description.
Always-on Skills
If you want the skill to always be active (without needing to be requested), set always: true:
---
name: project-rules
description: Rules the agent must always follow
always: true
---
Rules:
- Always respond in Portuguese
- Never delete files without confirming
Skills with Dependencies
If the skill requires an installed program or environment variable:
---
name: docker-helper
description: Helps with Docker
requires:
bins: [docker, docker-compose]
env: [DOCKER_HOST]
---
If dependencies are not satisfied, the skill appears as "unavailable" and is not used.
12. Agent Loop
The Agent Loop (_run_agent_loop() in openbotx/agent/loop.py) is the heart of the processing. It implements the ReAct pattern (Reason + Act) — the LLM thinks, acts (uses tools), observes the result, and repeats.
graph TD
A[Send messages + tool definitions to LLM] --> B{LLM responded with...}
B -->|Plain text| C[Return as final response - END]
B -->|Tool calls| D[For each tool call:]
D --> E[ToolRegistry finds the tool]
E --> F[Validate parameters]
F --> G[Execute tool]
G --> H["Add result to messages (role: tool)"]
H --> I{More tool calls?}
I -->|Yes| D
I -->|No| J{Hit iteration limit?}
J -->|No| A
J -->|Yes| K[Return limit reached message]
Real-time Broadcasting
During processing, the agent broadcasts events via the EventDispatcher so the frontend can show progress. All chat:* events include agent_name for multi-agent identification:
chat:tool_use— when a tool is executed. Includestool(human-readable name likeRead File,Web Search) anddescription(human-readable summary of the tool call with arguments)chat:thinking— when the LLM returns reasoning content (see below)chat:user_message— when a non-web message arrives (e.g., from Telegram or cron). Includeschannel,content, andmedia. This allows the web UI to display messages from other channels in real timechat:transcription— when audio media is transcribed. Includes the transcription text, which is also prepended to the user message contentsessions:updated— after saving each conversation turn (user + assistant messages) to the session. Triggers sidebar reload in the frontend
Extended Thinking (Reasoning)
Some LLM models support extended thinking — a feature where the model exposes its internal chain-of-thought alongside the final response. This is a model-specific capability, not something OpenBotX controls.
Which models support it:
- Anthropic Claude models with extended thinking enabled (e.g.,
claude-sonnet-4-20250514with thinking mode) - Other providers may add support via LiteLLM as they release thinking-capable models
How it flows through the system:
- The agent loop sends a request to the LLM via
provider.chat() - LiteLLM returns the response with an optional
reasoning_contentfield on the message object LiteLLMProviderextracts it:getattr(message, "reasoning_content", None)and includes it in theLLMResponsedataclass- The agent loop checks if
reasoning_contentexists. If it does, it broadcasts achat:thinkingWebSocket event:
if response.reasoning_content and self._dispatcher:
await self._dispatcher.broadcast("chat:thinking", {
"task_id": task_id,
"chat_id": chat_id,
"content": response.reasoning_content,
"agent_name": agent_name,
})
- The frontend receives this event and can display the model's thought process (e.g., in a collapsible "thinking" section above the response)
What happens with the reasoning content after broadcast:
- It is saved into the conversation messages via
ContextBuilder.add_assistant_message()(as thereasoning_contentfield) - However, it is stripped before sending to the LLM on subsequent turns —
_sanitize_messages()only keepsrole,content,tool_calls,tool_call_id, andname. The reasoning is not re-sent to the model - This means reasoning is broadcast once for real-time display, stored in the session for history, but never re-injected into future LLM calls
For models that don't support extended thinking, reasoning_content is simply None and no chat:thinking event is broadcast. The agent loop works identically in both cases.
Iteration Limit
The loop has a maximum iteration limit (default: 40, configurable via agent_params.max_iterations). If the agent hits this limit without reaching a final response, it returns:
"I've reached my processing limit. Please try again or simplify your request."
Output Limits
Each tool manages its own output limits internally. There is no global truncation applied by the agent loop or ContextBuilder — tool results are passed through as-is. This design ensures each tool controls the size and format of its results appropriately:
- read_file — 50 KB per read, with line-based pagination (offset/limit parameters)
- exec — 200 KB, tail-based truncation (keeps the last 200 KB, where errors and results typically appear)
- http_client — 50 KB response body truncation
- web_fetch — 50 KB content extraction limit
Other tools (file operations, web search, memory, etc.) return naturally small results and don't need truncation.
Practical Example
Imagine the user asks "List the files in the docs/ folder":
Iteration 1:
- LLM receives the message and decides to call the
list_dirtool with{"path": "docs/"} - The tool returns:
"architecture.md\nconfiguration.md\napi.md\nskills.md\ntools.md" - Result is added to messages
Iteration 2:
- LLM receives the result and generates text: "The files in the docs/ folder are: architecture.md, configuration.md, ..."
- Since it's plain text (no tool calls), the loop ends
- This is the final response sent to the user
13. AI Provider Selection
OpenBotX supports multiple AI providers via LiteLLM. The selection of which provider to use follows a cascade logic in Config.get_provider() (openbotx/config/schema.py):
graph TD
A["Configured model (e.g., anthropic/claude-sonnet-4-20250514)"] --> B{Has prefix with /?}
B -->|Yes| C["Extract prefix (e.g., 'anthropic')"]
C --> D{Prefix matches a provider whose credential has an API key?}
D -->|Yes| E[Use that provider]
D -->|No| F[Continue to keyword search]
B -->|No| F
F --> G{Any provider has a keyword matching the model name?}
G -->|Yes| H[Use that provider]
G -->|No| I{Any provider with a configured credential that has an API key?}
I -->|Yes| J[Use the first one found]
I -->|No| K[No provider available]
ServerFactory.create_provider() then resolves through provider -> credential -> concrete key to build a LiteLLMProvider.
Model Parameters Resolution
Each provider can define a default model_params dict with arbitrary key-value pairs (e.g., max_tokens, temperature, top_p). At startup, these are merged into each agent's model_params using simple dict merge. The resolution order is:
- Provider
model_params— default parameters for all agents using the provider - Agent
model_params— override provider defaults (agent keys always take precedence)
This merge happens once in ServerFactory.create_orchestrator before agent loops are created: {**provider_params, **agent_params}.
Model Name Resolution
The LiteLLMProvider (openbotx/providers/litellm_provider.py) transforms the model name before sending to LiteLLM:
- If there's a configured gateway (explicit model provider), it applies the gateway prefix and optionally strips the original prefix
- Otherwise, it looks up the provider specification by model name and adds the correct LiteLLM prefix
Prompt Caching
Cloud LLM providers like Anthropic charge tokens to process the system prompt and tool definitions on every API call. Since these rarely change between turns (the system prompt, .md files, memory, and tool schemas are the same throughout a conversation), re-processing them on every call wastes time and money.
Prompt caching tells the provider to keep these blocks in a temporary server-side cache. On subsequent calls within the same session, the provider recognizes the cached content and skips re-processing, reducing both latency and cost.
Currently, two providers support this: Anthropic and OpenRouter. The flag supports_prompt_caching is set per provider in openbotx/providers/registry.py.
When enabled, _apply_cache_control() in openbotx/providers/litellm_provider.py transforms the messages before sending to the API:
System message — the content is converted from a string to a list of content blocks, with cache_control on the last block:
# Before:
{"role": "system", "content": "You are OpenBotX..."}
# After:
{"role": "system", "content": [
{"type": "text", "text": "You are OpenBotX...", "cache_control": {"type": "ephemeral"}}
]}
Tool definitions — cache_control is added to the last tool in the list:
# Before:
[{"type": "function", "function": {"name": "read_file", ...}}, ...]
# After (last tool only):
[..., {"type": "function", "function": {"name": "generate_image", ...}, "cache_control": {"type": "ephemeral"}}]
The ephemeral type means the cache is temporary — the provider decides how long to keep it (typically 5 minutes of inactivity for Anthropic). Non-system messages (user, assistant, tool results) are not cached because they change every turn.
This transformation is applied automatically and transparently — the rest of the codebase works with regular strings and lists, unaware of caching.
LLM Error Recovery
- Malformed JSON: When the LLM returns tool call arguments with invalid JSON, the system uses
json_repair.loads()to attempt automatic correction - Empty content: Messages with empty content are replaced with
"(empty)"to avoid provider 400 errors - Sanitization: Only allowed keys (
role,content,tool_calls,tool_call_id,name) are sent to the LLM — extra fields are stripped
14. Tools
Tools are capabilities that the agent can use to interact with the world. Each tool is a Python class that inherits from Tool (openbotx/tools/base.py) and implements the execute() method.
Tool Registry
The ToolRegistry (openbotx/tools/registry.py) manages all tools:
class ToolRegistry:
def register(self, tool: Tool) # register a tool
def get_definitions(self) -> list[dict] # return schemas for the LLM
async def execute(name, params) -> str # execute a tool
When the AgentLoop calls the LLM, it sends tool definitions (name, description, parameters) so the LLM knows what's available.
Available Tools
| Tool | Description |
|---|---|
read_file |
Read file contents |
write_file |
Write content to a file |
edit_file |
Edit parts of an existing file |
list_dir |
List files and folders in a directory |
exec |
Execute shell commands |
web_search |
Search the web (Brave Search API) |
web_fetch |
Fetch content from a URL |
http_client |
Make HTTP requests (GET, POST, PUT, DELETE, PATCH, HEAD, OPTIONS) with download/upload support |
rss_reader |
Read RSS/Atom feeds and return latest entries |
message |
Send intermediate messages to the user |
spawn |
Create subagents for parallel tasks |
cron |
Schedule recurring or one-time tasks |
memory_save |
Save information to long-term memory |
memory_read |
Read MEMORY.md or HISTORY.md on demand |
memory_search |
Search across memory files for specific information |
browser |
Browser automation (Chrome/Chromium via CDP) |
generate_image |
Generate images (if configured) |
Parameter Validation
Before executing any tool, the ToolRegistry validates parameters using the JSON schema defined by the tool. Validation includes:
- Required field checking (
required) - Type validation (
string,integer,boolean,array,object) - Enum verification (allowed values)
- Minimum and maximum limits
Tool Error Handling
When a tool fails, the result includes the error message followed by a hint for the LLM:
Error: File not found: config.yml
[Analyze the error above and try a different approach.]
This hint is automatically appended by the ToolRegistry to any result starting with "Error" or when an exception occurs. This instructs the LLM to try a different approach instead of repeating the same error.
If the tool is not found, the error message lists all available tools, helping the LLM self-correct.
Workspace Restriction and PathResolver
By default, file tools (read_file, write_file, edit_file, list_dir), the HTTP client (http_client), and the shell (exec) are restricted to the agent's workspace and the shared public directory. The restriction is enforced by the PathResolver class (openbotx/helpers/path.py):
- File tools and HTTP client: All use a
PathResolverinstance for path resolution. The resolver:- Expands
~(home directory) viaPath.expanduser() - Resolves relative paths against the agent's workspace
- When
tools.general.restrict_to_workspaceis enabled, verifies the resolved path falls within one of the allowed directories (workspace or public). If not, raises aPermissionError
- Expands
- Per-agent isolation: Each agent has its own
PathResolverwithallowed_dirs = [agent_workspace, public_dir]. In multi-agent setups, agents cannot access each other's workspaces - Shell (exec): The working directory is set to the agent's workspace. The
PathResolver.is_restrictedproperty is used to determine whether workspace restriction is active for the exec tool
Shell Safety Guards (exec)
The exec tool (openbotx/tools/shell.py) has multiple layers of protection:
Blocked commands — Regex patterns that prevent execution:
rm -rf,del /f,rmdir /s(destructive deletion)format,mkfs,diskpart(disk formatting)dd if=,> /dev/sd(direct device writing)shutdown,reboot,poweroff(machine shutdown)- Fork bombs (
:(){ ... };:)
Path traversal protection — Blocks commands containing ../ or ..\
Absolute path restriction — When restrict_to_workspace=True, absolute paths in the command are verified to ensure they are within the working directory
Timeout — Commands exceeding the timeout (default: 60s) are automatically killed with process.kill()
Truncation — Output large
…(truncated)