# How OpenBotX Works

> This document explains the complete execution flow of the OpenBotX AI agent, from the moment a user sends a message to the final response delivery. Each step follows the actual execution sequence of the system.

- Skill: `tools-only/how-openbotx-works` (Agent Skill, multi-file: 3 files)
- Install (CLI): `npx skillmds@latest add tools-only/how-openbotx-works`
- Raw SKILL.md: https://api.skillmd.com/api/skills/tools-only/how-openbotx-works/raw
- Safety review: pending (external: skill-scanner PASS, skillspector WARNING)
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: Coding & Dev Tools
- Author: tools-only (https://skillmd.com/u/tools-only)
- Updated: 2026-09-29
- Page: https://skillmd.com/skills/tools-only/how-openbotx-works

---

# How OpenBotX Works

This document explains the complete execution flow of the OpenBotX AI agent, from the moment a user sends a message to the final response delivery. Each step follows the actual execution sequence of the system.

---

## Table of Contents

1. [Overview](#1-overview)
2. [Server Startup](#2-server-startup)
3. [Configuration Loading](#3-configuration-loading)
4. [Authentication](#4-authentication)
5. [Message Entry](#5-message-entry)
6. [Message Bus](#6-message-bus)
7. [Message Processing](#7-message-processing)
8. [Context Building](#8-context-building)
9. [The .md Files](#9-the-md-files)
10. [Skills](#10-skills)
11. [Creating New Skills](#11-creating-new-skills)
12. [Agent Loop](#12-agent-loop)
13. [AI Provider Selection](#13-ai-provider-selection)
14. [Tools](#14-tools)
15. [Subagents](#15-subagents)
16. [Memory and Consolidation](#16-memory-and-consolidation)
17. [Sessions](#17-sessions)
18. [Task Management](#18-task-management)
19. [WebSocket](#19-websocket)
20. [Heartbeat Service](#20-heartbeat-service)
21. [Scheduler (Cron)](#21-scheduler-cron)
22. [Output Routing](#22-output-routing)
23. [Media Pipeline](#23-media-pipeline)
24. [Server Shutdown](#24-server-shutdown)
25. [Complete Cycle](#25-complete-cycle)

---

## 1. Overview

OpenBotX is an AI assistant platform that runs as a web server. When you start the system with `openbotx start`, a FastAPI server boots up and creates all the necessary components.

```mermaid
graph TD
    A[User - Browser / Telegram] -->|Sends message| B[Input Channel]
    B -->|WebSocket or Telegram| C[Message Bus - Inbound Queue]
    C --> D[Agent Loop]
    D -->|Queries| E[LLM - AI Model]
    D -->|Uses| F[Tools]
    F -->|Result| D
    E -->|Response| D
    D -->|Final response| G[Message Bus - Outbound Queue]
    G --> H[Channel Manager]
    H -->|WebSocket| A
    H -->|Telegram API| A
```

Everything is wired together in `openbotx/server/app.py`. The `ServerFactory` class handles dependency creation, and the `lifespan()` function orchestrates the startup/shutdown sequence.

---

## 2. Server Startup

When you run `openbotx start`, the CLI (`openbotx/cli/commands.py`) does the following:

1. Starts a Uvicorn (ASGI) server with the configured host and port (default: `0.0.0.0:8000`)
2. After 1.5 seconds, automatically opens the browser (unless `--no-browser` is passed):
   - If `server.public_url` is configured (e.g., `https://my-domain.com`), opens that URL
   - Otherwise, opens `http://localhost:{port}/app/`

The FastAPI server uses an async **lifespan** that manages the entire lifecycle. During startup, components are created in this order:

```mermaid
graph TD
    A[load_config] -->|Loads config.yml + .env| B[Create ServerFactory]
    B --> C[Create Workspace + Setup Logging]
    C --> D[Generate JWT Secret if needed]
    D --> E[Create WebSocketManager + EventDispatcher]
    E --> F[Create MessageBus]
    F --> G[Create SessionManager]
    G --> H[Create TaskManager]
    H --> I[Create SkillsLoader]
    I --> J[Create CronService]
    J --> K[Create public/ directory structure]
    K --> L[Create Storage backend]
    L --> M["ServerFactory.create_orchestrator()"]
    M --> M0["Create ProjectContext (paths, tool configs, storage)"]
    M0 --> M1["For each agent: Create Provider"]
    M1 --> M2["Resolve workspace directory"]
    M2 --> M3["Create SubagentManager + AgentLoop (with AgentConfig + ProjectContext)"]
    M3 --> M4["Create AgentClassifier (if multi-agent)"]
    M4 --> M5["Return Orchestrator"]
    M5 --> N[Create ChannelManager]
    N --> O[Create HeartbeatService]
    O --> P["Start Orchestrator (background task)"]
    P --> Q["Start CronService (background task)"]
    Q --> R["Start ChannelManager (channels + dispatch)"]
    R --> S["Start HeartbeatService (background task)"]
    S --> T[Re-queue recovered tasks]
    T --> U[Server ready]
```

The `ServerFactory` class encapsulates all dependency creation logic. It receives the `Config` object and provides methods to create providers, storage backends, cron callbacks, and the full orchestrator graph.

Each agent gets its own `AgentLoop` with:
- Its own `LiteLLMProvider` (configured for the agent's model)
- Its own workspace directory (resolved via `AgentConfig.resolve_workspace()`)
- A shared `ProjectContext` (project paths, tool configs, storage)
- Its own `SubagentManager` (receives `AgentConfig` + `ProjectContext`)
- Its own `ToolRegistry` built via `build_registry()`, which creates a `PathResolver` per agent via `ProjectContext.create_resolver()` and filters by the agent's `tools` whitelist if set

After all services are started, the server re-queues any tasks that were interrupted by a previous shutdown. Tasks that were in the DOING state are reset to TODO during `TaskManager` initialization, and then published back to the MessageBus so the agent can re-execute them (see section 18).

---

## 3. Configuration Loading

The `load_config()` in `openbotx/config/loader.py` does the following:

1. **Loads `.env`**: Calls `load_dotenv()` to load environment variables from the `.env` file at the project root
2. **Reads `config.yml`**: Parses the YAML file with `yaml.safe_load()`
3. **Expands environment variables**: Substitutes `${VAR}` patterns with actual environment variable values, recursively across strings, dicts, and lists:

```yaml
# config.yml
credentials:
  anthropic:
    type: simple
    key: ${ANTHROPIC_API_KEY}  # will be replaced with the actual value
```

4. **Validates with Pydantic**: The resulting dictionary is validated by the `Config` model (Pydantic), which applies defaults for missing fields

### Important Defaults

| Setting | Default |
|---|---|
| Server | `host: 0.0.0.0`, `port: 8000`, `public_url: ""` |
| Authentication | `username: ""`, `password: ""` (must be configured) |
| Model | `""` (must be configured per agent) |
| Model params | `model_params: {}` (empty dict — set via provider or agent config) |
| Agent params | `agent_params.max_iterations: 40`, `agent_params.memory_window: 100` |
| Shell timeout | `exec.timeout: 60` seconds |
| Workspace restriction | `general.restrict_to_workspace: true` |
| WebSocket progress | `send_progress: true` |
| Tool hints | `send_tool_hints: false` |
| Heartbeat | `enabled: true`, `interval: 1800` (30 minutes) |

### Saving

`save_config()` uses `model_dump(exclude_defaults=True)` — meaning it only saves values that differ from defaults, keeping `config.yml` clean and minimal.

---

## 4. Authentication

The authentication system (`openbotx/server/auth.py`) uses **JWT (JSON Web Tokens)** to protect API routes.

```mermaid
sequenceDiagram
    participant U as User (Browser)
    participant S as Server

    U->>S: POST /api/auth/login {username, password}
    S->>S: Validate credentials
    S->>U: {token: "eyJ..."} (JWT valid for 24h)

    U->>S: GET /api/tasks (Header: Bearer eyJ...)
    S->>S: Verify JWT
    S->>U: [task list]

    U->>S: WebSocket /ws?token=eyJ...
    S->>S: Verify JWT via query param
    S->>U: Connection accepted
```

### Flow Details

1. **Login**: The user sends `POST /api/auth/login` with username and password. If valid, receives a JWT token signed with HS256, valid for **24 hours**
2. **API routes**: All routes under `/api/` require the `Authorization: Bearer {token}` header. The middleware extracts and validates the token
3. **WebSocket**: The token is passed as a query parameter (`/ws?token=...`) and validated before accepting the connection. If invalid, the connection is closed with code `4001`
4. **Public routes**: The following routes do not require authentication:
   - `/api/auth/login` and `/api/health` (public routes)
   - `/app/*` (static frontend)
   - `/ws` (authentication handled internally)

If no `server.credential` is configured, the server automatically creates a `simple` credential with a random key for JWT signing. If no `web_client.credential` is configured, a default `login` credential with admin/admin credentials is auto-created.

---

## 5. Message Entry

A message can enter the system through three paths:

### Via WebSocket (Browser)

1. The user opens the browser at `http://localhost:8000`
2. The Vue.js frontend connects via **WebSocket** at `/ws?token={jwt}`
3. When the user sends a message, the frontend sends a JSON event:

```json
{
  "type": "chat:send",
  "data": {
    "message": "Hello, how are you?",
    "session_id": "direct",
    "metadata": {}
  }
}
```

4. The `websocket_endpoint()` (`openbotx/server/websocket.py`) creates an `InboundMessage` and publishes it to the bus inbound queue:

```python
msg = InboundMessage(
    channel="web",
    sender_id="web_user",
    chat_id=session_id,
    content=content,
    metadata=data.get("data", {}).get("metadata", {}),
)
await bus.publish_inbound(msg)
```

### Via REST API

In addition to WebSocket, there is a REST route for sending messages:

```
POST /api/chat
{
  "message": "Hello",
  "session_id": "direct"
}
```

The route (`openbotx/server/routes/chat.py`) does the following synchronously before returning:

1. Creates a task and injects the `task_id` into the message metadata
2. Persists the user message to the session immediately (creates the session if needed)
3. Broadcasts `sessions:updated` so connected clients see the new session right away
4. Publishes the `InboundMessage` to the bus with `message_saved: true` in metadata

Returns `{"task_id": "abc123", "session_id": "direct"}`. Because the session is saved before the response, a page refresh always shows the session and the user's message — even if the agent hasn't started processing yet.

### Via Telegram

1. The `TelegramChannel` performs periodic **polling** on the Telegram API to fetch new messages
2. Upon receiving a message, it creates an `InboundMessage` with `channel="telegram"` and Telegram metadata (`user_id`, `username`, `first_name`, `is_group`, `message_id`)
3. If the message contains media (photos, documents), the file is downloaded to a date-organized subdirectory under `public/media/` (e.g., `public/media/2026/02/26/abc123.jpg`) and the **relative path** is added to the `media` list
4. The message is published to the same bus inbound queue

### InboundMessage Structure

In all cases, the message becomes an `InboundMessage` with these fields:

| Field | Description |
|---|---|
| `channel` | Source: `"web"`, `"telegram"`, `"cron"`, or `"heartbeat"` |
| `sender_id` | Who sent it (e.g., `"web_user"`, `"telegram_12345"`) |
| `chat_id` | Conversation identifier |
| `content` | Message text |
| `media` | List of attached file paths |
| `metadata` | Extra data (task_id, message_id, etc.) |
| `session_key` | Session key: `channel:chat_id` (e.g., `web:direct`) |
| `session_key_override` | If set, overrides the session key |

---

## 6. Message Bus

The Message Bus (`openbotx/bus/queue.py`) is the heart of the communication system. It works as an internal mailbox with two queues:

```mermaid
graph LR
    WEB[Web - WebSocket] -->|InboundMessage| IN[Inbound Queue]
    TEL[Telegram] -->|InboundMessage| IN
    CRON[Cron Service] -->|InboundMessage| IN
    IN --> AGENT[Agent Loop]
    AGENT -->|OutboundMessage| OUT[Outbound Queue]
    OUT --> CM[Channel Manager]
    CM -->|WebSocket| WEB2[Browser]
    CM -->|Telegram API| TEL2[Telegram]
```

```python
class MessageBus:
    def __init__(self):
        self.inbound: asyncio.Queue[InboundMessage] = asyncio.Queue()
        self.outbound: asyncio.Queue[OutboundMessage] = asyncio.Queue()
```

The bus uses Python async queues (`asyncio.Queue`). This means the producer doesn't need to wait for the consumer — they are independent processes.

The bus **decouples** channels from the agent. Telegram knows nothing about the AgentLoop, and the AgentLoop knows nothing about Telegram. They only know the bus.

---

## 7. Message Processing

The `Orchestrator` (`openbotx/agent/orchestrator.py`) runs in the background, waiting for messages in the inbound queue:

```python
async def run(self):
    while not self._stop_event.is_set():
        try:
            msg = await asyncio.wait_for(self._bus.consume_inbound(), timeout=1.0)
            agent_name = await self._route(msg)
            agent = self._agents[agent_name]
            await agent.process_message(msg, agent_name=agent_name)
        except TimeoutError:
            continue  # check stop_event again
        except Exception as e:
            logger.error("orchestrator error: %s", e)  # single message failure doesn't crash
```

The 1-second timeout on `consume_inbound()` ensures the `_stop_event` is checked regularly, enabling graceful shutdown. Exceptions from individual message processing are caught and logged — a single bad message does not crash the orchestrator.

**Message routing:** When multiple agents are configured, the `AgentClassifier` uses an LLM call to determine the best agent. It analyzes the user's message plus the last 20 messages of conversation history and calls a `route(agent_name, confidence)` tool. The classifier uses hardcoded `max_tokens=256` and `temperature=0.0` for deterministic, fast classification. Assistant messages in the history are prefixed with `[Agent: name]` so the classifier can see which agent handled previous turns and maintain conversation continuity (it keeps the same agent unless the topic clearly changes). If the classifier returns an unknown agent name, it falls back to the default (first) agent. For single-agent setups, messages go directly to the default agent with no classification overhead.

When a message arrives, the selected `AgentLoop.process_message()` does the following in sequence:

```mermaid
graph TD
    A[Message arrives from Bus] --> B{Has task_id in metadata?}
    B -->|Yes| C[Retrieve existing Task]
    B -->|No| D[Create new Task]
    C --> E[Mark Task as DOING]
    D --> E
    E --> E1{Non-web channel?}
    E1 -->|Yes| E2["Broadcast chat:user_message"]
    E1 -->|No| F
    E2 --> F
    F{Is it a special command?}
    F -->|/new| G["Clear session, broadcast sessions:updated, respond"]
    F -->|/help| H[Return command list]
    F -->|No| I[Configure tool contexts]
    I --> I1{Has audio media?}
    I1 -->|Yes| I2["Transcribe audio, broadcast chat:transcription"]
    I1 -->|No| J
    I2 --> J
    J{User message already saved?}
    J -->|Yes - web/REST| J1[Skip user save]
    J -->|No - channel/cron| J2[Save user message to session]
    J2 --> J3["Broadcast sessions:updated"]
    J3 --> J1
    J1 --> K[Build System Prompt]
    K --> K1["Retrieve session history (exclude current turn)"]
    K1 --> L[Assemble messages]
    L --> M[Execute Agent Loop]
    M --> N[Save assistant response to session]
    N --> N1["Broadcast sessions:updated"]
    N1 --> O[Publish response to Bus]
    O --> P[Mark Task as DONE]
    P --> Q{Need memory consolidation?}
    Q -->|Yes| R[Run consolidation]
    Q -->|No| S[End]
    R --> S
```

### 7.1. Task Creation/Recovery

Each user message becomes a **task**. The `process_message()` method accepts an optional `agent_name` parameter (set by the Orchestrator after classification). If provided, this overrides the agent loop's default name, and the task's `agent_name` is updated to match. If the message already has a `task_id` in its metadata (as with the REST API), it retrieves the existing task. Otherwise, it creates a new task with the title being the first 50 characters of the message. The task starts in the **DOING** state.

### 7.2. Special Commands

Before sending to the LLM, the agent checks if the message is a command:

- **`/new`** — Clears the session history and responds "Conversation cleared. How can I help you?"
- **`/help`** — Returns the list of available commands

If it's a command, the task is marked as **DONE** and the response is sent directly via the bus, without going through the LLM.

### 7.3. Tool Contexts

Before calling the LLM, the agent configures three tools with the current message context:

- **message tool**: receives `channel` and `chat_id` to know where to send intermediate messages. Also calls `start_turn()` which resets the `_sent_in_turn` flag — this tracks whether the tool has already sent a message in this turn
- **spawn tool**: receives `channel`, `chat_id`, and `parent_task_id` (to link subagents to the parent task)
- **cron tool**: receives `channel` and `chat_id` (so scheduled jobs know where they originated)

### 7.4. Error Handling

If the Agent Loop throws an exception, the error is caught and:
1. The task is marked as **ERROR** with the error message
2. An error response is sent to the user: `"I encountered an error: {error}"`

---

## 8. Context Building

The `ContextBuilder` (`openbotx/agent/context.py`) is responsible for assembling the "system prompt" — the text that tells the LLM who it is and what it knows. The prompt is built in this order:

```mermaid
graph TD
    A[Base identity] --> A1[Agent name]
    A1 --> B[Current date and time]
    B --> B1[Public URL]
    B1 --> B2[Project Structure - workspace and public paths]
    B2 --> C[AGENTS.md]
    C --> D[SOUL.md]
    D --> E[USER.md]
    E --> F[TOOLS.md]
    F --> G[Memory - MEMORY.md]
    G --> H[Always-on skills]
    H --> I[Available skills summary]
    I --> I1[Agent instructions]
    I1 --> J[Complete System Prompt]
```

### 8.1. Base Identity

```
"You are OpenBotX, a personal AI assistant."
```

If the agent has a name (multi-agent setup), this is followed by:

```
"You are acting as the **crypto** agent."
```

### 8.2. Current Date and Time

```
"Current date and time: 2025-01-15 14:30:00."
```

### 8.2a. Public URL and Project Structure

If a public URL is configured, it is included:

```
"Public URL: https://my-domain.com"
```

Then the directory context, which tells the agent about its workspace and the public directory:

```
# Project Structure
You have access to your workspace and the public directory. Always use absolute paths.

Workspace: /home/user/myproject/workspace
  Internal files: reports, data, drafts.

Public: /home/user/myproject/public
  Web-accessible at /public/ URL. Use for anything the user needs to access:
  - /home/user/myproject/public/media — images, audio, video (organized by date: YYYY/MM/DD)
  - /home/user/myproject/public/documents — PDFs, spreadsheets, exports
```

This explicitly provides the agent with the absolute paths for its workspace and the public directory, guiding it to use the correct directories for different file types.

### 8.3. Bootstrap Files (.md)

The system reads 4 files from the **project root** (not the workspace), in this fixed order:
1. `AGENTS.md`
2. `SOUL.md`
3. `USER.md`
4. `TOOLS.md`

Each is added as a section in the prompt (details in step 9).

### 8.4. Memory

If a `memory/MEMORY.md` file exists in the workspace, its content is added as a "Memory" section in the prompt.

### 8.5. Always-on Skills

Skills marked with `always: true` are automatically loaded and added to the prompt as individual sections.

### 8.6. Available Skills Summary

An XML list of all available skills is added, so the LLM knows what it can request. Each skill includes its name, description, file location (so the LLM can read it if needed), and availability status:

```xml
<available_skills>
  <skill><name>code-review</name><description>Review code for issues</description><location>/path/to/skills/code-review/SKILL.md</location><status>available</status></skill>
  <skill always="true"><name>git</name><description>Git operations</description><location>/path/to/skills/git/SKILL.md</location><status>available</status></skill>
</available_skills>
```

### 8.7. Agent Instructions

If the agent has an `instructions` field configured, it is appended as a dedicated section:

```
# Agent Instructions
You are a market analyst specializing in cryptocurrency. Always include disclaimers.
```

This is the last section in the system prompt, so it takes highest priority when instructions conflict with earlier context.

### 8.8. Final Message Assembly

After building the system prompt, `build_messages()` assembles the message list:

```python
[
    {"role": "system", "content": system_prompt},
    # ... session history (up to 500 messages) ...
    {"role": "user", "content": "current user message"}
]
```

If the message contains media (images), the agent loop resolves relative paths to base64 data URIs via the storage provider (see section 23). The user content is then assembled as multimodal:

```python
{"role": "user", "content": [
    {"type": "text", "text": "describe this image"},
    {"type": "image", "url": "data:image/jpeg;base64,/9j/4AAQ..."}
]}
```

Before sending to the LLM API, the provider converts this to the OpenAI-compatible format:

```python
{"type": "image_url", "image_url": {"url": "data:image/jpeg;base64,/9j/4AAQ..."}}
```

Using data URIs instead of HTTP URLs ensures compatibility with all cloud LLM providers, which cannot access `localhost` URLs.

---

## 9. The .md Files

These files reside at the project root (where `config.yml` lives) and define the agent's personality and behavior:

### SOUL.md — The Agent's "Soul"

Defines personality, tone of voice, and general behavior rules. Example:

```markdown
You are a helpful AI assistant.
- Always be polite and professional
- Respond in the same language as the user
- When unsure, ask for clarification
```

### USER.md — User Information

Contains information about the user that helps the agent personalize responses:

```markdown
- Name: Paulo
- Preferred language: Portuguese
- Timezone: America/Sao_Paulo
```

### AGENTS.md — Agent Descriptions

Describes the available agents and their capabilities:

```markdown
## Main Agent
General-purpose assistant capable of file operations, web search, and code execution.
```

### TOOLS.md — Tools Documentation

Provides additional instructions on how to use tools:

```markdown
## File Operations
When reading files, always check if the file exists first.
```

All these files are **optional**. If they don't exist, the agent works with the base identity. They are read **every time** a message is processed, so you can edit them at any moment and the effect is immediate.

---

## 10. Skills

Skills are extra capabilities you can give to the agent. Each skill is a Markdown file with specific instructions.

### How They Are Organized

Skills reside in two directories:
- **Built-in**: `openbotx/skills/` (shipped with the system)
- **Workspace**: `workspace/skills/` (created by the user)

Each skill is a folder with a `SKILL.md` file inside:

```
skills/
  code-review/
    SKILL.md
  git/
    SKILL.md
  summarize/
    SKILL.md
```

### SKILL.md Format

Each file has a **YAML frontmatter** at the top (between `---`) and the content below:

```markdown
---
name: code-review
description: Reviews code for bugs and improvements
always: false
requires:
  bins: []
  env: []
---

When asked to review code, analyze it for:
1. Bugs and logic errors
2. Security vulnerabilities
3. Performance issues
```

### Frontmatter Fields

| Field | Description |
|---|---|
| `name` | Skill name (used as identifier) |
| `description` | Short description (appears in the skills summary) |
| `always` | If `true`, the content is automatically included in every message's prompt |
| `requires.bins` | Programs that must be installed (e.g., `git`, `npm`). Checked via `shutil.which()` |
| `requires.env` | Environment variables that must exist (e.g., `BRAVE_API_KEY`). Checked via `os.environ.get()` |

The `bins` and `env` fields accept both a single string and a list.

### How They Are Loaded

The `SkillsLoader` (`openbotx/agent/skills.py`) does the following:

1. Scans skill directories (**built-in first**, then workspace)
2. Workspace skills with the same name **override** built-in ones
3. For each skill, reads the `SKILL.md` and parses the frontmatter using a regex `^---\n...\n---\n`
4. Checks if dependencies are satisfied (binaries installed, env vars present)
5. Tags each skill with a `source` field: `"builtin"` for skills from `openbotx/skills/`, `"project"` for skills from `workspace/skills/`. This is exposed via the REST API and displayed in the web UI as a visual tag on each skill card
6. Skills with `always: true` are included in the prompt automatically — **but only if their requirements are satisfied**. An `always: true` skill with unsatisfied requirements (missing binary or env var) is silently excluded from the prompt, preventing broken instructions from being injected
7. Skills with `always: false` appear in the available skills list for the LLM to use when needed
8. Skills with unsatisfied dependencies appear as `status="unavailable"` in the skills summary and are not loaded when requested

---

## 11. Creating New Skills

To create a new skill, follow these steps:

### Step 1: Create the Folder

```bash
mkdir -p workspace/skills/my-skill
```

### Step 2: Create the SKILL.md

```markdown
---
name: my-skill
description: Does something very useful
always: false
requires:
  bins: []
  env: []
---

Instructions for the agent on how to use this skill.

When the user asks to use "my-skill", do the following:
1. First step
2. Second step
3. Third step
```

### Step 3: Use It

The skill automatically appears in the available skills list. The agent can use it when the user asks for something related to its description.

### Always-on Skills

If you want the skill to always be active (without needing to be requested), set `always: true`:

```yaml
---
name: project-rules
description: Rules the agent must always follow
always: true
---

Rules:
- Always respond in Portuguese
- Never delete files without confirming
```

### Skills with Dependencies

If the skill requires an installed program or environment variable:

```yaml
---
name: docker-helper
description: Helps with Docker
requires:
  bins: [docker, docker-compose]
  env: [DOCKER_HOST]
---
```

If dependencies are not satisfied, the skill appears as "unavailable" and is not used.

---

## 12. Agent Loop

The Agent Loop (`_run_agent_loop()` in `openbotx/agent/loop.py`) is the heart of the processing. It implements the **ReAct** pattern (Reason + Act) — the LLM thinks, acts (uses tools), observes the result, and repeats.

```mermaid
graph TD
    A[Send messages + tool definitions to LLM] --> B{LLM responded with...}
    B -->|Plain text| C[Return as final response - END]
    B -->|Tool calls| D[For each tool call:]
    D --> E[ToolRegistry finds the tool]
    E --> F[Validate parameters]
    F --> G[Execute tool]
    G --> H["Add result to messages (role: tool)"]
    H --> I{More tool calls?}
    I -->|Yes| D
    I -->|No| J{Hit iteration limit?}
    J -->|No| A
    J -->|Yes| K[Return limit reached message]
```

### Real-time Broadcasting

During processing, the agent broadcasts events via the `EventDispatcher` so the frontend can show progress. All `chat:*` events include `agent_name` for multi-agent identification:

- **`chat:tool_use`** — when a tool is executed. Includes `tool` (human-readable name like `Read File`, `Web Search`) and `description` (human-readable summary of the tool call with arguments)
- **`chat:thinking`** — when the LLM returns reasoning content (see below)
- **`chat:user_message`** — when a non-web message arrives (e.g., from Telegram or cron). Includes `channel`, `content`, and `media`. This allows the web UI to display messages from other channels in real time
- **`chat:transcription`** — when audio media is transcribed. Includes the transcription text, which is also prepended to the user message content
- **`sessions:updated`** — after saving each conversation turn (user + assistant messages) to the session. Triggers sidebar reload in the frontend

### Extended Thinking (Reasoning)

Some LLM models support **extended thinking** — a feature where the model exposes its internal chain-of-thought alongside the final response. This is a model-specific capability, not something OpenBotX controls.

**Which models support it:**
- Anthropic Claude models with extended thinking enabled (e.g., `claude-sonnet-4-20250514` with thinking mode)
- Other providers may add support via LiteLLM as they release thinking-capable models

**How it flows through the system:**

1. The agent loop sends a request to the LLM via `provider.chat()`
2. LiteLLM returns the response with an optional `reasoning_content` field on the message object
3. `LiteLLMProvider` extracts it: `getattr(message, "reasoning_content", None)` and includes it in the `LLMResponse` dataclass
4. The agent loop checks if `reasoning_content` exists. If it does, it broadcasts a `chat:thinking` WebSocket event:

```python
if response.reasoning_content and self._dispatcher:
    await self._dispatcher.broadcast("chat:thinking", {
        "task_id": task_id,
        "chat_id": chat_id,
        "content": response.reasoning_content,
        "agent_name": agent_name,
    })
```

5. The frontend receives this event and can display the model's thought process (e.g., in a collapsible "thinking" section above the response)

**What happens with the reasoning content after broadcast:**
- It is saved into the conversation messages via `ContextBuilder.add_assistant_message()` (as the `reasoning_content` field)
- However, it is **stripped before sending to the LLM** on subsequent turns — `_sanitize_messages()` only keeps `role`, `content`, `tool_calls`, `tool_call_id`, and `name`. The reasoning is not re-sent to the model
- This means reasoning is broadcast once for real-time display, stored in the session for history, but never re-injected into future LLM calls

For models that don't support extended thinking, `reasoning_content` is simply `None` and no `chat:thinking` event is broadcast. The agent loop works identically in both cases.

### Iteration Limit

The loop has a maximum iteration limit (default: **40**, configurable via `agent_params.max_iterations`). If the agent hits this limit without reaching a final response, it returns:

```
"I've reached my processing limit. Please try again or simplify your request."
```

### Output Limits

Each tool manages its own output limits internally. There is no global truncation applied by the agent loop or `ContextBuilder` — tool results are passed through as-is. This design ensures each tool controls the size and format of its results appropriately:

- **read_file** — 50 KB per read, with line-based pagination (offset/limit parameters)
- **exec** — 200 KB, tail-based truncation (keeps the last 200 KB, where errors and results typically appear)
- **http_client** — 50 KB response body truncation
- **web_fetch** — 50 KB content extraction limit

Other tools (file operations, web search, memory, etc.) return naturally small results and don't need truncation.

### Practical Example

Imagine the user asks "List the files in the docs/ folder":

**Iteration 1:**
- LLM receives the message and decides to call the `list_dir` tool with `{"path": "docs/"}`
- The tool returns: `"architecture.md\nconfiguration.md\napi.md\nskills.md\ntools.md"`
- Result is added to messages

**Iteration 2:**
- LLM receives the result and generates text: "The files in the docs/ folder are: architecture.md, configuration.md, ..."
- Since it's plain text (no tool calls), the loop ends
- This is the final response sent to the user

---

## 13. AI Provider Selection

OpenBotX supports multiple AI providers via **LiteLLM**. The selection of which provider to use follows a cascade logic in `Config.get_provider()` (`openbotx/config/schema.py`):

```mermaid
graph TD
    A["Configured model (e.g., anthropic/claude-sonnet-4-20250514)"] --> B{Has prefix with /?}
    B -->|Yes| C["Extract prefix (e.g., 'anthropic')"]
    C --> D{Prefix matches a provider whose credential has an API key?}
    D -->|Yes| E[Use that provider]
    D -->|No| F[Continue to keyword search]
    B -->|No| F
    F --> G{Any provider has a keyword matching the model name?}
    G -->|Yes| H[Use that provider]
    G -->|No| I{Any provider with a configured credential that has an API key?}
    I -->|Yes| J[Use the first one found]
    I -->|No| K[No provider available]
```

`ServerFactory.create_provider()` then resolves through provider -> credential -> concrete key to build a `LiteLLMProvider`.

### Model Parameters Resolution

Each provider can define a default `model_params` dict with arbitrary key-value pairs (e.g., `max_tokens`, `temperature`, `top_p`). At startup, these are merged into each agent's `model_params` using simple dict merge. The resolution order is:

1. **Provider `model_params`** — default parameters for all agents using the provider
2. **Agent `model_params`** — override provider defaults (agent keys always take precedence)

This merge happens once in `ServerFactory.create_orchestrator` before agent loops are created: `{**provider_params, **agent_params}`.

### Model Name Resolution

The `LiteLLMProvider` (`openbotx/providers/litellm_provider.py`) transforms the model name before sending to LiteLLM:

1. If there's a configured **gateway** (explicit model provider), it applies the gateway prefix and optionally strips the original prefix
2. Otherwise, it looks up the provider specification by model name and adds the correct LiteLLM prefix

### Prompt Caching

Cloud LLM providers like Anthropic charge tokens to process the system prompt and tool definitions on every API call. Since these rarely change between turns (the system prompt, .md files, memory, and tool schemas are the same throughout a conversation), re-processing them on every call wastes time and money.

**Prompt caching** tells the provider to keep these blocks in a temporary server-side cache. On subsequent calls within the same session, the provider recognizes the cached content and skips re-processing, reducing both latency and cost.

Currently, two providers support this: **Anthropic** and **OpenRouter**. The flag `supports_prompt_caching` is set per provider in `openbotx/providers/registry.py`.

When enabled, `_apply_cache_control()` in `openbotx/providers/litellm_provider.py` transforms the messages before sending to the API:

**System message** — the content is converted from a string to a list of content blocks, with `cache_control` on the last block:

```python
# Before:
{"role": "system", "content": "You are OpenBotX..."}

# After:
{"role": "system", "content": [
    {"type": "text", "text": "You are OpenBotX...", "cache_control": {"type": "ephemeral"}}
]}
```

**Tool definitions** — `cache_control` is added to the last tool in the list:

```python
# Before:
[{"type": "function", "function": {"name": "read_file", ...}}, ...]

# After (last tool only):
[..., {"type": "function", "function": {"name": "generate_image", ...}, "cache_control": {"type": "ephemeral"}}]
```

The `ephemeral` type means the cache is temporary — the provider decides how long to keep it (typically 5 minutes of inactivity for Anthropic). Non-system messages (user, assistant, tool results) are **not** cached because they change every turn.

This transformation is applied automatically and transparently — the rest of the codebase works with regular strings and lists, unaware of caching.

### LLM Error Recovery

- **Malformed JSON**: When the LLM returns tool call arguments with invalid JSON, the system uses `json_repair.loads()` to attempt automatic correction
- **Empty content**: Messages with empty content are replaced with `"(empty)"` to avoid provider 400 errors
- **Sanitization**: Only allowed keys (`role`, `content`, `tool_calls`, `tool_call_id`, `name`) are sent to the LLM — extra fields are stripped

---

## 14. Tools

Tools are capabilities that the agent can use to interact with the world. Each tool is a Python class that inherits from `Tool` (`openbotx/tools/base.py`) and implements the `execute()` method.

### Tool Registry

The `ToolRegistry` (`openbotx/tools/registry.py`) manages all tools:

```python
class ToolRegistry:
    def register(self, tool: Tool)           # register a tool
    def get_definitions(self) -> list[dict]  # return schemas for the LLM
    async def execute(name, params) -> str   # execute a tool
```

When the AgentLoop calls the LLM, it sends tool definitions (name, description, parameters) so the LLM knows what's available.

### Available Tools

| Tool | Description |
|---|---|
| `read_file` | Read file contents |
| `write_file` | Write content to a file |
| `edit_file` | Edit parts of an existing file |
| `list_dir` | List files and folders in a directory |
| `exec` | Execute shell commands |
| `web_search` | Search the web (Brave Search API) |
| `web_fetch` | Fetch content from a URL |
| `http_client` | Make HTTP requests (GET, POST, PUT, DELETE, PATCH, HEAD, OPTIONS) with download/upload support |
| `rss_reader` | Read RSS/Atom feeds and return latest entries |
| `message` | Send intermediate messages to the user |
| `spawn` | Create subagents for parallel tasks |
| `cron` | Schedule recurring or one-time tasks |
| `memory_save` | Save information to long-term memory |
| `memory_read` | Read MEMORY.md or HISTORY.md on demand |
| `memory_search` | Search across memory files for specific information |
| `browser` | Browser automation (Chrome/Chromium via CDP) |
| `generate_image` | Generate images (if configured) |

### Parameter Validation

Before executing any tool, the `ToolRegistry` validates parameters using the JSON schema defined by the tool. Validation includes:
- Required field checking (`required`)
- Type validation (`string`, `integer`, `boolean`, `array`, `object`)
- Enum verification (allowed values)
- Minimum and maximum limits

### Tool Error Handling

When a tool fails, the result includes the error message followed by a hint for the LLM:

```
Error: File not found: config.yml

[Analyze the error above and try a different approach.]
```

This hint is automatically appended by the `ToolRegistry` to any result starting with "Error" or when an exception occurs. This instructs the LLM to try a different approach instead of repeating the same error.

If the tool is not found, the error message lists all available tools, helping the LLM self-correct.

### Workspace Restriction and PathResolver

By default, file tools (`read_file`, `write_file`, `edit_file`, `list_dir`), the HTTP client (`http_client`), and the shell (`exec`) are restricted to the agent's workspace and the shared public directory. The restriction is enforced by the `PathResolver` class (`openbotx/helpers/path.py`):

- **File tools and HTTP client**: All use a `PathResolver` instance for path resolution. The resolver:
  1. Expands `~` (home directory) via `Path.expanduser()`
  2. Resolves relative paths against the agent's workspace
  3. When `tools.general.restrict_to_workspace` is enabled, verifies the resolved path falls within one of the allowed directories (workspace or public). If not, raises a `PermissionError`
- **Per-agent isolation**: Each agent has its own `PathResolver` with `allowed_dirs = [agent_workspace, public_dir]`. In multi-agent setups, agents cannot access each other's workspaces
- **Shell (exec)**: The working directory is set to the agent's workspace. The `PathResolver.is_restricted` property is used to determine whether workspace restriction is active for the exec tool

### Shell Safety Guards (exec)

The `exec` tool (`openbotx/tools/shell.py`) has multiple layers of protection:

**Blocked commands** — Regex patterns that prevent execution:
- `rm -rf`, `del /f`, `rmdir /s` (destructive deletion)
- `format`, `mkfs`, `diskpart` (disk formatting)
- `dd if=`, `> /dev/sd` (direct device writing)
- `shutdown`, `reboot`, `poweroff` (machine shutdown)
- Fork bombs (`:(){ ... };:`)

**Path traversal protection** — Blocks commands containing `../` or `..\`

**Absolute path restriction** — When `restrict_to_workspace=True`, absolute paths in the command are verified to ensure they are within the working directory

**Timeout** — Commands exceeding the timeout (default: 60s) are automatically killed with `process.kill()`

**Truncation** — Output large

…(truncated)
