Open Executive self-hosted virtual executive team
Open Executive is a self-hosted application, not a library, CLI package, or MCP
server you install into another project. You clone it, configure providers, and
run two services:
user message -> Executive orchestrator -> parallel specialist calls
-> ChromaDB retrieval -> one synthesized executive answer
Eight specialists sit behind one voice: CSO, CFO, CHRO, General Counsel, COO,
CMO, CPO, and Board Communications. The internal agent structure is
deliberately never exposed to end users.
This skill tracks upstream commit 3a48f77a35e6980335553b9bdd02724e00f6f239
(2026-08-27) on main. The repository has no Git tags or GitHub releases even
though CHANGELOG.md documents 0.1.0, so pin a commit rather than a tag.
When to use this skill
- Decide whether Open Executive fits a request, and pick local, Docker, or Fly.io
- Preflight Python 3.11+, uv, Node 22+, and provider configuration before a first run
- Choose between Anthropic, OpenRouter, and local Ollama, LM Studio, or vLLM models
- Control paid spend, including the web-search default that contradicts
.env.example
- Preserve Anthropic prompt caching while changing prompts or tools
- Connect Slack, Discord, Telegram, Google Chat, or Gmail without unwanted outbound messages
- Load or unload demo fixtures without triggering the irreversible factory reset
- Operate the single-instance scheduler, episodic memory, and Fly.io deploys
- Add a specialist agent or workflow with the tests, evals, and docs upstream requires
- Debug 401 test failures, a stale
make eval path, or first-boot slowness
Do not use this skill for:
- Generic multi-agent framework selection: use
microsoft-agent-framework, openai-agents-python, or goalflow
- Provider-neutral LLM gateway routing: use
amrouter
- LLM tracing and offline eval platforms: use
langsmith or opik
- Generic FastAPI, Next.js, or ChromaDB questions unrelated to this codebase
- Business strategy advice itself; this skill operates the tool, it is not the advisor
- Fly.io-independent deployment strategy: use
deployment-automation
Instructions
Step 0: Enforce the safety contract
- Treat every turn as billable. Each message can fan out to several
specialists,
claude-opus-4-7 runs deep reasoning for CSO, CFO, GC, and
Board, and a background claude-haiku-4-5 pass extracts episodic memory
after every response. Confirm before running paid traffic on someone's key.
- Verify the web-search posture explicitly.
config.py defaults
enable_web_search to True while .env.example claims it is off. Each
search is billed and the cost scales with specialist fan-out. Never assume
it is off because the sample env file says so.
- Integrations send real messages. Slack, Discord, Telegram, Google Chat,
and Gmail deliver to real people. The outbound guard is anti-spam only, it
fails open on internal errors, and it is not an approval gate. Get explicit
approval before enabling a channel or triggering a proactive send.
- Never run the destructive fixture and cleanup operations casually.
POST /fixtures/reset is an irreversible factory reset of live state and
the snapshot. make stop kills any process on ports 8000 and 3000.
make clean deletes the virtualenv, node_modules, and .next.
- Keep company data out of Git.
packages/core/company/ and .env are
gitignored on purpose. Never commit a profile, uploaded document, database,
or key.
- Never print secret or personal values. Report provider keys, tokens, and
recipient addresses as set or unset only.
- Respect the single-instance rule. The scheduler claims jobs with
UPDATE ... RETURNING; a second API machine double-fires scheduled actions.
Run the read-only helpers before touching a real deployment:
bash .agent-skills/openexecutive/scripts/openexecutive.sh doctor /path/to/OpenExecutive
python3 .agent-skills/openexecutive/scripts/audit-config.py /path/to/OpenExecutive/.env
Step 1: Pick exactly one operating mode
| Mode |
Choose it when |
First action |
fit-check |
It is unclear whether this project fits |
Read the architecture summary in references/upstream-and-architecture.md |
preflight |
Host or provider readiness is unknown |
Run openexecutive.sh doctor |
run-local |
The app must start on this machine |
Confirm a provider, then make dev or Docker |
provider-cost |
Spend, models, caching, or local models matter |
Run audit-config.py, read references/setup-and-providers.md |
integrations |
A messaging or email channel is involved |
Read the outbound rules in references/operations-and-safety.md |
operate |
Fixtures, scheduler, memory, or Fly.io work is needed |
Classify the operation risk tier first |
contribute |
Code, prompts, or a new agent will change |
Read references/contributing-and-evals.md |
troubleshoot |
A concrete failure exists |
Identify the failing layer before retrying |
Do not blend provider setup, paid runs, integration enablement, and deployment
into one unreviewable shell block.
Step 2: Preflight the host and repository
bash .agent-skills/openexecutive/scripts/openexecutive.sh doctor /path/to/OpenExecutive
The helper checks Python 3.11+, uv, Node 22+, npm, Docker, flyctl, Git, and
Make, detects a checkout, and reports provider and integration variables as set
or unset without printing values. It installs nothing and starts nothing.
Missing pieces are lane facts, not blockers for every lane. Docker replaces uv
and Node; flyctl only matters for a Fly deployment.
Step 3: Configure a provider before the first run
The app refuses to start with no provider. Pick one path:
- Anthropic: set
ANTHROPIC_API_KEY, keep the default model trio.
- OpenRouter: set
OPENROUTER_ENABLED=true and OPENROUTER_API_KEY to bill
through OpenRouter and unlock non-Anthropic models per agent.
- Local: set
LOCAL_MODELS_ENABLED=true, LOCAL_BASE_URL including the
version path, and LOCAL_MODELS, then point DEFAULT_MODEL,
DEEP_REASONING_MODEL, and ROUTING_MODEL at local slugs to run with no
Anthropic key. Server-side web search has no local equivalent.
Audit the resulting file before any run:
python3 .agent-skills/openexecutive/scripts/audit-config.py /path/to/OpenExecutive/.env
It reports provider coverage, the effective web-search posture, outbound-guard
posture, and access-control gaps. Only an allowlist of non-sensitive settings,
such as feature flags and model names, is echoed; credentials, hostnames, and
email addresses are shown as present or absent. Details are in
references/setup-and-providers.md.
Step 4: Run it locally
cp .env.example .env # then edit; .env is gitignored
make dev # API on 8000, UI on 3000
The first run pulls heavy ML dependencies and downloads a roughly 90 MB
embedding model, so it takes minutes before the UI is usable. Docker Compose is
the alternative. Use make stop only when you accept that it kills every
process on ports 8000 and 3000, including unrelated dev servers.
Complete onboarding in the UI to build the company profile, or use the CLI:
openexecutive onboard
openexecutive chat
openexecutive ask "How should we price the new tier?"
Step 5: Control cost and preserve prompt caching
Prompt caching is load-bearing; upstream states that breaking it multiplies
cost roughly tenfold. When editing prompts or tools:
- keep tool definitions sorted by name;
- keep the Executive persona a constant, never an f-string;
- keep dynamic content out of any block carrying
cache_control;
- inject retrieval context into the user turn, not the cached system prompt.
Cap search spend with ENABLE_WEB_SEARCH=false or a low WEB_SEARCH_MAX_USES,
and leave XCRAWL_ENABLED off unless the user asked for it.
Step 6: Treat integrations and outbound sends as external actions
Enable a channel only when the user asks. Before enabling, confirm the roster
model: Email, Telegram, and Discord access is driven by non-archived Person
rows, and the deployed UI is gated by Google sign-in plus ALLOWED_EMAILS.
Keep the anti-spam knobs conservative and remember they are best-effort
suppression, not authorization. references/operations-and-safety.md lists the
send chokepoint, guard behavior, and the questions to answer before turning on
Gmail, Slack, Discord, Telegram, or Google Chat.
Step 7: Classify every operation by risk before running it
- Read-only: health, listing fixtures, reading status, viewing audit rows.
- Stateful but recoverable: loading a fixture, snapshotting, uploading a
document, running a chat turn that spends money.
- Destructive:
POST /fixtures/reset, POST /fixtures/unload,
DELETE /fixtures/{name}, openexecutive consolidate-initiatives --apply,
make clean, make stop, and any flyctl secrets change that restarts a
live app.
Preview merges before applying them, since --apply deletes rows and takes an
immediate SQLite write lock:
openexecutive consolidate-initiatives # dry run
openexecutive consolidate-initiatives --apply # only after review
Pause the API before a large consolidation. Never scale the API beyond one
machine.
Step 8: Contribute with the checks upstream actually enforces
Run unit tests with the shared secret unset, because a leftover value makes the
full-app tests return 401 instead of their expected status:
env -u BACKEND_SHARED_SECRET uv run pytest tests/unit/ -v
uv run ruff check openexecutive/ && uv run mypy openexecutive/
make eval passes a scenarios path that does not exist in the repository. The
runner's own default is also relative to the current directory and only
resolves from evals/, so pass an explicit scenario path instead. Adding an
agent requires prompts, registry and tool-enum updates, knowledge, evals, and
architecture-doc updates in the same pull request. See
references/contributing-and-evals.md.
Examples
Example 1: Check readiness without installing anything
bash .agent-skills/openexecutive/scripts/openexecutive.sh doctor ~/src/OpenExecutive
Resolve only the lane you need, then configure exactly one provider.
Example 2: Find the real spend posture of an existing config
python3 .agent-skills/openexecutive/scripts/audit-config.py ~/src/OpenExecutive/.env --json
Treat an unset ENABLE_WEB_SEARCH as billable searches enabled, because the
code default is on.
Example 3: Preview an initiative merge before deleting rows
openexecutive consolidate-initiatives
Only after reviewing the proposed clusters, re-run with --apply.
Example 4: Run the eval suite on the path that exists
cd packages/core
uv run python ../../evals/run_evals.py \
--scenarios openexecutive/evals/_scenarios/ \
--output ../../evals/results/
Pass the scenario path explicitly. Both the Makefile target and the runner's
relative default resolve to a missing directory from this working directory,
and a missing directory yields zero scenarios instead of an error.
Best practices
- Confirm the provider and the spend before the first paid turn.
- Set
ENABLE_WEB_SEARCH explicitly instead of trusting the sample env file.
- Keep the API at one machine so scheduled actions fire once.
- Enable a messaging channel only on request, and verify the roster first.
- Preview fixture and consolidation operations; never reach for the reset route to fix a smaller problem.
- Keep company profiles, uploads, databases, and keys out of Git.
- Preserve prompt-cache structure when editing prompts or tools.
- Prefer a local or OpenRouter provider when the user wants to avoid Anthropic billing.
- Verify claims against the code, since the sample env file and
make eval are known to disagree with it.
- Re-read current upstream before asserting latest behavior; this skill is pinned to one commit.
References
references/upstream-and-architecture.md - pinned metadata, layout, agent and memory architecture
references/setup-and-providers.md - prerequisites, provider choices, env keys, cost controls
references/operations-and-safety.md - risk tiers, destructive operations, integrations, scheduler, deploys
references/contributing-and-evals.md - tests, evals, agent-addition checklist, known repo discrepancies
scripts/openexecutive.sh - read-only host and repository doctor, safety summary, pinned URLs
scripts/audit-config.py - offline env posture audit that never prints values
- Open Executive repository
- Pinned upstream source
1---2name: openexecutive3description: Operate SenteLabsAI/OpenExecutive, the Apache-2.0 self-hosted virtual executive team that answers business questions through one Executive persona backed by eight specialist Claude agents over FastAPI, Next.js, ChromaDB, and SQLite. Route one request to one mode: fit-check the project; preflight Python 3.11, uv, and Node 22; choose an Anthropic, OpenRouter, or local model provider; run it with `make dev` or Docker; control paid spend and Anthropic prompt caching; connect Slack, Discord, Telegram, Google Chat, or Gmail without sending unwanted messages; operate fixtures, the single-instance scheduler, and Fly.io deploys; or contribute an agent with tests and evals. Use when a user wants to run, configure, extend, or debug Open Executive. Triggers on: openexecutive, Open Executive, SenteLabs, virtual executive team, virtual CFO or CSO agent, consult_specialist, openexec-api, episodic memory executive, fixtures reset, single-instance scheduler.4---56# Open Executive self-hosted virtual executive team78Open Executive is a self-hosted application, not a library, CLI package, or MCP9server you install into another project. You clone it, configure providers, and10run two services:1112```text13user message -> Executive orchestrator -> parallel specialist calls14 -> ChromaDB retrieval -> one synthesized executive answer15```1617Eight specialists sit behind one voice: CSO, CFO, CHRO, General Counsel, COO,18CMO, CPO, and Board Communications. The internal agent structure is19deliberately never exposed to end users.2021This skill tracks upstream commit `3a48f77a35e6980335553b9bdd02724e00f6f239`22(2026-08-27) on `main`. The repository has no Git tags or GitHub releases even23though `CHANGELOG.md` documents `0.1.0`, so pin a commit rather than a tag.2425## When to use this skill2627- Decide whether Open Executive fits a request, and pick local, Docker, or Fly.io28- Preflight Python 3.11+, uv, Node 22+, and provider configuration before a first run29- Choose between Anthropic, OpenRouter, and local Ollama, LM Studio, or vLLM models30- Control paid spend, including the web-search default that contradicts `.env.example`31- Preserve Anthropic prompt caching while changing prompts or tools32- Connect Slack, Discord, Telegram, Google Chat, or Gmail without unwanted outbound messages33- Load or unload demo fixtures without triggering the irreversible factory reset34- Operate the single-instance scheduler, episodic memory, and Fly.io deploys35- Add a specialist agent or workflow with the tests, evals, and docs upstream requires36- Debug 401 test failures, a stale `make eval` path, or first-boot slowness3738Do not use this skill for:3940- Generic multi-agent framework selection: use `microsoft-agent-framework`, `openai-agents-python`, or `goalflow`41- Provider-neutral LLM gateway routing: use `amrouter`42- LLM tracing and offline eval platforms: use `langsmith` or `opik`43- Generic FastAPI, Next.js, or ChromaDB questions unrelated to this codebase44- Business strategy advice itself; this skill operates the tool, it is not the advisor45- Fly.io-independent deployment strategy: use `deployment-automation`4647## Instructions4849### Step 0: Enforce the safety contract50511. **Treat every turn as billable.** Each message can fan out to several52 specialists, `claude-opus-4-7` runs deep reasoning for CSO, CFO, GC, and53 Board, and a background `claude-haiku-4-5` pass extracts episodic memory54 after every response. Confirm before running paid traffic on someone's key.552. **Verify the web-search posture explicitly.** `config.py` defaults56 `enable_web_search` to `True` while `.env.example` claims it is off. Each57 search is billed and the cost scales with specialist fan-out. Never assume58 it is off because the sample env file says so.593. **Integrations send real messages.** Slack, Discord, Telegram, Google Chat,60 and Gmail deliver to real people. The outbound guard is anti-spam only, it61 fails open on internal errors, and it is not an approval gate. Get explicit62 approval before enabling a channel or triggering a proactive send.634. **Never run the destructive fixture and cleanup operations casually.**64 `POST /fixtures/reset` is an irreversible factory reset of live state and65 the snapshot. `make stop` kills any process on ports 8000 and 3000.66 `make clean` deletes the virtualenv, `node_modules`, and `.next`.675. **Keep company data out of Git.** `packages/core/company/` and `.env` are68 gitignored on purpose. Never commit a profile, uploaded document, database,69 or key.706. **Never print secret or personal values.** Report provider keys, tokens, and71 recipient addresses as set or unset only.727. **Respect the single-instance rule.** The scheduler claims jobs with73 `UPDATE ... RETURNING`; a second API machine double-fires scheduled actions.7475Run the read-only helpers before touching a real deployment:7677```bash78bash .agent-skills/openexecutive/scripts/openexecutive.sh doctor /path/to/OpenExecutive79python3 .agent-skills/openexecutive/scripts/audit-config.py /path/to/OpenExecutive/.env80```8182### Step 1: Pick exactly one operating mode8384| Mode | Choose it when | First action |85|---|---|---|86| `fit-check` | It is unclear whether this project fits | Read the architecture summary in `references/upstream-and-architecture.md` |87| `preflight` | Host or provider readiness is unknown | Run `openexecutive.sh doctor` |88| `run-local` | The app must start on this machine | Confirm a provider, then `make dev` or Docker |89| `provider-cost` | Spend, models, caching, or local models matter | Run `audit-config.py`, read `references/setup-and-providers.md` |90| `integrations` | A messaging or email channel is involved | Read the outbound rules in `references/operations-and-safety.md` |91| `operate` | Fixtures, scheduler, memory, or Fly.io work is needed | Classify the operation risk tier first |92| `contribute` | Code, prompts, or a new agent will change | Read `references/contributing-and-evals.md` |93| `troubleshoot` | A concrete failure exists | Identify the failing layer before retrying |9495Do not blend provider setup, paid runs, integration enablement, and deployment96into one unreviewable shell block.9798### Step 2: Preflight the host and repository99100```bash101bash .agent-skills/openexecutive/scripts/openexecutive.sh doctor /path/to/OpenExecutive102```103104The helper checks Python 3.11+, uv, Node 22+, npm, Docker, flyctl, Git, and105Make, detects a checkout, and reports provider and integration variables as set106or unset without printing values. It installs nothing and starts nothing.107108Missing pieces are lane facts, not blockers for every lane. Docker replaces uv109and Node; flyctl only matters for a Fly deployment.110111### Step 3: Configure a provider before the first run112113The app refuses to start with no provider. Pick one path:114115- **Anthropic**: set `ANTHROPIC_API_KEY`, keep the default model trio.116- **OpenRouter**: set `OPENROUTER_ENABLED=true` and `OPENROUTER_API_KEY` to bill117 through OpenRouter and unlock non-Anthropic models per agent.118- **Local**: set `LOCAL_MODELS_ENABLED=true`, `LOCAL_BASE_URL` including the119 version path, and `LOCAL_MODELS`, then point `DEFAULT_MODEL`,120 `DEEP_REASONING_MODEL`, and `ROUTING_MODEL` at local slugs to run with no121 Anthropic key. Server-side web search has no local equivalent.122123Audit the resulting file before any run:124125```bash126python3 .agent-skills/openexecutive/scripts/audit-config.py /path/to/OpenExecutive/.env127```128129It reports provider coverage, the effective web-search posture, outbound-guard130posture, and access-control gaps. Only an allowlist of non-sensitive settings,131such as feature flags and model names, is echoed; credentials, hostnames, and132email addresses are shown as present or absent. Details are in133`references/setup-and-providers.md`.134135### Step 4: Run it locally136137```bash138cp .env.example .env # then edit; .env is gitignored139make dev # API on 8000, UI on 3000140```141142The first run pulls heavy ML dependencies and downloads a roughly 90 MB143embedding model, so it takes minutes before the UI is usable. Docker Compose is144the alternative. Use `make stop` only when you accept that it kills every145process on ports 8000 and 3000, including unrelated dev servers.146147Complete onboarding in the UI to build the company profile, or use the CLI:148149```bash150openexecutive onboard151openexecutive chat152openexecutive ask "How should we price the new tier?"153```154155### Step 5: Control cost and preserve prompt caching156157Prompt caching is load-bearing; upstream states that breaking it multiplies158cost roughly tenfold. When editing prompts or tools:159160- keep tool definitions sorted by name;161- keep the Executive persona a constant, never an f-string;162- keep dynamic content out of any block carrying `cache_control`;163- inject retrieval context into the user turn, not the cached system prompt.164165Cap search spend with `ENABLE_WEB_SEARCH=false` or a low `WEB_SEARCH_MAX_USES`,166and leave `XCRAWL_ENABLED` off unless the user asked for it.167168### Step 6: Treat integrations and outbound sends as external actions169170Enable a channel only when the user asks. Before enabling, confirm the roster171model: Email, Telegram, and Discord access is driven by non-archived Person172rows, and the deployed UI is gated by Google sign-in plus `ALLOWED_EMAILS`.173174Keep the anti-spam knobs conservative and remember they are best-effort175suppression, not authorization. `references/operations-and-safety.md` lists the176send chokepoint, guard behavior, and the questions to answer before turning on177Gmail, Slack, Discord, Telegram, or Google Chat.178179### Step 7: Classify every operation by risk before running it180181- **Read-only**: health, listing fixtures, reading status, viewing audit rows.182- **Stateful but recoverable**: loading a fixture, snapshotting, uploading a183 document, running a chat turn that spends money.184- **Destructive**: `POST /fixtures/reset`, `POST /fixtures/unload`,185 `DELETE /fixtures/{name}`, `openexecutive consolidate-initiatives --apply`,186 `make clean`, `make stop`, and any `flyctl secrets` change that restarts a187 live app.188189Preview merges before applying them, since `--apply` deletes rows and takes an190immediate SQLite write lock:191192```bash193openexecutive consolidate-initiatives # dry run194openexecutive consolidate-initiatives --apply # only after review195```196197Pause the API before a large consolidation. Never scale the API beyond one198machine.199200### Step 8: Contribute with the checks upstream actually enforces201202Run unit tests with the shared secret unset, because a leftover value makes the203full-app tests return 401 instead of their expected status:204205```bash206env -u BACKEND_SHARED_SECRET uv run pytest tests/unit/ -v207uv run ruff check openexecutive/ && uv run mypy openexecutive/208```209210`make eval` passes a scenarios path that does not exist in the repository. The211runner's own default is also relative to the current directory and only212resolves from `evals/`, so pass an explicit scenario path instead. Adding an213agent requires prompts, registry and tool-enum updates, knowledge, evals, and214architecture-doc updates in the same pull request. See215`references/contributing-and-evals.md`.216217## Examples218219### Example 1: Check readiness without installing anything220221```bash222bash .agent-skills/openexecutive/scripts/openexecutive.sh doctor ~/src/OpenExecutive223```224225Resolve only the lane you need, then configure exactly one provider.226227### Example 2: Find the real spend posture of an existing config228229```bash230python3 .agent-skills/openexecutive/scripts/audit-config.py ~/src/OpenExecutive/.env --json231```232233Treat an unset `ENABLE_WEB_SEARCH` as billable searches enabled, because the234code default is on.235236### Example 3: Preview an initiative merge before deleting rows237238```bash239openexecutive consolidate-initiatives240```241242Only after reviewing the proposed clusters, re-run with `--apply`.243244### Example 4: Run the eval suite on the path that exists245246```bash247cd packages/core248uv run python ../../evals/run_evals.py \249 --scenarios openexecutive/evals/_scenarios/ \250 --output ../../evals/results/251```252253Pass the scenario path explicitly. Both the Makefile target and the runner's254relative default resolve to a missing directory from this working directory,255and a missing directory yields zero scenarios instead of an error.256257## Best practices2582591. Confirm the provider and the spend before the first paid turn.2602. Set `ENABLE_WEB_SEARCH` explicitly instead of trusting the sample env file.2613. Keep the API at one machine so scheduled actions fire once.2624. Enable a messaging channel only on request, and verify the roster first.2635. Preview fixture and consolidation operations; never reach for the reset route to fix a smaller problem.2646. Keep company profiles, uploads, databases, and keys out of Git.2657. Preserve prompt-cache structure when editing prompts or tools.2668. Prefer a local or OpenRouter provider when the user wants to avoid Anthropic billing.2679. Verify claims against the code, since the sample env file and `make eval` are known to disagree with it.26810. Re-read current upstream before asserting latest behavior; this skill is pinned to one commit.269270## References271272- `references/upstream-and-architecture.md` - pinned metadata, layout, agent and memory architecture273- `references/setup-and-providers.md` - prerequisites, provider choices, env keys, cost controls274- `references/operations-and-safety.md` - risk tiers, destructive operations, integrations, scheduler, deploys275- `references/contributing-and-evals.md` - tests, evals, agent-addition checklist, known repo discrepancies276- `scripts/openexecutive.sh` - read-only host and repository doctor, safety summary, pinned URLs277- `scripts/audit-config.py` - offline env posture audit that never prints values278- [Open Executive repository](https://github.com/SenteLabsAI/OpenExecutive)279- [Pinned upstream source](https://github.com/SenteLabsAI/OpenExecutive/tree/3a48f77a35e6980335553b9bdd02724e00f6f239)