Stata MCP Skill
Instructions
- Ensure the
stata MCP server is registered (see project README for config) and request it if not already active.
- When the user asks for Stata work:
- Use
run_command for ad-hoc syntax (trace=True for call stacks, raw=True for plain output).
- Use
load_data before analyses that require datasets.
- Use
get_data, describe, codebook, or get_variable_list to inspect data.
- Use
run_do_file for provided .do scripts.
- Use
export_graph/export_graphs_all for visualization requests.
- Use
get_help when the user wants Stata documentation.
- Use
get_stored_results to return r()/e() scalars/macros after commands for validation.
- Use
read_log to tail or retrieve output from long-running commands.
- Use
get_ui_channel to obtain a localhost HTTP endpoint for high-volume data browsing.
- Surface
rc/stderr info back to the user, referencing r()/e() codes.
- If Stata isn't auto-discovered, remind the user to set
STATA_PATH (examples in README).
Tool quick reference
Command Execution
run_command(code, echo=True, as_json=True, trace=False, raw=False, max_output_lines=None): Run Stata syntax.
code: The Stata command(s) to execute.
echo: Include the command itself in output (default: True).
as_json: Return JSON envelope with rc/stdout/stderr/error (default: True).
trace: Enable set trace on for deeper error diagnostics (default: False).
raw: Return plain stdout/error message instead of JSON (default: False).
max_output_lines: Truncate output to this many lines (default: None for no truncation).
- Note: Always writes output to a temporary log file and emits a
notifications/logMessage with {"event":"log_path","path":"..."} so the client can tail it locally.
run_do_file(path, echo=True, as_json=True, trace=False, raw=False, max_output_lines=None): Execute .do files.
path: Path to the .do file.
echo: Include commands in output (default: True).
as_json: Return JSON envelope (default: True).
trace: Enable trace mode for debugging (default: False).
raw: Return plain output instead of JSON (default: False).
max_output_lines: Truncate output to this many lines (default: None).
- Note: Always writes output to a temporary log file and emits incremental
notifications/progress when the client provides a progress token/callback.
read_log(path, offset=0, max_bytes=65536): Read a slice of a previously-provided log file.
path: Path to the log file (from notifications/logMessage).
offset: Byte offset to start reading from (default: 0).
max_bytes: Maximum bytes to read (default: 65536).
- Returns JSON:
path, offset, next_offset, data.
Data Loading & Inspection
load_data(source, clear=True, as_json=True, raw=False, max_output_lines=None): Load data using sysuse/webuse/use heuristics.
source: Dataset name, URL, or file path (e.g., "auto", "webuse nlsw88", "/path/to/file.dta").
clear: Append , clear to replace existing data (default: True).
as_json: Return JSON envelope (default: True).
raw: Return plain output (default: False).
max_output_lines: Truncate output to this many lines (default: None).
- Note: After loading, use UI channel for advanced filtering/sorting at scale.
get_data(start=0, count=50): Retrieve a slice of the active dataset as JSON.
start: Zero-based index of first observation (default: 0).
count: Number of observations to retrieve (default: 50, max: 500).
- Note: For advanced sorting/filtering at scale, use the UI channel endpoints (see
get_ui_channel()).
describe(): Return variable descriptions, storage types, and labels.
get_variable_list(): Return JSON list of all variables with names, labels, and types.
codebook(variable, as_json=True, trace=False, raw=False, max_output_lines=None): Return codebook/summary for a specific variable.
variable: Variable name to describe.
as_json: Return JSON envelope (default: True).
trace: Enable trace mode (default: False).
raw: Return plain output (default: False).
max_output_lines: Truncate output to this many lines (default: None).
Graph Management
list_graphs(): List all graphs in Stata's memory with active graph marked.
- Note: Graphs are automatically cached during command execution for instant exports.
export_graph(graph_name=None, format="pdf"): Export a stored graph to file.
graph_name: Name of graph to export (from list_graphs); if None, exports active graph.
format: Output format—"pdf" (default) or "png". Use "png" to view plots directly.
export_graphs_all(): Export all graphs in memory. Returns file paths.
Help & Results
get_help(topic, plain_text=False): Return Stata help text.
topic: Command or help topic (e.g., "regress", "graph").
plain_text: Return plain text instead of Markdown (default: False).
get_stored_results(): Return current r() and e() results as JSON after a command.
Session Management
create_session(session_id): Manually create a new Stata session.
list_sessions(): List all active sessions and their status (running, idle, etc.).
stop_session(session_id): Terminate and clean up a specific session.
break_session(session_id="default"): Interrupt the currently executing command in a session.
- Use this tool when a command is taking too long or you want to stop a long-running loop without losing data already in memory.
- Follow-up with
read_log to see where execution stopped.
UI Data Browser
get_ui_channel(): Return a short-lived localhost HTTP endpoint + bearer token for the UI-only data browser.
- Returns JSON with
baseUrl, token, expiresAt, and capabilities.
- Intended for VS Code extension UI to browse data at high volume (paging, filtering, sorting) without sending large payloads over MCP.
- Loopback only (binds to
127.0.0.1), requires bearer auth.
- Key endpoints (all require
Authorization: Bearer <token> header):
GET /v1/dataset: Dataset identity and state
GET /v1/vars: Variable metadata
POST /v1/page: Page data with optional sorting (sortBy parameter)
POST /v1/arrow: Binary Arrow IPC stream
POST /v1/views: Create filtered view
POST /v1/views/:viewId/page: Page within filtered view (supports sorting)
POST /v1/views/:viewId/arrow: Arrow stream from filtered view
DELETE /v1/views/:viewId: Delete view
POST /v1/filters/validate: Validate filter expression
- Sorting: Use
sortBy array in page requests (e.g., ["price"] for ascending, ["-price"] for descending, ["foreign", "-price"] for multi-level)
- Filtering: Filter expressions use Python boolean operators (
==, !=, <, >, and, or); Stata-style &/| also accepted
- Server limits: maxLimit=500, maxVars=32767, maxChars=500, maxRequestBytes=1000000, maxArrowLimit=1000000
- Dataset tracking:
datasetId used for cache invalidation; changing dataset invalidates view handles
Cancellation
- Clients may cancel an in-flight request by sending the MCP notification
notifications/cancelled with params.requestId set to the original tool call ID.
- Pass a
_meta.progressToken when invoking the tool if you want progress updates (optional).
- Cancellation is best-effort and depends on Stata surfacing
BreakError.
Error Reporting
- All tools executing Stata commands support JSON envelopes (
as_json=true) containing:
rc: Return code from r()/c(rc)
stdout: Standard output
stderr: Standard error (captures "red text")
message: Error message
line: Line number (when Stata reports it)
command: The command that was executed
log_path: Path to log file for streaming (when applicable)
snippet: Excerpt of error output
- Stata-specific error codes (
r(XXX)) are parsed and preserved
- Use
trace=true to enable set trace on for detailed program-defined error diagnostics
- Set
MCP_STATA_LOGLEVEL environment variable (e.g., DEBUG, INFO) to control server logging
MCP Resources
The server exposes these resources for MCP clients:
stata://data/summary → summarize
stata://data/metadata → describe
stata://graphs/list → graph list
stata://variables/list → variable list
stata://results/stored → stored r()/e() results
Graph review workflow
- Call
list_graphs() to see available plots and identify the active graph.
- Use
export_graphs_all() to fetch file paths for every graph; view them directly in the client.
- For a single plot, call
export_graph(graph_name="GraphName", format="png") to get a viewable file.
- Compare the rendered PNGs to the user spec (titles, axes labels, legends, colors, filters); state whether the graph matches and what to change.
Examples
Run a regression
# Load sample data and run regression
load_data("auto")
run_command("regress price mpg")
get_stored_results() # Retrieve coefficients and statistics
Export a histogram
# Create and export a graph
run_command("histogram price")
list_graphs() # Confirm graph exists
export_graph(graph_name="Graph", format="png") # Export for viewing
Debug a do-file
run_do_file("/path/to/analysis.do", trace=True)
Inspect data structure
load_data("nlsw88", clear=True)
describe()
get_variable_list()
codebook("wage")
get_data(start=0, count=10)
Read log output from long-running command
# After run_command emits a log_path notification
read_log("/tmp/stata_log_abc123.log", offset=0)
# Continue reading with next_offset for incremental output
read_log("/tmp/stata_log_abc123.log", offset=4096)
Advanced data browsing with sorting and filtering
# Get UI channel for high-volume data operations
get_ui_channel() # Returns baseUrl, token, expiresAt
# Example UI channel usage (requires HTTP client):
# POST {baseUrl}/v1/page with Authorization: Bearer {token}
# Body: {"datasetId":"...","offset":0,"limit":50,"vars":["price","mpg"],"sortBy":["-price"]}
# Create filtered view for price < 5000
# POST {baseUrl}/v1/views
# Body: {"datasetId":"...","frame":"default","filterExpr":"price < 5000"}
# Page through filtered view with sorting
# POST {baseUrl}/v1/views/{viewId}/page
# Body: {"offset":0,"limit":50,"vars":["price","mpg"],"sortBy":["-price"]}
1---2name: stata-mcp3description: Run or debug Stata workflows through the local io.github.tmonk/mcp-stata server. Use when users mention Stata commands, .do files, r()/e() results, dataset inspection, Stata graph exports, or data browsing with sorting/filtering.4---56# Stata MCP Skill78## Instructions91. Ensure the `stata` MCP server is registered (see project README for config) and request it if not already active.102. When the user asks for Stata work:11 - Use `run_command` for ad-hoc syntax (`trace=True` for call stacks, `raw=True` for plain output).12 - Use `load_data` before analyses that require datasets.13 - Use `get_data`, `describe`, `codebook`, or `get_variable_list` to inspect data.14 - Use `run_do_file` for provided `.do` scripts.15 - Use `export_graph`/`export_graphs_all` for visualization requests.16 - Use `get_help` when the user wants Stata documentation.17 - Use `get_stored_results` to return `r()`/`e()` scalars/macros after commands for validation.18 - Use `read_log` to tail or retrieve output from long-running commands.19 - Use `get_ui_channel` to obtain a localhost HTTP endpoint for high-volume data browsing.203. Surface `rc`/`stderr` info back to the user, referencing `r()`/`e()` codes.214. If Stata isn't auto-discovered, remind the user to set `STATA_PATH` (examples in README).2223## Tool quick reference2425### Command Execution26- `run_command(code, echo=True, as_json=True, trace=False, raw=False, max_output_lines=None)`: Run Stata syntax.27 - `code`: The Stata command(s) to execute.28 - `echo`: Include the command itself in output (default: True).29 - `as_json`: Return JSON envelope with rc/stdout/stderr/error (default: True).30 - `trace`: Enable `set trace on` for deeper error diagnostics (default: False).31 - `raw`: Return plain stdout/error message instead of JSON (default: False).32 - `max_output_lines`: Truncate output to this many lines (default: None for no truncation).33 - Note: Always writes output to a temporary log file and emits a `notifications/logMessage` with `{"event":"log_path","path":"..."}` so the client can tail it locally.3435- `run_do_file(path, echo=True, as_json=True, trace=False, raw=False, max_output_lines=None)`: Execute .do files.36 - `path`: Path to the .do file.37 - `echo`: Include commands in output (default: True).38 - `as_json`: Return JSON envelope (default: True).39 - `trace`: Enable trace mode for debugging (default: False).40 - `raw`: Return plain output instead of JSON (default: False).41 - `max_output_lines`: Truncate output to this many lines (default: None).42 - Note: Always writes output to a temporary log file and emits incremental `notifications/progress` when the client provides a progress token/callback.4344- `read_log(path, offset=0, max_bytes=65536)`: Read a slice of a previously-provided log file.45 - `path`: Path to the log file (from `notifications/logMessage`).46 - `offset`: Byte offset to start reading from (default: 0).47 - `max_bytes`: Maximum bytes to read (default: 65536).48 - Returns JSON: `path`, `offset`, `next_offset`, `data`.4950### Data Loading & Inspection51- `load_data(source, clear=True, as_json=True, raw=False, max_output_lines=None)`: Load data using sysuse/webuse/use heuristics.52 - `source`: Dataset name, URL, or file path (e.g., "auto", "webuse nlsw88", "/path/to/file.dta").53 - `clear`: Append `, clear` to replace existing data (default: True).54 - `as_json`: Return JSON envelope (default: True).55 - `raw`: Return plain output (default: False).56 - `max_output_lines`: Truncate output to this many lines (default: None).57 - Note: After loading, use UI channel for advanced filtering/sorting at scale.5859- `get_data(start=0, count=50)`: Retrieve a slice of the active dataset as JSON.60 - `start`: Zero-based index of first observation (default: 0).61 - `count`: Number of observations to retrieve (default: 50, max: 500).62 - Note: For advanced sorting/filtering at scale, use the UI channel endpoints (see `get_ui_channel()`).6364- `describe()`: Return variable descriptions, storage types, and labels.6566- `get_variable_list()`: Return JSON list of all variables with names, labels, and types.6768- `codebook(variable, as_json=True, trace=False, raw=False, max_output_lines=None)`: Return codebook/summary for a specific variable.69 - `variable`: Variable name to describe.70 - `as_json`: Return JSON envelope (default: True).71 - `trace`: Enable trace mode (default: False).72 - `raw`: Return plain output (default: False).73 - `max_output_lines`: Truncate output to this many lines (default: None).7475### Graph Management76- `list_graphs()`: List all graphs in Stata's memory with active graph marked.77 - Note: Graphs are automatically cached during command execution for instant exports.7879- `export_graph(graph_name=None, format="pdf")`: Export a stored graph to file.80 - `graph_name`: Name of graph to export (from `list_graphs`); if None, exports active graph.81 - `format`: Output format—"pdf" (default) or "png". Use "png" to view plots directly.8283- `export_graphs_all()`: Export all graphs in memory. Returns file paths.8485### Help & Results86- `get_help(topic, plain_text=False)`: Return Stata help text.87 - `topic`: Command or help topic (e.g., "regress", "graph").88 - `plain_text`: Return plain text instead of Markdown (default: False).8990- `get_stored_results()`: Return current `r()` and `e()` results as JSON after a command.9192### Session Management93- `create_session(session_id)`: Manually create a new Stata session.94- `list_sessions()`: List all active sessions and their status (running, idle, etc.).95- `stop_session(session_id)`: Terminate and clean up a specific session.96- `break_session(session_id="default")`: Interrupt the currently executing command in a session.97 - Use this tool when a command is taking too long or you want to stop a long-running loop without losing data already in memory.98 - Follow-up with `read_log` to see where execution stopped.99100### UI Data Browser101- `get_ui_channel()`: Return a short-lived localhost HTTP endpoint + bearer token for the UI-only data browser.102 - Returns JSON with `baseUrl`, `token`, `expiresAt`, and `capabilities`.103 - Intended for VS Code extension UI to browse data at high volume (paging, filtering, sorting) without sending large payloads over MCP.104 - Loopback only (binds to `127.0.0.1`), requires bearer auth.105 - **Key endpoints** (all require `Authorization: Bearer <token>` header):106 - `GET /v1/dataset`: Dataset identity and state107 - `GET /v1/vars`: Variable metadata108 - `POST /v1/page`: Page data with optional sorting (`sortBy` parameter)109 - `POST /v1/arrow`: Binary Arrow IPC stream110 - `POST /v1/views`: Create filtered view111 - `POST /v1/views/:viewId/page`: Page within filtered view (supports sorting)112 - `POST /v1/views/:viewId/arrow`: Arrow stream from filtered view113 - `DELETE /v1/views/:viewId`: Delete view114 - `POST /v1/filters/validate`: Validate filter expression115 - **Sorting**: Use `sortBy` array in page requests (e.g., `["price"]` for ascending, `["-price"]` for descending, `["foreign", "-price"]` for multi-level)116 - **Filtering**: Filter expressions use Python boolean operators (`==`, `!=`, `<`, `>`, `and`, `or`); Stata-style `&`/`|` also accepted117 - **Server limits**: maxLimit=500, maxVars=32767, maxChars=500, maxRequestBytes=1000000, maxArrowLimit=1000000118 - **Dataset tracking**: `datasetId` used for cache invalidation; changing dataset invalidates view handles119120## Cancellation121- Clients may cancel an in-flight request by sending the MCP notification `notifications/cancelled` with `params.requestId` set to the original tool call ID.122- Pass a `_meta.progressToken` when invoking the tool if you want progress updates (optional).123- Cancellation is best-effort and depends on Stata surfacing `BreakError`.124125## Error Reporting126- All tools executing Stata commands support JSON envelopes (`as_json=true`) containing:127 - `rc`: Return code from r()/c(rc)128 - `stdout`: Standard output129 - `stderr`: Standard error (captures "red text")130 - `message`: Error message131 - `line`: Line number (when Stata reports it)132 - `command`: The command that was executed133 - `log_path`: Path to log file for streaming (when applicable)134 - `snippet`: Excerpt of error output135- Stata-specific error codes (`r(XXX)`) are parsed and preserved136- Use `trace=true` to enable `set trace on` for detailed program-defined error diagnostics137- Set `MCP_STATA_LOGLEVEL` environment variable (e.g., `DEBUG`, `INFO`) to control server logging138139## MCP Resources140The server exposes these resources for MCP clients:141- `stata://data/summary` → `summarize`142- `stata://data/metadata` → `describe`143- `stata://graphs/list` → graph list144- `stata://variables/list` → variable list145- `stata://results/stored` → stored r()/e() results146147## Graph review workflow1481. Call `list_graphs()` to see available plots and identify the active graph.1492. Use `export_graphs_all()` to fetch file paths for every graph; view them directly in the client.1503. For a single plot, call `export_graph(graph_name="GraphName", format="png")` to get a viewable file.1514. Compare the rendered PNGs to the user spec (titles, axes labels, legends, colors, filters); state whether the graph matches and what to change.152153## Examples154155### Run a regression156```157# Load sample data and run regression158load_data("auto")159run_command("regress price mpg")160get_stored_results() # Retrieve coefficients and statistics161```162163### Export a histogram164```165# Create and export a graph166run_command("histogram price")167list_graphs() # Confirm graph exists168export_graph(graph_name="Graph", format="png") # Export for viewing169```170171### Debug a do-file172```173run_do_file("/path/to/analysis.do", trace=True)174```175176### Inspect data structure177```178load_data("nlsw88", clear=True)179describe()180get_variable_list()181codebook("wage")182get_data(start=0, count=10)183```184185### Read log output from long-running command186```187# After run_command emits a log_path notification188read_log("/tmp/stata_log_abc123.log", offset=0)189# Continue reading with next_offset for incremental output190read_log("/tmp/stata_log_abc123.log", offset=4096)191```192193### Advanced data browsing with sorting and filtering194```195# Get UI channel for high-volume data operations196get_ui_channel() # Returns baseUrl, token, expiresAt197198# Example UI channel usage (requires HTTP client):199# POST {baseUrl}/v1/page with Authorization: Bearer {token}200# Body: {"datasetId":"...","offset":0,"limit":50,"vars":["price","mpg"],"sortBy":["-price"]}201202# Create filtered view for price < 5000203# POST {baseUrl}/v1/views204# Body: {"datasetId":"...","frame":"default","filterExpr":"price < 5000"}205206# Page through filtered view with sorting207# POST {baseUrl}/v1/views/{viewId}/page208# Body: {"offset":0,"limit":50,"vars":["price","mpg"],"sortBy":["-price"]}209```