Windows-MCP Tool Tester
An automated testing skill that generates comprehensive test cases for a single Windows-MCP tool,
executes them, and produces a structured test report with pass/fail results, performance metrics,
and actionable recommendations.
Core Principles
- One tool per invocation. If the user doesn't specify which tool to test, ask them before proceeding.
- Black-box testing only. Derive test cases exclusively from the MCP tool description and parameter schema — never read source code. Silence in the schema is a documentation gap, not a testing hint.
- Auto-generate test cases from the tool's MCP description and parameter schema. Cover common scenarios, edge cases, parameter combinations, and error handling paths.
- Measure two dimensions: correctness (return value matches expectations) and response time (end-to-end, including MCP overhead).
- Mandatory side-effect verification: every tool call that may modify system state MUST be independently verified — no exceptions, no sampling.
- Safe cleanup: track process PIDs spawned during testing; only kill those specific PIDs during teardown, never kill by process name alone.
- Safety first: Windows-MCP has full system access with no sandboxing. Tests involving
destructive tools (FileSystem delete, Registry set/delete, Process kill, PowerShell) can
modify or destroy data. Running in a VM or Windows Sandbox is strongly recommended.
Before executing destructive test cases, confirm the user accepts the risk. See
SECURITY.md.
- Produce a structured report at the end (see Step 4).
Step 0: Identify the Target Tool
If the user hasn't specified a tool, present the full list and ask them to pick one:
App, PowerShell, Screenshot, Snapshot, Click, Type, Scroll, Move, Shortcut, Wait,
MultiSelect, MultiEdit, Clipboard, Process, Notification, FileSystem, Registry, Scrape
Once a tool is confirmed, proceed to Step 1. Do NOT test multiple tools in one session.
Step 1: Analyze the Tool
Read the tool's MCP description and parameter schema via the MCP server's tool listing. Identify:
- All parameters — name, type, required/optional, default value, allowed values (enums)
- All modes (if the tool is mode-based, e.g., FileSystem has read/write/copy/move/delete/list/search/info)
- Return value structure — what the tool returns on success vs. failure
- Side effects — does it modify system state? (important for test isolation)
- Dependencies — does it require a running app, an open window, existing files, etc.?
Use this analysis to inform test case generation. If the description is ambiguous or silent on a
behavior, note it as a documentation gap and design a test to probe it. The tool's response is
the ground truth.
Step 2: Generate Test Cases
Design test cases that cover the following categories. Not every category applies to every tool —
use judgment based on the tool's nature.
Category A: Basic Functionality (Required)
Test the tool's primary purpose with standard, well-formed inputs.
- One test per mode/operation type (e.g., for FileSystem: read, write, list, etc.)
- Use realistic parameter values
Category B: Parameter Variations (Required)
- Test each optional parameter individually to verify it takes effect
- Test enum parameters with every allowed value
- Test boolean parameters in both true and false states
- For
anyOf union types: The schema advertises all listed types as valid, so the tool must handle each correctly.
anyOf: [boolean, string] (e.g., drag, use_vision): test boolean true/false AND string "true"/"false" — both must produce the same behavior.
- Caveat: MCP transport layers may silently coerce
"true" (string) to true (boolean)
before the tool sees it. To genuinely probe the string path, also test non-standard truthy
strings like "yes" or "1" — if those fail while "true" passes, the tool likely only
receives booleans and the string branch is untested.
anyOf: [<type>, null] (nullable): test with a valid value AND explicit null. An unhandled TypeError on null is a FAIL.
- Test with default values (omit optional params) vs. explicit values
Category C: Edge Cases (Required)
- Empty strings, zero values, negative numbers where applicable
- Boundary values (e.g., very long text for Type, timeout=0 for PowerShell)
- Unicode / special characters in string parameters
- Very large or very small numeric inputs
Category D: Error Handling (Required)
- Missing required parameters
- Invalid parameter types or out-of-range values
- Referencing nonexistent resources (files, windows, processes, registry keys)
- Operations that should fail gracefully (e.g., deleting a non-existent file)
Category E: Parameter Interaction (When Applicable)
- Combinations of parameters that might interact (e.g., Click with both
loc and label)
- Mutually exclusive parameters
- Mode-specific parameter requirements
- Cross-mode parameter applicability: For mode-based tools, pass parameters meant for one
mode while calling another (e.g.,
window_loc in launch mode). Silently ignoring them is
a documentation gap worth reporting.
Category F: Idempotency & State (When Applicable)
- For tools marked
idempotentHint: true: call twice with same args, verify same result
- For destructive tools: verify cleanup or rollback is possible
- For stateful tools: verify state changes are reflected correctly
Test Case Format
For each test case, define:
ID: TC-{ToolName}-{Number}
Category: A/B/C/D/E/F
Description: What this test verifies
Parameters: The exact parameters to pass
Expected: What a correct result looks like (success/failure, key content in response)
Setup: Any prerequisite actions (create a file, open an app, etc.)
Teardown: Any cleanup actions after the test
Present the test plan to the user for confirmation before executing.
Aim for 10-20 test cases depending on tool complexity.
Include an estimated execution time: approximately 30-45 seconds per test case
(includes timing calls, tool execution, verification). App launches add 5-10s extra.
Present as a range, e.g., "Estimated execution time: 8-12 minutes (15 test cases)".
Step 3: Execute Tests
Pre-Test Step 1: Gather Environment Info
Before running test cases, collect the test environment details for the report. Use these
PowerShell commands for reliable results:
# OS version
(Get-CimInstance Win32_OperatingSystem).Caption + " " + (Get-CimInstance Win32_OperatingSystem).Version
# Display resolution (physical pixels)
Get-CimInstance Win32_VideoController | Select-Object CurrentHorizontalResolution, CurrentVerticalResolution
# Display count
(Get-CimInstance Win32_PnPEntity | Where-Object { $_.PNPClass -eq 'Monitor' -and $_.Status -eq 'OK' }).Count
# DPI scale factor (96 = 100%, 120 = 125%, 144 = 150%, 192 = 200%)
Get-ItemProperty 'HKCU:\Control Panel\Desktop\WindowMetrics' -Name AppliedDPI -ErrorAction SilentlyContinue | Select-Object -ExpandProperty AppliedDPI
Also call Screenshot once — its Screenshot Original Size cross-checks the DPI value,
and its output includes Active Desktop and All Desktops.
Pre-Test Step 2: Prepare Environment
For input tools (Type, Click, Scroll, Move, Shortcut, MultiSelect, MultiEdit), prepare the
environment before executing any test cases:
IME (Input Method) state: Check the Tray Input Indicator in the Snapshot output. If it
shows a non-English input mode (e.g., "Chinese Mode", "Japanese Mode"), switch to English
mode first using the Shortcut tool (typically shift to toggle). This is critical for
Type tool tests — an active IME will intercept keystrokes and produce incorrect characters.
Record the original IME state and restore it after testing.
Label availability check: Call Snapshot on the test target window and verify whether the
element you intend to use with label parameter is actually listed in the Interactive Elements.
Common pitfalls:
- Modern Windows 11 Notepad's text editing area is not exposed as an interactive element
in the UI tree — use
loc coordinates instead.
- Some complex controls (e.g., rich text editors, canvas-based UIs) may not enumerate child
elements.
If a planned
label-based test has no valid label target, adapt the test to use loc, or
pick a different element that does have a label (e.g., a search box, address bar).
Warm-up call: Execute 1-2 throwaway tool calls (not counted in test results) to warm up
the MCP connection, window focus, and UI tree cache. First calls are typically slower due to
cold start effects — excluding them gives more representative performance numbers.
Test Execution
Run each test case sequentially. For each test:
Label freshness rule: Snapshot labels are a point-in-time snapshot. If any action between
tests could change the UI state, call Snapshot again before using label parameters.
State reset between tests: Each test case should start from a known, clean state. For
input tools sharing a test window (e.g., Notepad), define a standard reset procedure and
execute it in Setup:
- Type tests: Shortcut (Ctrl+A) → Shortcut (Delete) to clear the text area
- Click/Move tests: Move cursor to a neutral position away from interactive elements
- Scroll tests: Reset scroll position to top (Ctrl+Home)
If a test's Setup includes clear=true in the tool call itself, you may skip the manual
reset — but verify in the teardown that the state is clean for the next test.
- Setup — perform any prerequisite actions (e.g., create a temp file for FileSystem read tests).
When spawning processes, record their PIDs for teardown (see Test Isolation Guidelines).
- Record start time — call PowerShell to capture a precise timestamp in milliseconds before calling the tool:
[long](([System.DateTime]::UtcNow - [System.DateTime]::UnixEpoch).TotalMilliseconds)
Save the returned integer as $t_start.
- Call the MCP tool with the specified parameters
- Record end time — immediately after the tool returns, call PowerShell again with the same command. Save as
$t_end.
- Compute elapsed time —
elapsed_ms = $t_end - $t_start. Note: this includes MCP
overhead from the timestamp calls themselves (~3-5s each). Use for relative comparison
between test cases only. When testing the PowerShell tool itself, timing is
self-referential — record times as N/A (self-referential) and rely on the PowerShell
tool's own timeout behavior and status codes for performance assessment instead.
- Capture the response — store the full return value and measure
response_size:
- Text-only tools: character count of the returned string.
- Mixed-content tools (Screenshot, Snapshot): character count of the text portion only,
note
+image in the report. Do not attempt to measure image byte size.
- Evaluate correctness — compare the response against expected behavior:
- Does the response indicate success/failure as expected?
- Does the response content match expected patterns?
- For error cases: does the error message make sense and provide useful information?
- Verify side effects (MANDATORY) — independently verify EVERY mutating tool call. Never
rely solely on the tool's return value. Never skip or sample.
Rule: verification MUST NOT use the same tool under test. Use a different tool
(preferably PowerShell) to cross-check. Verification methods:
- Move/Click: call Screenshot or Snapshot to verify the expected UI change occurred.
- Type: use Shortcut (Ctrl+A → Ctrl+C) then PowerShell
Get-Clipboard to capture
exact text for comparison.
- Drag (Move with drag=True): call Snapshot to verify the target window actually moved.
- App (launch/resize): call PowerShell (
Get-Process) to verify process exists, and
Screenshot/Snapshot to verify window position/size.
- FileSystem: verify with PowerShell (
Test-Path, Get-Content, Get-ChildItem)
— never with the FileSystem tool itself.
- Registry: verify with PowerShell (
Get-ItemProperty, reg query)
— never with the Registry tool itself.
- Clipboard (set): verify with PowerShell (
Get-Clipboard)
— never with the Clipboard tool itself.
- Process (kill): verify with PowerShell (
Get-Process -Id $pid)
— never with the Process tool itself.
- Shortcut: call Screenshot to verify the shortcut's expected effect occurred.
- For read-only tools (Screenshot, Snapshot, Scrape, Process list), this step is not needed.
- If verification fails but the tool reported success, mark the test as FAIL and note
"tool reported success but side effect not confirmed" in the root cause analysis.
- Teardown — clean up any side effects (delete temp files, close apps, etc.)
IMPORTANT: Never estimate response times. Always use the PowerShell measurement above.
If unavailable, record N/A and explain why.
Correctness Evaluation Criteria
| Result |
Meaning |
| PASS |
Response matches expected behavior exactly |
| SOFT PASS |
Response is acceptable but slightly different from ideal (e.g., extra whitespace, ordering) |
| FAIL |
Response doesn't match expected behavior — includes cases where the tool rejects schema-valid input. Never look up source code to explain away a failure. |
| ERROR |
Tool threw an unexpected exception or timed out |
| SKIP |
Test couldn't run due to missing prerequisites (document why) |
When to SKIP vs. Adapt
If a planned test case cannot execute as designed (e.g., the target element has no label in the
UI tree, or a required window state cannot be achieved), decide:
- SKIP if the prerequisite is truly missing and no workaround exists. Document the reason.
- Adapt if you can achieve the same test intent with a different approach (e.g., use
loc
instead of label, use a different target app). Update the test case description and note the
adaptation in the report. Adapting is preferred over skipping when the test intent is still
achievable.
Performance Tracking
For each test case, record response time (ms) via PowerShell timestamps and
response size (character count of the raw response text).
Warm-up effect: The first 1-2 tool calls in a session are typically slower due to MCP
connection warm-up, UI tree cache initialization, and window focus acquisition. If warm-up
calls were performed in Pre-Test, note this in the report. If not, flag the first test case's
timing as potentially inflated and exclude it from aggregate statistics (average, median, P95)
or mark it separately.
Step 4: Generate the Test Report
Localization: If the user specified a language, write the entire report in that language
(headings, tables, commentary, recommendations). Keep test case IDs (e.g., TC-Move-01) in
English. The template below is a structural reference — translate all prose while preserving
the markdown structure.
# Windows-MCP Tool Test Report: {ToolName}
**Date:** {timestamp}
**Tool:** {ToolName}
**Total Test Cases:** {N}
**PASS:** {P} | **SOFT PASS:** {SP} | **FAIL:** {F} | **ERROR:** {E} | **SKIP:** {S}
**Overall Pass Rate:** {(P+SP)/N * 100}%
---
## 1. Test Environment
{Record the environment to aid reproducibility. Data gathered in Pre-Test step.}
| Item | Value |
|--------------------------|----------------------------------------------------------|
| OS Version | {e.g., Windows 11 Pro 10.0.26200} |
| Display Resolution | {e.g., 2560x1440} |
| Screenshot Original Size | {e.g., 3840x2160 — this is resolution x scale factor} |
| Display Count | {e.g., 1} |
| Active Virtual Desktop | {e.g., Desktop 1} |
| MCP Transport | {e.g., SSE via http://localhost:8088/sse} |
| Scale Factor | {e.g., 150% (AppliedDPI=144)} |
---
## 2. Executive Summary
{2-3 sentences summarizing the overall health of the tool. Highlight critical failures if any.
Note any patterns — e.g., "all error-handling tests failed" or "basic functionality is solid
but Unicode support is incomplete." Also assess these dimensions when relevant:}
- **Error message quality**: descriptive and actionable, or cryptic?
- **Input validation**: does the tool validate params before executing, or fail deep with confusing errors?
- **Consistency**: do repeated calls with same params return consistent results?
- **Graceful degradation**: when prerequisites are missing, does the tool explain what's needed?
---
## 3. Failed & Error Test Cases
{For each non-passing test case, provide:}
### TC-{ID}: {Description}
- **Category:** {category}
- **Parameters:** `{params}`
- **Expected:** {what should have happened}
- **Actual:** {what actually happened}
- **Side-Effect Verification:** {what the independent verification revealed, if applicable}
- **Root Cause Analysis:** {your best assessment of why it failed}
- **Suggested Fix:** {actionable recommendation for the developer}
{If all tests passed, write: "All test cases passed. No issues to report."}
---
## 4. Performance Analysis
> **Note:** All times are end-to-end measurements including MCP transport overhead
> (serialization, network round-trip, SSE/stdio latency). They do NOT represent pure tool
> execution time. Use these numbers for **relative comparison** and outlier detection — not
> as absolute benchmarks. For pure execution time, check server-side logs with
> `WINDOWS_MCP_PROFILE_SNAPSHOT=1`.
### Response Time
| Test Case | Time (ms) | Assessment |
|-----------|-----------|------------|
| TC-XXX-01 | 6500 | Normal |
| TC-XXX-02 | 15200 | Slow |
| ... | ... | ... |
**Average:** {avg} ms | **Median:** {median} ms | **P95:** {p95} ms | **Max:** {max} ms
**Assessment thresholds (end-to-end including MCP overhead):**
- Fast: < 5000ms
- Normal: 5000ms – 10000ms
- Slow: 10000ms – 20000ms
- Very Slow: > 20000ms
{Commentary on any outliers or concerning patterns. When a test case is significantly slower
than peers, note possible causes: app launch wait, UI tree traversal, screenshot capture, etc.}
### Response Size
| Test Case | Response Size (chars) |
|-----------|-----------------------|
| TC-XXX-01 | 245 |
| ... | ... |
{Note any unexpectedly large or empty responses.}
---
## 5. Environmental Interference & Notes
{List any environmental factors that affected test execution but are not bugs in the tool itself.
These factors help future testers reproduce results and avoid false failures.}
| # | Factor | Impact | Mitigation |
|---|--------|--------|------------|
| 1 | {e.g., IME in Chinese mode} | {e.g., TC-Type-01 typed wrong characters} | {e.g., Switched IME to English before retesting} |
| ... | ... | ... | ... |
**Common environmental factors:**
- **IME state**: Active non-English input methods intercept keystrokes (affects Type, Shortcut)
- **Notification popups**: System or app notifications may steal focus mid-test
- **Background app focus changes**: Chat apps, update dialogs may overlay the test window
- **Screen lock / screensaver**: Can interrupt long-running test sessions
- **Clipboard managers**: Third-party clipboard tools may interfere with Clipboard tests
{If no environmental interference occurred, write: "No environmental interference observed."}
---
## 6. Documentation & Schema Gaps
{List any discrepancies between the tool's MCP parameter schema / description and its actual
behavior or environmental interactions. These are not necessarily bugs — they are places where
the documentation or schema could be improved to set correct expectations for callers.}
| # | Gap Type | Description | Recommendation |
|-----|-----------------------------------|-------------|----------------|
| 1 | {schema / description / behavior} | {desc} | {rec} |
| ... | ... | ... | ... |
**Gap Types:**
- **schema**: parameter schema (types, required/optional, allowed values) does not match actual behavior
- **description**: tool description is silent or ambiguous about a behavior that testing revealed
- **behavior**: tool behaves inconsistently with what the schema + description together imply
{If no gaps were found, write: "No documentation or schema gaps identified."}
---
## 7. All Test Cases
| ID | Category | Description | Result | Time (ms) | Response Size |
|------------|------------|-------------|--------|-----------|---------------|
| TC-XXX-01 | A - Basic | {desc} | PASS | 6500 | 245 |
| TC-XXX-02 | B - Params | {desc} | FAIL | 8200 | 310 |
| ... | ... | ... | ... | ... | ... |
Test Isolation Guidelines
To avoid polluting the system or interfering with user state:
- FileSystem tests: Use a dedicated temp directory (e.g.,
%TEMP%\wmcp-test-{timestamp}\).
Clean up after all tests complete.
- Registry tests: Use a dedicated test key under
HKCU:\Software\WMCP-Test-{timestamp}.
Delete the entire key after testing.
- Process tests: Only list processes (don't kill user processes). If testing kill, spawn a
sacrificial process first (e.g.,
notepad.exe) and record its PID.
- App tests: Use lightweight apps (Notepad, Calculator). Record PIDs of all processes
spawned during testing (use
(Start-Process notepad -PassThru).Id or query process list
before/after launch). In teardown, only kill processes by PID — NEVER by name (e.g.,
Stop-Process -Id $pid, not Stop-Process -Name notepad), because the user may have their
own instances of the same application running.
- Modern tabbed apps caveat: Windows 11 Notepad/Terminal may reuse a single process for
multiple tabs. Diff the process list before/after each launch to detect new PIDs. Only kill
PIDs that did not exist before testing began.
- Clipboard tests: Save and restore the original clipboard content.
- Input tools (Click, Type, Scroll, Move, Shortcut): Open a dedicated test window
(e.g., Notepad) to receive input. Don't interact with user's active work.
See also Pre-Test Step 2 for IME state handling — switch to English input mode before testing
and restore the original state in final teardown.
- Read-only tools (Screenshot, Snapshot, Scrape): Safe to run freely.
- Notification tests: User-visible (sends Windows toasts). Avoid repeated or unnecessary
notifications. Prefer a single clearly labeled test notification per test case.
- PowerShell tests: Use read-only commands where possible.
Tool-Specific Testing Guidance
Hints per tool. Always read the actual schema to discover additional scenarios beyond these.
App
- Modes: launch, resize, switch. Test each mode.
- Launch: test with known apps (notepad, calc), unknown app names
- Resize: test with valid window_loc/window_size, without an active app
- Switch: test switching to a running app, to a non-existent app
PowerShell
- Simple commands:
echo "hello", Get-Date, Get-Process | Select-Object -First 3
- Timeout behavior: set a very short timeout with a long-running command
- Encoding: commands with Unicode output
- Error output: commands that write to stderr
- Exit codes: commands that fail (e.g.,
Get-Item nonexistent)
Screenshot
- Default parameters (no args)
- With annotation enabled/disabled
- With reference lines
- With specific display index
- Verify return includes image data
Snapshot
- Various flag combinations: use_vision, use_dom, use_annotation, use_ui_tree
- All flags off vs. all flags on
- With/without reference lines
Click
- By coordinates (loc) vs. by label
- Different button types: left, right, middle
- clicks=0 (hover), clicks=1 (single), clicks=2 (double)
- Invalid coordinates (negative, off-screen)
- Invalid label (non-existent element ID)
Type
- Normal text, Unicode text, special characters
- With and without clear=true
- With and without press_enter=true
- Different caret_position values: start, idle, end
- By coordinates vs. by label
- IME sensitivity: Test with IME active to verify behavior (expect failure if tool uses
keystroke simulation rather than Unicode input). This is a high-value edge case because
many Windows machines have non-English IMEs installed.
- Emoji / surrogate pair characters: Test with characters outside the Basic Multilingual
Plane (e.g., 🌍, 😀) to verify supplementary plane Unicode support.
- Empty string: Test
text="" — this is a common edge case that may crash if the
implementation indexes into the string without a length check.
Scroll
- Vertical up/down, horizontal left/right
- Different wheel_times values (1, 5, 10)
- By coordinates vs. by label
Move
- Simple move to coordinates
- Drag mode (drag=true)
- By coordinates vs. by label
Shortcut
- Common shortcuts: ctrl+c, ctrl+v, ctrl+a, alt+tab
- Windows key shortcuts: win+r, win+d
- Multi-key combinations
- Invalid key names
Wait
- Short duration (1 second)
- Zero duration
- Verify actual elapsed time roughly matches requested duration
MultiSelect
- Select multiple items by coordinates
- Select by labels
- With and without press_ctrl
- Empty list of items
MultiEdit
- Edit multiple fields by coordinates
- Edit by labels
- Mixed valid and invalid targets
Clipboard
- get mode when clipboard has text
- get mode when clipboard is empty
- set mode with normal text
- set mode with Unicode text
- Roundtrip: set then get, verify content matches
Process
- list mode with default sort
- list mode with different sort_by values (memory, cpu, name)
- list mode with name filter
- list mode with different limit values
- kill mode with a sacrificial process (spawn notepad, then kill it)
Notification
- Valid notification with title, message, app_id
- Empty title or message
- Special characters in title/message
FileSystem
- Full mode coverage: read, write, copy, move, delete, list, search, info
- Read: existing file, non-existent file, offset/limit, different encodings
- Write: new file, overwrite, append
- List: with and without pattern, recursive, show_hidden
- Delete: file, empty dir, non-empty dir with recursive
Registry
- Full mode coverage: get, set, delete, list
- Set and get roundtrip
- Different value types (String, DWord, QWord)
- Non-existent key/value
- Use test-only registry path
Scrape
- With a URL (lightweight page)
- With and without query parameter
- With use_dom enabled (requires open browser)
- Invalid URL
1---2name: windows-mcp-tool-tester3description: Automated testing skill for Windows-MCP tools. Use this skill whenever the user wants to test, validate, benchmark, or evaluate any Windows-MCP tool (App, PowerShell, Screenshot, Snapshot, Click, Type, Scroll, Move, Shortcut, Wait, MultiSelect, MultiEdit, Clipboard, Process, Notification, FileSystem, Registry, Scrape). Triggers on phrases like "test the Click tool", "benchmark Screenshot", "validate FileSystem", "run QA on Registry", "check if PowerShell works", "evaluate tool performance", or any mention of testing/validating a Windows-MCP tool. Each invocation tests exactly ONE tool.4---5
6# Windows-MCP Tool Tester
7
8An automated testing skill that generates comprehensive test cases for a single Windows-MCP tool,
9executes them, and produces a structured test report with pass/fail results, performance metrics,
10and actionable recommendations.
11
12## Core Principles
13
14- **One tool per invocation.** If the user doesn't specify which tool to test, ask them before proceeding.
15- **Black-box testing only.** Derive test cases exclusively from the MCP tool description and parameter schema — never read source code. Silence in the schema is a documentation gap, not a testing hint.
16- **Auto-generate test cases** from the tool's MCP description and parameter schema. Cover common scenarios, edge cases, parameter combinations, and error handling paths.
17- **Measure two dimensions**: correctness (return value matches expectations) and response time (end-to-end, including MCP overhead).
18- **Mandatory side-effect verification**: every tool call that may modify system state MUST be independently verified — no exceptions, no sampling.
19- **Safe cleanup**: track process PIDs spawned during testing; only kill those specific PIDs during teardown, never kill by process name alone.
20- **Safety first**: Windows-MCP has full system access with no sandboxing. Tests involving
21 destructive tools (FileSystem delete, Registry set/delete, Process kill, PowerShell) can
22 modify or destroy data. Running in a **VM or Windows Sandbox** is strongly recommended.
23 Before executing destructive test cases, confirm the user accepts the risk. See `SECURITY.md`.
24- **Produce a structured report** at the end (see Step 4).
25
26---
27
28## Step 0: Identify the Target Tool
29
30If the user hasn't specified a tool, present the full list and ask them to pick one:
31
32> App, PowerShell, Screenshot, Snapshot, Click, Type, Scroll, Move, Shortcut, Wait,
33> MultiSelect, MultiEdit, Clipboard, Process, Notification, FileSystem, Registry, Scrape
34
35Once a tool is confirmed, proceed to Step 1. Do NOT test multiple tools in one session.
36
37---
38
39## Step 1: Analyze the Tool
40
41Read the tool's MCP description and parameter schema via the MCP server's tool listing. Identify:
42
431. **All parameters** — name, type, required/optional, default value, allowed values (enums)
442. **All modes** (if the tool is mode-based, e.g., FileSystem has read/write/copy/move/delete/list/search/info)
453. **Return value structure** — what the tool returns on success vs. failure
464. **Side effects** — does it modify system state? (important for test isolation)
475. **Dependencies** — does it require a running app, an open window, existing files, etc.?
48
49Use this analysis to inform test case generation. If the description is ambiguous or silent on a
50behavior, note it as a documentation gap and design a test to probe it. The tool's response is
51the ground truth.
52
53---
54
55## Step 2: Generate Test Cases
56
57Design test cases that cover the following categories. Not every category applies to every tool —
58use judgment based on the tool's nature.
59
60### Category A: Basic Functionality (Required)
61Test the tool's primary purpose with standard, well-formed inputs.
62- One test per mode/operation type (e.g., for FileSystem: read, write, list, etc.)
63- Use realistic parameter values
64
65### Category B: Parameter Variations (Required)
66- Test each optional parameter individually to verify it takes effect
67- Test enum parameters with every allowed value
68- Test boolean parameters in both true and false states
69- **For `anyOf` union types**: The schema advertises all listed types as valid, so the tool must handle each correctly.
70 - `anyOf: [boolean, string]` (e.g., `drag`, `use_vision`): test boolean `true`/`false` AND string `"true"`/`"false"` — both must produce the same behavior.
71 - **Caveat:** MCP transport layers may silently coerce `"true"` (string) to `true` (boolean)
72 before the tool sees it. To genuinely probe the string path, also test non-standard truthy
73 strings like `"yes"` or `"1"` — if those fail while `"true"` passes, the tool likely only
74 receives booleans and the string branch is untested.
75 - `anyOf: [<type>, null]` (nullable): test with a valid value AND explicit `null`. An unhandled TypeError on null is a FAIL.
76- Test with default values (omit optional params) vs. explicit values
77
78### Category C: Edge Cases (Required)
79- Empty strings, zero values, negative numbers where applicable
80- Boundary values (e.g., very long text for Type, timeout=0 for PowerShell)
81- Unicode / special characters in string parameters
82- Very large or very small numeric inputs
83
84### Category D: Error Handling (Required)
85- Missing required parameters
86- Invalid parameter types or out-of-range values
87- Referencing nonexistent resources (files, windows, processes, registry keys)
88- Operations that should fail gracefully (e.g., deleting a non-existent file)
89
90### Category E: Parameter Interaction (When Applicable)
91- Combinations of parameters that might interact (e.g., Click with both `loc` and `label`)
92- Mutually exclusive parameters
93- Mode-specific parameter requirements
94- **Cross-mode parameter applicability**: For mode-based tools, pass parameters meant for one
95 mode while calling another (e.g., `window_loc` in `launch` mode). Silently ignoring them is
96 a documentation gap worth reporting.
97
98### Category F: Idempotency & State (When Applicable)
99- For tools marked `idempotentHint: true`: call twice with same args, verify same result
100- For destructive tools: verify cleanup or rollback is possible
101- For stateful tools: verify state changes are reflected correctly
102
103### Test Case Format
104
105For each test case, define:
106
107```
108ID: TC-{ToolName}-{Number}
109Category: A/B/C/D/E/F
110Description: What this test verifies
111Parameters: The exact parameters to pass
112Expected: What a correct result looks like (success/failure, key content in response)
113Setup: Any prerequisite actions (create a file, open an app, etc.)
114Teardown: Any cleanup actions after the test
115```
116
117Present the test plan to the user for confirmation before executing.
118Aim for **10-20 test cases** depending on tool complexity.
119
120Include an **estimated execution time**: approximately **30-45 seconds** per test case
121(includes timing calls, tool execution, verification). App launches add 5-10s extra.
122Present as a range, e.g., "Estimated execution time: 8-12 minutes (15 test cases)".
123
124---
125
126## Step 3: Execute Tests
127
128### Pre-Test Step 1: Gather Environment Info
129
130Before running test cases, collect the test environment details for the report. Use these
131PowerShell commands for reliable results:
132
133```powershell
134# OS version
135(Get-CimInstance Win32_OperatingSystem).Caption + " " + (Get-CimInstance Win32_OperatingSystem).Version
136
137# Display resolution (physical pixels)
138Get-CimInstance Win32_VideoController | Select-Object CurrentHorizontalResolution, CurrentVerticalResolution
139
140# Display count
141(Get-CimInstance Win32_PnPEntity | Where-Object { $_.PNPClass -eq 'Monitor' -and $_.Status -eq 'OK' }).Count
142
143# DPI scale factor (96 = 100%, 120 = 125%, 144 = 150%, 192 = 200%)
144Get-ItemProperty 'HKCU:\Control Panel\Desktop\WindowMetrics' -Name AppliedDPI -ErrorAction SilentlyContinue | Select-Object -ExpandProperty AppliedDPI
145```
146
147Also call Screenshot once — its `Screenshot Original Size` cross-checks the DPI value,
148and its output includes `Active Desktop` and `All Desktops`.
149
150### Pre-Test Step 2: Prepare Environment
151
152For **input tools** (Type, Click, Scroll, Move, Shortcut, MultiSelect, MultiEdit), prepare the
153environment before executing any test cases:
154
1551. **IME (Input Method) state**: Check the Tray Input Indicator in the Snapshot output. If it
156 shows a non-English input mode (e.g., "Chinese Mode", "Japanese Mode"), switch to English
157 mode first using the `Shortcut` tool (typically `shift` to toggle). **This is critical for
158 Type tool tests** — an active IME will intercept keystrokes and produce incorrect characters.
159 Record the original IME state and restore it after testing.
160
1612. **Label availability check**: Call Snapshot on the test target window and verify whether the
162 element you intend to use with `label` parameter is actually listed in the Interactive Elements.
163 Common pitfalls:
164 - Modern Windows 11 Notepad's text editing area is **not** exposed as an interactive element
165 in the UI tree — use `loc` coordinates instead.
166 - Some complex controls (e.g., rich text editors, canvas-based UIs) may not enumerate child
167 elements.
168 If a planned `label`-based test has no valid label target, adapt the test to use `loc`, or
169 pick a different element that does have a label (e.g., a search box, address bar).
170
1713. **Warm-up call**: Execute 1-2 throwaway tool calls (not counted in test results) to warm up
172 the MCP connection, window focus, and UI tree cache. First calls are typically slower due to
173 cold start effects — excluding them gives more representative performance numbers.
174
175### Test Execution
176
177Run each test case sequentially. For each test:
178
179> **Label freshness rule:** Snapshot labels are a point-in-time snapshot. If any action between
180> tests could change the UI state, call Snapshot again before using `label` parameters.
181
182> **State reset between tests:** Each test case should start from a known, clean state. For
183> input tools sharing a test window (e.g., Notepad), define a standard reset procedure and
184> execute it in Setup:
185> - **Type tests**: Shortcut (Ctrl+A) → Shortcut (Delete) to clear the text area
186> - **Click/Move tests**: Move cursor to a neutral position away from interactive elements
187> - **Scroll tests**: Reset scroll position to top (Ctrl+Home)
188>
189> If a test's Setup includes `clear=true` in the tool call itself, you may skip the manual
190> reset — but verify in the teardown that the state is clean for the next test.
191
1921. **Setup** — perform any prerequisite actions (e.g., create a temp file for FileSystem read tests).
193 When spawning processes, **record their PIDs** for teardown (see Test Isolation Guidelines).
1942. **Record start time** — call PowerShell to capture a precise timestamp in milliseconds **before** calling the tool:
195 ```powershell
196 [long](([System.DateTime]::UtcNow - [System.DateTime]::UnixEpoch).TotalMilliseconds)
197 ```
198 Save the returned integer as `$t_start`.
1993. **Call the MCP tool** with the specified parameters
2004. **Record end time** — immediately after the tool returns, call PowerShell again with the same command. Save as `$t_end`.
2015. **Compute elapsed time** — `elapsed_ms = $t_end - $t_start`. Note: this includes MCP
202 overhead from the timestamp calls themselves (~3-5s each). Use for relative comparison
203 between test cases only. **When testing the PowerShell tool itself**, timing is
204 self-referential — record times as `N/A (self-referential)` and rely on the PowerShell
205 tool's own `timeout` behavior and status codes for performance assessment instead.
2066. **Capture the response** — store the full return value and measure `response_size`:
207 - **Text-only tools**: character count of the returned string.
208 - **Mixed-content tools** (Screenshot, Snapshot): character count of the **text portion only**,
209 note `+image` in the report. Do not attempt to measure image byte size.
2107. **Evaluate correctness** — compare the response against expected behavior:
211 - Does the response indicate success/failure as expected?
212 - Does the response content match expected patterns?
213 - For error cases: does the error message make sense and provide useful information?
2148. **Verify side effects (MANDATORY)** — independently verify EVERY mutating tool call. Never
215 rely solely on the tool's return value. Never skip or sample.
216 **Rule: verification MUST NOT use the same tool under test.** Use a different tool
217 (preferably PowerShell) to cross-check. Verification methods:
218 - **Move/Click**: call Screenshot or Snapshot to verify the expected UI change occurred.
219 - **Type**: use Shortcut (Ctrl+A → Ctrl+C) then PowerShell `Get-Clipboard` to capture
220 exact text for comparison.
221 - **Drag (Move with drag=True)**: call Snapshot to verify the target window actually moved.
222 - **App (launch/resize)**: call PowerShell (`Get-Process`) to verify process exists, and
223 Screenshot/Snapshot to verify window position/size.
224 - **FileSystem**: verify with PowerShell (`Test-Path`, `Get-Content`, `Get-ChildItem`)
225 — never with the FileSystem tool itself.
226 - **Registry**: verify with PowerShell (`Get-ItemProperty`, `reg query`)
227 — never with the Registry tool itself.
228 - **Clipboard (set)**: verify with PowerShell (`Get-Clipboard`)
229 — never with the Clipboard tool itself.
230 - **Process (kill)**: verify with PowerShell (`Get-Process -Id $pid`)
231 — never with the Process tool itself.
232 - **Shortcut**: call Screenshot to verify the shortcut's expected effect occurred.
233 - For read-only tools (Screenshot, Snapshot, Scrape, Process list), this step is not needed.
234 - If verification fails but the tool reported success, mark the test as **FAIL** and note
235 "tool reported success but side effect not confirmed" in the root cause analysis.
2369. **Teardown** — clean up any side effects (delete temp files, close apps, etc.)
237
238> **IMPORTANT:** Never estimate response times. Always use the PowerShell measurement above.
239> If unavailable, record `N/A` and explain why.
240
241### Correctness Evaluation Criteria
242
243| Result | Meaning |
244|-----------|-----------------------------------------------------------------------------------------------|
245| PASS | Response matches expected behavior exactly |
246| SOFT PASS | Response is acceptable but slightly different from ideal (e.g., extra whitespace, ordering) |
247| FAIL | Response doesn't match expected behavior — includes cases where the tool rejects schema-valid input. Never look up source code to explain away a failure. |
248| ERROR | Tool threw an unexpected exception or timed out |
249| SKIP | Test couldn't run due to missing prerequisites (document why) |
250
251### When to SKIP vs. Adapt
252
253If a planned test case cannot execute as designed (e.g., the target element has no `label` in the
254UI tree, or a required window state cannot be achieved), decide:
255
256- **SKIP** if the prerequisite is truly missing and no workaround exists. Document the reason.
257- **Adapt** if you can achieve the same test intent with a different approach (e.g., use `loc`
258 instead of `label`, use a different target app). Update the test case description and note the
259 adaptation in the report. Adapting is preferred over skipping when the test intent is still
260 achievable.
261
262### Performance Tracking
263
264For each test case, record **response time (ms)** via PowerShell timestamps and
265**response size** (character count of the raw response text).
266
267> **Warm-up effect:** The first 1-2 tool calls in a session are typically slower due to MCP
268> connection warm-up, UI tree cache initialization, and window focus acquisition. If warm-up
269> calls were performed in Pre-Test, note this in the report. If not, flag the first test case's
270> timing as potentially inflated and exclude it from aggregate statistics (average, median, P95)
271> or mark it separately.
272
273---
274
275## Step 4: Generate the Test Report
276
277**Localization:** If the user specified a language, write the entire report in that language
278(headings, tables, commentary, recommendations). Keep test case IDs (e.g., TC-Move-01) in
279English. The template below is a structural reference — translate all prose while preserving
280the markdown structure.
281
282````markdown
283# Windows-MCP Tool Test Report: {ToolName}
284
285**Date:** {timestamp}
286**Tool:** {ToolName}
287**Total Test Cases:** {N}
288**PASS:** {P} | **SOFT PASS:** {SP} | **FAIL:** {F} | **ERROR:** {E} | **SKIP:** {S}
289**Overall Pass Rate:** {(P+SP)/N * 100}%
290
291---
292
293## 1. Test Environment
294
295{Record the environment to aid reproducibility. Data gathered in Pre-Test step.}
296
297| Item | Value |
298|--------------------------|----------------------------------------------------------|
299| OS Version | {e.g., Windows 11 Pro 10.0.26200} |
300| Display Resolution | {e.g., 2560x1440} |
301| Screenshot Original Size | {e.g., 3840x2160 — this is resolution x scale factor} |
302| Display Count | {e.g., 1} |
303| Active Virtual Desktop | {e.g., Desktop 1} |
304| MCP Transport | {e.g., SSE via http://localhost:8088/sse} |
305| Scale Factor | {e.g., 150% (AppliedDPI=144)} |
306
307---
308
309## 2. Executive Summary
310
311{2-3 sentences summarizing the overall health of the tool. Highlight critical failures if any.
312Note any patterns — e.g., "all error-handling tests failed" or "basic functionality is solid
313but Unicode support is incomplete." Also assess these dimensions when relevant:}
314
315- **Error message quality**: descriptive and actionable, or cryptic?
316- **Input validation**: does the tool validate params before executing, or fail deep with confusing errors?
317- **Consistency**: do repeated calls with same params return consistent results?
318- **Graceful degradation**: when prerequisites are missing, does the tool explain what's needed?
319
320---
321
322## 3. Failed & Error Test Cases
323
324{For each non-passing test case, provide:}
325
326### TC-{ID}: {Description}
327- **Category:** {category}
328- **Parameters:** `{params}`
329- **Expected:** {what should have happened}
330- **Actual:** {what actually happened}
331- **Side-Effect Verification:** {what the independent verification revealed, if applicable}
332- **Root Cause Analysis:** {your best assessment of why it failed}
333- **Suggested Fix:** {actionable recommendation for the developer}
334
335{If all tests passed, write: "All test cases passed. No issues to report."}
336
337---
338
339## 4. Performance Analysis
340
341> **Note:** All times are end-to-end measurements including MCP transport overhead
342> (serialization, network round-trip, SSE/stdio latency). They do NOT represent pure tool
343> execution time. Use these numbers for **relative comparison** and outlier detection — not
344> as absolute benchmarks. For pure execution time, check server-side logs with
345> `WINDOWS_MCP_PROFILE_SNAPSHOT=1`.
346
347### Response Time
348
349| Test Case | Time (ms) | Assessment |
350|-----------|-----------|------------|
351| TC-XXX-01 | 6500 | Normal |
352| TC-XXX-02 | 15200 | Slow |
353| ... | ... | ... |
354
355**Average:** {avg} ms | **Median:** {median} ms | **P95:** {p95} ms | **Max:** {max} ms
356
357**Assessment thresholds (end-to-end including MCP overhead):**
358- Fast: < 5000ms
359- Normal: 5000ms – 10000ms
360- Slow: 10000ms – 20000ms
361- Very Slow: > 20000ms
362
363{Commentary on any outliers or concerning patterns. When a test case is significantly slower
364than peers, note possible causes: app launch wait, UI tree traversal, screenshot capture, etc.}
365
366### Response Size
367
368| Test Case | Response Size (chars) |
369|-----------|-----------------------|
370| TC-XXX-01 | 245 |
371| ... | ... |
372
373{Note any unexpectedly large or empty responses.}
374
375---
376
377## 5. Environmental Interference & Notes
378
379{List any environmental factors that affected test execution but are not bugs in the tool itself.
380These factors help future testers reproduce results and avoid false failures.}
381
382| # | Factor | Impact | Mitigation |
383|---|--------|--------|------------|
384| 1 | {e.g., IME in Chinese mode} | {e.g., TC-Type-01 typed wrong characters} | {e.g., Switched IME to English before retesting} |
385| ... | ... | ... | ... |
386
387**Common environmental factors:**
388- **IME state**: Active non-English input methods intercept keystrokes (affects Type, Shortcut)
389- **Notification popups**: System or app notifications may steal focus mid-test
390- **Background app focus changes**: Chat apps, update dialogs may overlay the test window
391- **Screen lock / screensaver**: Can interrupt long-running test sessions
392- **Clipboard managers**: Third-party clipboard tools may interfere with Clipboard tests
393
394{If no environmental interference occurred, write: "No environmental interference observed."}
395
396---
397
398## 6. Documentation & Schema Gaps
399
400{List any discrepancies between the tool's MCP parameter schema / description and its actual
401behavior or environmental interactions. These are not necessarily bugs — they are places where
402the documentation or schema could be improved to set correct expectations for callers.}
403
404| # | Gap Type | Description | Recommendation |
405|-----|-----------------------------------|-------------|----------------|
406| 1 | {schema / description / behavior} | {desc} | {rec} |
407| ... | ... | ... | ... |
408
409**Gap Types:**
410- **schema**: parameter schema (types, required/optional, allowed values) does not match actual behavior
411- **description**: tool description is silent or ambiguous about a behavior that testing revealed
412- **behavior**: tool behaves inconsistently with what the schema + description together imply
413
414{If no gaps were found, write: "No documentation or schema gaps identified."}
415
416---
417
418## 7. All Test Cases
419
420| ID | Category | Description | Result | Time (ms) | Response Size |
421|------------|------------|-------------|--------|-----------|---------------|
422| TC-XXX-01 | A - Basic | {desc} | PASS | 6500 | 245 |
423| TC-XXX-02 | B - Params | {desc} | FAIL | 8200 | 310 |
424| ... | ... | ... | ... | ... | ... |
425````
426
427---
428
429## Test Isolation Guidelines
430
431To avoid polluting the system or interfering with user state:
432
433- **FileSystem tests**: Use a dedicated temp directory (e.g., `%TEMP%\wmcp-test-{timestamp}\`).
434 Clean up after all tests complete.
435- **Registry tests**: Use a dedicated test key under `HKCU:\Software\WMCP-Test-{timestamp}`.
436 Delete the entire key after testing.
437- **Process tests**: Only list processes (don't kill user processes). If testing kill, spawn a
438 sacrificial process first (e.g., `notepad.exe`) and record its PID.
439- **App tests**: Use lightweight apps (Notepad, Calculator). **Record PIDs** of all processes
440 spawned during testing (use `(Start-Process notepad -PassThru).Id` or query process list
441 before/after launch). In teardown, only kill processes by PID — NEVER by name (e.g.,
442 `Stop-Process -Id $pid`, not `Stop-Process -Name notepad`), because the user may have their
443 own instances of the same application running.
444 - **Modern tabbed apps caveat:** Windows 11 Notepad/Terminal may reuse a single process for
445 multiple tabs. Diff the process list before/after each launch to detect new PIDs. Only kill
446 PIDs that did not exist before testing began.
447- **Clipboard tests**: Save and restore the original clipboard content.
448- **Input tools (Click, Type, Scroll, Move, Shortcut)**: Open a dedicated test window
449 (e.g., Notepad) to receive input. Don't interact with user's active work.
450 See also Pre-Test Step 2 for IME state handling — switch to English input mode before testing
451 and restore the original state in final teardown.
452- **Read-only tools (Screenshot, Snapshot, Scrape)**: Safe to run freely.
453- **Notification tests**: User-visible (sends Windows toasts). Avoid repeated or unnecessary
454 notifications. Prefer a single clearly labeled test notification per test case.
455- **PowerShell tests**: Use read-only commands where possible.
456
457---
458
459## Tool-Specific Testing Guidance
460
461Hints per tool. Always read the actual schema to discover additional scenarios beyond these.
462
463### App
464- Modes: launch, resize, switch. Test each mode.
465- Launch: test with known apps (notepad, calc), unknown app names
466- Resize: test with valid window_loc/window_size, without an active app
467- Switch: test switching to a running app, to a non-existent app
468
469### PowerShell
470- Simple commands: `echo "hello"`, `Get-Date`, `Get-Process | Select-Object -First 3`
471- Timeout behavior: set a very short timeout with a long-running command
472- Encoding: commands with Unicode output
473- Error output: commands that write to stderr
474- Exit codes: commands that fail (e.g., `Get-Item nonexistent`)
475
476### Screenshot
477- Default parameters (no args)
478- With annotation enabled/disabled
479- With reference lines
480- With specific display index
481- Verify return includes image data
482
483### Snapshot
484- Various flag combinations: use_vision, use_dom, use_annotation, use_ui_tree
485- All flags off vs. all flags on
486- With/without reference lines
487
488### Click
489- By coordinates (loc) vs. by label
490- Different button types: left, right, middle
491- clicks=0 (hover), clicks=1 (single), clicks=2 (double)
492- Invalid coordinates (negative, off-screen)
493- Invalid label (non-existent element ID)
494
495### Type
496- Normal text, Unicode text, special characters
497- With and without clear=true
498- With and without press_enter=true
499- Different caret_position values: start, idle, end
500- By coordinates vs. by label
501- **IME sensitivity**: Test with IME active to verify behavior (expect failure if tool uses
502 keystroke simulation rather than Unicode input). This is a high-value edge case because
503 many Windows machines have non-English IMEs installed.
504- **Emoji / surrogate pair characters**: Test with characters outside the Basic Multilingual
505 Plane (e.g., 🌍, 😀) to verify supplementary plane Unicode support.
506- **Empty string**: Test `text=""` — this is a common edge case that may crash if the
507 implementation indexes into the string without a length check.
508
509### Scroll
510- Vertical up/down, horizontal left/right
511- Different wheel_times values (1, 5, 10)
512- By coordinates vs. by label
513
514### Move
515- Simple move to coordinates
516- Drag mode (drag=true)
517- By coordinates vs. by label
518
519### Shortcut
520- Common shortcuts: ctrl+c, ctrl+v, ctrl+a, alt+tab
521- Windows key shortcuts: win+r, win+d
522- Multi-key combinations
523- Invalid key names
524
525### Wait
526- Short duration (1 second)
527- Zero duration
528- Verify actual elapsed time roughly matches requested duration
529
530### MultiSelect
531- Select multiple items by coordinates
532- Select by labels
533- With and without press_ctrl
534- Empty list of items
535
536### MultiEdit
537- Edit multiple fields by coordinates
538- Edit by labels
539- Mixed valid and invalid targets
540
541### Clipboard
542- get mode when clipboard has text
543- get mode when clipboard is empty
544- set mode with normal text
545- set mode with Unicode text
546- Roundtrip: set then get, verify content matches
547
548### Process
549- list mode with default sort
550- list mode with different sort_by values (memory, cpu, name)
551- list mode with name filter
552- list mode with different limit values
553- kill mode with a sacrificial process (spawn notepad, then kill it)
554
555### Notification
556- Valid notification with title, message, app_id
557- Empty title or message
558- Special characters in title/message
559
560### FileSystem
561- Full mode coverage: read, write, copy, move, delete, list, search, info
562- Read: existing file, non-existent file, offset/limit, different encodings
563- Write: new file, overwrite, append
564- List: with and without pattern, recursive, show_hidden
565- Delete: file, empty dir, non-empty dir with recursive
566
567### Registry
568- Full mode coverage: get, set, delete, list
569- Set and get roundtrip
570- Different value types (String, DWord, QWord)
571- Non-existent key/value
572- Use test-only registry path
573
574### Scrape
575- With a URL (lightweight page)
576- With and without query parameter
577- With use_dom enabled (requires open browser)
578- Invalid URL