Run Tests via Copilot Studio Kit
Run a batch test suite against a published Copilot Studio agent using the Power CAT Copilot Studio Kit.
Prerequisites
The user must have:
- The Copilot Studio Kit installed in their Power Platform environment
- Published their agent in the Copilot Studio UI
- Created a test set in the Copilot Studio Kit
- An Azure App Registration with Dataverse permissions
Phase 1: Configure Settings
Read tests/settings.json (relative to the user's project CWD) and check for missing or placeholder values (containing YOUR_).
If the file doesn't exist, create it from the template:
cp ${CLAUDE_SKILL_DIR}/../../tests/settings-example.json ./tests/settings.json
If values are missing, ask the user for each missing value. Explain where to find each one:
- Environment URL (
dataverse.environmentUrl): "What is your Dataverse environment URL? Find it in Power Platform admin center or Copilot Studio > Settings > Session Details. It looks like https://orgXXXXXX.crm.dynamics.com"
- Tenant ID (
dataverse.tenantId): "What is your Azure tenant ID? Find it in Azure Portal > Microsoft Entra ID > Overview. It's a GUID like c87f36f7-fc65-453c-9019-0d724f21bc42"
- Client ID (
dataverse.clientId): "What is your App Registration client ID? Find it in Azure Portal > App Registrations > your app > Application (client) ID. It's a GUID."
- Agent Configuration ID (
testRun.agentConfigurationId): "What is your agent configuration ID? In Copilot Studio, go to your agent > Tests tab. The ID is a GUID found in the URL or test configuration."
- Test Set ID (
testRun.agentTestSetId): "What is your test set ID? In Copilot Studio, go to your agent > Tests tab > select your test set. The ID is a GUID found in the URL."
Ask for ALL missing values at once (don't ask one at a time).
Write tests/settings.json with the collected values:
{
"dataverse": {
"environmentUrl": "<value>",
"tenantId": "<value>",
"clientId": "<value>"
},
"testRun": {
"agentConfigurationId": "<value>",
"agentTestSetId": "<value>"
}
}
If all values are already configured and valid, proceed to Phase 2.
Phase 2: Run Tests
Ensure tests/package.json exists in the user's project. If not, copy it:
cp ${CLAUDE_SKILL_DIR}/../../tests/package.json ./tests/package.json
Install dependencies if tests/node_modules/ doesn't exist:
npm install --prefix tests
Run the test script in the background with a 100-minute timeout (6000000ms):
node ${CLAUDE_SKILL_DIR}/../../tests/run-tests.js --config-dir ./tests
Use run_in_background: true for this command. Save the returned task ID.
Wait 10 seconds, then check the background task output (non-blocking check).
Detect the authentication state from the output:
If the output contains "Using cached token": Authentication succeeded automatically. Tell the user: "Authentication successful (cached credentials). Tests are running, this may take several minutes..."
If the output contains "use a web browser to open the page": Extract the URL and device code from the message. Present this prominently to the user:
Authentication Required
Open your browser to: https://microsoft.com/devicelogin
Enter the code: XXXXXXXXX (extract the actual code from the output)
After signing in, the tests will continue automatically.
If the output contains an error: Report the error to the user and stop.
If the output is empty or incomplete: Wait another 10 seconds and check again (retry up to 3 times).
Wait for the background task to complete (blocking). The script polls every 20 seconds until all tests finish and downloads results as a CSV.
Read the final output to get the success rate and CSV filename.
Proceed to Phase 3.
Phase 3: Analyze Results
Get the results: Glob: tests/test-results-*.csv — read the most recent CSV file (newest by modification time).
Parse the CSV columns:
| Column |
Meaning |
| Test Utterance |
The user message that was tested |
| Expected Response |
What the test expected |
| Response |
What the agent actually responded |
| Latency (ms) |
Response time |
| Result |
Success, Failed, Unknown, Error, or Pending |
| Test Type |
Response Match, Topic Match, Generative Answers, Multi-turn, Plan Validation, or Attachments |
| Result Reason |
Why the test passed or failed |
Focus on failed tests (Result = Failed or Error). For each failure, analyze:
- Test Type = Topic Match: The wrong topic was triggered, or no topic matched. Check trigger phrases and model descriptions.
- Test Type = Response Match: The response didn't match expected. Check
SendActivity messages, instructions, or generative answer config.
- Test Type = Generative Answers: The generative answer was incorrect or missing. Check knowledge sources,
SearchAndSummarizeContent, and agent instructions.
- Test Type = Plan Validation: The orchestrator's plan was wrong. Check topic descriptions and agent-level instructions.
- Test Type = Multi-turn: A multi-turn conversation failed. Check topic flow, variable handling, and conditions.
Proceed to Phase 4 (Propose Fixes).
Phase 4: Propose Fixes
For each failure, identify the relevant YAML file(s):
- Auto-discover the agent:
Glob: **/agent.mcs.yml
- Find the relevant topic by matching the test utterance against trigger phrases and model descriptions
- Read the topic file to understand the current flow
Propose specific YAML changes to fix each failure. Present them to the user as a summary:
- Which test(s) failed and why
- Which file(s) need changes
- What the proposed change is (show the diff)
Wait for user decision. The user can:
- Accept all — apply all proposed changes
- Accept partially — apply only some changes (ask which ones)
- Reject — discard proposed changes and discuss alternative approaches
Apply accepted changes using the Edit tool. After applying, remind the user to push and publish again before re-running tests.
Test Result Codes Reference
Result: 1=Success, 2=Failed, 3=Unknown, 4=Error, 5=Pending
Test Type: 1=Response Match, 2=Topic Match, 3=Attachments, 4=Generative Answers, 5=Multi-turn, 6=Plan Validation
Run Status: 1=Not Run, 2=Running, 3=Complete, 4=Not Available, 5=Pending, 6=Error
1---2name: run-tests-kit3description: Run a batch test suite via the Copilot Studio Kit (Dataverse API). Uses the Power CAT Copilot Studio Kit to execute test cases against a published agent and produces pass/fail results with latencies. Requires the Kit installed in the environment, an App Registration with Dataverse permissions, and a published agent.4---56# Run Tests via Copilot Studio Kit78Run a batch test suite against a **published** Copilot Studio agent using the [Power CAT Copilot Studio Kit](https://github.com/microsoft/Power-CAT-Copilot-Studio-Kit).910## Prerequisites1112The user must have:131. The **[Copilot Studio Kit](https://github.com/microsoft/Power-CAT-Copilot-Studio-Kit)** installed in their Power Platform environment142. **Published** their agent in the Copilot Studio UI153. **Created a test set** in the Copilot Studio Kit164. An **Azure App Registration** with Dataverse permissions1718## Phase 1: Configure Settings19201. **Read `tests/settings.json`** (relative to the user's project CWD) and check for missing or placeholder values (containing `YOUR_`).21222. **If the file doesn't exist**, create it from the template:23 ```bash24 cp ${CLAUDE_SKILL_DIR}/../../tests/settings-example.json ./tests/settings.json25 ```26273. **If values are missing**, ask the user for each missing value. Explain where to find each one:2829 - **Environment URL** (`dataverse.environmentUrl`): "What is your Dataverse environment URL? Find it in Power Platform admin center or Copilot Studio > Settings > Session Details. It looks like `https://orgXXXXXX.crm.dynamics.com`"30 - **Tenant ID** (`dataverse.tenantId`): "What is your Azure tenant ID? Find it in Azure Portal > Microsoft Entra ID > Overview. It's a GUID like `c87f36f7-fc65-453c-9019-0d724f21bc42`"31 - **Client ID** (`dataverse.clientId`): "What is your App Registration client ID? Find it in Azure Portal > App Registrations > your app > Application (client) ID. It's a GUID."32 - **Agent Configuration ID** (`testRun.agentConfigurationId`): "What is your agent configuration ID? In Copilot Studio, go to your agent > Tests tab. The ID is a GUID found in the URL or test configuration."33 - **Test Set ID** (`testRun.agentTestSetId`): "What is your test set ID? In Copilot Studio, go to your agent > Tests tab > select your test set. The ID is a GUID found in the URL."3435 Ask for ALL missing values at once (don't ask one at a time).36374. **Write `tests/settings.json`** with the collected values:38 ```json39 {40 "dataverse": {41 "environmentUrl": "<value>",42 "tenantId": "<value>",43 "clientId": "<value>"44 },45 "testRun": {46 "agentConfigurationId": "<value>",47 "agentTestSetId": "<value>"48 }49 }50 ```51525. If all values are already configured and valid, proceed to Phase 2.5354## Phase 2: Run Tests55561. **Ensure `tests/package.json` exists** in the user's project. If not, copy it:57 ```bash58 cp ${CLAUDE_SKILL_DIR}/../../tests/package.json ./tests/package.json59 ```60612. **Install dependencies** if `tests/node_modules/` doesn't exist:62 ```bash63 npm install --prefix tests64 ```65663. **Run the test script in the background** with a 100-minute timeout (6000000ms):67 ```bash68 node ${CLAUDE_SKILL_DIR}/../../tests/run-tests.js --config-dir ./tests69 ```70 Use `run_in_background: true` for this command. Save the returned task ID.71724. **Wait 10 seconds**, then check the background task output (non-blocking check).73745. **Detect the authentication state** from the output:7576 - If the output contains **"Using cached token"**: Authentication succeeded automatically. Tell the user: "Authentication successful (cached credentials). Tests are running, this may take several minutes..."7778 - If the output contains **"use a web browser to open the page"**: Extract the URL and device code from the message. **Present this prominently to the user**:7980 > **Authentication Required**81 >82 > Open your browser to: https://microsoft.com/devicelogin83 > Enter the code: **XXXXXXXXX** (extract the actual code from the output)84 >85 > After signing in, the tests will continue automatically.8687 - If the output contains an **error**: Report the error to the user and stop.8889 - If the output is empty or incomplete: Wait another 10 seconds and check again (retry up to 3 times).90916. **Wait for the background task to complete** (blocking). The script polls every 20 seconds until all tests finish and downloads results as a CSV.92937. **Read the final output** to get the success rate and CSV filename.94958. Proceed to **Phase 3**.9697## Phase 3: Analyze Results98991. **Get the results**: `Glob: tests/test-results-*.csv` — read the most recent CSV file (newest by modification time).1001012. **Parse the CSV columns**:102 | Column | Meaning |103 |--------|---------|104 | Test Utterance | The user message that was tested |105 | Expected Response | What the test expected |106 | Response | What the agent actually responded |107 | Latency (ms) | Response time |108 | Result | `Success`, `Failed`, `Unknown`, `Error`, or `Pending` |109 | Test Type | `Response Match`, `Topic Match`, `Generative Answers`, `Multi-turn`, `Plan Validation`, or `Attachments` |110 | Result Reason | Why the test passed or failed |1111123. **Focus on failed tests** (Result = `Failed` or `Error`). For each failure, analyze:113 - **Test Type = Topic Match**: The wrong topic was triggered, or no topic matched. Check trigger phrases and model descriptions.114 - **Test Type = Response Match**: The response didn't match expected. Check `SendActivity` messages, instructions, or generative answer config.115 - **Test Type = Generative Answers**: The generative answer was incorrect or missing. Check knowledge sources, `SearchAndSummarizeContent`, and agent instructions.116 - **Test Type = Plan Validation**: The orchestrator's plan was wrong. Check topic descriptions and agent-level instructions.117 - **Test Type = Multi-turn**: A multi-turn conversation failed. Check topic flow, variable handling, and conditions.1181194. Proceed to **Phase 4** (Propose Fixes).120121## Phase 4: Propose Fixes1221231. **For each failure, identify the relevant YAML file(s)**:124 - Auto-discover the agent: `Glob: **/agent.mcs.yml`125 - Find the relevant topic by matching the test utterance against trigger phrases and model descriptions126 - Read the topic file to understand the current flow1271282. **Propose specific YAML changes** to fix each failure. Present them to the user as a summary:129 - Which test(s) failed and why130 - Which file(s) need changes131 - What the proposed change is (show the diff)1321333. **Wait for user decision**. The user can:134 - **Accept all** — apply all proposed changes135 - **Accept partially** — apply only some changes (ask which ones)136 - **Reject** — discard proposed changes and discuss alternative approaches1371384. **Apply accepted changes** using the Edit tool. After applying, remind the user to push and publish again before re-running tests.139140## Test Result Codes Reference141142```143Result: 1=Success, 2=Failed, 3=Unknown, 4=Error, 5=Pending144Test Type: 1=Response Match, 2=Topic Match, 3=Attachments, 4=Generative Answers, 5=Multi-turn, 6=Plan Validation145Run Status: 1=Not Run, 2=Running, 3=Complete, 4=Not Available, 5=Pending, 6=Error146```