Trigger: Use when managing Alert Configuration on Google Cloud's Agent Platform — Google Cloud AI and agent infrastructure.
Agent Platform Alert Configuration
Critical Steps
1. Safety & Confirmation Tiers (CRITICAL)
Before executing any commands or writing configurations on behalf of the user,
you MUST adhere to the following safety tiers based on the action requested:
- Tier R: Read-only (
check_telemetry.py)
- Rule: No confirmation needed. You may execute these scripts
immediately to inspect the telemetry status of the Reasoning Engine.
- Tier B: Billing & Resource Creation (
create_online_monitor.py /
provisioning)
- Rule: Explicit User Confirmation Required. These actions incur
additional billing charges and create cloud resources. The agent MUST
ALWAYS warn the user explicitly about the potential extra billing costs
of BOTH the Online Monitor (specifically mentioning LLM evaluations)
and Telemetry (specifically mentioning Cloud Trace/Logging export).
You MUST STOP and ask for explicit approval before proceeding with
provisioning or providing setup commands.
2. Prerequisites & Dependencies
Agent Telemetry
- Disclaimer: For Reliability, Cost, Safety, and Security alerts to
function, the underlying agent MUST be instrumented to emit OpenTelemetry
(OTel) metrics. If the agent does not emit these metrics, the alerting
policies will have no data stream to evaluate.
Python Environment
Before executing any python script in this skill you MUST install the required
dependencies in your environment. Run this command first:
pip install -r scripts/requirements.txt
3. Input Assumptions
- Explicit Project Adherence: You must ONLY configure alerts, query
telemetry, or interact with the Google Cloud Project(s) explicitly provided
by the user in the prompt. Do NOT assume or use other projects from your
environment or history unless the user explicitly directs you to do so.
- Sequential File Transformations: If the user explicitly asks to copy a
file and then modify it, you MUST perform these actions sequentially (copy
first, then modify) rather than writing the final content directly.
4. Execution Steps
Mandatory Prerequisite Execution Protocol (SEQUENTIAL): Before
generating or writing ANY configuration, you MUST execute these steps in
order:
- Step 1: Metric Scope Check: Determine where to deploy policies.
- Action A (CLI): Run
gcloud beta monitoring metrics-scopes list projects/{project_id}. If a scoping project is returned, you MUST
deploy policies there.
- Action B (Code Scan): Search Terraform configurations for
google_monitoring_monitored_project resources to extract the
scoping project.
- Action C (Fallback): If ambiguous, ASK the user: "Are you using
a multi-project Cloud Monitoring Metric Scope? If so, what is the
scoping project ID?"
- Step 2: Pre-existing Policies Check: Avoid duplicates.
- Action: Scan the target directory to see if aggregated policies
already exist targeting the same metrics (grouped by
reasoning_engine_id or gen_ai_agent_name). Use
validate_config.py --directory to verify.
Alert Policy Type Resource Files: You MUST list and read files under
references/ with names ending in _alert_policies.md to learn how to
configure alert policies based on type. By default you should configure all
of the following alert types UNLESS the user requests to generate explicit
alert policies and/or types. Follow their tables of content to help you find
the reference sections you need to read:
| Alert Type |
Reference File |
| Reliability |
reliability_alert_policies.md |
| Quality |
quality_alert_policies.md |
| Cost |
cost_alert_policies.md |
| Safety |
safety_alert_policies.md |
| Security |
security_alert_policies.md |
5. Outputs & Formats
- Always configure the supported alerting policies for the target agent:
- For Reliability Monitoring: You MUST configure exactly five alerting
policies:
- Latency (anomaly monitoring)
- Error Rate - Fast Burn SLO (1-Hour Window)
- Error Rate - Slow Burn SLO (3-Day Window)
- Model Call Error Rate (SQL-based Log Analytics Alerting)
- Tool Call Error Rate (SQL-based Log Analytics Alerting)
- For Quality Monitoring: You MUST configure exactly three alerting
policies (Requires Vertex AI Online Monitors):
- Final Response Quality
- Tool Use Quality
- Hallucination
- For Cost Monitoring: You MUST configure exactly one cost alerting
policy:
- Rapid Token Burn Rate (anomaly monitoring)
- For Safety Monitoring: You MUST configure exactly one safety
alerting policy:
- High Model Armor Safety Policy Trigger Rate (SQL-based Log
Analytics Alerting)
- For Security Monitoring: You MUST configure exactly one security
alerting policy:
- High IAM Permission Denied Trigger Rate (SQL-based Log Analytics
Alerting)
- Terraform Only: Write the generated observability configuration ONLY as
Terraform (
.tf) files (e.g., alerts.tf, variables.tf).
- You ONLY need to install Terraform if you're asked to deploy the
alerts AND there is no valid Terraform install. SQL-based alerting using
condition_sql requires the provider version >= 6.0.0 (or late 5.x
versions supporting the feature).
- If you are NOT asked to deploy the alerts you do not need to install
terraform.
- Dynamic Multi-Resource Alerting (No Single-Resource Pinning): You MUST
NOT hardcode specific agent IDs or resource name filters (e.g.,
{gen_ai_agent_name="{agent_name}"} or
metric.labels.agent_resource_name="{agent_name}") in alerting conditions
unless explicitly requested (e.g., "ONLY for this agent"). Merely mentioning
a specific agent name or ID in the request does NOT constitute an explicit
request to pin/filter; you MUST still default to dynamic grouping to cover
all agents. To cover all active agents in the project dynamically:
- For Reliability Metrics using PromQL: ALWAYS use grouping
aggregations. Group by
gen_ai_agent_name (e.g., by (gen_ai_agent_name)). Avoid filtering to a single ID/Name unless
requested.
- For Quality Metrics using Standard Threshold Filters: Omit the
agent_resource_name filter entirely. Configure the condition filter to
only target the monitored resource type
(aiplatform.googleapis.com/OnlineEvaluator) and metric type
(aiplatform.googleapis.com/online_evaluator/scores) globally for the
project.
- For Downstream Calls using SQL: Omit the
ENDS_WITH filter
targeting a specific agent name. Instead, extract the agent identifier
(e.g., JSON_VALUE(resource.attributes, '$."cloud.resource_id"')) and
add it to the GROUP BY clause alongside the model or tool name.
- Directory Inference: Prefer the path explicitly provided by the user (if
any). Otherwise, deploy configuration files to target Terraform or SRE
folders (e.g.
monitoring/, ops/, sre/). Use tools to locate where
alert policies or state pointers exist in the project, rather than blindly
writing to the root.
- Notification Channels: By default, never configure any notification
channels without user input. If the user explicitly provides a notification
channel in their prompt, configure the alerts to use it. If no notification
channel is provided, you MUST explicitly ask the user in your final response
if they would like to configure notification channels. This is a mandatory
question and you MUST NOT omit it from your response. IMPORTANT Do NOT
make assumptions about notification channels. If you search the codebase for
a notification channel you must ALWAYS confirm with the user before using
it.
- Plain English Response: You MUST include a plain English explanation for
what the alerts do in your response. This must explain in plain English what
the alert measures, how the algorithm works, and what a trigger indicates.
6. Output Verification
- Background Task Cleanup: You MUST check the status of all background
tasks that you spawn. Before completing your execution and returning your
final response, you MUST terminate or kill any active or hanging background
tasks (using the
manage_task tool with action kill).
- Validate Configuration: Run the Config Linting tool to make sure all
the output files are written with the correct grammar and structure. See
details about the tool in the
Tooling Scripts section below.
Tooling Scripts
Use the following scripts to resolve duplicates and validate configs before
presenting or applying Terraform changes:
- Duplicate Check & Merge: Checks for pre-existing alerts in the target
folder to ensure changes are merged in-place rather than appended:
- Command:
python3 scripts/validate_config.py --directory {target_tf_dir} --engine-var '${var.gen_ai_agent_name}'
- Config Linting: Validates PromQL grammar, matching engine labels, and
HCL structure:
- Command:
python3 scripts/validate_config.py --file {path_to_tf_file}
- Self-Correction Loop: If validation fails (exits non-zero or outputs
errors), you MUST read the command output, locate the line/file
containing the lint error, analyze the PromQL syntax or Terraform HCL
issue, apply adjustments in-place, and re-run the
validate_config.py --file validation. Repeat this loop until the validation script passes
successfully.
Gotchas & Behavioral Corrections
- Raw Error Boundaries: Explain that raw error counts or absolute failed
request count boundaries do not scale under changing traffic throughput.
Recommend ratio-based error rate alerts instead.
- Safe Threshold Modulation E2E Validation: When verifying a dynamic
metric threshold policy end-to-end, do NOT attempt to force real platform
errors. Instead, deploy the alert policy with standard safe bounds (Z-score
multiplier > 15), then temporarily update standard deviation Z-score limits
to a negative value (e.g. > -3) to trigger/verify the "Firing" state before
reverting. Always get confirmation before taking this action proactively.
- Expected Script Failures:
validate_config.py --directory exiting with code 1: Parse the JSON
output for duplicate resource targets. Perform in-place upgrade edits,
then re-check until it passes with 0.
- Script Execution Failures & Self-Correction: If script execution
fails unexpectedly, you MUST read and inspect the stdout/stderr logs or
error output. Analyze the error message and attempt to dynamically
correct parameters and retry execution before escalating or
falling back to manual plans. Consult the relevant domain-specific
reference file for detailed troubleshooting steps for specific scripts.
- Distribution Metric Aligner Constraint: Standard
ALIGN_MEAN cannot be
applied to DELTA distribution metrics like online_evaluator/scores. You
MUST use percentile-based aligners (like ALIGN_PERCENTILE_50) to reduce
the score distribution into a comparable numeric stream.
- HCL Heredoc Interpolation: When referencing Terraform variables inside
PromQL or SQL queries (which are defined as strings), you MUST use the
${var.variable_name} syntax. Bare references like var.variable_name will
fail at deployment time.
- Avoid Recursive Directory Operations: You MUST NOT run recursive listing
or search commands (such as
ls -R, find ., or raw recursive grep) from
the repository root if it contains a very large number of files, as this
will freeze your session. Always target specific subdirectories.
Supporting Links
1---2name: agent-platform-alert-configuration3description: **Trigger**: Use when managing Alert Configuration on Google Cloud's Agent Platform — Google Cloud AI and agent infrastructure.4---56**Trigger**: Use when managing Alert Configuration on Google Cloud's Agent Platform — Google Cloud AI and agent infrastructure.78# Agent Platform Alert Configuration910## Critical Steps1112### 1. Safety & Confirmation Tiers (CRITICAL)1314Before executing any commands or writing configurations on behalf of the user,15you MUST adhere to the following safety tiers based on the action requested:16171. **Tier R: Read-only (`check_telemetry.py`)**18 * **Rule**: No confirmation needed. You may execute these scripts19 immediately to inspect the telemetry status of the Reasoning Engine.202. **Tier B: Billing & Resource Creation (`create_online_monitor.py` /21 provisioning)**22 * **Rule**: **Explicit User Confirmation Required**. These actions incur23 additional billing charges and create cloud resources. The agent MUST24 ALWAYS warn the user explicitly about the potential extra billing costs25 of BOTH the Online Monitor (specifically mentioning **LLM evaluations**)26 and Telemetry (specifically mentioning **Cloud Trace/Logging export**).27 You MUST STOP and ask for explicit approval before proceeding with28 provisioning or providing setup commands.2930### 2. Prerequisites & Dependencies3132#### Agent Telemetry3334* **Disclaimer**: For Reliability, Cost, Safety, and Security alerts to35 function, the underlying agent MUST be instrumented to emit OpenTelemetry36 (OTel) metrics. If the agent does not emit these metrics, the alerting37 policies will have no data stream to evaluate.3839#### Python Environment4041Before executing any python script in this skill you MUST install the required42dependencies in your environment. Run this command first:4344```bash45pip install -r scripts/requirements.txt46```4748### 3. Input Assumptions4950* **Explicit Project Adherence**: You must ONLY configure alerts, query51 telemetry, or interact with the Google Cloud Project(s) explicitly provided52 by the user in the prompt. Do NOT assume or use other projects from your53 environment or history unless the user explicitly directs you to do so.54* **Sequential File Transformations**: If the user explicitly asks to copy a55 file and then modify it, you MUST perform these actions sequentially (copy56 first, then modify) rather than writing the final content directly.5758### 4. Execution Steps59601. **Mandatory Prerequisite Execution Protocol (SEQUENTIAL)**: Before61 generating or writing ANY configuration, you MUST execute these steps in62 order:63 1. **Step 1: Metric Scope Check**: Determine where to deploy policies.64 * **Action A (CLI)**: Run `gcloud beta monitoring metrics-scopes list65 projects/{project_id}`. If a scoping project is returned, you MUST66 deploy policies there.67 * **Action B (Code Scan)**: Search Terraform configurations for68 `google_monitoring_monitored_project` resources to extract the69 scoping project.70 * **Action C (Fallback)**: If ambiguous, ASK the user: "Are you using71 a multi-project Cloud Monitoring Metric Scope? If so, what is the72 scoping project ID?"73 2. **Step 2: Pre-existing Policies Check**: Avoid duplicates.74 * **Action**: Scan the target directory to see if aggregated policies75 already exist targeting the same metrics (grouped by76 `reasoning_engine_id` or `gen_ai_agent_name`). Use77 `validate_config.py --directory` to verify.782. **Alert Policy Type Resource Files**: You MUST list and read files under79 `references/` with names ending in `_alert_policies.md` to learn how to80 configure alert policies based on type. By default you should configure all81 of the following alert types UNLESS the user requests to generate explicit82 alert policies and/or types. Follow their tables of content to help you find83 the reference sections you need to read:8485 Alert Type | Reference File86 :-------------- | :-------------87 **Reliability** | [reliability_alert_policies.md](references/reliability_alert_policies.md)88 **Quality** | [quality_alert_policies.md](references/quality_alert_policies.md)89 **Cost** | [cost_alert_policies.md](references/cost_alert_policies.md)90 **Safety** | [safety_alert_policies.md](references/safety_alert_policies.md)91 **Security** | [security_alert_policies.md](references/security_alert_policies.md)9293### 5. Outputs & Formats9495* **Always configure the supported alerting policies** for the target agent:96 * **For Reliability Monitoring**: You MUST configure exactly five alerting97 policies:98 1. **Latency** (anomaly monitoring)99 2. **Error Rate - Fast Burn SLO** (1-Hour Window)100 3. **Error Rate - Slow Burn SLO** (3-Day Window)101 4. **Model Call Error Rate** (SQL-based Log Analytics Alerting)102 5. **Tool Call Error Rate** (SQL-based Log Analytics Alerting)103 * **For Quality Monitoring**: You MUST configure exactly three alerting104 policies (Requires Vertex AI Online Monitors):105 1. **Final Response Quality**106 2. **Tool Use Quality**107 3. **Hallucination**108 * **For Cost Monitoring**: You MUST configure exactly one cost alerting109 policy:110 1. **Rapid Token Burn Rate** (anomaly monitoring)111 * **For Safety Monitoring**: You MUST configure exactly one safety112 alerting policy:113 1. **High Model Armor Safety Policy Trigger Rate** (SQL-based Log114 Analytics Alerting)115 * **For Security Monitoring**: You MUST configure exactly one security116 alerting policy:117 1. **High IAM Permission Denied Trigger Rate** (SQL-based Log Analytics118 Alerting)119* **Terraform Only**: Write the generated observability configuration ONLY as120 Terraform (`.tf`) files (e.g., `alerts.tf`, `variables.tf`).121 - You **ONLY** need to install Terraform if you're asked to deploy the122 alerts AND there is no valid Terraform install. SQL-based alerting using123 `condition_sql` requires the provider version **>= 6.0.0** (or late 5.x124 versions supporting the feature).125 - If you are **NOT** asked to deploy the alerts you do not need to install126 terraform.127* **Dynamic Multi-Resource Alerting (No Single-Resource Pinning)**: You MUST128 NOT hardcode specific agent IDs or resource name filters (e.g.,129 `{gen_ai_agent_name="{agent_name}"}` or130 `metric.labels.agent_resource_name="{agent_name}"`) in alerting conditions131 unless explicitly requested (e.g., "ONLY for this agent"). Merely mentioning132 a specific agent name or ID in the request does NOT constitute an explicit133 request to pin/filter; you MUST still default to dynamic grouping to cover134 all agents. To cover all active agents in the project dynamically:135 * **For Reliability Metrics using PromQL**: ALWAYS use grouping136 aggregations. Group by `gen_ai_agent_name` (e.g., `by137 (gen_ai_agent_name)`). Avoid filtering to a single ID/Name unless138 requested.139 * **For Quality Metrics using Standard Threshold Filters**: Omit the140 `agent_resource_name` filter entirely. Configure the condition filter to141 only target the monitored resource type142 (`aiplatform.googleapis.com/OnlineEvaluator`) and metric type143 (`aiplatform.googleapis.com/online_evaluator/scores`) globally for the144 project.145 * **For Downstream Calls using SQL**: Omit the `ENDS_WITH` filter146 targeting a specific agent name. Instead, extract the agent identifier147 (e.g., `JSON_VALUE(resource.attributes, '$."cloud.resource_id"')`) and148 add it to the `GROUP BY` clause alongside the model or tool name.149* **Directory Inference**: Prefer the path explicitly provided by the user (if150 any). Otherwise, deploy configuration files to target Terraform or SRE151 folders (e.g. `monitoring/`, `ops/`, `sre/`). Use tools to locate where152 alert policies or state pointers exist in the project, rather than blindly153 writing to the root.154* **Notification Channels**: By default, never configure any notification155 channels without user input. If the user explicitly provides a notification156 channel in their prompt, configure the alerts to use it. If no notification157 channel is provided, you MUST explicitly ask the user in your final response158 if they would like to configure notification channels. **This is a mandatory159 question and you MUST NOT omit it from your response.** **IMPORTANT** Do NOT160 make assumptions about notification channels. If you search the codebase for161 a notification channel you must ALWAYS confirm with the user before using162 it.163* **Plain English Response**: You MUST include a plain English explanation for164 what the alerts do in your response. This must explain in plain English what165 the alert measures, how the algorithm works, and what a trigger indicates.166167### 6. Output Verification168169* **Background Task Cleanup**: You MUST check the status of all background170 tasks that you spawn. Before completing your execution and returning your171 final response, you MUST terminate or kill any active or hanging background172 tasks (using the `manage_task` tool with action `kill`).173* **Validate Configuration**: Run the **Config Linting** tool to make sure all174 the output files are written with the correct grammar and structure. See175 details about the tool in the `Tooling Scripts` section below.176177## Tooling Scripts178179Use the following scripts to resolve duplicates and validate configs before180presenting or applying Terraform changes:1811821. **Duplicate Check & Merge**: Checks for pre-existing alerts in the target183 folder to ensure changes are merged in-place rather than appended:184 * Command: `python3 scripts/validate_config.py --directory {target_tf_dir}185 --engine-var '${var.gen_ai_agent_name}'`1862. **Config Linting**: Validates PromQL grammar, matching engine labels, and187 HCL structure:188 * Command: `python3 scripts/validate_config.py --file {path_to_tf_file}`189 * **Self-Correction Loop**: If validation fails (exits non-zero or outputs190 errors), you MUST read the command output, locate the line/file191 containing the lint error, analyze the PromQL syntax or Terraform HCL192 issue, apply adjustments in-place, and re-run the `validate_config.py193 --file` validation. Repeat this loop until the validation script passes194 successfully.195196## Gotchas & Behavioral Corrections197198* **Raw Error Boundaries**: Explain that raw error counts or absolute failed199 request count boundaries do not scale under changing traffic throughput.200 Recommend ratio-based error rate alerts instead.201* **Safe Threshold Modulation E2E Validation**: When verifying a dynamic202 metric threshold policy end-to-end, do NOT attempt to force real platform203 errors. Instead, deploy the alert policy with standard safe bounds (Z-score204 multiplier > 15), then temporarily update standard deviation Z-score limits205 to a negative value (e.g. > -3) to trigger/verify the "Firing" state before206 reverting. Always get confirmation before taking this action proactively.207* **Expected Script Failures**:208 * `validate_config.py --directory` exiting with code 1: Parse the JSON209 output for duplicate resource targets. Perform in-place upgrade edits,210 then re-check until it passes with 0.211 * **Script Execution Failures & Self-Correction**: If script execution212 fails unexpectedly, you MUST read and inspect the stdout/stderr logs or213 error output. Analyze the error message and attempt to dynamically214 correct parameters and retry execution before escalating or215 falling back to manual plans. Consult the relevant domain-specific216 reference file for detailed troubleshooting steps for specific scripts.217* **Distribution Metric Aligner Constraint**: Standard `ALIGN_MEAN` cannot be218 applied to `DELTA` distribution metrics like `online_evaluator/scores`. You219 MUST use percentile-based aligners (like `ALIGN_PERCENTILE_50`) to reduce220 the score distribution into a comparable numeric stream.221* **HCL Heredoc Interpolation**: When referencing Terraform variables inside222 PromQL or SQL queries (which are defined as strings), you MUST use the223 ${var.variable_name} syntax. Bare references like var.variable_name will224 fail at deployment time.225* **Avoid Recursive Directory Operations**: You MUST NOT run recursive listing226 or search commands (such as `ls -R`, `find .`, or raw recursive `grep`) from227 the repository root if it contains a very large number of files, as this228 will freeze your session. Always target specific subdirectories.229230## Supporting Links231232* [Continuous evaluation with online monitors](https://docs.cloud.google.com/gemini-enterprise-agent-platform/optimize/evaluation/evaluate-online)233* [Agent Platform Quality Metrics](https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/rubric-metric-details)234* [Google Cloud Alerting Policies Guide](https://cloud.google.com/monitoring/alerts)235* [Google Cloud Monitoring PromQL Documentation](https://cloud.google.com/monitoring/promql)