Quality Flywheel

Evaluate and improve GenAI models and agents using the Google GenAI Evaluation SDK. Creates eval datasets (from session traces or synthetic generation), selects and configures metrics (RubricMetric, LLMMetric, CodeExecutionMetric), executes evals via client.evals.evaluate(), and analyzes results to suggest concrete fixes. Supports both single-turn model evaluation and multi-turn agent trajectory evaluation. Use when asked to "evaluate my agent", "evaluate my model", "create eval dataset", "run evals", "analyze eval results", "which metrics should I use", "generate test data", or "improve quality".

GoogleCloudPlatform Updated

File contents

GoogleCloudPlatform/vertex-ai-samples/tree/main/skills/quality-flywheel commit aa6f11425f

Frequently asked questions

npx skillmds@latest add googlecloudplatform-vertex-ai-samples/quality-flywheel