Agent Observability Auto Experiment

Run an iterative code-improvement hill-climb against real Datadog LLM-Obs data, locally, with Claude Code as the agent. Establishes a baseline eval, makes one focused change, re-scores with the same harness, keeps the change only if it beats the best, and repeats. Use when the user says "run an auto experiment", "hill-climb this code", "iteratively improve X and measure the delta", "optimize this prompt/file against my traces", "auto-optimize against LLM-Obs", or wants the local equivalent of the auto_experiments worker. Works from an ml_app, a dataset_id, or a list of trace_ids.

om-scogo c76ce4b 4 files · 41.9 KB Updated 0 repo stars

File contents

om-scogo/skillsh-scraper/tree/main/data/skills/datadog-labs-agent-skills/agent-skills commit c76ce4b049

Frequently asked questions

npx skillmds add om-scogo/agent-observability-auto-experiment