# Trace and evaluate agent runs with MLflow

> Instrument LLM and agent applications with MLflow tracing, evaluation, prompt tracking, and monitoring so operators can debug behavior before and after deployment.

- Skill: `agentskillexchange/trace-and-evaluate-agent-runs-with-mlflow` (Agent Skill)
- Install (CLI): `npx skillmds@latest add agentskillexchange/trace-and-evaluate-agent-runs-with-mlflow`
- Raw SKILL.md: https://api.skillmd.com/api/skills/agentskillexchange/trace-and-evaluate-agent-runs-with-mlflow/raw
- Safety review: pending (external: skill-scanner PASS, skillspector PASS)
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: AI & ML
- Author: agentskillexchange (https://skillmd.com/u/agentskillexchange)
- Updated: 2026-09-08
- Page: https://skillmd.com/skills/agentskillexchange/trace-and-evaluate-agent-runs-with-mlflow

---


# Trace and evaluate agent runs with MLflow

Instrument LLM and agent applications with MLflow tracing, evaluation, prompt tracking, and monitoring so operators can debug behavior before and after deployment.

## Prerequisites

Python, uv or pip, MLflow tracking server, LLM or agent application

## Installation

Requirements and caveats from upstream:
- [![Python SDK](https://img.shields.io/pypi/v/mlflow)](https://pypi.org/project/mlflow/)
- python
- MLflow provides everything you need to build, debug, evaluate, and deploy production-quality LLM applications and AI agents. Supports Python, TypeScript/JavaScript, Java and any other programming language. MLflow also...

Basic usage or getting-started notes:
- **3. Run Your Code**
- <a href="https://mlflow.org/docs/latest/genai/tracing/quickstart/">Getting Started →</a>
- <div>Run systematic evaluations, track quality metrics over time, and catch regressions before they reach production. Choose from 50+ built-in metrics and LLM judges, or define your own.</div><br>

- Source: https://github.com/mlflow/mlflow
- Extracted from upstream docs: https://raw.githubusercontent.com/mlflow/mlflow/HEAD/README.md

## Documentation

- https://mlflow.org/docs/latest/genai/

## Source

- [Agent Skill Exchange](https://agentskillexchange.com/skills/trace-and-evaluate-agent-runs-with-mlflow/)

