# Capture local screen and audio context so agents can search what happened on your device

> Use Screenpipe when an agent needs private, local-first memory of what you saw or heard on your computer, including searchable screen text, app context, and transcripts, instead of relying on a chat-only memory layer.

- Skill: `agentskillexchange/capture-local-screen-and-audio-context-so-agents-can-search-` (Agent Skill)
- Install (CLI): `npx skillmds@latest add agentskillexchange/capture-local-screen-and-audio-context-so-agents-can-search-`
- Raw SKILL.md: https://api.skillmd.com/api/skills/agentskillexchange/capture-local-screen-and-audio-context-so-agents-can-search-/raw
- Safety review: pending (external: skill-scanner PASS, skillspector PASS)
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: AI & ML
- Author: agentskillexchange (https://skillmd.com/u/agentskillexchange)
- Updated: 2026-09-08
- Page: https://skillmd.com/skills/agentskillexchange/capture-local-screen-and-audio-context-so-agents-can-search-

---


# Capture local screen and audio context so agents can search what happened on your device

Use Screenpipe when an agent needs private, local-first memory of what you saw or heard on your computer, including searchable screen text, app context, and transcripts, instead of relying on a chat-only memory layer.

## Prerequisites

Screenpipe desktop app or source build, local screen and audio permissions, sufficient local storage, and optionally an MCP-compatible agent client

## Installation

Use the upstream install or setup path that matches your environment:
- Make sure to understand the main branch is moving fast and breaking things, if you're looking for a stable version check app releases https://github.com/screenpipe/screenpipe/releases and use the git commit accordingl...

Basic usage or getting-started notes:
- Minimum requirements: 8 GB RAM recommended. ~5–10 GB disk space per month. CPU usage typically 5–10% on modern hardware thanks to event-driven capture.
- Typical CPU usage is 5–10% on modern hardware. Event-driven capture only processes frames when something changes, and accessibility tree extraction is much lighter than OCR.

- Source: https://github.com/screenpipe/screenpipe
- Extracted from upstream docs: https://raw.githubusercontent.com/screenpipe/screenpipe/HEAD/README.md

## Documentation

- https://docs.screenpi.pe

## Source

- [Agent Skill Exchange](https://agentskillexchange.com/skills/capture-local-screen-and-audio-context-so-agents-can-search-what-happened-on-your-device/)

