# File Management

> Search, get, create, upload (single & multipart), update, and delete files in Spuree projects with checksum-verified upload flow

- Skill: `cheehoolabs/file-management` (Agent Skill)
- Install (CLI): `npx skillmds@latest add cheehoolabs/file-management`
- Raw SKILL.md: https://api.skillmd.com/api/skills/cheehoolabs/file-management/raw
- Safety review: pending
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: Coding & Dev Tools
- Author: cheehoolabs (https://skillmd.com/u/cheehoolabs)
- Updated: 2026-09-17
- Page: https://skillmd.com/skills/cheehoolabs/file-management

---


# File Management

## Overview

Spuree is an agent-friendly cloud storage. Projects contain folders (nestable) and files at any level. This skill manages files — they can live directly under a project or inside any folder. In the API, projects and folders are both called **sessions** (`sessionType`: `creative_project` = project, `session` = folder).

Use this skill when an agent needs to:

- Search for files and folders by name across all projects
- Get a file by ID with download URL
- Upload new files to a project or folder
- Update file metadata (rename, move) or content (with conflict detection)
- Delete files (soft delete)

## Authentication

```
Authorization: Bearer $SPUREE_ACCESS_TOKEN
```

Or: `X-API-Key: $SPUREE_API_KEY`. See the **authentication** skill.

## Base URLs

| Base URL | Endpoints |
| --- | --- |
| `https://data.spuree.com/api/v1/files` | File CRUD |
| `https://data.spuree.com/api/v1/search` | Cross-project search |

## Endpoints

### GET /v1/search

Full-text search across files, folders, projects, and assets. Backed by the `content_index` collection (Atlas `$search`), kept in sync via MongoDB Change Streams.

Search is Atlas analyzer-based rather than substring matching. It returns one
grouped item per source object, with the matching indexed rows in `matches[]`,
and uses an opaque cursor for pagination.

**Query Parameters:**

| Parameter | Type | Required | Default | Description |
| --- | --- | --- | --- | --- |
| `q` | string | Yes | — | Search query (1–255 chars). Full-text token match. |
| `type` | string | No | — | Filter by source type: `file` \| `folder` \| `project` \| `asset`. Deprecated `workspace` remains accepted for compatibility and returns an empty page because workspace rows are not indexed. Omit for all. |
| `searchIn` | string | No | `all` | Indexed rows to search: `name` \| `content` (body + annotation) \| `all`. |
| `matchMode` | string | No | `any` | Term matching inside each indexed row: `any` \| `all` \| `phrase`. `any` matches at least one analyzed term, `all` requires every analyzed term, and `phrase` requires ordered adjacent terms. |
| `format` | string | No | — | Comma-separated file formats, e.g. `txt,md`. Applies to `file` type only. |
| `entityType` | string | No | — | Comma-separated entity types, e.g. `character,prop`. Applies to `asset` type only. |
| `workspaceId` | string | No | — | Restrict to a single workspace (ObjectId). Must be one the caller is a member of, else empty page; malformed id → 422. |
| `projectId` | string | No | — | Restrict to a single project (ObjectId). Must be one the caller can access, else empty page; malformed id → 422. |
| `createdAfter` | string | No | — | ISO 8601 lower bound on source `createdAt`. **Timezone required** (`Z` or `±HH:MM`) — naive datetimes → 422. |
| `createdBefore` | string | No | — | ISO 8601 upper bound. Timezone required. |
| `limit` | integer | No | 50 | Page size (1–200). |
| `cursor` | string | No | — | Opaque pagination token from a previous response. Replay the original `q` and every corpus-shaping filter. Newly issued cursors use HMAC-signed `v1` envelopes and bind that corpus. Unsigned pre-v1 cursors remain accepted as unbound `any`-mode compatibility tokens only until 2026-09-07 00:00 UTC and must never be reused with another corpus. |
| `sortBy` | string | No | `relevance` | Result ordering: `relevance` \| `createdAt`. |
| `sortOrder` | string | No | `desc` | Direction for `sortBy=createdAt`: `asc` \| `desc`. Ignored for relevance, which is always descending. |
| `includePreview` | boolean | No | `false` | When `true`, **`asset`** results add a short-lived signed `previewUrl` (~1h) and `previewFileFormat` for the cover image/video. Other source types are unaffected; the raw bucket/key pointer is never returned. |

**Response:** `{ "data": [...], "count": N, "cursor": "<opaque token or null>" }`

Continue pagination whenever `cursor` is non-null, even when `count` is smaller
than the requested `limit`; canonical validation can suppress stale hits and
produce a short page that still has another page.

Each `data` item is one matched source object. Common fields:

| Field | Type | Description |
| --- | --- | --- |
| `sourceType` | string | `file` \| `folder` \| `project` \| `asset` |
| `sourceId` | string | ObjectId of the source object (the file/folder/project/asset id) |
| `score` | number | Top relevance score across this object's matches |
| `matchCount` | integer | Number of matching rows |
| `matches` | array | Per-row matches (see below) |
| `workspaceId` | string? | Workspace ObjectId |
| `projectId` | string? | Project ObjectId |
| `sourceCreatedAt` | datetime \| null | Source object's creation timestamp when indexed. |
| `container` | object \| null | Canonical object with `id`, `name`, and `kind`; `kind` is `folder` or `project`. File: immediate owner. Folder: immediate parent. |
| `project` | object \| null | Canonical creative-project owner (`id`, `name`). |
| `breadcrumb` | array \| null | Canonical project-root-to-container path (`id`, `name`, `sessionType`). |

Canonical context objects use these exact wire shapes:

| Object | Exact shape |
| --- | --- |
| `container` | `{ id: string, name: string, kind: "folder" \| "project" }` |
| `project` | `{ id: string, name: string }` |
| `breadcrumb[]` | `{ id: string, name: string, sessionType: "creative_project" \| "session" }` |

The `container`, `project`, and `breadcrumb` keys are required on every search
item but nullable as one all-or-nothing context unit. Project and asset results
use null context. File and folder results always carry a complete canonical
context unit; hits whose current lineage is missing, deleted, malformed, stale,
or unauthorized are omitted rather than returned with null or partial hierarchy
metadata. For file and folder results, use these canonical names and IDs instead
of deriving hierarchy from snippets or denormalized index fields.

`container` has the exact shape `{ id, name, kind }`, where `kind` is `folder`
or `project`. Only a `folder` container ID belongs in a
`/folders/{containerId}` URL; a `project` container is the project root and
belongs in `/projects/{containerId}`.

**File / folder / project** items may additionally carry `fileName`,
`fileFormat` (file only), `sessionId`, and `entitySessionId` (owning entity, if
any). For a file, `sessionId` is its current canonical containing folder or
project. For a folder or project result, `sessionId` identifies the matched
session itself; determine parentage only from `container` and `breadcrumb`.
**Folder / project** items also carry `sessionName` — the clean display name.
Use it for display rather than deriving a name from `snippets`/row text, which
mix the name with tags (folders) or description + tags (projects).

**Asset** items instead carry asset-specific fields (no `fileName`/`fileFormat`):

| Field | Type | Description |
| --- | --- | --- |
| `assetName` | string | Asset (entity session) name |
| `entityType` | string | e.g. `character`, `prop` |
| `visibility` | string | e.g. `workspace` \| `public` |
| `previewUrl` | string? | Signed cover URL — only when `includePreview=true` and the asset has a cover |
| `previewFileFormat` | string? | Cover format — `jpg`/`png` (image) or `mp4`/`mov` (video). Render by this: the signed URL carries no extension. Present only with `includePreview=true` |

Each entry in `matches[]`:

| Field | Type | Description |
| --- | --- | --- |
| `rowKind` | string | `name` \| `body` \| `annotation` |
| `score` | number | Relevance score for this row |
| `snippets` | string[] | Pre-rendered HTML strings — hits wrapped in `<mark>`, surrounding context HTML-escaped. Safe to inject directly into an HTML renderer (e.g. React `dangerouslySetInnerHTML`); no client-side offset math. |
| `chunkIndex` | integer? | Body chunk index (content rows) |
| `charOffset` | integer? | Char offset of the chunk within the file (content rows) |
| `lineStart` | integer? | Starting line number of the chunk (content rows) |

```bash
# Find files with any analyzed term from "hero"
curl "https://data.spuree.com/api/v1/search?q=hero&type=file" \
  -H "Authorization: Bearer $SPUREE_ACCESS_TOKEN"

# Search file names for an exact analyzed phrase, restricted to one project
curl "https://data.spuree.com/api/v1/search?q=quarterly%20report&type=file&searchIn=name&matchMode=phrase&projectId=...&format=txt,md&limit=50" \
  -H "Authorization: Bearer $SPUREE_ACCESS_TOKEN"

# Next page: replay the original query and corpus-shaping parameters
curl "https://data.spuree.com/api/v1/search?q=hero&type=file&cursor=<token>" \
  -H "Authorization: Bearer $SPUREE_ACCESS_TOKEN"
```

Newly issued HMAC-signed `v1` cursors also bind the caller's current permission
scope. Treat a cursor as opaque; do not edit or reuse it with a different query,
source type, filter set, or caller. A malformed, tampered, or mismatched bound
cursor returns 422. If a previously valid cursor returns 422 after the caller's
permission scope changes, discard it and restart from page one with the same
query and filters; do not retry the rejected cursor.

Unsigned pre-v1 cursors are treated as unbound `any`-mode compatibility tokens,
regardless of a re-sent `matchMode`, only until **2026-09-07 00:00 UTC**. During
that bounded migration window, a fully stripped and re-encoded modern payload
is statelessly indistinguishable from a genuine v0 token, so keep any unsigned
cursor only within its original pagination sequence. At and after the sunset,
every unsigned cursor is rejected; only a valid signed `v1` envelope is
accepted.

**Status Codes:**

| Code | Description |
| --- | --- |
| 200 | One canonical search page returned. |
| 400 | `search_query_too_broad`: query would match at least the bounded threshold; refine the terms. |
| 401 | Invalid or expired credential. |
| 403 | OAuth credential lacks the `read` scope. |
| 422 | Missing/invalid parameter, timezone-less date, malformed narrowing ID, malformed/tampered/mismatched cursor, or an unsigned cursor at/after the migration sunset. |
| 500 | Non-transient search or canonical-storage failure. |
| 503 | `search_context_unavailable`: canonical validation could not safely advance a stale-only page. |

---

### GET /v1/files/{fileId}

Get file metadata and presigned download URL.

**Response:** `{ "data": { id, fileName, fileFormat, mimeType, size, workspaceId, sessionId, entitySessionId, downloadUrl, createdAt, updatedAt } }`

```bash
curl "https://data.spuree.com/api/v1/files/{fileId}" \
  -H "Authorization: Bearer $SPUREE_ACCESS_TOKEN"
```

---

### GET /v1/files/{fileId}/content

<!-- spuree-agent
surfaces: ["local", "desktop", "backend", "hosted-web"]
webSafe: true
-->

Get inline UTF-8 text content for a small text-like file. Use this when an agent needs to read markdown, text, JSON, CSV, YAML, XML, subtitles, logs, or source files directly instead of following a short-lived CloudFront download URL.

This endpoint is not a replacement for file downloads. Binary files and large files should still use `GET /v1/files/{fileId}` and its `downloadUrl`.

**Response:** `{ "data": { id, fileName, fileFormat, mimeType, size, encoding, content, truncated } }`

`truncated` is reserved: oversize files return `413` rather than a partial body, so it is currently always `false`.

| Code | Description |
| --- | --- |
| 200 | Text content returned |
| 413 | File is too large to return inline |
| 415 | Unsupported/binary format or invalid UTF-8 |

```bash
curl "https://data.spuree.com/api/v1/files/{fileId}/content" \
  -H "Authorization: Bearer $SPUREE_ACCESS_TOKEN"
```

---

### POST /v1/files

Create a file record and get presigned upload URL(s). Automatically selects upload mode based on file size:

- **< 100 MB** → single presigned PUT URL (`mode: "single"`)
- **≥ 100 MB** → multipart upload with per-part presigned URLs (`mode: "multipart"`)

The caller **must** then upload to S3 **and** call `POST /v1/files/{fileId}/upload/complete`. Without the complete call, the file remains pending and unusable.

**Request Body:**

| Field | Type | Required | Description |
| --- | --- | --- | --- |
| `fileName` | string | Yes | File name without extension |
| `fileFormat` | string | Yes | Extension in lowercase (e.g., `fbx`, `png`) |
| `fileSize` | integer | Yes | File size in bytes (accepts string for backward compat) |
| `sessionId` | string | Yes | Target project or folder ObjectId |
| `checksum` | string | No | CRC32 base64 (8 chars) for end-to-end verification. When provided, the presigned URL is signed with the checksum and S3 verifies the upload matches. |

**Response (single mode):** `{ messageCode, fileId, mode: "single", uploadUrl, checksumBase64, contentType }`

**Response (multipart mode):** `{ messageCode, fileId, mode: "multipart", parts: [{partNumber, startByte, endByte, url}, ...], totalParts, expiresAt }`

| Code | Description |
| --- | --- |
| 200 | Record created, presigned URL(s) returned |
| 404 | Session not found |
| 409 | File already exists |

```bash
curl -X POST "https://data.spuree.com/api/v1/files" \
  -H "Authorization: Bearer $SPUREE_ACCESS_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"fileName":"hero_walk","fileFormat":"fbx","fileSize":5242880,"sessionId":"...","checksum":"AAAAAA=="}'
```

---

### POST /v1/files/{fileId}/upload/complete

Mark a file upload as completed. Works for both single and multipart uploads — the server auto-detects by checking whether the file has an active multipart `uploadId`. For multipart, the server calls S3 `ListParts` to collect ETags and then `CompleteMultipartUpload`.

**Request Body:**

| Field | Type | Required | Description |
| --- | --- | --- | --- |
| `fileId` | string | Yes | Must match path parameter |
| `clientChecksumCRC32` | string | No | Client-computed CRC32 for end-to-end verification |

**Response:** `{ messageCode, fileId, checksum, checksumCRC32 }`

- `checksum`: S3-verified CRC32 (base64) for both single and multipart uploads
- `checksumCRC32`: S3 FULL_OBJECT CRC32 (multipart only)

```bash
curl -X POST "https://data.spuree.com/api/v1/files/{fileId}/upload/complete" \
  -H "Authorization: Bearer $SPUREE_ACCESS_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"fileId": "..."}'
```

---

### PATCH /v1/files/{fileId}

Update file metadata. At least one field required.

| Field | Type | Description |
| --- | --- | --- |
| `fileName` | string | New name (without extension) |
| `fileFormat` | string | New extension (lowercase) |
| `checksum` | string | CRC32 base64 (write-once backfill, no edit permission needed) |
| `sessionId` | string | Move to target project or folder |

**Move:** Only `creative_project` (project) and `session` (folder) targets supported.

**Response:** `{ messageCode, fileId }`

| Code | Description |
| --- | --- |
| 200 | Updated |
| 409 | Filename conflict in target |

```bash
curl -X PATCH "https://data.spuree.com/api/v1/files/{fileId}" \
  -H "Authorization: Bearer $SPUREE_ACCESS_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"fileName": "hero_run_cycle"}'
```

---

### PUT /v1/files/{fileId}

Prepare a content update with optimistic concurrency. Acquires an upload lock. Like `POST /v1/files`, automatically selects single or multipart mode based on `fileSize`.

**Request Body:**

| Field | Type | Required | Description |
| --- | --- | --- | --- |
| `expectedChecksum` | string | Yes | CRC32 base64 of current content (for conflict detection) |
| `newChecksum` | string | Yes | CRC32 base64 of new content to upload |
| `fileSize` | integer | No | Size in bytes. If ≥ 100 MB, returns multipart mode. |

**Response (single):** `{ messageCode, fileId, mode: "single", uploadUrl, checksumBase64, contentType }`

**Response (multipart):** `{ messageCode, fileId, mode: "multipart", parts: [{partNumber, startByte, endByte, url}, ...], totalParts, expiresAt }`

After receiving the response, upload to S3 then call `POST /v1/files/{fileId}/upload/complete`.

**Upload lock:** Single uploads lock for 1 hour. Multipart uploads lock for 6 hours. Same user can re-acquire. Expired locks auto-release.

| Code | Description |
| --- | --- |
| 200 | Lock acquired, presigned URL(s) returned |
| 409 | Checksum mismatch or locked by another user |

```bash
curl -X PUT "https://data.spuree.com/api/v1/files/{fileId}" \
  -H "Authorization: Bearer $SPUREE_ACCESS_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"expectedChecksum": "a1b2...","newChecksum": "f6e5..."}'
```

---

### GET /v1/files/{fileId}/upload

Resume a multipart upload. Returns which parts are already in S3 and fresh presigned URLs for remaining parts.

**Query Parameters:**

| Parameter | Type | Required | Description |
| --- | --- | --- | --- |
| `uploadId` | string | No | Client's expected uploadId for mismatch detection |
| `checksum` | string | No | Client's expected pendingChecksum for conflict detection |

**Response:** `{ messageCode, fileId, completedParts: [{partNumber, etag, size}, ...], remainingParts: [{partNumber, startByte, endByte, url}, ...], totalParts }`

| Code | Description |
| --- | --- |
| 200 | Success |
| 400 | No active multipart upload |
| 409 | Not in pending status or locked by another user |

```bash
curl "https://data.spuree.com/api/v1/files/{fileId}/upload" \
  -H "Authorization: Bearer $SPUREE_ACCESS_TOKEN"
```

---

### POST /v1/files/{fileId}/upload/urls

Refresh presigned URLs for specific parts of a multipart upload. Use when URLs have expired before upload completes.

**Request Body:** `{ "partNumbers": [1, 3, 5] }` (1-indexed, min 1 part)

**Response:** `{ messageCode, urls: [{partNumber, startByte, endByte, url}, ...] }`

| Code | Description |
| --- | --- |
| 200 | Success |
| 400 | No active multipart upload or invalid part numbers |

```bash
curl -X POST "https://data.spuree.com/api/v1/files/{fileId}/upload/urls" \
  -H "Authorization: Bearer $SPUREE_ACCESS_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"partNumbers": [1, 3, 5]}'
```

---

### DELETE /v1/files/{fileId}/upload

Abort a multipart upload. Cleans up S3 parts and marks the file as failed. Idempotent — safe to call on already-aborted uploads.

**Response:** `{ messageCode, fileId, message }`

| Code | Description |
| --- | --- |
| 200 | Aborted |
| 400 | No active multipart upload |

```bash
curl -X DELETE "https://data.spuree.com/api/v1/files/{fileId}/upload" \
  -H "Authorization: Bearer $SPUREE_ACCESS_TOKEN"
```

---

### DELETE /v1/files/{fileId}

Soft-delete a file.

**Response:** `{ messageCode, fileId }`

```bash
curl -X DELETE "https://data.spuree.com/api/v1/files/{fileId}" \
  -H "Authorization: Bearer $SPUREE_ACCESS_TOKEN"
```

## Common Patterns

### Single Upload Flow (< 100 MB)

1. **Compute CRC32 checksum (base64):**
   ```bash
   CHECKSUM=$(python3 -c "import base64, zlib, sys; print(base64.b64encode(zlib.crc32(open('$FILE_PATH','rb').read()).to_bytes(4,'big')).decode())")
   ```

2. **Create file record** — `sessionId` is the target project or folder:
   ```
   POST /v1/files { fileName, fileFormat, fileSize, sessionId, checksum }
   → { fileId, mode: "single", uploadUrl, checksumBase64, contentType }
   ```

3. **Upload binary to S3:**
   ```
   PUT {uploadUrl}
   Content-Type: {contentType}
   x-amz-checksum-crc32: {checksumBase64}
   Body: <file binary>
   ```

4. **Complete the upload — REQUIRED (file is not visible until this is called):**
   ```
   POST /v1/files/{fileId}/upload/complete
   ```

> **Important:** All three steps (create → S3 upload → complete) must be performed. Skipping complete leaves the file in a pending state.

### Multipart Upload Flow (≥ 100 MB)

1. **Create file record** with `fileSize` ≥ 100 MB:
   ```
   POST /v1/files { fileName, fileFormat, fileSize, sessionId, checksum }
   → { fileId, mode: "multipart", parts: [{partNumber, startByte, endByte, url}, ...], totalParts, expiresAt }
   ```

2. **Upload each part** to its presigned URL:
   ```
   PUT {part.url}
   Content-Length: {part.endByte - part.startByte + 1}
   Body: <file slice from startByte to endByte>
   ```

3. **If interrupted**, resume with `GET /v1/files/{fileId}/upload` to get completed parts and fresh URLs for remaining parts.

4. **If URLs expire**, refresh with `POST /v1/files/{fileId}/upload/urls` for specific parts.

5. **Complete** — server auto-detects multipart and calls S3 `CompleteMultipartUpload`:
   ```
   POST /v1/files/{fileId}/upload/complete
   ```

6. **To abort**, call `DELETE /v1/files/{fileId}/upload` to clean up S3 parts.

### Content Update Flow

1. `PUT /v1/files/{fileId}` with checksums (and `fileSize` for multipart) → get upload URL(s)
2. Upload new content to S3 (single or multipart, same as above)
3. `POST /v1/files/{fileId}/upload/complete` — **REQUIRED**

### Viewing / Previewing a File

When a user wants to **view**, **preview**, or **see** a file:

1. If the agent has browser tools (e.g., Chrome DevTools MCP), **open the Studio preview URL directly**:
   ```
   navigate_page → https://studio.spuree.com/files/{fileId}
   ```

2. Otherwise, **return the Studio preview URL** for the user to click:
   ```
   https://studio.spuree.com/files/{fileId}
   ```

Do **not** download the file or attempt to open it locally. The Studio URL is permanent, permission-aware, and renders images, video, 3D, and other supported formats inline in the browser.

| Use case | URL | Properties |
| --- | --- | --- |
| **Share with user / preview** | `https://studio.spuree.com/files/{fileId}` | Permanent, permission-aware, renders preview UI |
| **Programmatic download** | `downloadUrl` from `GET /v1/files/{fileId}` | Short-lived presigned S3 URL — do not share, bypasses permissions and expires |

**Rule of thumb:** after `POST /v1/files` (upload) build the Studio URL from the returned `fileId`; after `GET /v1/search` use the result's `sourceId` (the file id, for `sourceType: "file"` items). Return that URL to the user. Only fetch `downloadUrl` when the agent itself needs the bytes.

### Search Then Download

1. `GET /v1/search?q=hero_walk&type=file` → take `sourceId` from a `sourceType: "file"` result
2. `GET /v1/files/{sourceId}` → get `downloadUrl`
3. Download from the presigned URL

### File Organization

**Projects and folders are both sessions.** `sessionId` always refers to a session. Session types: `creative_project` (project), `session` (folder). Folders nest inside projects or other folders.

S3 key: `works_{workspaceId}/sess_{sessionId}/file_{fileId}`

### Studio URLs

| Resource | URL Pattern |
| --- | --- |
| Project | `https://studio.spuree.com/projects/{projectId}` |
| Folder | `https://studio.spuree.com/folders/{folderId}` |
| File | `https://studio.spuree.com/files/{fileId}` |

## Error Handling

| Code | Cause | Resolution |
| --- | --- | --- |
| 400 | Invalid checksum format, missing fields, bad ID | Fix input format |
| 403 | No workspace access or edit permission | Check permissions |
| 404 | File or session not found | Verify IDs |
| 409 (checksum mismatch) | File modified by another user | Re-fetch checksum, retry |
| 409 (upload lock) | Another user is uploading | Wait (lock expires in 1 hour for single, 6 hours for multipart) |
| 409 (filename conflict) | Same name exists in target | Use different name or target |
| 401 | Invalid or expired token | Refresh via **authentication** skill |

