Tool Specification: mshtools-read_file
Overview
Universal file reader supporting text, images, video, and complex binary files (converting Office/PDF to markdown). Provides direct filesystem access for content ingestion and analysis.
JSON Schema
{
"type": "object",
"properties": {
"file_path": {
"type": "string",
"description": "Absolute path to file (must be absolute, not relative)"
},
"offset": {
"type": "integer",
"description": "Line offset for partial reading (optional, default: 0)"
},
"limit": {
"type": "integer",
"description": "Maximum lines to read (optional, default: 1000)"
}
},
"required": ["file_path"]
}
Streaming Mechanism
- Transport: Direct filesystem read via kernel
- Processing Pipeline:
- Text files: Line-numbered output (cat -n format)
- Images: Direct pixel data display in chat interface
- Videos: Metadata extraction + frame preview (for MP4, MOV, WEBM, MKV, AVI, M4V ≤100MB)
- Binary files (Office/PDF ≤20MB): Conversion to markdown via parsing engines
- Truncation: Lines >2000 characters truncated; Text files >100MB rejected
- Output Format: Contextual based on MIME type detection
Integration Architecture
File Type Handlers
| Type |
Extension |
Handler |
Output Format |
| Text |
.txt, .md, .py, etc |
Direct read |
Line-numbered text |
| Image |
.png, .jpg, etc |
Native display |
Inline image |
| Video |
.mp4, .mov, etc |
FFmpeg metadata |
Metadata + thumbnail |
| Word |
.docx |
OpenXML parser |
Markdown |
| Excel |
.xlsx |
OpenXML parser |
Markdown table |
| PDF |
.pdf |
pdfplumber/pikepdf |
Markdown |
| Archive |
.zip, .tar |
Listing |
File tree |
Path Resolution
- Absolute Required: Must start with
/ (e.g., /mnt/kimi/upload/file.txt)
- Automatic Expansion:
~ not supported (use explicit paths)
- Symlink Resolution: Followed automatically
- Access Control: Respects filesystem permissions
Performance Characteristics
- Large Files: Partial reading supported via offset/limit
- Binary Conversion: CPU-intensive for Office/PDF (runs in isolated process)
- Image Loading: Direct memory mapping for fast display
- Caching: No caching; re-reads file on each call
Usage Patterns
Sequential Reading
Step 1: read_file(/path/to/large.txt, offset=0, limit=1000) # Lines 1-1000
Step 2: read_file(/path/to/large.txt, offset=1000, limit=1000) # Lines 1001-2000
Image Analysis
read_file(/mnt/kimi/upload/chart.png) # Displays image for vision analysis
Document Ingestion
read_file(/mnt/kimi/upload/report.docx) # Converts to markdown for text analysis
Error Handling
- Non-existent: Returns error with "file not found"
- Permission Denied: Returns error if read access blocked
- Size Exceeded: Returns error for files exceeding limits
- Malformed Binary: Best-effort parsing with partial content warnings
Skill Integration
- docx skill: Reads SKILL.md, C# templates, validation schemas
- xlsx skill: Reads Excel files for data analysis
- pdf skill: Reads existing PDFs for processing route
- webapp skill: Reads source code files during development
1---2name: tool-specification-mshtools-read-file3description: Universal file reader supporting text, images, video, and complex binary files (converting Office/PDF to markdown). Provides direct filesystem access for content ingestion and analysis.4---5# Tool Specification: mshtools-read_file67## Overview8Universal file reader supporting text, images, video, and complex binary files (converting Office/PDF to markdown). Provides direct filesystem access for content ingestion and analysis.910## JSON Schema11```json12{13 "type": "object",14 "properties": {15 "file_path": {16 "type": "string",17 "description": "Absolute path to file (must be absolute, not relative)"18 },19 "offset": {20 "type": "integer",21 "description": "Line offset for partial reading (optional, default: 0)"22 },23 "limit": {24 "type": "integer",25 "description": "Maximum lines to read (optional, default: 1000)"26 }27 },28 "required": ["file_path"]29}30```3132## Streaming Mechanism33- **Transport**: Direct filesystem read via kernel34- **Processing Pipeline**:35 - Text files: Line-numbered output (cat -n format)36 - Images: Direct pixel data display in chat interface37 - Videos: Metadata extraction + frame preview (for MP4, MOV, WEBM, MKV, AVI, M4V ≤100MB)38 - Binary files (Office/PDF ≤20MB): Conversion to markdown via parsing engines39- **Truncation**: Lines >2000 characters truncated; Text files >100MB rejected40- **Output Format**: Contextual based on MIME type detection4142## Integration Architecture4344### File Type Handlers45| Type | Extension | Handler | Output Format |46|------|-----------|---------|---------------|47| Text | .txt, .md, .py, etc | Direct read | Line-numbered text |48| Image | .png, .jpg, etc | Native display | Inline image |49| Video | .mp4, .mov, etc | FFmpeg metadata | Metadata + thumbnail |50| Word | .docx | OpenXML parser | Markdown |51| Excel | .xlsx | OpenXML parser | Markdown table |52| PDF | .pdf | pdfplumber/pikepdf | Markdown |53| Archive | .zip, .tar | Listing | File tree |5455### Path Resolution56- **Absolute Required**: Must start with `/` (e.g., `/mnt/kimi/upload/file.txt`)57- **Automatic Expansion**: `~` not supported (use explicit paths)58- **Symlink Resolution**: Followed automatically59- **Access Control**: Respects filesystem permissions6061### Performance Characteristics62- **Large Files**: Partial reading supported via offset/limit63- **Binary Conversion**: CPU-intensive for Office/PDF (runs in isolated process)64- **Image Loading**: Direct memory mapping for fast display65- **Caching**: No caching; re-reads file on each call6667## Usage Patterns6869### Sequential Reading70```71Step 1: read_file(/path/to/large.txt, offset=0, limit=1000) # Lines 1-100072Step 2: read_file(/path/to/large.txt, offset=1000, limit=1000) # Lines 1001-200073```7475### Image Analysis76```77read_file(/mnt/kimi/upload/chart.png) # Displays image for vision analysis78```7980### Document Ingestion81```82read_file(/mnt/kimi/upload/report.docx) # Converts to markdown for text analysis83```8485## Error Handling86- **Non-existent**: Returns error with "file not found"87- **Permission Denied**: Returns error if read access blocked88- **Size Exceeded**: Returns error for files exceeding limits89- **Malformed Binary**: Best-effort parsing with partial content warnings9091## Skill Integration92- **docx skill**: Reads SKILL.md, C# templates, validation schemas93- **xlsx skill**: Reads Excel files for data analysis94- **pdf skill**: Reads existing PDFs for processing route95- **webapp skill**: Reads source code files during development