# Implementation Guide

> Use soup.get_text(separator=' ', strip=True) to ensure words don't stick together after tag removal.

- Skill: `tools-only/implementation-guide-6` (Agent Skill, multi-file: 3 files)
- Install (CLI): `npx skillmds@latest add tools-only/implementation-guide-6`
- Raw SKILL.md: https://api.skillmd.com/api/skills/tools-only/implementation-guide-6/raw
- Safety review: pending (external: skill-scanner PASS, skillspector PASS)
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: Coding & Dev Tools
- Author: tools-only (https://skillmd.com/u/tools-only)
- Updated: 2026-09-29
- Page: https://skillmd.com/skills/tools-only/implementation-guide-6

---


# Implementation Guide

## Code Structure
**How is the code organized?**

- `pdf_to_epub/core/epub_extractor.py`

## Implementation Notes
**Key technical details to remember:**

### EbookLib Snippet
```python
import ebooklib
from ebooklib import epub

book = epub.read_epub(path)
for item in book.get_items():
    if item.get_type() == ebooklib.ITEM_DOCUMENT:
        # process html
```

### BeautifulSoup Cleaning
Use `soup.get_text(separator=' ', strip=True)` to ensure words don't stick together after tag removal.

## Error Handling
**How do we handle failures?**

- Catch errors if the file is not a valid ZIP or missing `mimetype`.

