Summary
Reduced redundant token estimation in tool-output pruning while preserving pruning behavior.
Context
The pruning loop in prune_old_tool_outputs() was estimating token counts twice for pruned parts and recalculating placeholder tokens for every mutation.
Changes
- Added
PRUNE_PLACEHOLDER_TOKENSto cache placeholder token counts. - Introduced
get_part_content_text()to normalize part content before estimation/pruning. - Passed precomputed token counts into
prune_part_content()to avoid re-estimation. - Stopped token estimation after
protect + minimumis satisfied while still scanning for parts to prune.
Behavioral Impact
Pruning decisions and output remain unchanged; token estimation work is reduced and placeholder token counting is cached.
Related Cards
- [[token-pruning-efficiency-plan]]