Backlog
Tracked features, ideas, and deferred work for grooming and future sessions.
P0 - Must Have
(Empty)
P1 - Should Have
gitlab-skill: Remove hardcoded corporate URL
Source: Plugin code review session 2026-02-21
Added: 2026-02-21
Completed: 2026-02-22
Priority: P1
Status: DONE — replaced https://sourcery.assaabloy.net with generic placeholder https://gitlab.example.com in validate_glfm.py (default arg + help text) and gitlab-ci-local-guide.md (example URL)
Description: validate_glfm.py lines 152-153 hardcode https://sourcery.assaabloy.net as the default GitLab instance URL. gitlab-ci-local-guide.md line 51 also references this URL. This leaks a corporate internal URL into a public repository. Replace with a generic placeholder (e.g., https://gitlab.example.com) or make the URL a required argument with no default.
Files:
plugins/gitlab-skill/skills/gitlab-skill/scripts/validate_glfm.py(lines 152-153)plugins/gitlab-skill/skills/gitlab-skill/references/gitlab-ci-local-guide.md(line 51)
bash-development: Fix bash-53-features inaccuracies and task_output bug
Source: Plugin code review session 2026-02-21
Added: 2026-02-21
Completed: 2026-02-22
Priority: P1
Status: DONE — Fixed GLOBSORT syntax (:asc/:desc → +/- prefix, date → mtime, added missing specifiers), added required space in ${ command; } and ${| command; } examples, added REPLY-is-local note, corrected C23 claim to "C standard conformance improvements", added interface disclaimer for kv/strptime. Fixed task_output() bug in log_functions.sh — ${task_output} → ${raw_task_output} on lines 1410/1412.
Description: Two issues:
- bash-53-features/SKILL.md inaccuracies (features exist but details wrong):
- GLOBSORT:
:asc/:descsuffixes fabricated — actual syntax is+/-prefix;datenot a valid specifier (should bemtime); missing specifiers:blocks,atime,ctime,numeric,nosort ${ command; }examples: missing required space after{(bash.1 requires space/tab/newline/|after{)${| command; }REPLY examples: misleading — REPLY is local within substitution, restored after completion- C23 claim overstated: build minimum is C90, not C23; C23 changes are about conformance
kv/strptimeusage examples: unverifiable from official sources (existence confirmed, interface speculative)
- GLOBSORT:
- log_functions.sh bug (line 1401):
task_output()function references${task_output}variable on lines 1410/1412, but onlyraw_task_outputis assigned (line 1406). Output will be empty.
Citations:
- Bash CHANGES: https://tiswww.case.edu/php/chet/bash/CHANGES lines 850-860 (accessed 2026-02-21)
- Bash NEWS: https://ftp.gnu.org/gnu/bash/bash-5.3.tar.gz extracted NEWS lines 48-57 (accessed 2026-02-21)
- Bash manpage:
bash-5.3/doc/bash.1lines 2525-2566 (GLOBSORT), 4146+ (command substitution) (accessed 2026-02-21) Files: plugins/bash-development/skills/bash-53-features/SKILL.md(GLOBSORT, examples, C23)plugins/bash-development/skills/bash-logging/scripts/log_functions.sh(lines 1401, 1410, 1412)
commitlint: Verify --last flag and exit codes against primary sources RESOLVED
Source: Plugin code review session 2026-02-21
Added: 2026-02-21
Resolved: 2026-02-21
Priority: P1
Status: FACT-CHECKED 2026-02-21 — --last flag VERIFIED across 4 independent sources. Flag exists in source code (cli.ts with alias -l), official CLI reference docs, commitlint help output, and raw documentation. Exit codes verified against ExitCode enum in cli-error.ts. Original review agent claim that --last was fabricated was wrong.
Citations:
- commitlint source:
@commitlint/cli/src/cli.ts(accessed 2026-02-21) - commitlint docs: https://commitlint.js.org/reference/cli.html (accessed 2026-02-21)
- commitlint source:
@commitlint/cli/src/cli-error.tsExitCodeenum (accessed 2026-02-21)
clang-format: Fix broken YAML frontmatter
Source: Plugin code review session 2026-02-21
Added: 2026-02-21
Completed: 2026-02-22
Priority: P1
Status: DONE — Fixed description:"Configure..." → description: "Configure..." in plugins/clang-format/skills/clang-format/SKILL.md line 3.
Description: SKILL.md line 3 has description:"Configure clang-format..." — missing the required space after the description: key. This causes YAML parsing failures. The frontmatter should be description: "Configure clang-format...".
Files:
plugins/clang-format/skills/clang-format/SKILL.md(line 3)
agent-orchestration: Remove phantom /is-it-done command references
Source: Plugin code review session 2026-02-21
Added: 2026-02-21
Completed: 2026-02-22
Priority: P1
Status: DONE — Replaced all 9 actionable /is-it-done slash command calls in SKILL.md and how-to-delegate/SKILL.md with /am-i-complete (the actual existing command). Source attribution references in post-completion-validation-protocol.md and synthesis-improvements-from-research.md left as historical metadata. Removed Clavix references from clear-framework.md — replaced with generic imperative descriptions of the CLEAR framework operations.
perl-development: Fix shell injection vulnerability in example template
Source: Plugin code review session 2026-02-21
Added: 2026-02-21
Completed: 2026-02-22
Priority: P1
Status: DONE — Added #!/usr/bin/env perl shebang. Replaced undefined App::Logger import with inline logging functions (print_success, print_error, print_info, print_debug, print_warning). Added use File::Spec, use IPC::Open3, use Symbol 'gensym' for safe system calls. Fixed command_exists to use File::Spec->path() instead of backtick command -v shell call. Fixed run_command to accept list args and use IPC::Open3::open3 instead of backtick string injection. Fixed system("stty ...") calls in query_terminal_safely to use list form.
Files:
plugins/perl-development/skills/perl-development/references/perl_example_file.pl
hallucination-detector: Fix dead backtick evidence marker and multi-occurrence "because" bug
Source: Plugin code review session 2026-02-21
Added: 2026-02-21
Completed: 2026-02-22
Priority: P1
Status: DONE — Fixed both bugs in hallucination-audit-stop.js: (1) Removed dead backtick regex from EVIDENCE_MARKERS (was never matched since stripLowSignalRegions removes backtick spans before scanning); replaced with a BACKTICK_RE check against the original pre-strip text via new rawText param in hasEvidenceNearby. (2) Changed because loop from lower.indexOf() (first occurrence only) to a while loop scanning all occurrences — each unsuppressed because is now flagged independently.
Files:
plugins/hallucination-detector/scripts/hallucination-audit-stop.js
python3-development: Fix 3 malformed frontmatter files
Source: Plugin code review session 2026-02-21
Added: 2026-02-21
Completed: 2026-02-22
Priority: P1
Status: DONE — Fixed 2 malformed frontmatter files (backlog said 3, validator found 2): skills/development/add-new-feature/SKILL.md and skills/development/complete-implementation/SKILL.md — both had description:"..." missing the space after the colon. Also unquoted the description in complete-implementation (was double-escaped "\"..."\""). No ghost agent reference found after inspection.
the-rewrite-room: Fix nonexistent script reference and 6 broken links
Source: Plugin code review session 2026-02-21
Added: 2026-02-21
Completed: 2026-02-22
Priority: P1
Status: DONE — Fixed validators.yaml line 12: changed validate_frontmatter.py (no longer exists) to plugin_validator.py. Fixed 6 broken links in registry-guide.md (lines 9, 10, 11, 35, 57, 83) — updated ./workflows.yaml, ./validators.yaml, ./routing-rules.yaml to ../registry/workflows.yaml, ../registry/validators.yaml, ../registry/routing-rules.yaml. Added --json flag to file_metrics.py (count and scan commands) to support machine-readable output as referenced in research-utilities.md. Also fixed validators.yaml line 65 false-positive link pattern.
Files:
plugins/the-rewrite-room/skills/the-rewrite-room/registry/validators.yamlplugins/the-rewrite-room/skills/the-rewrite-room/references/registry-guide.mdplugins/the-rewrite-room/skills/the-rewrite-room/workflows/research-utilities.mdplugins/the-rewrite-room/skills/the-rewrite-room/scripts/file_metrics.py
fastmcp-creator: Add citations for 1200+ lines of FastMCP 3.x API documentation
Source: Plugin code review session 2026-02-21
Added: 2026-02-21
Priority: P1
Status: FACT-CHECKED 2026-02-21 — require_auth hallucination flag was WRONG. The string require_auth does not appear anywhere in the plugin. The actual auth APIs documented (require_scopes, restrict_tag, AuthContext, AuthMiddleware) are all VERIFIED against installed FastMCP 3.0.0rc2 source code (fastmcp/server/auth/authorization.py, fastmcp/server/middleware/authorization.py). Citation need still valid — documentation lacks source attribution.
Description: The plugin contains over 1200 lines of FastMCP 3.x API documentation derived from a release candidate version, with no source citations. Per CLAUDE.md citation requirements, all factual claims must have cited sources. Auth API claims verified but citations still needed for all documentation.
Citations:
- FastMCP auth:
fastmcp/server/auth/__init__.pylines 8-13,authorization.pylines 48, 78, 106 (accessed 2026-02-21) - FastMCP middleware:
fastmcp/server/middleware/authorization.pyline 51 (accessed 2026-02-21) Files: plugins/fastmcp-creator/skills/fastmcp-creator/references/(multiple files)
Consolidate validate_frontmatter.py into plugin_validator.py
Source: Frontmatter validation bug-fix session 2026-02-20
Added: 2026-02-20
Completed: 2026-02-20
Status: DONE — commit bfb03f1 on branch claude/fix-frontmatter-validation-u2iV2
Priority: P1
Description: plugin_validator.py (4190 lines) copy-pastes frontmatter
validation logic from validate_frontmatter.py (1341 lines) instead of
importing it. Comments acknowledge this: "PYDANTIC FRONTMATTER MODELS (from
validate_frontmatter.py)" and "Complexity preserved from validate_frontmatter.py
for behavioral parity." This creates two maintenance surfaces: any change must
be applied to both scripts (as seen when reversing the name-field bug workaround).
Required work:
- Audit both scripts for all divergences (validation checks, Pydantic models, helper functions, CLI flags, scan patterns).
- Extract shared code (Pydantic models,
_fix_skill_name*,extract_frontmatter,detect_file_type,validate_and_normalize) into a shared moduleplugins/plugin-creator/scripts/frontmatter_core.py. - Refactor
validate_frontmatter.pyandplugin_validator.pyto import fromfrontmatter_core.py. - Port any validation steps present in
validate_frontmatter.pybut missing fromplugin_validator.py(e.g. skill-directory-name check, name-mismatch warning added in 2026-02-20). - Update
frontmatter_utils.pyif overlap exists. - Update all documentation: CLAUDE.md, reference files, script docstrings.
- Update tests to import from the correct locations.
- Run full test suite; verify pre-commit hooks still pass.
Suggested location: plugins/plugin-creator/scripts/
Validate and verify orchestrator-discipline plugin hooks and processes
Source: Plugin creation session 2026-02-19
Added: 2026-02-19
Completed: 2026-02-21
Status: DONE — commits 49e0ae5, ea3e737 on branch claude/bulk-backlog-grooming-LIQDi
Plan: plan/tasks-4-validate-orchestrator-discipline.md
Description: T1-T4 complete. Plugin passes claude plugin validate. Hook directory detection added. user-invocable: true added to SKILL.md. All 5 hook behavior tests pass.
Suggested location: plugins/orchestrator-discipline/
SAM: Error Recovery / Rollback Procedures
Source: Gap analysis of SAM framework
Added: 2026-02-01
Description: Define explicit procedure when a task fails irrecoverably. How to undo artifact changes? How to restore artifact plane to consistent state after failure?
Research first: How do GSD, BMAD-METHOD, AutoGPT, and traditional CI/CD handle rollback? What patterns exist for transactional artifact updates?
Suggested location: stateless-software-engineering-framework.md (new Appendix or Part 6 addition)
SAM: Human Escalation Criteria
Source: Gap analysis of SAM framework
Added: 2026-02-01
Description: Define explicit triggers for escalating to human at each stage (not just Discovery). When should agents block and ask vs attempt repair vs fail?
Research first: How do GSD deviation rules work? How does BMAD-METHOD handle human checkpoints? What escalation patterns exist in agent frameworks?
Suggested location: stateless-software-engineering-framework.md (each Agent Specification section)
SAM: Timeout/Stall Detection
Source: Gap analysis of SAM framework
Added: 2026-02-01
Description: Define mechanism to detect when an agent is stuck or has stalled. Include timeout thresholds per stage, health check patterns, and recovery actions.
Research first: How do orchestration frameworks (Temporal, Prefect, Airflow) handle task timeouts? What heartbeat patterns exist? How does Gas Town handle session recycling?
Suggested location: stateless-software-engineering-framework.md (Orchestrator section 3.8)
SAM: Artifact Schema Validation
Source: Gap analysis of SAM framework
Added: 2026-02-01
Description: Define formal validation rules or JSON schemas for artifact formats. Currently only templates provided. Enable automated validation at stage boundaries.
Research first: How do GSD artifacts (STATE.md, ROADMAP.md) enforce structure? What validation approaches exist in BMAD-METHOD? JSON Schema vs YAML validation vs custom parsers?
Suggested location: sam-artifact-schemas/ (new directory with schema files)
SAM: Scope Creep Detection
Source: Gap analysis of SAM framework
Added: 2026-02-01
Description: Define mechanism to detect when execution diverges from plan. How does Forensic Review detect that the execution agent solved a different problem than planned?
Research first: How does GSD plan-checker detect deviation? What diff/comparison techniques exist? How do code review tools detect scope creep in PRs?
Suggested location: stateless-software-engineering-framework.md (section 3.6 Forensic Review)
Meta-Process Capture — Expert Panel Dataset Builder
Source: ARL expert panel process (sessions 2026-02-12 to 2026-02-13) Added: 2026-02-13 Description: Document the multi-agent expert panel methodology as a reusable system for building datasets that inform skills and systems. The process — assign framework experts to repositories, ask structured questions, cross-examine, synthesize, map to requirements, validate — produced high-quality sourced findings. Capture this as a repeatable pattern. Key elements to document:
- Expert assignment protocol (one agent per source repo, source-code-only evidence standard)
- Question group design (themed questions → cross-examination → synthesis)
- Cross-examination as adversarial validation (experts challenge each other's claims)
- Phased output (discussion → requirement mapping → synthesis → validation)
- Traceability chain (claim → expert citation → file:line evidence)
- Session continuity handling (state file enables cross-session resumption)
Input artifacts:
plugins/plugin-creator/skills/assessor/references/ARL/ARL-agent-instructions.md,qa-expert-panel.mdSuggested location: New skill or methodology document — captures the meta-process, not the ARL content
SAM Extension — Integrate ARL General Theory
Source: ARL expert panel Phase 3 output
Added: 2026-02-13
Description: Integrate the 7 universal principles from synthesis-general-theory.md into SAM methodology documents. These principles (structure over instruction, front-loading reduces gates, AI cannot self-evaluate, compression is architectural, iteration-aware state required, parallelism enables independent verification, failure paths need more compression) extend SAM's scope to cover autonomous refinement loops.
Input artifacts: plugins/plugin-creator/skills/assessor/references/ARL/synthesis-general-theory.md
Target files: stateless-agent-methodology.md, stateless-software-engineering-framework.md
Dependencies: None — general theory is framework-agnostic
Related backlog items: SAM gap items (error recovery, human escalation, scope creep detection) — the general theory findings inform several of these
ARL Skill Development
Source: ARL expert panel Phase 3 output
Added: 2026-02-13
Description: Build the Autonomous Refinement Loop as a skill using synthesis-arl-applicable.md as its reference foundation. The ARL is a logical process (Assess → Plan → Implement → Review → Repeat) for autonomous skill refinement with R1-R10 requirement gates.
Input artifacts: plugins/plugin-creator/skills/assessor/references/ARL/synthesis-arl-applicable.md, synthesis-general-theory.md
Scope: Logical process design — what gates fire when, what each gate checks, success/failure criteria. NOT implementation artifacts (schemas, thresholds, pseudocode).
Dependencies: Benefits from SAM extension being done first (ARL would be a SAM-based skill)
Suggested location: New skill under plugins/plugin-creator/skills/ or standalone plugin
Extract claude-plugin-lint to standalone PyPI package
Source: Gap analysis - no existing Claude Code plugin linters exist
Added: 2026-02-01
Description: Extract and enhance validate_frontmatter.py into a standalone open-source project. First dedicated linter for Claude Code plugin frontmatter (SKILL.md, agents/.md, commands/.md). Official claude plugin validate only checks plugin.json structure.
Features to include:
- YAML frontmatter schema validation with Pydantic models
- Auto-fix capabilities (arrays → comma-separated, multiline → single-line)
- Token-based complexity metrics (tiktoken) instead of line counts
- Cross-reference validation (agent references non-existent skill)
- Marketplace readiness scoring
- Pre-commit hook integration
- CLI with
--fixand--reportmodes Current source:plugins/plugin-creator/scripts/validate_frontmatter.pySuggested repo name:claude-plugin-lintorcc-plugin-validator
P2 - Could Have
Add ty support alongside mypy in distributed plugins COMPLETED
Source: Plugin code review session 2026-02-21 Added: 2026-02-21 Completed: 2026-02-21 Description: Completed as full mypy-to-ty migration across all active documentation, skills, agents, and reference files. All stale mypy references updated to ty (Astral ecosystem). Third-party reference docs (mypy-docs/, rules/mypy/) retained as-is. Historical plan/ files left unchanged.
conventional-commits: Fix CHANGELOG references to nonexistent files
Source: Plugin code review session 2026-02-21
Added: 2026-02-21
Description: CHANGELOG references files that do not exist in the repository. Additionally, related skills referenced in the plugin do not exist. All dead references need to be either created or removed.
Files: plugins/conventional-commits/ (CHANGELOG and skill cross-references)
dasel: Reconcile 265 -f flag occurrences with reference documentation
Source: Plugin code review session 2026-02-21
Added: 2026-02-21
Description: Reference documentation states dasel uses a specific flag pattern, but 265 occurrences of the -f flag across the skill contradict this documentation. The hook file exists on disk but is not registered in the plugin manifest (plugin.json). Run auto_sync_manifests.py --reconcile to fix manifest drift, and audit -f flag usage against official dasel documentation.
Files:
plugins/dasel/(skill files with-fflag usage)plugins/dasel/.claude-plugin/plugin.json(missing hook registration)
litellm: Remove private API documentation and update verification date RESOLVED
Source: Plugin code review session 2026-02-21
Added: 2026-02-21
Resolved: 2026-02-21
Status: FACT-CHECKED 2026-02-21 — Review claim REFUTED. litellm._should_retry() is NOT a private/internal API. It is explicitly documented in official litellm Exception Mapping docs (https://docs.litellm.ai/docs/exception_mapping) with code examples showing litellm._should_retry(e.status_code). Despite the underscore prefix, this is an intentionally public utility function exported in __init__.py. The function exists at utils.py:6506.
Description: Original review flagged _should_retry() as private API. Fact-checking verified it is documented public API in official docs. Stale verification date may still need updating.
Citations:
- litellm docs: https://docs.litellm.ai/docs/exception_mapping (accessed 2026-02-21)
- litellm source:
litellm/utils.pyline 6506 (accessed 2026-02-21) - litellm source:
litellm/__init__.pytype stub exports_should_retry(accessed 2026-02-21) Files: plugins/litellm/skills/litellm/SKILL.md(line 257)
verification-gate: Remove unsubstantiated 95% confidence claim
Source: Plugin code review session 2026-02-21
Added: 2026-02-21
Description: The skill contains an unsubstantiated claim about "95% confidence" with no supporting data, citation, or methodology. Per CLAUDE.md verification protocol, claims must be cited. Either add supporting evidence or remove the specific percentage. Also fix missing code fence language specifiers and stale dates.
Files: plugins/verification-gate/ (SKILL.md and reference files)
development-harness: Remove hardcoded machine path and fix version drift
Source: Plugin code review session 2026-02-21
Added: 2026-02-21
Description: Contains a hardcoded machine-specific path (likely /home/user/... or similar) that won't work on other systems. Version references have drifted from actual tool versions. Role table has inconsistencies between documented and actual roles.
Files: plugins/development-harness/ (SKILL.md and reference files)
llamafile: Fix HuggingFace model URLs (wrong org name + fabricated repos)
Source: Plugin code review session 2026-02-21
Added: 2026-02-21
Priority: P2
Status: FACT-CHECKED 2026-02-21 — SourceForge URL VERIFIED (legitimate auto-mirror with explicit disclaimer). HuggingFace model URLs REFUTED — wrong org name (Mozilla should be mozilla-ai) and model repo names (gemma-3-3b-it-gguf, Qwen3-0.6B-gguf, Mistral-7B-gguf, Llama-3.1-8B-gguf) all return 404. GitHub URLs using mozilla-ai/llamafile are correct (old Mozilla-Ocho redirects).
Description: SourceForge mirror is confirmed legitimate (auto-mirror with disclaimer). HuggingFace URLs need two fixes: (1) org name Mozilla must be changed to mozilla-ai, (2) model repo names need verification against actual mozilla-ai org repos — current names appear fabricated. The llava-v1.5-7b-llamafile URL works only via redirect.
Citations:
- SourceForge mirror: https://sourceforge.net/projects/llamafile.mirror/files/0.9.3/ — disclaimer states "exact mirror of the llamafile project" (accessed 2026-02-21)
- HuggingFace 404s:
Mozilla/Qwen3-0.6B-gguf,Mozilla/Mistral-7B-gguf,Mozilla/gemma-3-3b-it-gguf,Mozilla/Llama-3.1-8B-gguf(accessed 2026-02-21) - HuggingFace redirect:
Mozilla/llava-v1.5-7b-llamafileredirects tomozilla-ai/llava-v1.5-7b-llamafile(accessed 2026-02-21) - GitHub:
mozilla-ai/llamafileresolves correctly;Mozilla-Ocho/llamafileredirects (accessed 2026-02-21) Files: plugins/llamafile/skills/llamafile/SKILL.md(lines 78, 87-92 model URLs; line 69 SourceForge OK)
prompt-optimization: Fix unreachable reference files and raw JSX/MDX markup
Source: Plugin code review session 2026-02-21
Added: 2026-02-21
Description: Two reference files are unreachable (not linked from SKILL.md or any other file). Contains raw Anthropic JSX/MDX markup that should be converted to standard markdown for compatibility with Claude Code's markdown rendering.
Files: plugins/prompt-optimization-claude-45/ (reference files)
brainstorming-skill: Remove orphaned bibliography entry and cross-reference headings
Source: Plugin code review session 2026-02-21
Added: 2026-02-21
Description: Contains an orphaned bibliography entry (referenced nowhere) and orphaned cross-reference section headings that point to removed or renamed content.
Files: plugins/brainstorming-skill/ (SKILL.md and reference files)
uv: Fix incorrect script paths in README
Source: Plugin code review session 2026-02-21
Added: 2026-02-21
Description: README contains incorrect script paths that don't match the actual file locations. The README is also disproportionately large for what is essentially a thin wrapper plugin.
Files: plugins/uv/ (README.md)
plugin-creator: Remove dead code and triplicated regex
Source: Plugin code review session 2026-02-21
Added: 2026-02-21
Description: Contains triplicated regex patterns (same regex defined 3 times), a dead skipped list that is populated but never read, an unused sum() call, and HK005 warning is incorrectly treated as an error in certain code paths. Also has a noqa BLE001 suppression that should be addressed per CLAUDE.md linting policy.
Files: plugins/plugin-creator/ (scripts and skill files)
Add PR003/PR004 test coverage to plugin registration validator
Source: Code review session 2026-02-21
Added: 2026-02-21
Description: PluginRegistrationValidator defines PR003 (missing metadata fields: repository, homepage, author) and PR004 (repository URL mismatches git remote URL) at lines 276-277 of plugin_validator.py, and emits them at lines 2815 and 2834. Tests exist for PR001 (unregistered) and PR002 (missing file), but not PR003/PR004. Add tests to plugins/plugin-creator/tests/test_plugin_registration_validator.py covering: (1) PR003 emitted when metadata fields absent; (2) PR004 emitted when repo URL mismatches remote.
kaizen: MCP consolidation analysis
Source: Design session 2026-02-20
Added: 2026-02-20
Description: The plugin currently runs two MCP servers (kaizen-duckdb via mcp-server-motherduck, kaizen-analysis via server.py) plus a standalone CLI script (sentiment-score.py). Investigate: (1) What does each MCP server provide that the other cannot? Can they be merged into a single server? (2) Why is sentiment-score.py a standalone script rather than an MCP tool inside server.py? What would be gained or lost by moving scoring into the MCP server (always-on scoring, no manual invocation, lock ownership)? (3) Is there a clean boundary between "batch processing" (script) and "query/serve" (MCP) that should be preserved?
Decision needed: Consolidate vs. keep separate, with rationale.
Suggested location: plugins/agentskill-kaizen/
SAM: Parser regex false positive on "## Task Summary Statistics"
Source: Migration proof-of-concept (2026-02-13)
Added: 2026-02-13
Description: The widened task header regex ^#{2,3}\s+Task:?\s+([A-Za-z0-9.]+)[:\s-]+(.+)$ in implementation_manager.py matches ## Task Summary Statistics as task ID "Summary" with title "Statistics". The regex needs a negative lookahead or post-parse filter to exclude non-task sections. Observed when parsing plan/tasks-1-plugin-linter.md.
File: plugins/python3-development/skills/implementation-manager/scripts/implementation_manager.py line 645
SAM: Replace validate-task-file.sh with Python validator
Source: Task format standardization plan (2026-02-13)
Added: 2026-02-13
Description: The bash validator at plugins/plugin-creator/scripts/validate-task-file.sh validates a different schema (tasks-refactor-*.md) and doesn't understand YAML frontmatter. Replace with Python validator that uses the shared task_format.py module.
File: plugins/plugin-creator/scripts/validate-task-file.sh
SAM: Parallel Execution Details
Source: Gap analysis of SAM framework
Added: 2026-02-01
Description: Detail safe parallelization within SAM pipeline. When can tasks run in parallel? How to handle merge conflicts? Reference GSD wave execution pattern.
Research first: How does GSD wave execution work in detail? How do task orchestrators (Temporal, Prefect) handle parallel dependencies? What conflict resolution patterns exist?
Suggested location: stateless-software-engineering-framework.md (new section 2.4 or Appendix)
SAM: Multi-Model Strategy
Source: Gap analysis of SAM framework
Added: 2026-02-01
Description: Define guidance for using different models for different agent types. E.g., cheaper/faster models for simple verification, stronger models for planning.
Research first: How do agent frameworks handle model selection? What cost/quality tradeoffs exist? How does Claude Code's haiku/sonnet/opus selection work?
Suggested location: stateless-software-engineering-framework.md (Implementation Roadmap or new Appendix)
SAM: Audit Trail / Observability
Source: Gap analysis of SAM framework
Added: 2026-02-01
Description: Beyond artifacts, define logging/metrics/tracing guidance. How to diagnose pipeline issues? What telemetry to capture?
Research first: How do GSD and BMAD-METHOD handle logging? What observability patterns exist in agent frameworks? OpenTelemetry for LLM workflows?
Suggested location: stateless-software-engineering-framework.md (new Appendix I)
SAM: Partial Success Handling
Source: Gap analysis of SAM framework
Added: 2026-02-01
Description: Define how to represent and handle partial task success. Task completes some DoD items but not all. How is this state represented in artifacts?
Research first: How do GSD checkpoints represent partial progress? How do CI/CD systems handle partial test passes? What state machine patterns exist?
Suggested location: stateless-software-engineering-framework.md (section 3.5 Execution Agent output)
SAM: Context Size Management
Source: Gap analysis of SAM framework
Added: 2026-02-01
Description: Define explicit guidance for measuring and managing context size per agent. What's the target token budget? How to detect context pressure?
Research first: How do agent frameworks measure context usage? What token counting approaches exist? How does Claude Code handle context limits internally?
Suggested location: stateless-software-engineering-framework.md (section 2.1 or Appendix C)
SAM: Conflicting Review Findings
Source: Gap analysis of SAM framework
Added: 2026-02-01
Description: Define protocol when forensic review and self-verification disagree. Which takes precedence? How to adjudicate conflicts?
Research first: How do code review systems handle conflicting reviewers? What adjudication patterns exist in multi-agent systems? How does GSD handle verification disagreements?
Suggested location: stateless-software-engineering-framework.md (section 3.6 Forensic Review)
Multi-session build state lost during context compaction
Source: agentskill-kaizen plugin build (2026-02-18), 3 sessions with 1 compaction
Added: 2026-02-18
Description: During the agentskill-kaizen build (8-phase /plugin-dev:create-plugin workflow), context compaction mid-build converted structured task state into a narrative summary. The resuming session had to reconstruct "what's done vs pending" from prose rather than a checklist. Background agent results that were already consumed and applied reappeared as late notifications after compaction, requiring manual deduplication ("did I already handle this?"). No persistent artifact tracked phase completion, commit SHAs per phase, deferred items, or agent result consumption status.
Observed symptoms:
- Phase completion status existed only in ephemeral context — lost on compaction
- Background agent notifications arrived after their findings were already applied (3 duplicate notifications)
- Plan committed early (
87a0b93) diverged from actual implementation but was never updated - No mechanism to mark agent results as "consumed" — each notification required re-evaluation
/plugin-dev:create-plugin workflow lacks intra-phase parallelism tracking
Source: agentskill-kaizen plugin build (2026-02-18) Added: 2026-02-18 Description: The 8-phase create-plugin workflow treats each phase as a serial step, but Phase 5 (Implementation) actually consisted of 6 parallel sub-tasks and Phase 6 (Validation) spawned 3 parallel review agents. The workflow provides no structure for tracking parallel work within a phase — no task dependencies, no completion gates, no way to know which sub-tasks are done after compaction. Batching validation fixes by file rather than by finding would also have been more efficient (SKILL.md was edited 3 separate times when one pass would have sufficed).
Background agent result deduplication after compaction
Source: agentskill-kaizen plugin build sessions 2-3 (2026-02-18)
Added: 2026-02-18
Description: Background agents launched in session N may complete after context compaction or session restart. The system delivers their results as <task-notification> messages, but there is no mechanism to mark results as already consumed. During the kaizen build, 3 review agents (plugin-validator, 2x skill-reviewer) completed during Phase 6 and their findings were applied in commit 0d61480. After compaction, all 3 re-delivered their notifications in session 3, requiring manual evaluation each time ("was this already handled?"). A persistent state file (e.g., .planning/kaizen/agent-results.json tracking task IDs → consumed/pending) would eliminate this waste.
/plugin-dev:create-plugin Phase 6 validation should batch fixes by file, not by finding
Source: agentskill-kaizen plugin build Phase 6 (2026-02-18) Added: 2026-02-18 Description: Phase 6 collected findings from 3 parallel review agents, then applied fixes one finding at a time. This resulted in SKILL.md being edited 3 separate times (description rewrite, SQL removal, MCP server name fix) when a single pass through the file would have applied all fixes together. The workflow should group all findings by file, then make one editing pass per file. Reduces Edit tool calls and context consumed by repeated reads.
Plan artifact diverges from implementation without update mechanism
Source: agentskill-kaizen plugin build (2026-02-18), plan committed as 87a0b93
Added: 2026-02-18
Description: The research/plan from Phase 2-3 was committed early as a markdown file. During implementation (Phase 5), decisions changed — the MCP server grew from planned scope, analysis dimensions were rebalanced, hook patterns shifted. The plan was never updated to reflect actual implementation. After compaction, the stale plan became a potential source of confusion. Two options to address: (1) update the plan artifact after each phase, or (2) treat the plan as disposable and track only the living state (what's done, what's pending, what deferred).
Evaluate scikit-learn dependency weight for agentskill-kaizen cluster_sessions tool
Source: agentskill-kaizen MCP server review (2026-02-18)
Added: 2026-02-18
Description: The cluster_sessions tool in plugins/agentskill-kaizen/mcp/server.py uses scikit-learn (KMeans, CountVectorizer) for session clustering. scikit-learn pulls in ~40MB of transitive dependencies (numpy, scipy, joblib, threadpoolctl). The typical use case is clustering dozens of sessions, not thousands. The Phase 1-2 research produced no durable artifact evaluating library choices — scikit-learn was assumed from the backlog without validation. Investigate whether a lighter alternative (e.g., pyclustering, stdlib-based implementation, or just numpy directly) would suffice for this scale, and whether the KMeans-on-bag-of-words approach is even appropriate for tool-call sequence similarity.
Research first: What clustering approaches work for short categorical sequences? Is cosine similarity on bag-of-words tool vectors meaningful for workflow comparison? What lightweight Python clustering libraries exist that don't pull in scipy?
github_project_setup.py: add milestone close command
Source: PR #149 follow-up — start-milestone automation (2026-02-22)
Added: 2026-02-22
Completed: 2026-02-22
Status: DONE — Added milestone close subcommand to github_project_setup.py: validates milestone is open, lists open issues, transitions status:in-progress → status:done, closes milestone, prints summary. Added status:done to LABELS taxonomy. Added _transition_to_done() helper. Updated complete-milestone/SKILL.md Step 4 to reference the script command (consistent with start-milestone).
Description: github_project_setup.py now has milestone start which bulk-transitions status:needs-grooming → status:in-progress. Add a symmetric milestone close command for the complete-milestone skill. The command should: (1) validate the milestone is open, (2) list all still-open issues (warn if any remain), (3) transition open issues from status:in-progress → status:done or close them, (4) close the milestone itself via milestone.edit(state="closed"), (5) print a completion summary. Update complete-milestone/SKILL.md to reference the script (consistent with start-milestone).
Suggested location: .claude/skills/gh/scripts/github_project_setup.py — add milestone close subcommand under milestone_app; update .claude/skills/complete-milestone/SKILL.md
work-backlog-item: accept #N in close and resolve routing
**Sour
…(truncated)