Watching AI Format Evolution
Overview
Run a manual format-watch pass that prioritizes official evidence, extracts only parser-relevant structural signals, and produces a bounded maintenance report.
Core principle: separate confirmed upstream evidence from local inference.
Required Output Contract
Unless the user explicitly asks for a different structure, the response MUST be organized as a separate section for each assistant checked or considered.
This applies to general monthly watch passes, monthly summaries, watch updates, exploratory checks, and triage requests.
For each assistant section, use these exact headings in this exact order:
Sources checked
Confirmed deltas
Likely deltas
Unknowns
Confidence
Risk
Inspect next
Recommended action
Do not rename, merge, reorder, omit, or replace these headings.
If a heading has no content, write an explicit placeholder such as None confirmed, Unknown, or Not checked.
The exact heading contract is mandatory. Treat it like a schema, not a stylistic preference.
Use ## <Assistant name> for the assistant section heading, then the exact eight ### headings below it.
Do not answer this skill as freeform prose.
Format Violations
The following are violations of this skill:
- Combining multiple assistants into one shared report
- Using only some of the required headings
- Replacing headings with close variants such as
Evidence, Findings, Next steps, or Action
- Writing prose summaries instead of the required per-assistant template
- Using bold assistant names instead of
## <Assistant name> headings
- Omitting headings because the answer feels obvious or evidence is thin
Partial compliance is non-compliance.
When to Use
- User asks to watch, monitor, track, or review AI assistant format evolution.
- User wants a periodic manual check on supported or candidate assistants.
- User wants to know whether docs, fixtures, tests, or parser assumptions may be stale.
- User explicitly does not want GitHub Actions or an external bot/service.
When NOT to Use
- User wants parser implementation now.
- User is debugging one concrete parser failure with a reproducible fixture.
- User wants broad market research rather than storage/transcript format evidence.
Hard Rules
- Official source code, migrations, and typed protocol definitions are primary evidence.
- Official docs, changelogs, releases, and package metadata are secondary evidence.
- Local session capture and fixture diffing are supporting evidence by default.
- For closed-source assistants such as Claude Code, a fresh real session diff can justify docs, fixtures, and tests updates, but not parser edits by itself.
- Community posts are context only.
- For assistants with public persistence source, inspect sensitive source files before release notes.
- Label every conclusion as
confirmed, likely, or unknown.
- Do not recommend parser edits without concrete format evidence.
- Do not default to GitHub Actions, OpenClaw-style services, or autonomous monitoring.
Watch Loop
- Build the watchlist.
Critical: already-supported assistants only.
In this repo that means Claude Code, OpenCode, Codex, and Mistral Vibe.
Exploratory: candidate assistants only if the user asks.
- Collect authoritative evidence per assistant.
- Extract parser-relevant signals only: storage paths, file naming, top-level fields, enum/variant values, SQLite tables/columns, tool call linkage, subagent markers.
- Compare those signals to local docs, fixtures, tests, and parser assumptions.
- Produce a bounded report with confidence, risk, impacted local paths, and smallest justified next step.
Assistant Matrix
| Assistant | Watch level | Check first | Then check | Watch for |
|---|---|---|---|
| OpenCode | critical | session types, Drizzle schema, storage paths in official source | release notes | SQLite tables/columns, part types, path changes |
| Codex | critical | Rust protocol/session types, rollout recorder, session path code | release notes | enum variants, JSONL layout, storage naming |
| Mistral Vibe | critical | Python session/message classes, session logging config | release notes/docs | roles, fields, filename pattern |
| Claude Code | critical | official changelog/releases, package or SDK session types, local sessions | repo docs | path changes, event field drift, partial/incomplete transcript behavior |
| Gemini CLI | exploratory only | official docs plus chat/session persistence source | release notes | JSON structure, file layout, schema stability |
Output Format
For each assistant, report exactly these sections:
Sources checked: official URLs, packages, commits, or versions
Confirmed deltas: only evidence-backed structural changes
Likely deltas: plausible impact not yet proven
Unknowns: what still needs direct evidence
Confidence: low, medium, or high
Risk: low, medium, or high
Inspect next: existing local paths only
Recommended action: docs, fixtures, tests, parser, or no action yet
Before sending the final answer, perform a heading check:
- count assistant sections
- verify all 8 headings are present in each section
- verify heading text matches exactly
- if any heading is missing, add it before responding
Example
## Claude Code
### Sources checked
- official changelog vX.Y.Z
- npm package @anthropic-ai/claude-code vX.Y.Z
- fresh local session sample from ~/.claude/projects/...
### Confirmed deltas
- changelog mentions storage migration
### Likely deltas
- local docs about storage paths may be stale
### Unknowns
- whether event field names changed or only storage location changed
### Confidence
medium
### Risk
medium
### Inspect next
- docs/session-formats/claude-code.md
- tests/fixtures/claude_sessions/
- src/parsers/claude_code.rs
### Recommended action
- refresh docs and fixtures first
- hold parser edits until a fresh fixture or source/type diff shows parser-impacting drift
Common Mistakes
| Mistake |
Fix |
| Treating release notes as schema proof |
For open-source assistants, inspect source/types first |
| Mixing upstream facts with repo assumptions |
Label repo-side impact as inference, not evidence |
| Recommending parser edits too early |
Update docs/fixtures/tests first unless parser breakage is confirmed |
| Naming local files that were not verified |
Only cite existing repository paths |
| Treating one fresh session as a stable contract |
Use it as confirmation, not sole authority |
Rationalization Table
| Excuse |
Reality |
| "Release notes are enough" |
They are secondary evidence for assistants with public source. |
| "Our parser assumptions imply upstream drift" |
They show impact surface, not proof of a format change. |
| "One local session proves the contract changed" |
It proves one sample changed, not that the upstream format is stable or general. |
| "I should suggest CI or a bot" |
This skill is for manual watch passes unless the user expands scope. |
| "I can recommend parser edits now to be safe" |
Parser edits need concrete structural evidence, not precaution alone. |
Red Flags - Stop and Correct
- Claiming
confirmed without a cited official source, package, or checked file
- Recommending parser edits without a specific field, variant, path, or schema delta
- Using community chatter as primary evidence
- Listing impacted local paths that do not exist
- Defaulting to GitHub Actions or external agent services
1---2name: watching-ai-format-evolution3description: Use when manually monitoring, watching, tracking, or reviewing AI assistant storage, session, transcript, JSONL, or SQLite format drift after official upstream repository, changelog, or package updates, especially when fixtures, parser docs, tests, or parsers may be stale.4---5
6# Watching AI Format Evolution
7
8## Overview
9
10Run a manual format-watch pass that prioritizes official evidence, extracts only parser-relevant structural signals, and produces a bounded maintenance report.
11Core principle: separate confirmed upstream evidence from local inference.
12
13## Required Output Contract
14
15Unless the user explicitly asks for a different structure, the response MUST be organized as a separate section for each assistant checked or considered.
16This applies to general monthly watch passes, monthly summaries, watch updates, exploratory checks, and triage requests.
17
18For each assistant section, use these exact headings in this exact order:
19
201. `Sources checked`
212. `Confirmed deltas`
223. `Likely deltas`
234. `Unknowns`
245. `Confidence`
256. `Risk`
267. `Inspect next`
278. `Recommended action`
28
29Do not rename, merge, reorder, omit, or replace these headings.
30If a heading has no content, write an explicit placeholder such as `None confirmed`, `Unknown`, or `Not checked`.
31The exact heading contract is mandatory. Treat it like a schema, not a stylistic preference.
32Use `## <Assistant name>` for the assistant section heading, then the exact eight `###` headings below it.
33Do not answer this skill as freeform prose.
34
35## Format Violations
36
37The following are violations of this skill:
38
39- Combining multiple assistants into one shared report
40- Using only some of the required headings
41- Replacing headings with close variants such as `Evidence`, `Findings`, `Next steps`, or `Action`
42- Writing prose summaries instead of the required per-assistant template
43- Using bold assistant names instead of `## <Assistant name>` headings
44- Omitting headings because the answer feels obvious or evidence is thin
45
46Partial compliance is non-compliance.
47
48## When to Use
49
50- User asks to watch, monitor, track, or review AI assistant format evolution.
51- User wants a periodic manual check on supported or candidate assistants.
52- User wants to know whether docs, fixtures, tests, or parser assumptions may be stale.
53- User explicitly does not want GitHub Actions or an external bot/service.
54
55## When NOT to Use
56
57- User wants parser implementation now.
58- User is debugging one concrete parser failure with a reproducible fixture.
59- User wants broad market research rather than storage/transcript format evidence.
60
61## Hard Rules
62
63- Official source code, migrations, and typed protocol definitions are primary evidence.
64- Official docs, changelogs, releases, and package metadata are secondary evidence.
65- Local session capture and fixture diffing are supporting evidence by default.
66- For closed-source assistants such as Claude Code, a fresh real session diff can justify docs, fixtures, and tests updates, but not parser edits by itself.
67- Community posts are context only.
68- For assistants with public persistence source, inspect sensitive source files before release notes.
69- Label every conclusion as `confirmed`, `likely`, or `unknown`.
70- Do not recommend parser edits without concrete format evidence.
71- Do not default to GitHub Actions, OpenClaw-style services, or autonomous monitoring.
72
73## Watch Loop
74
751. Build the watchlist.
76Critical: already-supported assistants only.
77In this repo that means Claude Code, OpenCode, Codex, and Mistral Vibe.
78Exploratory: candidate assistants only if the user asks.
792. Collect authoritative evidence per assistant.
803. Extract parser-relevant signals only: storage paths, file naming, top-level fields, enum/variant values, SQLite tables/columns, tool call linkage, subagent markers.
814. Compare those signals to local docs, fixtures, tests, and parser assumptions.
825. Produce a bounded report with confidence, risk, impacted local paths, and smallest justified next step.
83
84## Assistant Matrix
85
86| Assistant | Watch level | Check first | Then check | Watch for |
87|---|---|---|---|
88| OpenCode | critical | session types, Drizzle schema, storage paths in official source | release notes | SQLite tables/columns, part types, path changes |
89| Codex | critical | Rust protocol/session types, rollout recorder, session path code | release notes | enum variants, JSONL layout, storage naming |
90| Mistral Vibe | critical | Python session/message classes, session logging config | release notes/docs | roles, fields, filename pattern |
91| Claude Code | critical | official changelog/releases, package or SDK session types, local sessions | repo docs | path changes, event field drift, partial/incomplete transcript behavior |
92| Gemini CLI | exploratory only | official docs plus chat/session persistence source | release notes | JSON structure, file layout, schema stability |
93
94## Output Format
95
96For each assistant, report exactly these sections:
97
98- `Sources checked:` official URLs, packages, commits, or versions
99- `Confirmed deltas:` only evidence-backed structural changes
100- `Likely deltas:` plausible impact not yet proven
101- `Unknowns:` what still needs direct evidence
102- `Confidence:` low, medium, or high
103- `Risk:` low, medium, or high
104- `Inspect next:` existing local paths only
105- `Recommended action:` docs, fixtures, tests, parser, or no action yet
106
107Before sending the final answer, perform a heading check:
108
109- count assistant sections
110- verify all 8 headings are present in each section
111- verify heading text matches exactly
112- if any heading is missing, add it before responding
113
114## Example
115
116```markdown
117## Claude Code
118
119### Sources checked
120- official changelog vX.Y.Z
121- npm package @anthropic-ai/claude-code vX.Y.Z
122- fresh local session sample from ~/.claude/projects/...
123
124### Confirmed deltas
125- changelog mentions storage migration
126
127### Likely deltas
128- local docs about storage paths may be stale
129
130### Unknowns
131- whether event field names changed or only storage location changed
132
133### Confidence
134medium
135
136### Risk
137medium
138
139### Inspect next
140- docs/session-formats/claude-code.md
141- tests/fixtures/claude_sessions/
142- src/parsers/claude_code.rs
143
144### Recommended action
145- refresh docs and fixtures first
146- hold parser edits until a fresh fixture or source/type diff shows parser-impacting drift
147```
148
149## Common Mistakes
150
151| Mistake | Fix |
152|---|---|
153| Treating release notes as schema proof | For open-source assistants, inspect source/types first |
154| Mixing upstream facts with repo assumptions | Label repo-side impact as inference, not evidence |
155| Recommending parser edits too early | Update docs/fixtures/tests first unless parser breakage is confirmed |
156| Naming local files that were not verified | Only cite existing repository paths |
157| Treating one fresh session as a stable contract | Use it as confirmation, not sole authority |
158
159## Rationalization Table
160
161| Excuse | Reality |
162|---|---|
163| "Release notes are enough" | They are secondary evidence for assistants with public source. |
164| "Our parser assumptions imply upstream drift" | They show impact surface, not proof of a format change. |
165| "One local session proves the contract changed" | It proves one sample changed, not that the upstream format is stable or general. |
166| "I should suggest CI or a bot" | This skill is for manual watch passes unless the user expands scope. |
167| "I can recommend parser edits now to be safe" | Parser edits need concrete structural evidence, not precaution alone. |
168
169## Red Flags - Stop and Correct
170
171- Claiming `confirmed` without a cited official source, package, or checked file
172- Recommending parser edits without a specific field, variant, path, or schema delta
173- Using community chatter as primary evidence
174- Listing impacted local paths that do not exist
175- Defaulting to GitHub Actions or external agent services