Uberassess
Core rule
uberassess turns external signals, plan artifacts, idea seeds, and implementation questions into source-grounded recommendations. It owns pre-planning research: clarify what question is really being asked, inspect enough outside and local evidence to avoid stale guesses, and recommend adopt/watch/archive/reject/eval/plan. It does not implement, mutate project state, or launch work without explicit approval.
Use the lightest tier that can answer: is this idea worth adopting, watching, archiving, rejecting, or turning into an eval seed? Most interesting sources should not become implementation plans.
Direct-use only when explicitly named, when the user explicitly asks to assess/evaluate/research a source, plan, artifact, idea, or implementation question for adoption, or when routed by ubergoal.
Relationship to Ubergoal
uberassess is a pre-planning assessment phase.
- If the recommendation is approved for code/skill/workflow changes, hand off to
$ubergoal/$uberplan with the packet as evidence.
- If the source suggests only memory, watchlist, or eval work, do not escalate to implementation.
- Reflective readers may consume or propose assessment packets as read-only candidate signals; assessment does not grant mutation authority.
Tiers
| Tier |
Use for |
Output |
| 0 |
Quick triage of low-stakes links |
save / ignore / needs deeper assessment |
| 1 |
Normal source, plan, or contained idea assessment |
templates/assessment-packet.md |
| 2 |
Important idea/question likely to affect systems, architecture, skills, tools, or project direction |
packet + research frame/source map + project context freshness + alternatives/prior art + benefit >> cost + eval seed |
| 3 |
Likely code/skill/workflow/agentic-system change |
Tier 2 + Agent Advocate RCA + $ubergoal handoff + evidence plan |
Escalate only for concrete risk. Do not spend Tier 3 effort on every bookmark.
Procedure
- Frame the assessment question. Restate the seed in one sentence, identify whether this is source-first, plan-artifact, idea-first, implementation-question, or mixed, and name the research mode: quick source assessment or Deep Research Assessment. If the input is vague, define the smallest useful question instead of asking the user to over-specify.
- Clarify before broad research when it matters. Ask one to three targeted questions when the answer would materially change research direction, source lanes, cost, privacy/side-effect boundaries, or adoption criteria. Do not send a generic questionnaire. If the ambiguity is low-risk or the user asked for speed, state assumptions and proceed; record unanswered clarifications as coverage gaps.
- Resolve the source or seed. Capture raw source handles before summarizing. For source-first work, use source-specific tools/skills when available, e.g. bookmark resolvers, GitHub tooling, arXiv/PDF extraction, web fetch/browser, or transcript/OCR sidecars. For plan-artifact work, capture the draft plan as the raw source and label it
plan_artifact. For idea-first work, record the user's wording as the raw seed and label it idea_seed, implementation_question, or open_research_question. Record failures and uninspected media.
- Create a Source Packet. Use
templates/source-packet.md or embed its fields in the assessment packet. Separate raw source/seed, linked sources, retrieval limits, synthesis, and uncertainty.
- Build the research map. For Deep Research Assessment, intentionally choose lanes: local codebase/source-of-truth, primary docs/specs, GitHub alternatives/prior art, issues/forums/practitioner discussion, papers/blogs where relevant, and contradiction search. State coverage claimed and not claimed; never imply exhaustive research when only a slice was inspected.
- Check project context. Inspect the relevant project/source-of-truth paths/cards/docs enough to avoid stale or duplicate recommendations. Use local adapter references when present, but do not force unrelated local projects into portable assessments. Record context freshness and gaps.
- Extract ideas and claims. Distinguish user/source/author claims from your inferences. Record evidence quality, contradiction risk, local fit, and what is actually actionable.
- Compare alternatives. For implementation questions and architecture/tool ideas, produce an alternatives/prior-art matrix. Include the "do nothing / note / eval-only / improve context or skill" option alongside new machinery.
- Run benefit >> cost. Ask what to delete or simplify first. Prefer notes, eval seeds, or tool/context fixes over new machinery unless benefit is clearly much greater than hidden cost.
- Run Agent Advocate when agent systems are involved. Ask the human counterfactual: would a competent human with the same context/tools have made the error? If not, identify the missing context, source authority, tool feedback, memory, affordance, or approval boundary.
- Decide. Choose Adopt now, Watch, Archive, Reject, Needs more research, or Convert to eval only.
- Set approval boundary. State
Implementation before approval: no. If approved, provide a compact $ubergoal/$uberplan handoff and evidence plan.
- Leave a learning trail. For accepted/rejected high-signal assessments, note how outcome should feed
$uberskillevolver or future read-only review later.
Deep Research Assessment mode
Use this mode when the user says "boil the ocean," asks for deep research, requests alternatives or state-of-the-art, gives only a scrap of an idea, or asks an implementation/architecture question before planning. The job is not to gather every possible source; it is to make the research boundary honest and broad enough that a later plan is not built on vibes.
Minimum lanes to consider:
- Seed/source capture — exact user wording, source URL, linked artifacts, raw media/transcripts, and retrieval limits.
- Question framing and clarification — what decision this assessment must support, what would be overreach, and what targeted user feedback would materially improve the research.
- Local project context — current codebase, skills, docs, tests, policies, and duplicate/prior attempts.
- Primary docs/specs — official docs and source-of-truth references when relevant.
- Alternatives/prior art — GitHub repos, libraries, patterns, architecture options, and why they do or do not transfer.
- Practitioner evidence — issues, discussions, forums, postmortems, or threads that reveal real pain and edge cases.
- Contradiction search — evidence the idea is stale, solved, unsafe, too expensive, or better handled by deletion/simplification.
- Adoption fit — whether the output should be a memory note, watch item, eval seed, rejected idea, or
$ubergoal/$uberplan handoff.
Stop research when additional sources are unlikely to change the recommendation, the coverage gap itself is the recommendation, or approval is needed for paid/private/side-effecting access.
Plan Artifact Assessment mode
Use this mode when the user asks uberassess to assess a plan, or when uberplan needs a review of a draft plan before hardening or implementation. Treat the plan as the source artifact, but also capture the operator original instruction verbatim or by exact artifact path when available. Assess whether the plan matches the operator-original intent, asks the right clarifying questions, has enough research coverage, names assumptions and gaps honestly, maps risks to evidence, and deserves adoption, revision, or rejection. Do not assess only the plan author's summary.
A plan assessment should usually answer:
- Intent fit — does the plan solve the user's actual problem, or did it drift into a nearby process?
- Clarification need — what one to three questions would materially improve the plan before broad work?
- Research sufficiency — does the plan need source/code/docs/forum/GitHub research before implementation?
- Alternatives and deletion — what simpler/no-build/eval-only route should be considered?
- Evidence and adoption — what proof is required before the plan can move to implementation?
- Revision decision — approve as-is, revise then proceed, run deeper assessment, convert to eval/watch, or reject.
Output contract
For Tier 1+, produce an assessment packet with:
- source URL/type and raw capture status
- assessment mode, research question, and clarification checkpoint
- for plan artifacts: intent fit against the operator-original instruction, agent-interpreted scope, proposed narrowed scope, explicit deferrals/non-goals, approval evidence, evidence gaps, Scope fidelity verdict, and revision decision
- research/source map with coverage claimed and not claimed
- linked sources/media inspected and retrieval limitations
- key ideas, author claims, and model inferences
- alternatives/prior-art matrix when implementation or architecture choices are implicated
- project relevance matrix and project context freshness
- source authority, uncertainty, contradictions, and freshness
- benefit >> complexity cost analysis and simpler alternatives
- decision and rationale
- approval boundary and side effects
- if implementation is likely:
$ubergoal handoff, evidence plan, rollback/stop condition
- if agentic systems are implicated: Agent Advocate human counterfactual and affordance gap
- outcome-learning trail for future evolver or read-only review
Use scripts/validate_assessment_packet.py before treating a Tier 1+ packet as complete.
Do not build yet by default
Do not create MCP servers, scrapers, vector indexes, scheduled automation, model swaps, or new persistent state during assessment. Recommend them only after repeated usage proves a stable interface and benefit >> cost.
Architecture stepback route
Before recommending implementation for a system-scale question involving concurrency, scaling, queues, workers, long-running jobs, gateway stalls, orchestration, workflow durability, backpressure, repeated timeouts, or a pattern of symptom patches, route to $uberarchitect or embed its Architecture Stepback Packet. The recommendation is incomplete until it names the system class, normal industry architecture, fresh-start architecture, current mismatch, symptom patches demoted, smallest transition path, proof gate, and human counterfactual.
Optional Claude adversary
Contract: ../references/claude-adversary.md (opt-in only when the operator explicitly requests Claude by name; reconciliation + frame-independence rules there).
For uberassess, ask exactly:
- Source-lane sufficiency. Causal layer: source authority. Did assessment consult the relevant source lane, or stop at the first confirming source? Evidence: source map plus coverage gap. Minimum impact: inspect the missing lane or limit the claim.
- Actionability boundary. Causal layer: ownership/approval. Is the recommendation directly actionable by the agent, or does it require human escalation? Evidence: next action plus owner. Minimum impact: change decision to watch/archive/escalate if not actionable.
- 90-day falsifier. Causal layer: freshness. What would change this recommendation in 90 days? Evidence: named trigger/source. Minimum impact: add watch trigger or reduce confidence.
Helpful resources
templates/assessment-packet.md — canonical recommendation packet.
templates/source-packet.md — source capture subtemplate.
templates/project-context-card.md — lightweight project context card.
references/source-resolvers.md — source-type handling and limitations.
references/project-routing.md — routing destinations and non-goals.
references/hermes-and-approval.md — read-only reflection and approval policy.
scripts/validate_assessment_packet.py — packet validator.
evals/golden_skill_invocations.json — trigger/non-trigger examples for routing changes; load only when tuning assessment triggers.
1---2name: uberassess3description: Do not auto-trigger from task similarity. Use only when explicitly asked to assess or deeply research a source, plan artifact, idea seed, open research question, or implementation question for possible adoption, or routed by ubergoal: X/Twitter posts, bookmarked links, articles, GitHub repos, arXiv/papers, videos, Hermes/bookmark signals, internal artifacts, draft plans, planning packets, scraps of ideas, alternatives research, and codebase/docs/forum/GitHub reconnaissance. Produces a source-grounded recommendation packet, not implementation.4---56# Uberassess78## Core rule910`uberassess` turns external signals, plan artifacts, idea seeds, and implementation questions into **source-grounded recommendations**. It owns pre-planning research: clarify what question is really being asked, inspect enough outside and local evidence to avoid stale guesses, and recommend adopt/watch/archive/reject/eval/plan. It does not implement, mutate project state, or launch work without explicit approval.1112Use the lightest tier that can answer: **is this idea worth adopting, watching, archiving, rejecting, or turning into an eval seed?** Most interesting sources should not become implementation plans.1314Direct-use only when explicitly named, when the user explicitly asks to assess/evaluate/research a source, plan, artifact, idea, or implementation question for adoption, or when routed by `ubergoal`.1516## Relationship to Ubergoal1718- `uberassess` is a pre-planning assessment phase.19- If the recommendation is approved for code/skill/workflow changes, hand off to `$ubergoal`/`$uberplan` with the packet as evidence.20- If the source suggests only memory, watchlist, or eval work, do not escalate to implementation.21- Reflective readers may consume or propose assessment packets as read-only candidate signals; assessment does not grant mutation authority.2223## Tiers2425| Tier | Use for | Output |26|---|---|---|27| 0 | Quick triage of low-stakes links | save / ignore / needs deeper assessment |28| 1 | Normal source, plan, or contained idea assessment | `templates/assessment-packet.md` |29| 2 | Important idea/question likely to affect systems, architecture, skills, tools, or project direction | packet + research frame/source map + project context freshness + alternatives/prior art + benefit >> cost + eval seed |30| 3 | Likely code/skill/workflow/agentic-system change | Tier 2 + Agent Advocate RCA + `$ubergoal` handoff + evidence plan |3132Escalate only for concrete risk. Do not spend Tier 3 effort on every bookmark.3334## Procedure35360. **Frame the assessment question.** Restate the seed in one sentence, identify whether this is source-first, plan-artifact, idea-first, implementation-question, or mixed, and name the research mode: quick source assessment or Deep Research Assessment. If the input is vague, define the smallest useful question instead of asking the user to over-specify.371. **Clarify before broad research when it matters.** Ask one to three targeted questions when the answer would materially change research direction, source lanes, cost, privacy/side-effect boundaries, or adoption criteria. Do not send a generic questionnaire. If the ambiguity is low-risk or the user asked for speed, state assumptions and proceed; record unanswered clarifications as coverage gaps.382. **Resolve the source or seed.** Capture raw source handles before summarizing. For source-first work, use source-specific tools/skills when available, e.g. bookmark resolvers, GitHub tooling, arXiv/PDF extraction, web fetch/browser, or transcript/OCR sidecars. For plan-artifact work, capture the draft plan as the raw source and label it `plan_artifact`. For idea-first work, record the user's wording as the raw seed and label it `idea_seed`, `implementation_question`, or `open_research_question`. Record failures and uninspected media.393. **Create a Source Packet.** Use `templates/source-packet.md` or embed its fields in the assessment packet. Separate raw source/seed, linked sources, retrieval limits, synthesis, and uncertainty.404. **Build the research map.** For Deep Research Assessment, intentionally choose lanes: local codebase/source-of-truth, primary docs/specs, GitHub alternatives/prior art, issues/forums/practitioner discussion, papers/blogs where relevant, and contradiction search. State coverage claimed and not claimed; never imply exhaustive research when only a slice was inspected.415. **Check project context.** Inspect the relevant project/source-of-truth paths/cards/docs enough to avoid stale or duplicate recommendations. Use local adapter references when present, but do not force unrelated local projects into portable assessments. Record context freshness and gaps.426. **Extract ideas and claims.** Distinguish user/source/author claims from your inferences. Record evidence quality, contradiction risk, local fit, and what is actually actionable.437. **Compare alternatives.** For implementation questions and architecture/tool ideas, produce an alternatives/prior-art matrix. Include the "do nothing / note / eval-only / improve context or skill" option alongside new machinery.448. **Run benefit >> cost.** Ask what to delete or simplify first. Prefer notes, eval seeds, or tool/context fixes over new machinery unless benefit is clearly much greater than hidden cost.459. **Run Agent Advocate when agent systems are involved.** Ask the human counterfactual: would a competent human with the same context/tools have made the error? If not, identify the missing context, source authority, tool feedback, memory, affordance, or approval boundary.4610. **Decide.** Choose Adopt now, Watch, Archive, Reject, Needs more research, or Convert to eval only.4711. **Set approval boundary.** State `Implementation before approval: no`. If approved, provide a compact `$ubergoal`/`$uberplan` handoff and evidence plan.4812. **Leave a learning trail.** For accepted/rejected high-signal assessments, note how outcome should feed `$uberskillevolver` or future read-only review later.4950## Deep Research Assessment mode5152Use this mode when the user says "boil the ocean," asks for deep research, requests alternatives or state-of-the-art, gives only a scrap of an idea, or asks an implementation/architecture question before planning. The job is not to gather every possible source; it is to make the research boundary honest and broad enough that a later plan is not built on vibes.5354Minimum lanes to consider:5556- **Seed/source capture** — exact user wording, source URL, linked artifacts, raw media/transcripts, and retrieval limits.57- **Question framing and clarification** — what decision this assessment must support, what would be overreach, and what targeted user feedback would materially improve the research.58- **Local project context** — current codebase, skills, docs, tests, policies, and duplicate/prior attempts.59- **Primary docs/specs** — official docs and source-of-truth references when relevant.60- **Alternatives/prior art** — GitHub repos, libraries, patterns, architecture options, and why they do or do not transfer.61- **Practitioner evidence** — issues, discussions, forums, postmortems, or threads that reveal real pain and edge cases.62- **Contradiction search** — evidence the idea is stale, solved, unsafe, too expensive, or better handled by deletion/simplification.63- **Adoption fit** — whether the output should be a memory note, watch item, eval seed, rejected idea, or `$ubergoal`/`$uberplan` handoff.6465Stop research when additional sources are unlikely to change the recommendation, the coverage gap itself is the recommendation, or approval is needed for paid/private/side-effecting access.6667## Plan Artifact Assessment mode6869Use this mode when the user asks `uberassess` to assess a plan, or when `uberplan` needs a review of a draft plan before hardening or implementation. Treat the plan as the source artifact, but also capture the operator original instruction verbatim or by exact artifact path when available. Assess whether the plan matches the operator-original intent, asks the right clarifying questions, has enough research coverage, names assumptions and gaps honestly, maps risks to evidence, and deserves adoption, revision, or rejection. Do not assess only the plan author's summary.7071A plan assessment should usually answer:7273- **Intent fit** — does the plan solve the user's actual problem, or did it drift into a nearby process?74- **Clarification need** — what one to three questions would materially improve the plan before broad work?75- **Research sufficiency** — does the plan need source/code/docs/forum/GitHub research before implementation?76- **Alternatives and deletion** — what simpler/no-build/eval-only route should be considered?77- **Evidence and adoption** — what proof is required before the plan can move to implementation?78- **Revision decision** — approve as-is, revise then proceed, run deeper assessment, convert to eval/watch, or reject.7980## Output contract8182For Tier 1+, produce an assessment packet with:8384- source URL/type and raw capture status85- assessment mode, research question, and clarification checkpoint86- for plan artifacts: intent fit against the operator-original instruction, agent-interpreted scope, proposed narrowed scope, explicit deferrals/non-goals, approval evidence, evidence gaps, Scope fidelity verdict, and revision decision87- research/source map with coverage claimed and not claimed88- linked sources/media inspected and retrieval limitations89- key ideas, author claims, and model inferences90- alternatives/prior-art matrix when implementation or architecture choices are implicated91- project relevance matrix and project context freshness92- source authority, uncertainty, contradictions, and freshness93- benefit >> complexity cost analysis and simpler alternatives94- decision and rationale95- approval boundary and side effects96- if implementation is likely: `$ubergoal` handoff, evidence plan, rollback/stop condition97- if agentic systems are implicated: Agent Advocate human counterfactual and affordance gap98- outcome-learning trail for future evolver or read-only review99100Use `scripts/validate_assessment_packet.py` before treating a Tier 1+ packet as complete.101102## Do not build yet by default103104Do not create MCP servers, scrapers, vector indexes, scheduled automation, model swaps, or new persistent state during assessment. Recommend them only after repeated usage proves a stable interface and benefit >> cost.105106107## Architecture stepback route108109Before recommending implementation for a system-scale question involving concurrency, scaling, queues, workers, long-running jobs, gateway stalls, orchestration, workflow durability, backpressure, repeated timeouts, or a pattern of symptom patches, route to `$uberarchitect` or embed its Architecture Stepback Packet. The recommendation is incomplete until it names the system class, normal industry architecture, fresh-start architecture, current mismatch, symptom patches demoted, smallest transition path, proof gate, and human counterfactual.110111## Optional Claude adversary112113Contract: `../references/claude-adversary.md` (opt-in only when the operator explicitly requests Claude by name; reconciliation + frame-independence rules there).114115For `uberassess`, ask exactly:1161171. **Source-lane sufficiency.** Causal layer: source authority. Did assessment consult the relevant source lane, or stop at the first confirming source? Evidence: source map plus coverage gap. Minimum impact: inspect the missing lane or limit the claim.1182. **Actionability boundary.** Causal layer: ownership/approval. Is the recommendation directly actionable by the agent, or does it require human escalation? Evidence: next action plus owner. Minimum impact: change decision to watch/archive/escalate if not actionable.1193. **90-day falsifier.** Causal layer: freshness. What would change this recommendation in 90 days? Evidence: named trigger/source. Minimum impact: add watch trigger or reduce confidence.120121## Helpful resources122123- `templates/assessment-packet.md` — canonical recommendation packet.124- `templates/source-packet.md` — source capture subtemplate.125- `templates/project-context-card.md` — lightweight project context card.126- `references/source-resolvers.md` — source-type handling and limitations.127- `references/project-routing.md` — routing destinations and non-goals.128- `references/hermes-and-approval.md` — read-only reflection and approval policy.129- `scripts/validate_assessment_packet.py` — packet validator.130- `evals/golden_skill_invocations.json` — trigger/non-trigger examples for routing changes; load only when tuning assessment triggers.