Systems Engineering
Principle expression
Primary: P03
Supporting: P04, P13, P15
Scope
Own one judgment: under the actual goal, operating conditions, failure
consequences, and acceptable residual risk, how should fallible parts be related,
observed, corrected, and governed so the whole behaves reliably enough to use
and maintain?
Reliability here is approximate and evidence-relative. Humans, models, tools,
and reviewers will make mistakes. Engineering does not make each part perfect;
it makes material error observable, bounded, correctable, or explicitly
accepted at a cost proportionate to consequence.
This Skill designs the system relation. It does not execute a Work Cell or
Swarm, prove a model primitive, invent domain acceptance, choose every local
task packet, or approve its own result. It is not a mandatory preflight and does
not turn every task into a control diagram.
Principle source
Use a host Sequence and matching interpretations when the host declares them.
Otherwise use this package's read-only fallback in references/sequence.md.
Read only P03, P04, P13, and P15. A live task may select a different current
lead without changing this Skill's stable lineage.
Start from the actual system
Recover only what changes the engineering decision:
Desired whole behavior and acceptance owner:
Operating conditions, source state, and hard constraints:
Failure consequence and acceptable residual risk:
Actual path from input through effects to acceptance:
Observed failure, disturbance, variance, or near miss:
Available signals, controls, recovery, and resource margins:
Observation that would show the present system is already sufficient:
If the task is low-consequence, reversible, locally observable, and already has
an adequate correction path, retain the present form. Do not manufacture a
System Case.
Read concepts only when terms such as feedback,
stability, redundancy, independence, or residual risk affect the decision.
Read evaluation only when testing a new control
structure or claiming improved system reliability.
Specificity discipline
A system requirement can be correct while its implementation is still
undecided. False precision hides that boundary and can silently take another
owner's authority. Use the most concrete form supported by the supplied source:
- if a field, threshold, duration, batch size, sample, protocol, or role is
explicitly governed, preserve it and name its source;
- if only a property is governed, state the property and its acceptance signal;
use
[owning domain/runtime to determine] for the representation; and
- if even the property lacks evidence, return discovery rather than a plausible
mechanism.
Before submitting, audit every concrete name and number introduced by the
answer. Remove or bound any that came from the Agent rather than the case or a
governing source. Do not contradict this audit by naming a proposed field and
then claiming only to have stated a property.
Self-audit is not verification. For a consequential System Case, treat the
first result as a candidate until a designated independent verifier compares
each introduced mechanism, threshold, duration, sample, protocol, and role with
the supplied sources and ownership boundary. Give the verifier this narrow
claim/source task and only the governing sources plus accepted candidate
payload; do not inject the full reasoning trace or ask it to redesign the
system. Revise only the rejected specifics while preserving supported control
requirements.
Core method
Define sufficient behavior, not perfection. State the outcome, operating
range, acceptance condition, failure consequence, and tolerable residual
risk. “No mistakes” is not an engineering specification. Make risk tolerance
specific enough to distinguish harmless variation, recoverable failure, and
unacceptable escape.
Model the whole path. Follow the real movement from inputs and source
state through judgments, effects, verification, acceptance, and later use.
Name actors and components only where their relation changes behavior. A
list of Agents, tools, or workflow stages is not yet a system model.
Identify disturbances and failure paths. Include semantic error, stale
context, omitted obligations, correlated model error, protocol failure,
external change, unsafe effect, reviewer miss, false-positive burden, and
resource exhaustion when relevant. Do not invent an exhaustive hazard
register for a reversible low-risk task.
Find the principal control gap. Select the failure path whose containment
or correction most changes the whole. Preserve secondary hard constraints,
but do not answer every risk with more stages, more prompts, or more Cells.
Make the material state observable. Decide which signal can expose the
failure in time to act: source-linked evidence, a deterministic check,
structured status, disagreement, usage variance, production outcome, or a
later practice result. A control that cannot observe its target is ceremony.
Agent self-report and terminal success remain claims until their designated
verification and settlement. State the semantic property that must be
observed before selecting its representation. A whole-source claim must be
resolved against whole-source evidence before acceptance, but this Skill
does not decide which domain field or runtime protocol carries that evidence.
Assign authority and response. State who may observe, propose, verify,
intervene, commit effects, roll back, accept residual risk, and reopen the
case. Keep decision and execution authority distinct where consequence
requires it. Do not make a worker or review ensemble accept its own output.
Choose the minimum sufficient control structure. Select only mechanisms
that interrupt the named failure path:
- prevent by improving input, context, boundaries, or task form;
- detect through independent evidence, tests, comparison, or observation;
- contain with limited authority, blast radius, staging, or reversible effects;
- recover through retry, repair, rollback, reuse, or escalation; and
- tolerate through justified redundancy or resource margin.
Match control cost to failure consequence. A reversible documentation edit
may need one meaningful check; an irreversible production or strategic
decision may justify diverse independent review, repeated labor,
deterministic gates, prepared recovery, and explicit human acceptance.
More controls are not automatically more reliable: they can add delay,
correlated error, coordination failure, and review burden.
At a boundary owned by another method or runtime, specify the required
control property, observable signal, authority, and failure response, then
hand it off. Do not invent a terminal-tool field, domain packet, deployment
threshold, sample size, time window, or organization role without source
evidence and the owner's authority.
Engineer independence, not vote count. Repetition helps only when the
compared attempts can expose different failure modes. Identical context,
model, prompt, and method may reproduce one systematic error. Vary evidence
boundaries, method, role, model, verifier, or deterministic surface only
where correlated failure matters. Resolve factual disagreement against
sources and designated verification, never by majority alone.
Fit Agent nodes after the system relation is known. Use task-shaping
when one required contribution must be compiled into a stable Agent unit.
Use model-evaluation when its capability is unevidenced. Use a domain Skill
to define semantic partitions and acceptance. The Work Cell or orchestration
runtime carries prepared execution; none of them decides the whole system
relation by default.
Budget for completion and recovery. Begin from necessary work and risk,
then use work-estimation to convert it into profile-specific resources.
Prefer a realistic estimate, explicit margin for important uncertainty, a
high emergency ceiling only when required by the carrier, and post-run
audit of expected versus actual use. Do not treat low token use as quality
or a routine context ceiling as an engineering budget. Estimate variance
is normally a post-run correction signal; it does not itself authorize
interruption, replanning, or escalation during useful work. Add an in-run
resource control only when exhaustion is a named safety or availability
failure and its owner authorizes that response.
Close the loop in operation. Observe end-to-end outcomes, material
defects caught and escaped, false-positive repair burden, recovery, latency,
cost per useful result, and estimate variance. Compare with the unchanged
system when claiming improvement. Revise the system model, control, or risk
acceptance from what actually happened; do not merely append another rule.
Before making any control concrete, apply this authority-and-evidence gate:
Is the mechanism, threshold, duration, schema, or role fixed by source evidence
and owned here?
yes -> specify it and cite or name the governing source
no -> state only the required property, mark the representation
`[owning domain/runtime to determine]`, and route it
“Independent verification,” “staged effects,” and “recoverable rollback” may be
valid system requirements while their exact implementation remains a domain or
runtime decision. Unknown specificity is a discovery item, not an invitation to
invent plausible numbers.
System Case
Return the smallest handoff the consequence requires:
Desired behavior, operating range, and acceptance owner:
System boundary and effect path:
Principal disturbance or failure path:
Observable signal and evidence source:
Control action, authority, and recovery:
Required component contributions and local owners:
Expected work, margin, and audit signal:
Residual risk and who accepts it:
Operational measure and reopening condition:
For several interacting failure paths, a compact control map may help:
| Failure or disturbance |
Observable signal |
Response |
Authority |
Recovery |
Residual risk |
Use neither artifact when a direct explanation or one boundary change is
sufficient. The System Case is a decision carrier, not a universal manifest or
runtime schema.
Boundaries and routing
| Need |
Owner |
| Decide whether the whole is sufficiently reliable under disturbance |
systems-engineering |
| Determine whether one Agent operation fits an evidenced execution envelope |
task-shaping |
| Establish reusable model/provider/harness capability evidence |
model-evaluation |
| Form domain-specific review, refactoring, cognition, or other semantic packets |
owning domain Skill |
| Deliver authoritative context to one prepared unit |
context-engineering |
| Estimate necessary work and convert it into time, tokens, or money |
work-estimation |
| Run Cells, queues, retries, concurrency, and providers |
Work Cell or orchestration runtime |
| Verify domain facts and accept residual risk or effects |
designated verifier and human or host authority |
| Select the next bounded practice after observing operation |
practice-cycle |
Do not use systems-engineering as a synonym for architecture, project
planning, generic rigor, or “use more reviewers.” Do not hide a domain judgment
inside a control label. Do not claim end-to-end reliability from component
success, schema validity, test count, reviewer count, or one clean run.
Do not repair a semantic control gap by changing a generic runtime contract
unless runtime evidence identifies that contract as the owner and the change is
separately authorized.
Completion standard
The engineering judgment is ready when it defines sufficient whole behavior,
models the real effect and acceptance path, identifies the principal material
failure/control gap, connects an observable signal to an authorized response
and recovery, selects controls proportional to consequence, names residual risk
and its acceptance owner, and states an operational observation capable of
showing the design is wrong or incomplete.
For a consequential case, readiness also requires independent source-and-owner
verification of newly introduced specifics. Without that verification, report
a candidate. Without representative operation evidence, report a proposed
system design or trial—not a reliable system.
1---2name: systems-engineering3description: Design or revise a sufficiently reliable whole from fallible human, Agent, software, and organizational parts under concrete constraints and accepted residual risk. Use when a workflow keeps failing despite locally reasonable tasks, when deciding where verification, redundancy, retries, rollback, budget margin, or human acceptance belong, when a Swarm or multi-stage Agent process needs an end-to-end reliability model, when component quality is being confused with system quality, or when asking "how can unreliable agents form a reliable system?", "where should the feedback loop close?", "怎么把会犯错的 Agent 组成可靠系统?" / "如何从流程上保证近似可靠?". Do not use for an ordinary bounded task, generic planning, proving model capability, converting work into token/time/cost estimates, or operating an already prepared queue.4---56# Systems Engineering78## Principle expression910**Primary:** P031112**Supporting:** P04, P13, P151314## Scope1516Own one judgment: **under the actual goal, operating conditions, failure17consequences, and acceptable residual risk, how should fallible parts be related,18observed, corrected, and governed so the whole behaves reliably enough to use19and maintain?**2021Reliability here is approximate and evidence-relative. Humans, models, tools,22and reviewers will make mistakes. Engineering does not make each part perfect;23it makes material error observable, bounded, correctable, or explicitly24accepted at a cost proportionate to consequence.2526This Skill designs the system relation. It does not execute a Work Cell or27Swarm, prove a model primitive, invent domain acceptance, choose every local28task packet, or approve its own result. It is not a mandatory preflight and does29not turn every task into a control diagram.3031## Principle source3233Use a host Sequence and matching interpretations when the host declares them.34Otherwise use this package's read-only fallback in `references/sequence.md`.35Read only P03, P04, P13, and P15. A live task may select a different current36lead without changing this Skill's stable lineage.3738## Start from the actual system3940Recover only what changes the engineering decision:4142```text43Desired whole behavior and acceptance owner:44Operating conditions, source state, and hard constraints:45Failure consequence and acceptable residual risk:46Actual path from input through effects to acceptance:47Observed failure, disturbance, variance, or near miss:48Available signals, controls, recovery, and resource margins:49Observation that would show the present system is already sufficient:50```5152If the task is low-consequence, reversible, locally observable, and already has53an adequate correction path, retain the present form. Do not manufacture a54System Case.5556Read [concepts](references/concepts.md) only when terms such as feedback,57stability, redundancy, independence, or residual risk affect the decision.58Read [evaluation](references/evaluation.md) only when testing a new control59structure or claiming improved system reliability.6061## Specificity discipline6263A system requirement can be correct while its implementation is still64undecided. False precision hides that boundary and can silently take another65owner's authority. Use the most concrete form supported by the supplied source:6667- if a field, threshold, duration, batch size, sample, protocol, or role is68 explicitly governed, preserve it and name its source;69- if only a property is governed, state the property and its acceptance signal;70 use `[owning domain/runtime to determine]` for the representation; and71- if even the property lacks evidence, return discovery rather than a plausible72 mechanism.7374Before submitting, audit every concrete name and number introduced by the75answer. Remove or bound any that came from the Agent rather than the case or a76governing source. Do not contradict this audit by naming a proposed field and77then claiming only to have stated a property.7879Self-audit is not verification. For a consequential System Case, treat the80first result as a candidate until a designated independent verifier compares81each introduced mechanism, threshold, duration, sample, protocol, and role with82the supplied sources and ownership boundary. Give the verifier this narrow83claim/source task and only the governing sources plus accepted candidate84payload; do not inject the full reasoning trace or ask it to redesign the85system. Revise only the rejected specifics while preserving supported control86requirements.8788## Core method89901. **Define sufficient behavior, not perfection.** State the outcome, operating91 range, acceptance condition, failure consequence, and tolerable residual92 risk. “No mistakes” is not an engineering specification. Make risk tolerance93 specific enough to distinguish harmless variation, recoverable failure, and94 unacceptable escape.952. **Model the whole path.** Follow the real movement from inputs and source96 state through judgments, effects, verification, acceptance, and later use.97 Name actors and components only where their relation changes behavior. A98 list of Agents, tools, or workflow stages is not yet a system model.993. **Identify disturbances and failure paths.** Include semantic error, stale100 context, omitted obligations, correlated model error, protocol failure,101 external change, unsafe effect, reviewer miss, false-positive burden, and102 resource exhaustion when relevant. Do not invent an exhaustive hazard103 register for a reversible low-risk task.1044. **Find the principal control gap.** Select the failure path whose containment105 or correction most changes the whole. Preserve secondary hard constraints,106 but do not answer every risk with more stages, more prompts, or more Cells.1075. **Make the material state observable.** Decide which signal can expose the108 failure in time to act: source-linked evidence, a deterministic check,109 structured status, disagreement, usage variance, production outcome, or a110 later practice result. A control that cannot observe its target is ceremony.111 Agent self-report and terminal success remain claims until their designated112 verification and settlement. State the semantic property that must be113 observed before selecting its representation. A whole-source claim must be114 resolved against whole-source evidence before acceptance, but this Skill115 does not decide which domain field or runtime protocol carries that evidence.1166. **Assign authority and response.** State who may observe, propose, verify,117 intervene, commit effects, roll back, accept residual risk, and reopen the118 case. Keep decision and execution authority distinct where consequence119 requires it. Do not make a worker or review ensemble accept its own output.1207. **Choose the minimum sufficient control structure.** Select only mechanisms121 that interrupt the named failure path:122123 - prevent by improving input, context, boundaries, or task form;124 - detect through independent evidence, tests, comparison, or observation;125 - contain with limited authority, blast radius, staging, or reversible effects;126 - recover through retry, repair, rollback, reuse, or escalation; and127 - tolerate through justified redundancy or resource margin.128129 Match control cost to failure consequence. A reversible documentation edit130 may need one meaningful check; an irreversible production or strategic131 decision may justify diverse independent review, repeated labor,132 deterministic gates, prepared recovery, and explicit human acceptance.133 More controls are not automatically more reliable: they can add delay,134 correlated error, coordination failure, and review burden.135 At a boundary owned by another method or runtime, specify the required136 control property, observable signal, authority, and failure response, then137 hand it off. Do not invent a terminal-tool field, domain packet, deployment138 threshold, sample size, time window, or organization role without source139 evidence and the owner's authority.1408. **Engineer independence, not vote count.** Repetition helps only when the141 compared attempts can expose different failure modes. Identical context,142 model, prompt, and method may reproduce one systematic error. Vary evidence143 boundaries, method, role, model, verifier, or deterministic surface only144 where correlated failure matters. Resolve factual disagreement against145 sources and designated verification, never by majority alone.1469. **Fit Agent nodes after the system relation is known.** Use `task-shaping`147 when one required contribution must be compiled into a stable Agent unit.148 Use `model-evaluation` when its capability is unevidenced. Use a domain Skill149 to define semantic partitions and acceptance. The Work Cell or orchestration150 runtime carries prepared execution; none of them decides the whole system151 relation by default.15210. **Budget for completion and recovery.** Begin from necessary work and risk,153 then use `work-estimation` to convert it into profile-specific resources.154 Prefer a realistic estimate, explicit margin for important uncertainty, a155 high emergency ceiling only when required by the carrier, and post-run156 audit of expected versus actual use. Do not treat low token use as quality157 or a routine context ceiling as an engineering budget. Estimate variance158 is normally a post-run correction signal; it does not itself authorize159 interruption, replanning, or escalation during useful work. Add an in-run160 resource control only when exhaustion is a named safety or availability161 failure and its owner authorizes that response.16211. **Close the loop in operation.** Observe end-to-end outcomes, material163 defects caught and escaped, false-positive repair burden, recovery, latency,164 cost per useful result, and estimate variance. Compare with the unchanged165 system when claiming improvement. Revise the system model, control, or risk166 acceptance from what actually happened; do not merely append another rule.167168Before making any control concrete, apply this authority-and-evidence gate:169170```text171Is the mechanism, threshold, duration, schema, or role fixed by source evidence172and owned here?173 yes -> specify it and cite or name the governing source174 no -> state only the required property, mark the representation175 `[owning domain/runtime to determine]`, and route it176```177178“Independent verification,” “staged effects,” and “recoverable rollback” may be179valid system requirements while their exact implementation remains a domain or180runtime decision. Unknown specificity is a discovery item, not an invitation to181invent plausible numbers.182183## System Case184185Return the smallest handoff the consequence requires:186187```text188Desired behavior, operating range, and acceptance owner:189System boundary and effect path:190Principal disturbance or failure path:191Observable signal and evidence source:192Control action, authority, and recovery:193Required component contributions and local owners:194Expected work, margin, and audit signal:195Residual risk and who accepts it:196Operational measure and reopening condition:197```198199For several interacting failure paths, a compact control map may help:200201| Failure or disturbance | Observable signal | Response | Authority | Recovery | Residual risk |202|---|---|---|---|---|---|203204Use neither artifact when a direct explanation or one boundary change is205sufficient. The System Case is a decision carrier, not a universal manifest or206runtime schema.207208## Boundaries and routing209210| Need | Owner |211|---|---|212| Decide whether the whole is sufficiently reliable under disturbance | `systems-engineering` |213| Determine whether one Agent operation fits an evidenced execution envelope | `task-shaping` |214| Establish reusable model/provider/harness capability evidence | `model-evaluation` |215| Form domain-specific review, refactoring, cognition, or other semantic packets | owning domain Skill |216| Deliver authoritative context to one prepared unit | `context-engineering` |217| Estimate necessary work and convert it into time, tokens, or money | `work-estimation` |218| Run Cells, queues, retries, concurrency, and providers | Work Cell or orchestration runtime |219| Verify domain facts and accept residual risk or effects | designated verifier and human or host authority |220| Select the next bounded practice after observing operation | `practice-cycle` |221222Do not use `systems-engineering` as a synonym for architecture, project223planning, generic rigor, or “use more reviewers.” Do not hide a domain judgment224inside a control label. Do not claim end-to-end reliability from component225success, schema validity, test count, reviewer count, or one clean run.226Do not repair a semantic control gap by changing a generic runtime contract227unless runtime evidence identifies that contract as the owner and the change is228separately authorized.229230## Completion standard231232The engineering judgment is ready when it defines sufficient whole behavior,233models the real effect and acceptance path, identifies the principal material234failure/control gap, connects an observable signal to an authorized response235and recovery, selects controls proportional to consequence, names residual risk236and its acceptance owner, and states an operational observation capable of237showing the design is wrong or incomplete.238239For a consequential case, readiness also requires independent source-and-owner240verification of newly introduced specifics. Without that verification, report241a candidate. Without representative operation evidence, report a proposed242system design or trial—not a reliable system.