Procházet zdrojové kódy

用户令: ① 修 /detail/v2#tab=vibration 无内容(根因=清产物后缺无生成端产物; 并从旧含产物包补齐 567 件+台账重算); ② 门户逐页识别: 补页面侧产物登记(CMS 页面件/振动清单)+新增台账覆盖规则与 mtime 陈旧判据; ③ 文档新增 §12 门户逐页识别、§13 输入→功能→算法→产物 关系与实现逻辑

zhouyang.xie před 3 týdny
rodič
revize
c5202054e6
3 změnil soubory, kde provedl 826 přidání a 623 odebrání
  1. 271 245
      configs/portal_pages.yaml
  2. 110 0
      docs/系统设计说明.md
  3. 445 378
      scripts/pages_audit.py

+ 271 - 245
configs/portal_pages.yaml

@@ -1,245 +1,271 @@
-# 门户页面与子页的"内容归口"登记表  (2026-09-17, 用户令 1)
-#
-# 回答两个问题:
-#   ① 这一页/这一件 **该不该随输入数据(data/raw)变化**?
-#   ② 它 **算不算产物**? 算的话归到哪、由谁生成、怎么查它陈旧没陈旧?
-#
-# 五个 kind (每个都有机器可查的规则, 见 scripts/pages_audit.py):
-#   static          受管静态: 外壳正文(方法论/架构/案例叙述)。不随输入数据变, 进 git, 不引用产物。
-#   live            实时取数: iframe/链接到组件服务或 API。永远与产物一致, 不需要快照, 也不进产物管理。
-#   data-citing     静态正文 + **产物引用脚注**(如"判据全量见 outputs/<场>/…/findings.json")。
-#                   正文不随数据变; 但被引用的产物路径必须**存在**, 否则引用悬空(挂羊头)。
-#   data-derived    数据派生快照: 把产物/输入烘进页面(内嵌 JSON / 指纹 / 生成时间)。
-#                   ★**属于产物**: 必须登记 source + (source_sha256 或 generated_at), 并做陈旧检测。
-#   frozen-delivery 冻结交付件: 带版本号与日期的客户交付件(治理清单/报告/资料包)。
-#                   按**交付版本**变, 不随数据自动变; 必须能从文件名或正文读到版本号与日期。
-#
-# 判定依据(evidence)一律写清"凭什么这么判", 可复核: 看的是页面里有没有 iframe/fetch/内嵌数据/产物引用,
-# 以及内容是不是由 data/raw 算出来的。要改分类, 先改证据。
-version: 1
-farm: rudong
-
-# 产物路径写法统一: 一律以 outputs/<场>/… 记录(相对安装根), 便于陈旧检测直接定位文件
-pages:
-
-  # ── 门户外壳里的静态叙述页 (受管, 不随数据变) ───────────────────────────────
-  - id: index
-    title: 总览
-    kind: static
-    changes_with_data: false
-    file: release/portal_src/shell.html
-    evidence: "外壳 #index 段 2,519 B: 0 iframe / 0 fetch / 0 产物引用; 正文是产品承诺与入口导航"
-  - id: architecture
-    title: 系统架构
-    kind: static
-    changes_with_data: false
-    file: release/portal_src/shell.html
-    evidence: "外壳 #architecture 段 3,037 B: 无数据引用; 描述分层架构"
-  - id: method
-    title: 方法
-    kind: static
-    changes_with_data: false
-    file: release/portal_src/shell.html
-    evidence: "外壳 #method 段 2,959 B: 无数据引用; 描述方法论"
-  - id: findings
-    title: 经验发现
-    kind: static
-    changes_with_data: false
-    file: release/portal_src/shell.html
-    evidence: "★用户问的三页之一。外壳 #findings 段 2,322 B: 0 iframe / 0 fetch / 0 产物引用; 正文是公司级经验叙述(59 项内部检查 → 8 条公开教训, 60+ 个风场), 不是本场站数据的函数 ⇒ 不该随输入数据变, 也不是产物。若将来要放本场站计数, 必须改成 data-derived 并登记 source"
-  - id: case_hydraulic
-    title: 案例·液压
-    kind: static
-    changes_with_data: false
-    file: release/portal_src/shell.html
-    evidence: "外壳 #case_hydraulic 段 6,187 B: 无数据引用; 案例叙述"
-
-  # ── 实时取数页 (永远与产物一致) ──────────────────────────────────────────
-  - id: cms
-    title: 振动·CMS
-    kind: live
-    changes_with_data: true
-    file: release/portal_src/shell.html
-    live_target: http://127.0.0.1:18020/
-    evidence: "外壳 #cms 段里 1 个 iframe + 2 个链接指向组件 :18020 (src/windcms/serve.py), 组件按需读产物 ⇒ 页面本身不存快照, 数据变了刷新即变"
-  - id: recalc
-    title: 数据重算
-    kind: live
-    changes_with_data: true
-    file: release/portal_src/shell.html
-    live_target: http://127.0.0.1:28084/ops
-    evidence: "外壳 #recalc 段只有 1 个到 /ops 的链接(网关提供, 后端即真实状态)"
-  - id: login
-    title: 登录
-    kind: live
-    changes_with_data: true
-    file: release/portal_src/shell.html
-    live_target: http://127.0.0.1:18033/v2
-    evidence: "外壳 #login 段链接到工作台 :18033/v2"
-  - id: admin
-    title: 系统状态
-    kind: live
-    changes_with_data: true
-    file: release/portal_src/shell.html
-    live_target: http://127.0.0.1:28084/healthz
-    evidence: "外壳 #admin 段由外壳 JS 拉网关 /healthz 渲染(段内无静态数字)"
-
-  # ── #sim 仿真与回放 ────────────────────────────────────────────────────
-  - id: sim
-    title: 仿真与回放
-    kind: static
-    changes_with_data: false
-    file: release/portal_src/shell.html
-    evidence: "★用户问的三页之一。外壳 #sim 段 50,767 B 是**静态叙述**(机理/图纸说明), 段内 0 iframe / 0 fetch; 6 个链接指向 :18792 仿真台页与 :64292 三维工作台。仿真台的输入是**图纸/机型参数**(configs/machine_packs), 不是 data/raw ⇒ 页面本体不该随输入数据变"
-    children:
-      - name: 控制律仿真台资料包 (5 页)
-        file: release/如东SWT40_控制律仿真台_20260906.zip
-        members_in_zip: ["0_四系统合页.html", "1_偏航系统.html", "2_传动链与功率.html", "3_发电机与热管理.html", "4_变桨轮毂液压.html"]
-        served_by: release/sim_sys_server.py
-        kind: frozen-delivery
-        changes_with_data: false
-        version: "20260906"
-        evidence: "release/sim_sys_server.py 直接从 zip 里读这 5 页 ⇒ 冻结资料包(日期在包名里); 内容为图纸/机理仿真, 不随输入数据变"
-      - name: 三维拆装工作台
-        file: release/viewer/unit-workbench.html
-        kind: static
-        changes_with_data: false
-        evidence: "release/viewer/** 由图纸/三维模型构建(build-*.py + esbuild), 与 data/raw 无关; rev=hub-review-v1 是评审版本号"
-      - name: 控制律仪表台面板 (门户内嵌)
-        file: release/portal_src/templates/tpl-swt40_控制律仪表台_外发版.html.html
-        kind: data-citing
-        changes_with_data: false
-        cites: [outputs/rudong/yaw_verify/findings.json]
-        known_gap:
-          missing: [outputs/rudong/yaw_verify/findings.json]
-          why: 引用的产物目录 yaw_verify/ 不在本包(振动线/偏航专项产物未随包); 本包同类台账在 outputs/<场>/sop/findings.json。不改交付件正文(那是客户手里的冻结版本), 但按「引用悬空」如实记账, 见 docs §7
-        evidence: "面板正文是固定的判据/仪表说明, 只把产物当**脚注引用**('判据与证伪条件全量见 outputs/rudong/yaw_verify/findings.json') ⇒ 正文不随数据变, 但被引用的产物必须在位"
-      - name: 四系统判据面板 (门户内嵌, 4 份)
-        file: release/portal_src/templates/tpl-swt-subsystem-*.html.html
-        kind: data-citing
-        changes_with_data: false
-        cites: [outputs/rudong/yaw_verify/findings.json]
-        known_gap:
-          missing: [outputs/rudong/yaw_verify/findings.json]
-          why: "同仪表台: 引用 yaw_verify/ 专项产物, 该目录不在本包"
-        evidence: "同仪表台: 静态正文 + findings 引用脚注"
-
-  # ── #documents 交付文档 ───────────────────────────────────────────────
-  - id: documents
-    title: 交付文档
-    kind: static
-    changes_with_data: false
-    file: release/portal_src/shell.html
-    evidence: "★用户问的三页之一。外壳 #documents 段 5,690 B 是导航页(0 fetch), 1 个 iframe 指向网关下发的治理清单 HTML(release/如东/**)。段内不含数据快照"
-    children:
-      - name: 治理清单交付包 (分册 + 全册 + 正式报告)
-        file: "release/如东/如东治理清单_交付_20260901/**/*.html"
-        kind: frozen-delivery
-        changes_with_data: false
-        evidence: "文件名自带版本号与日期(如 如东_液压系统治理清单_v1.2_2026-09-01.html / 全册_v4.5_2026-09-06.html) ⇒ 客户交付件, 随交付版本变; 内容虽由分析产出, 但**冻结发布**, 不该随输入数据自动改"
-      - name: 全场状态一览 (脱敏)
-        file: release/如东/全场状态一览_脱敏_20260825.html
-        kind: data-derived
-        changes_with_data: true
-        source: null
-        known_gap:
-          why: "内容是实际运行状态的函数(脱敏快照), 属产物性质; 但页面里没有任何指纹/生成时间/来源登记, 全仓也搜不到生成端 ⇒ 现在**无法做陈旧检测**。处置: 要么补生成端并改成带 source_sha256 的产物(推荐), 要么在登记表里把它降级为 frozen-delivery 并接受它会过期(需用户裁定)"
-        evidence: "★这是问题所在: 1.5 MB 的场站状态快照(脱敏版), 内容是实际运行状态的函数 ⇒ **属于产物**, 但页面里**没有任何指纹/生成时间/来源登记**, 全仓也搜不到生成端 ⇒ 目前游离在产物管理之外(无溯源、无陈旧检测、无重算入口)"
-      - name: 如东取数单
-        file: release/如东/如东取数单_2026-08-21.html
-        kind: data-derived
-        changes_with_data: true
-        cites:
-          - outputs/rudong/cleaned/native_scturbine.parquet
-          - outputs/rudong/m5_cms_tcm/tcm_index.parquet
-          - outputs/rudong/structured/events/fastlog.parquet
-        known_gap:
-          missing:
-            - outputs/rudong/cleaned/native_scturbine.parquet
-            - outputs/rudong/structured/events/fastlog.parquet
-            - outputs/rudong/structured/events/fault_scada_xml.parquet
-          why: "取数单里列的 cleaned/、structured/ 两棵产物树不在本包(属缺失的六层链/结构化产物线, 见 docs §7); 本包在位的同类是 outputs/<场>/windscada/ 与 m5_cms_tcm/"
-        evidence: 取数单是「要哪些数据」的清单页, 明列产物路径 ⇒ 随产物结构变; 现在只有日期没有指纹, 也没有生成端随包
-
-  # ── 门户内嵌的数据派生页 (28 个模板里的) ──────────────────────────────
-  - id: embed_replay_chain
-    title: 传动链实际运行回放 (门户内嵌单文件版)
-    kind: data-derived
-    changes_with_data: true
-    file: release/portal_src/templates/tpl-如东传动链实际运行诊断_单文件版.html.html
-    embedded_json_var: RUDONG_ACTUAL_REPLAY
-    source_key: source
-    source_sha_key: source_sha256
-    json_escaped: true
-    source: outputs/rudong/cleaned/turbine_1min.parquet
-    known_gap:
-      missing:
-        - outputs/rudong/cleaned/turbine_1min.parquet
-        - outputs/rudong/cleaned/turbine_1min_fingerprint.json
-        - outputs/rudong/gearbox_life/load_spectrum_summary.csv
-        - outputs/rudong/gearbox_life/tcm_dynamic_severity.csv
-      why: "页面烘入的 cleaned/ 与 gearbox_life/ 两棵产物树不在本包(六层链未随包, docs §7) ⇒ 内嵌快照**无法与当前产物比对**(拿不到当前值)。这恰恰说明它属于产物: 有 source+sha256 的机制已经齐了, 缺的是那份产物本身"
-    evidence: "内嵌 window.RUDONG_ACTUAL_REPLAY 带 source=outputs/rudong/cleaned/turbine_1min.parquet 与 source_sha256 ⇒ 明明白白的产物烘入快照; **已可做陈旧检测**(产物在位时)"
-  - id: embed_standard_panel
-    title: 标准面板 (门户内嵌)
-    kind: data-derived
-    changes_with_data: true
-    file: release/portal_src/templates/tpl-standard_panel_zh.html.html
-    generated_at_key: 生成
-    source: outputs/rudong/paradigm_r1/releases/E1_measure_replay_20260905/release.yaml
-    cites:
-      - outputs/rudong/structured/reference/turbine_master.parquet
-    known_gap:
-      missing:
-        - outputs/rudong/paradigm_r1/releases/E1_measure_replay_20260905/release.yaml
-        - outputs/rudong/structured/reference/turbine_master.parquet
-      why: "引用的 paradigm 发布件与 structured/ 参考表不在本包(paradigm_r1 目录在, 但没有该 release 子目录) ⇒ 只能按正文里的生成时间判断新旧"
-    evidence: "正文带 '生成 2026-09-05 16:28' 与产物引用 ⇒ 某次生成的数据快照; 缺 source_sha ⇒ 只能按生成时间+引用在位判断"
-  - id: embed_coverage_ch0
-    title: 覆盖度报告 ch0 (门户内嵌)
-    kind: data-derived
-    changes_with_data: true
-    file: release/portal_src/templates/tpl-coverage_ch0_zh.html.html
-    generated_at_key: 生成
-    source: null
-    known_gap:
-      why: "正文写 '源: e1_base.parquet @e3b531b8eddf8197' —— 那份 e1_base.parquet 不在本包 ⇒ 指纹虽有, 但对不上任何在位文件, 溯源链断在包外; 只能按生成时间 2026-09-05 16:28 判新旧"
-    evidence: "正文写 '生成 2026-09-05 16:28 · 源: e1_base.parquet @e3b531b8eddf8197 (10-min, 2,973,638 格, 38 台)' ⇒ 数据派生; 但 e1_base.parquet **不在本包** ⇒ 溯源链断在包外"
-  - id: embed_u6_sim
-    title: U6 变桨液压仿真台 (门户内嵌)
-    kind: static
-    changes_with_data: false
-    file: release/portal_src/templates/tpl-U6_变桨液压仿真台.html.html
-    evidence: "图纸派生的物理回路仿真台面板; 无内嵌数据/指纹/产物引用"
-  - id: embed_u6_health
-    title: U6 液压公共站健康报告 (客户版, 门户内嵌)
-    kind: frozen-delivery
-    changes_with_data: false
-    file: release/portal_src/templates/tpl-U6_液压公共站健康报告_客户版.html.html
-    version: "客户版"
-    evidence: "984 KB 客户交付报告; 文件名标 '客户版'; 无数据指纹 ⇒ 按交付件冻结(缺版本号/日期, 见 audit 提示)"
-  - id: embed_gearbox_report
-    title: 整机综合诊断与风险评估 (主轴冲击深挖 V2.1, 门户内嵌)
-    kind: frozen-delivery
-    changes_with_data: false
-    file: release/portal_src/templates/tpl-如东海上风电场_整机综合诊断与风险评估_主轴冲击深挖修订版V2.1.html.html
-    version: "V2.1"
-    evidence: "3.7 MB 综合诊断报告, 文件名带修订版号 V2.1; 报告类交付件按版本发布"
-  - id: embed_reports_4
-    title: 分系统评估报告 (主轴承/齿轮箱/发电机/变桨, 门户内嵌 4 份)
-    kind: frozen-delivery
-    changes_with_data: false
-    file: release/portal_src/templates/tpl-report-*.html.html
-    evidence: "4 份系统评估报告(0.7–3.4 MB), 无内嵌数据/指纹; 属交付报告(缺版本号/日期登记)"
-  - id: embed_governance_12
-    title: 治理清单分册 (门户内嵌 12 份)
-    kind: frozen-delivery
-    changes_with_data: false
-    file: release/portal_src/templates/tpl-governance-*.html.html
-    evidence: "12 份治理分册(0–11), 与 release/如东 下的分册同源; 冻结交付件"
-  - id: embed_sc1_local
-    title: SC1 本地占位 (门户内嵌)
-    kind: static
-    changes_with_data: false
-    file: release/portal_src/templates/tpl-sc1-local.html.html
-    evidence: "0.4 KB 占位面板"
+# 门户页面与子页的"内容归口"登记表  (2026-09-17, 用户令 1)
+#
+# 回答两个问题:
+#   ① 这一页/这一件 **该不该随输入数据(data/raw)变化**?
+#   ② 它 **算不算产物**? 算的话归到哪、由谁生成、怎么查它陈旧没陈旧?
+#
+# 五个 kind (每个都有机器可查的规则, 见 scripts/pages_audit.py):
+#   static          受管静态: 外壳正文(方法论/架构/案例叙述)。不随输入数据变, 进 git, 不引用产物。
+#   live            实时取数: iframe/链接到组件服务或 API。永远与产物一致, 不需要快照, 也不进产物管理。
+#   data-citing     静态正文 + **产物引用脚注**(如"判据全量见 outputs/<场>/…/findings.json")。
+#                   正文不随数据变; 但被引用的产物路径必须**存在**, 否则引用悬空(挂羊头)。
+#   data-derived    数据派生快照: 把产物/输入烘进页面(内嵌 JSON / 指纹 / 生成时间)。
+#                   ★**属于产物**: 必须登记 source + (source_sha256 或 generated_at), 并做陈旧检测。
+#   frozen-delivery 冻结交付件: 带版本号与日期的客户交付件(治理清单/报告/资料包)。
+#                   按**交付版本**变, 不随数据自动变; 必须能从文件名或正文读到版本号与日期。
+#
+# 判定依据(evidence)一律写清"凭什么这么判", 可复核: 看的是页面里有没有 iframe/fetch/内嵌数据/产物引用,
+# 以及内容是不是由 data/raw 算出来的。要改分类, 先改证据。
+version: 1
+farm: rudong
+
+# 产物路径写法统一: 一律以 outputs/<场>/… 记录(相对安装根), 便于陈旧检测直接定位文件
+pages:
+
+  # ── 门户外壳里的静态叙述页 (受管, 不随数据变) ───────────────────────────────
+  - id: index
+    title: 总览
+    kind: static
+    changes_with_data: false
+    file: release/portal_src/shell.html
+    evidence: "外壳 #index 段 2,519 B: 0 iframe / 0 fetch / 0 产物引用; 正文是产品承诺与入口导航"
+  - id: architecture
+    title: 系统架构
+    kind: static
+    changes_with_data: false
+    file: release/portal_src/shell.html
+    evidence: "外壳 #architecture 段 3,037 B: 无数据引用; 描述分层架构"
+  - id: method
+    title: 方法
+    kind: static
+    changes_with_data: false
+    file: release/portal_src/shell.html
+    evidence: "外壳 #method 段 2,959 B: 无数据引用; 描述方法论"
+  - id: findings
+    title: 经验发现
+    kind: static
+    changes_with_data: false
+    file: release/portal_src/shell.html
+    evidence: "★用户问的三页之一。外壳 #findings 段 2,322 B: 0 iframe / 0 fetch / 0 产物引用; 正文是公司级经验叙述(59 项内部检查 → 8 条公开教训, 60+ 个风场), 不是本场站数据的函数 ⇒ 不该随输入数据变, 也不是产物。若将来要放本场站计数, 必须改成 data-derived 并登记 source"
+  - id: case_hydraulic
+    title: 案例·液压
+    kind: static
+    changes_with_data: false
+    file: release/portal_src/shell.html
+    evidence: "外壳 #case_hydraulic 段 6,187 B: 无数据引用; 案例叙述"
+
+  # ── 实时取数页 (永远与产物一致) ──────────────────────────────────────────
+  - id: cms
+    title: 振动·CMS
+    kind: live
+    changes_with_data: true
+    file: release/portal_src/shell.html
+    live_target: http://127.0.0.1:18020/
+    evidence: "外壳 #cms 段里 1 个 iframe + 2 个链接指向组件 :18020 (src/windcms/serve.py), 组件按需读产物 ⇒ 页面本身不存快照, 数据变了刷新即变"
+  - id: recalc
+    title: 数据重算
+    kind: live
+    changes_with_data: true
+    file: release/portal_src/shell.html
+    live_target: http://127.0.0.1:28084/ops
+    evidence: "外壳 #recalc 段只有 1 个到 /ops 的链接(网关提供, 后端即真实状态)"
+  - id: login
+    title: 登录
+    kind: live
+    changes_with_data: true
+    file: release/portal_src/shell.html
+    live_target: http://127.0.0.1:18033/v2
+    evidence: "外壳 #login 段链接到工作台 :18033/v2"
+  - id: admin
+    title: 系统状态
+    kind: live
+    changes_with_data: true
+    file: release/portal_src/shell.html
+    live_target: http://127.0.0.1:28084/healthz
+    evidence: "外壳 #admin 段由外壳 JS 拉网关 /healthz 渲染(段内无静态数字)"
+
+  # ── #sim 仿真与回放 ────────────────────────────────────────────────────
+  - id: sim
+    title: 仿真与回放
+    kind: static
+    changes_with_data: false
+    file: release/portal_src/shell.html
+    evidence: "★用户问的三页之一。外壳 #sim 段 50,767 B 是**静态叙述**(机理/图纸说明), 段内 0 iframe / 0 fetch; 6 个链接指向 :18792 仿真台页与 :64292 三维工作台。仿真台的输入是**图纸/机型参数**(configs/machine_packs), 不是 data/raw ⇒ 页面本体不该随输入数据变"
+    children:
+      - name: 控制律仿真台资料包 (5 页)
+        file: release/如东SWT40_控制律仿真台_20260906.zip
+        members_in_zip: ["0_四系统合页.html", "1_偏航系统.html", "2_传动链与功率.html", "3_发电机与热管理.html", "4_变桨轮毂液压.html"]
+        served_by: release/sim_sys_server.py
+        kind: frozen-delivery
+        changes_with_data: false
+        version: "20260906"
+        evidence: "release/sim_sys_server.py 直接从 zip 里读这 5 页 ⇒ 冻结资料包(日期在包名里); 内容为图纸/机理仿真, 不随输入数据变"
+      - name: 三维拆装工作台
+        file: release/viewer/unit-workbench.html
+        kind: static
+        changes_with_data: false
+        evidence: "release/viewer/** 由图纸/三维模型构建(build-*.py + esbuild), 与 data/raw 无关; rev=hub-review-v1 是评审版本号"
+      - name: 控制律仪表台面板 (门户内嵌)
+        file: release/portal_src/templates/tpl-swt40_控制律仪表台_外发版.html.html
+        kind: data-citing
+        changes_with_data: false
+        cites: [outputs/rudong/yaw_verify/findings.json]
+        known_gap:
+          missing: [outputs/rudong/yaw_verify/findings.json]
+          why: 引用的产物目录 yaw_verify/ 不在本包(振动线/偏航专项产物未随包); 本包同类台账在 outputs/<场>/sop/findings.json。不改交付件正文(那是客户手里的冻结版本), 但按「引用悬空」如实记账, 见 docs §7
+        evidence: "面板正文是固定的判据/仪表说明, 只把产物当**脚注引用**('判据与证伪条件全量见 outputs/rudong/yaw_verify/findings.json') ⇒ 正文不随数据变, 但被引用的产物必须在位"
+      - name: 四系统判据面板 (门户内嵌, 4 份)
+        file: release/portal_src/templates/tpl-swt-subsystem-*.html.html
+        kind: data-citing
+        changes_with_data: false
+        cites: [outputs/rudong/yaw_verify/findings.json]
+        known_gap:
+          missing: [outputs/rudong/yaw_verify/findings.json]
+          why: "同仪表台: 引用 yaw_verify/ 专项产物, 该目录不在本包"
+        evidence: "同仪表台: 静态正文 + findings 引用脚注"
+
+  # ── #documents 交付文档 ───────────────────────────────────────────────
+  - id: documents
+    title: 交付文档
+    kind: static
+    changes_with_data: false
+    file: release/portal_src/shell.html
+    evidence: "★用户问的三页之一。外壳 #documents 段 5,690 B 是导航页(0 fetch), 1 个 iframe 指向网关下发的治理清单 HTML(release/如东/**)。段内不含数据快照"
+    children:
+      - name: 治理清单交付包 (分册 + 全册 + 正式报告)
+        file: "release/如东/如东治理清单_交付_20260901/**/*.html"
+        kind: frozen-delivery
+        changes_with_data: false
+        evidence: "文件名自带版本号与日期(如 如东_液压系统治理清单_v1.2_2026-09-01.html / 全册_v4.5_2026-09-06.html) ⇒ 客户交付件, 随交付版本变; 内容虽由分析产出, 但**冻结发布**, 不该随输入数据自动改"
+      - name: 全场状态一览 (脱敏)
+        file: release/如东/全场状态一览_脱敏_20260825.html
+        kind: data-derived
+        changes_with_data: true
+        source: null
+        known_gap:
+          why: "内容是实际运行状态的函数(脱敏快照), 属产物性质; 但页面里没有任何指纹/生成时间/来源登记, 全仓也搜不到生成端 ⇒ 现在**无法做陈旧检测**。处置: 要么补生成端并改成带 source_sha256 的产物(推荐), 要么在登记表里把它降级为 frozen-delivery 并接受它会过期(需用户裁定)"
+        evidence: "★这是问题所在: 1.5 MB 的场站状态快照(脱敏版), 内容是实际运行状态的函数 ⇒ **属于产物**, 但页面里**没有任何指纹/生成时间/来源登记**, 全仓也搜不到生成端 ⇒ 目前游离在产物管理之外(无溯源、无陈旧检测、无重算入口)"
+      - name: 如东取数单
+        file: release/如东/如东取数单_2026-08-21.html
+        kind: data-derived
+        changes_with_data: true
+        cites:
+          - outputs/rudong/cleaned/native_scturbine.parquet
+          - outputs/rudong/m5_cms_tcm/tcm_index.parquet
+          - outputs/rudong/structured/events/fastlog.parquet
+        known_gap:
+          missing:
+            - outputs/rudong/cleaned/native_scturbine.parquet
+            - outputs/rudong/structured/events/fastlog.parquet
+            - outputs/rudong/structured/events/fault_scada_xml.parquet
+          why: "取数单里列的 cleaned/、structured/ 两棵产物树不在本包(属缺失的六层链/结构化产物线, 见 docs §7); 本包在位的同类是 outputs/<场>/windscada/ 与 m5_cms_tcm/"
+        evidence: 取数单是「要哪些数据」的清单页, 明列产物路径 ⇒ 随产物结构变; 现在只有日期没有指纹, 也没有生成端随包
+
+  # ── 页面侧产物 (页面读的就是它们; 用户令 2026-09-17: 是产物就要纳入产物管理) ──────────
+  - id: page_product_cms
+    title: CMS 评估报告页/件 (门户「振动·CMS」页读的就是它们)
+    kind: data-derived
+    changes_with_data: true
+    file: outputs/rudong/windcms/index.html
+    source: outputs/rudong/m5_cms_tcm/windows/w0316/index.parquet
+    generator: scripts/windcms.py report
+    generated_at_source: mtime
+    evidence: >-
+      门户 #cms 的 iframe 指向组件 :18020, 组件页读的就是 outputs/<场>/windcms 下的这些页面件
+      (index.html / 逐台页 / workbench.html / report_*.md) —— 它们是**产物**(由 windcms 报告生成器读
+      TCM 窗索引写成), 不是程序。本包缺六层链的 model_run/fusion 产物, 重生成会掉内容 (report.md −97%,
+      docs §7), 故**不在重算链里重建**, 交付时按随包件补齐并登记为 shipped。
+      管理动作: 登记(本条) + 溯源(_provenance.json 的 shipped) + 陈旧检测(源=TCM 窗索引的 mtime/sha)。
+  - id: page_product_vib_raw_manifest
+    title: 振动摄入清单 (数据层「CMS 振动评估报告」页引用)
+    kind: data-derived
+    changes_with_data: true
+    file: outputs/rudong/m5_cms_tcm/vib_raw_manifest.json
+    source: outputs/rudong/m5_cms_tcm/windows/w0316/index.parquet
+    generator: scripts/vib_raw_build.py
+    generated_at_source: mtime
+    evidence: >-
+      每次振动摄入重写: 记本窗索引/谱库的件数、时间跨度与**未随包的六层链缺口**(missing_chain)。
+      数据层页面与 docs §7 都引用它 ⇒ 属产物, 且已在台账中自登记。
+
+  - id: embed_replay_chain
+    title: 传动链实际运行回放 (门户内嵌单文件版)
+    kind: data-derived
+    changes_with_data: true
+    file: release/portal_src/templates/tpl-如东传动链实际运行诊断_单文件版.html.html
+    embedded_json_var: RUDONG_ACTUAL_REPLAY
+    source_key: source
+    source_sha_key: source_sha256
+    json_escaped: true
+    source: outputs/rudong/cleaned/turbine_1min.parquet
+    known_gap:
+      missing:
+        - outputs/rudong/cleaned/turbine_1min.parquet
+        - outputs/rudong/cleaned/turbine_1min_fingerprint.json
+        - outputs/rudong/gearbox_life/load_spectrum_summary.csv
+        - outputs/rudong/gearbox_life/tcm_dynamic_severity.csv
+      why: "页面烘入的 cleaned/ 与 gearbox_life/ 两棵产物树不在本包(六层链未随包, docs §7) ⇒ 内嵌快照**无法与当前产物比对**(拿不到当前值)。这恰恰说明它属于产物: 有 source+sha256 的机制已经齐了, 缺的是那份产物本身"
+    evidence: "内嵌 window.RUDONG_ACTUAL_REPLAY 带 source=outputs/rudong/cleaned/turbine_1min.parquet 与 source_sha256 ⇒ 明明白白的产物烘入快照; **已可做陈旧检测**(产物在位时)"
+  - id: embed_standard_panel
+    title: 标准面板 (门户内嵌)
+    kind: data-derived
+    changes_with_data: true
+    file: release/portal_src/templates/tpl-standard_panel_zh.html.html
+    generated_at_key: 生成
+    source: outputs/rudong/paradigm_r1/releases/E1_measure_replay_20260905/release.yaml
+    cites:
+      - outputs/rudong/structured/reference/turbine_master.parquet
+    known_gap:
+      missing:
+        - outputs/rudong/paradigm_r1/releases/E1_measure_replay_20260905/release.yaml
+        - outputs/rudong/structured/reference/turbine_master.parquet
+      why: "引用的 paradigm 发布件与 structured/ 参考表不在本包(paradigm_r1 目录在, 但没有该 release 子目录) ⇒ 只能按正文里的生成时间判断新旧"
+    evidence: "正文带 '生成 2026-09-05 16:28' 与产物引用 ⇒ 某次生成的数据快照; 缺 source_sha ⇒ 只能按生成时间+引用在位判断"
+  - id: embed_coverage_ch0
+    title: 覆盖度报告 ch0 (门户内嵌)
+    kind: data-derived
+    changes_with_data: true
+    file: release/portal_src/templates/tpl-coverage_ch0_zh.html.html
+    generated_at_key: 生成
+    source: null
+    known_gap:
+      why: "正文写 '源: e1_base.parquet @e3b531b8eddf8197' —— 那份 e1_base.parquet 不在本包 ⇒ 指纹虽有, 但对不上任何在位文件, 溯源链断在包外; 只能按生成时间 2026-09-05 16:28 判新旧"
+    evidence: "正文写 '生成 2026-09-05 16:28 · 源: e1_base.parquet @e3b531b8eddf8197 (10-min, 2,973,638 格, 38 台)' ⇒ 数据派生; 但 e1_base.parquet **不在本包** ⇒ 溯源链断在包外"
+  - id: embed_u6_sim
+    title: U6 变桨液压仿真台 (门户内嵌)
+    kind: static
+    changes_with_data: false
+    file: release/portal_src/templates/tpl-U6_变桨液压仿真台.html.html
+    evidence: "图纸派生的物理回路仿真台面板; 无内嵌数据/指纹/产物引用"
+  - id: embed_u6_health
+    title: U6 液压公共站健康报告 (客户版, 门户内嵌)
+    kind: frozen-delivery
+    changes_with_data: false
+    file: release/portal_src/templates/tpl-U6_液压公共站健康报告_客户版.html.html
+    version: "客户版"
+    evidence: "984 KB 客户交付报告; 文件名标 '客户版'; 无数据指纹 ⇒ 按交付件冻结(缺版本号/日期, 见 audit 提示)"
+  - id: embed_gearbox_report
+    title: 整机综合诊断与风险评估 (主轴冲击深挖 V2.1, 门户内嵌)
+    kind: frozen-delivery
+    changes_with_data: false
+    file: release/portal_src/templates/tpl-如东海上风电场_整机综合诊断与风险评估_主轴冲击深挖修订版V2.1.html.html
+    version: "V2.1"
+    evidence: "3.7 MB 综合诊断报告, 文件名带修订版号 V2.1; 报告类交付件按版本发布"
+  - id: embed_reports_4
+    title: 分系统评估报告 (主轴承/齿轮箱/发电机/变桨, 门户内嵌 4 份)
+    kind: frozen-delivery
+    changes_with_data: false
+    file: release/portal_src/templates/tpl-report-*.html.html
+    evidence: "4 份系统评估报告(0.7–3.4 MB), 无内嵌数据/指纹; 属交付报告(缺版本号/日期登记)"
+  - id: embed_governance_12
+    title: 治理清单分册 (门户内嵌 12 份)
+    kind: frozen-delivery
+    changes_with_data: false
+    file: release/portal_src/templates/tpl-governance-*.html.html
+    evidence: "12 份治理分册(0–11), 与 release/如东 下的分册同源; 冻结交付件"
+  - id: embed_sc1_local
+    title: SC1 本地占位 (门户内嵌)
+    kind: static
+    changes_with_data: false
+    file: release/portal_src/templates/tpl-sc1-local.html.html
+    evidence: "0.4 KB 占位面板"

+ 110 - 0
docs/系统设计说明.md

@@ -398,6 +398,8 @@
 | &nbsp;&nbsp;└ `治理清单交付包 (分册 + 全册 + 正式报告)` | `frozen-delivery` | 否 | 文件名自带版本号与日期(如 如东_液压系统治理清单_v1.2_2026-09-01.html / 全册_v4.5_2026-09-06.html) ⇒ 客户交付件, 随交付版本变; 内容虽由分析产出, 但**冻结发布**, 不该随输入数据自动改 |
 | &nbsp;&nbsp;└ `全场状态一览 (脱敏)` | `data-derived` | **是** | ★这是问题所在: 1.5 MB 的场站状态快照(脱敏版), 内容是实际运行状态的函数 ⇒ **属于产物**, 但页面里**没有任何指纹/生成时间/来源登记**, 全仓也搜不到生成端 ⇒ 目前游离在产物管理之外(无溯源、无陈旧检测、无重算入口) |
 | &nbsp;&nbsp;└ `如东取数单` | `data-derived` | **是** | 取数单是「要哪些数据」的清单页, 明列产物路径 ⇒ 随产物结构变; 现在只有日期没有指纹, 也没有生成端随包 |
+| `page_product_cms` CMS 评估报告页/件 (门户「振动·CMS」页读的就是它们) | `data-derived` | **是** | 门户 #cms 的 iframe 指向组件 :18020, 组件页读的就是 outputs/<场>/windcms 下的这些页面件 (index.html / 逐台页 / workbench.html / report_*.md) —— 它们是**产物**(由 windcms 报告生成器读 TCM 窗索引写成), 不是程序。本包缺六层链的 model_run/fusion 产物, 重生成会掉内容 (report.md −97%, docs §7), 故**不在重算链里重建**, 交付时按随包件补齐并登记为 shipped。 管理动作: 登记(本条) + 溯源(_provenance.json 的 shipped) + 陈旧检测(源=TCM 窗索引的 mtime/sha)。 |
+| `page_product_vib_raw_manifest` 振动摄入清单 (数据层「CMS 振动评估报告」页引用) | `data-derived` | **是** | 每次振动摄入重写: 记本窗索引/谱库的件数、时间跨度与**未随包的六层链缺口**(missing_chain)。 数据层页面与 docs §7 都引用它 ⇒ 属产物, 且已在台账中自登记。 |
 | `embed_replay_chain` 传动链实际运行回放 (门户内嵌单文件版) | `data-derived` | **是** | 内嵌 window.RUDONG_ACTUAL_REPLAY 带 source=outputs/rudong/cleaned/turbine_1min.parquet 与 source_sha256 ⇒ 明明白白的产物烘入快照; **已可做陈旧检测**(产物在位时) |
 | `embed_standard_panel` 标准面板 (门户内嵌) | `data-derived` | **是** | 正文带 '生成 2026-09-05 16:28' 与产物引用 ⇒ 某次生成的数据快照; 缺 source_sha ⇒ 只能按生成时间+引用在位判断 |
 | `embed_coverage_ch0` 覆盖度报告 ch0 (门户内嵌) | `data-derived` | **是** | 正文写 '生成 2026-09-05 16:28 · 源: e1_base.parquet @e3b531b8eddf8197 (10-min, 2,973,638 格, 38 台)' ⇒ 数据派生; 但 e1_base.parquet **不在本包** ⇒ 溯源链断在包外 |
@@ -540,4 +542,112 @@ R7 每个域必须声明 consumers 或写明"无人读";R8 `configs/farms/` 
    → 页面 **5/5 可用**(`/`·`/detail/`·`/cms/`·`/ops`·`/healthz`;`/cms/` 因为有新的"无产物"页也成了 200)
    → 收尾只停副本自己的进程;主实例 `/detail/v2` 验证后仍 **200 / 229,643 B**,副本端口无一残留。
 
+---
+
+## 12. 门户逐页识别:内容是什么、是不是产物、输出在哪
+
+判据三条(与 §10 同一套,登记表在 `configs/portal_pages.yaml`,检查器 `scripts/pages_audit.py`):
+**① 内容是不是由 `data/raw` 算出来的**(是 ⇒ 产物性质);**② 页面是现场读产物还是烘成快照**(现场读 ⇒ 永远一致,不用管);
+**③ 交付件按版本冻结还是随数据滚**(冻结 ⇒ 不是运行时产物,但要版本登记)。
+
+| 页面 / 子页 | 内容是什么 | 由输入数据重算? | 是产物? | 若是产物: 输出路径 + 生成端 | 现状与管理 |
+|---|---|---|---|---|---|
+| `#index` 总览 | 产品承诺 + 入口导航(外壳静态文本) | 否 | **不是** | — | 受管外壳(`release/portal_src/shell.html`) |
+| `#architecture` 系统架构 | 分层架构叙述 | 否 | **不是** | — | 同上 |
+| `#method` 方法 | 方法论叙述 | 否 | **不是** | — | 同上 |
+| `#findings` 经验发现 | 公司级经验("59 项内部检查 / 60+ 风场")+ 八条教训 | 否 | **不是** | — | 静态;若将来放本场站计数,必须改 `data-derived` 并登记 `source` |
+| `#case_hydraulic` 案例·液压 | 案例叙述 | 否 | **不是** | — | 静态 |
+| `#cms` 振动·CMS | iframe → 组件 `:18020`(CMS 诊断页 + 逐台页 + 工作台) | **是**(组件现场读产物) | 页面本体不是;**它读的 CMS 页面件是产物** | `outputs/<场>/windcms/*.html`、`report_CMS振动状态评估报告_*.md`、`report.md`、`workbench.html` | **产物管理**: 这些件由 `windcms.py report` 生成;本包缺六层链产物故**不在重算链里重建**(`vib_raw_build.py --with-report` 会掉内容,见 §7),交付时按"随包件"补齐并登记为 `shipped` |
+| `#sim` 仿真与回放 | 机理/图纸叙述 + 6 个链接 | 否 | **不是** | — | 输入是图纸/机型参数(`configs/machine_packs`、`release/viewer`),与 `data/raw` 无关 |
+| └ 控制律仿真台(5 页) | 图纸派生的物理回路仿真台 | 否 | **不是** | — | 冻结资料包 `release/如东SWT40_控制律仿真台_20260906.zip`(`release/sim_sys_server.py` 从 zip 现读) |
+| └ 三维拆装工作台 | 三维模型/装配(three.js) | 否 | **不是** | — | 交付资产 `release/viewer/**`(由图纸 + esbuild 构建,`rev=` 是评审版本号) |
+| └ 判据/仪表台面板(内嵌 5 块) | 固定判据说明 + **产物脚注引用** | 否(正文) | 正文不是;**被引用的产物是** | 引用 `outputs/<场>/yaw_verify/findings.json`(**本包无此目录**,见 §7 缺口) | `data-citing`:正文不随数据变,但引用必须在位(审计查悬空) |
+| └ **实际运行回放面板**(内嵌) | 把实际 1 分钟记录与证据包**烘进页面** | **是** | **是(快照)** | 源 `outputs/<场>/cleaned/turbine_1min.parquet` + 内嵌 `source_sha256` | **产物管理**: 已登记 `data-derived` + **陈旧检测**(sha 与当前产物比;源不在包,故现在只能按生成时间判) |
+| `#documents` 交付文档 | 交付件导航(0 fetch) | 否 | **不是** | — | 静态导航 |
+| └ 治理清单分册(12 + 全册 + 正式报告) | 分析结论的冻结交付版 | 否(按交付版本) | **不是**(冻结交付件) | — | 文件名自带版本号+日期(`v1.2_2026-09-01` 等);`frozen-delivery` 规则校验版本/日期 |
+| └ **全场状态一览(脱敏)** | 场站实际状态快照 | **是** | **是** | 无生成端、无指纹(`release/如东/全场状态一览_脱敏_20260825.html`) | **待裁定**: 补生成端并加 `source_sha256`(推荐),或降级为冻结交付件并接受过期(现登记 `known_gap`) |
+| └ **如东取数单** | "要哪些数据"的清单,明列产物路径 | **是** | **是** | 无生成端(`release/如东/如东取数单_2026-08-21.html`) | 同上;审计另查出它引用的 `cleaned/`、`structured/` 两棵树不在本包 |
+| `#admin` 系统状态 | 外壳 JS 拉 `/healthz` 渲染 | **是**(实时) | 不是(无快照) | — | `live` |
+| `#recalc` 数据重算 | 门户内嵌 `/ops/recalc` | **是**(实时) | 不是 | — | `live` |
+| `#login` 登录 | 链接到工作台 `:18033/v2` | **是**(实时) | 不是 | — | `live` |
+| 内嵌交付件(28 个模板) | 报告/清单/面板正文 | 少数是(3 件带数据指纹/时间戳) | 1 件是(回放面板) | 见上 | 其余按 `frozen-delivery`/`static` 登记 |
+
+**"纳入产物管理"的四件事**(§10 已述,这里补页面侧产物):① 登记(`portal_pages.yaml`);② 溯源(`source` + `source_sha256`
+或生成时间);③ **陈旧检测**(`pages_audit.py` rc=7);④ 进链(`rebuild_all.py` ⑧b 重装门户 + ⑧c 页面审计)。
+**页面侧产物(`outputs/<场>/windcms/*`)另有一条硬要求**:它们必须出现在产物台账 `_provenance.json` 里(见 §13.4 的检查)。
+
+---
 
+## 13. 输入数据 → 功能 → 算法 → 产物(关系与实现逻辑)
+
+### 13.1 关系总表(每条都能指到脚本与产物)
+
+| 输入数据(`data/raw/<场>/…`) | 功能 | 实现(脚本) | 算法/实现逻辑要点 | 产物(`outputs/<场>/…`) | 验收锚点 |
+|---|---|---|---|---|---|
+| `故障报警/*.xls`(实为 **SpreadsheetML XML**) | 报警台账 | `scripts/windscada_alarms_ingest.py` | 逐行解析 XML 报警事件 → 统一列(t_on/机组/文本/类型);列映射按随包件锁定 | `windscada/alarms.parquet` | **39,211 行** |
+| `风机故障记录/**/*.xls(x)`(78 张月度/年度同模板表) | 检修工单台账 | `scripts/windscada_workorder_ingest.py` | 前两行是说明与分组表头、**第 3 行列名**、数据从第 4 行起;`str(cell)` 原样字符串化(与随包件一致) | `windscada/workorders.parquet` | **5,876 行** |
+| `油样报告/<台号>/<部件>/*.pdf` | 油液化验索引 | `scripts/windscada_watch_channels_build.py` | 从**文件名**抽全部关键字段(`DDMMYYYY_sampleid_项目_台号_部件`),不读 PDF 正文 | `windscada/oil_samples_index.parquet` | **404 行**(缺 2026-07 批 102 行,源件不在包) |
+| `scada_10min/<机组>.csv`(14 GB) | L0 标准仓 + 10 个 SCADA 构建器 | `scripts/rebuild_from_raw.py --scada` | 逐台流式读 CSV → 统一列 → 按窗聚合/派生(温度、功率曲线、损失、控制、偏航、热链、停机事件…) | `windscada/{temp_bins,temp_monthly,powercurve_bins,powercurve_dev,loss_monthly,control_*,yaw_daily,stop_events,thermal_chain,system_aux,curve_*}.parquet` | `temp_monthly` **19,494 行** |
+| 同上(月度聚合) | 月度派生件(15 件无生成端的补上) | `scripts/windscada_monthly_build.py` | 从 10min 件**重新聚合**并**逐值对齐随包件**(随包件当标准答案,`--verify` 报命中率);规则是反推+全量比对确认 | `windscada/temp_monthly`、`loss_monthly`、`ctrl_monthly` 等同族 | 逐值相等才落盘 |
+| `windcms/**/*_decode.json`(Brande TCM 导出,25,680 件) | 振动窗索引 | `scripts/rudong_tcm_index.py` | 解析 `body.body[<ISO时间戳>]=[{Record:…}]`;机组取 `Record.Location/LocationName`;一行一条记录 | `m5_cms_tcm/windows/<窗>/index.parquet` | 窗 `w0316` **2,066,686 行** |
+| 同上 | 振动谱库 | `scripts/rudong_tcm_spectra.py` | 每条谱写进 `npz` 分片 + `spectra_meta` 指针(`shard/shard_row/x_offset/x_delta`);取数 `x = x_offset + arange(n)*x_delta` | `m5_cms_tcm/windows/<窗>/spectra/*.npz` + `spectra_meta.parquet` | 窗 `w0316` **420,742 条谱 / 1,712 片** |
+| 同上 + 厂家报告 docx | 振动侧一键摄入 | `scripts/vib_raw_build.py`(可 `--with-report`) | 编排"索引→谱→清单";**默认不重生成报告**(缺六层链,重生成会掉内容 −97%,见 §7) | `m5_cms_tcm/{index,spectra_meta,vib_raw_manifest}.…` | `vib_raw_manifest.json` 自述缺口 |
+| 厂家报告 docx(PDF 是扫描件,无文本层) | 报告转录 | `scripts/vib_reports_build.py` | docx 有文本层 ⇒ **逐字转录**判级为结构化件;扫描件 PDF 只登记归档、不进判级 | `windcms/报告_CMS*.md`、`windcms/*.json` | 转录件与报告逐条对齐 |
+| `西门子4.0技术资料/**`(含 4 个台账 xlsx) | 本体知识层 | `python -m src.ontology.kb_ingest` | 四条链(预防/抢修/故障树/计划)图谱化:手册→AlarmCode+作业指导;对译表→中文名;维护指导书→预防性任务;每条声明出处 | `ontology/objects.json` 等 | 码覆盖 263/322 |
+| 上述全部产物 + 判级矩阵 | 本体铺开 | `python -m src.ontology.populate` | **零判级权**:只转录 L1 正本(`taxonomy.system_matrix` / `fusion_table`);每条 Verdict 带审级与 `claim_window`,Evidence 指回 L1 产物;同部件多源判级**并存不合并** | `ontology/objects.json`(9,613 对象) | `ontology.audit` **0 问题** |
+| 链盘(运行时 `/api/fleet`) | 决策链进度 | `python -m src.ontology.chain_ingest` | 每台一个 `Decision` 对象记六步状态/卡点/下一步(不是 Evidence——它是系统推出来的处置状态) | `ontology/objects.json` | 需服务在跑 |
+| `component_history.json` + 链盘 | 趋势/闭环证据 | `python -m src.ontology.trend_ingest` | 在升=末三窗/早三窗 ≥1.3(未闭环);闭环=检出→换件→复测回落;两类都建 Evidence 并挂 `about→turbine`、`affects→component` | `ontology/objects.json` | 幂等覆盖 |
+| `sop/findings.json`(119 条) | 事实契约 + 派生件 | `scripts/guanlan_facts_contract.py` | 每条结论 = claim_id + 脱敏来源引用(`文件@sha16` + 章节)+ 本条 sha256 + 时间窗 + 聚合层级 + 裁决(六枚举) | `guanlan/facts_contract_v0.json` → 派生 `guanlan/derived/{portal_claims,detail_cards,qa_refs,report_summary}` | 探针:改一条 claim → 四个消费者全变 |
+| 门户外壳 + 28 个内嵌件 + 契约派生件 | 门户装配 | `scripts/portal_build.py`(+ `guanlan_portal_inject_claims.py`) | 外壳受管、内嵌件与合成结果不入库;装配**先写临时文件再原子替换**,失败不破坏在服务的门户;行尾强制 LF | `release/portal.html`(20.2 MB) | `--verify` 逐字节一致 |
+| 上述全部 | 产物清点 + 呼应校验 | `scripts/inventory_products.py --check` | 对每类输入算跨度/条数,与它喂出来的产物逐项对拍(产物跨度必须落在输入跨度内) | `docs/系统设计说明.md` §2(自动块) | rc=0 |
+
+### 13.2 端到端顺序(为什么是这个顺序)
+
+`scripts/rebuild_all.py` 把 14 步写死(各步幂等,可整条重跑):
+
+```
+① 放数据 place_raw_data.py            原样搬 现场包 → data/raw/<场>/(同尺寸跳过,冲突拦下)
+② 三门台账 rebuild_from_raw.py         报警/工单/油样 ← 台账类源件
+③ SCADA 侧 rebuild_from_raw.py --scada 10 个构建器 ← 14 GB 10min CSV(约 15 分钟)
+④ 月度派生件 windscada_monthly_build.py 逐值对齐随包件才落盘(缺基线返回 5 只跳过"等价验收")
+④b 振动摄入 vib_raw_build.py           索引/谱库 ← TCM 导出(约 14 分钟)
+⑤ 补齐随包件 products_restore_missing.py 无生成端的那批(源 = 交付包 zip;缺源返回 6 跳过)
+⑥ 重启组件服务                         只重启组件、保留网关(chain_ingest 要吃重新加载的产物)
+⑦ 本体六步 kb_ingest → populate → chain_ingest → trend_ingest → 检索索引 → 实机参数表
+⑧ 审计 ontology.audit(期望 0 问题) + ⑧b 重装门户(把新产物灌进结论段) + ⑧c 页面归口审计
+```
+
+**依赖为什么不能换**:⑦ 的 `populate` 要吃 ②③④④b⑤ 的产物;⑤ 必须在 ⑥ 之前(本体层要读补齐后的 L1 件);
+⑧b 必须在产物全部就位之后(门户的结论段是产物的渲染结果)。
+
+### 13.3 前端取数与"即时进页面"
+
+| 服务 | 读法 | 产物更新后 |
+|---|---|---|
+| 工作台 `:18033`(`scripts/windscada_serve.py`) | 首请求把产物读进内存 `_CACHE`;每次取数前比**产物指纹**(相关文件的 `(路径, mtime_ns, size)` 摘要,TTL 2 s) | 指纹一变自动重载,**不用重启** |
+| CMS `:18020`(`src/windcms/serve.py`) | 直接读产物页面件 + `spectra_meta` 指纹(`src/windcms/data.py::_mtime_stamp`) | 同上 |
+| 门户 `:28084`(`scripts/guanlan_gateway.py`) | 门户按 `(mtime,size)` 缓存;契约结论段由 §13.1 的事实契约派生 | 需 `portal_build.py` 重装(⑧b) |
+| 运维控制台 `/ops`、`/ops/recalc` | 后端即真实状态;`Cache-Control: no-store` | 刷新即最新 |
+
+### 13.4 产物管理的四条硬要求(机器执行)
+
+1. **路径单一真源**:产物只落 `outputs/<场>/…`,代码只经 `src/paths.py` 取(§3)。
+2. **逐件来源台账**:`outputs/<场>/_provenance.json` 记 `raw-derived`(由 `data/raw` 重算,含验证依据)
+   或 `shipped`(包内无生成端,用随包件补齐);`raw-derived` 另由 `_derived_manifest.json` 自登记。
+3. **呼应校验**:`inventory_products.py --check` —— 产物跨度必须落在输入跨度内,rc≠0 即不符。
+4. **页面侧产物同样要登记**:`configs/portal_pages.yaml` 里 `data-derived` 的页面件,其 `source`(产物路径)
+   必须在 `_provenance.json` 里出现,否则 `pages_audit.py` 报"未纳入产物台账"。
+
+### 13.5 没有生成端的那批(`shipped`)与补齐办法
+
+`pitch_daily`、`pc_monthly_bins`、`duty_monthly`、`thermal_monthly`、`sector_power`、`yaw_*`、`genbearing_monthly`、
+`mblub_monthly`、`structure.parquet`、`watch_channels_monthly`、振动 `handoff_vibration_v2.json` /
+`component_history.json` / `baseline_38.json`、`sop/**`、`paradigm_r1/**`、`guanlan/**`、CMS 页面件 —— 全库只有读取方、
+0 处写入方(§7)。**唯一来源是"含产物的交付包"**:
+```
+python scripts/products_restore_missing.py --stash <含产物的包.zip>    # 只补缺件, 不覆盖 raw 重算件
+```
+2026-09-17 实测:用旧版含产物包 `guanlan-rudong-v2_0.2.0_test_win64.zip` 一次补齐 **567 件**(其中 39 件是日志,
+已按 §11.2 归位到 `logs/build/`),补齐后与旧包**逐字节相同 538 件**。
+**教训(写进 §9 的口径里)**:交付包默认不含产物 + 清除产物不留备份 ⇒ 这批件**一旦清掉就没有来源**;
+要留退路就用 `scripts/pack_dist.py --with-products` 打一份带产物的包存档。

+ 445 - 378
scripts/pages_audit.py

@@ -1,378 +1,445 @@
-#!/usr/bin/env python3
-# -*- coding: utf-8 -*-
-r"""页面归口审计 —— 把"这一页该不该随输入数据变 / 算不算产物"变成机器每天能查的事 (2026-09-17, 用户令 1)。
-
-## 背景(用户的问题)
-
-    http://127.0.0.1:28084/#findings  #sim  #documents 及它们包含的子页,
-    "应否随着输入数据的变化而变化"、"是否也属于产物"、"如是产物应纳入产物管理"。
-
-答案不是一句话能给的: 三页里既有**受管静态叙述**(方法论/架构), 也有**冻结交付件**(治理清单/报告),
-还有**真·数据派生快照**(把 outputs 里的数据烘进页面) —— 后者才是"属于产物、必须纳入产物管理"的那批。
-本模块把这份判断落成 `configs/portal_pages.yaml` 登记表 + 五条机器可查的规则, 于是:
-
-  · static            —— 正文不得引用产物(setting 变了就该改分类);
-  · live              —— iframe/链接指向的端口必须是 `configs/serve.json` 里的已知服务;
-  · data-citing       —— 引用的产物路径**必须存在**(否则是悬空引用, 挂羊头卖狗肉);
-  · data-derived      —— **必须有 source + (source_sha256 或 generated_at)**; 有 sha 的还要跟当前产物
-                         逐字节比对 ⇒ **陈旧检测**(数据换了、页面没换 = 报出来);
-  · frozen-delivery   —— 文件名或正文里必须能读到版本号与日期(客户拿到手才知道是哪一版)。
-
-## 为什么要有"陈旧检测"
-
-页面把产物烘进去, 就**脱离**了"产物即时进页面"的机制(§5.1 的指纹热重载只对实时读产物的服务生效)。
-数据重算后, 这种页面**不会自己变**; 没有检测, 它就一直挂着旧数字 —— 这正是用户担心的情形。
-
-## 退出码 (给 rebuild_all / check 用)
-
-    0 全部一致 · 5 登记缺项/文件缺失/未登记页 · 6 溯源缺失 · 7 页面陈旧(数据已变, 页面没变) · 8 交付件缺版本号 · 9 分类错误
-
-## 用法
-
-    python scripts/pages_audit.py                 # 表 + 检查(默认)
-    python scripts/pages_audit.py --list          # 只打表
-    python scripts/pages_audit.py --check         # 只检查(安静模式, 只打结论)
-    python scripts/pages_audit.py --write-doc     # 把表写进 docs/系统设计说明.md (标记区间内)
-    python scripts/pages_audit.py --json          # 机器可读结果
-"""
-from __future__ import annotations
-
-import argparse
-import fnmatch
-import glob as globmod
-import hashlib
-import html as htmllib
-import json
-import pathlib
-import re
-import sys
-
-ROOT = pathlib.Path(__file__).resolve().parents[1]
-sys.path.insert(0, str(ROOT))
-from src import paths as P                                          # noqa: E402
-
-REG = P.CONFIGS / 'portal_pages.yaml'
-DOC = P.ROOT / 'docs' / '系统设计说明.md'
-DOC_BEGIN, DOC_END = '<!-- PAGES:BEGIN -->', '<!-- PAGES:END -->'
-
-RC_OK, RC_MISS, RC_PROV, RC_STALE, RC_VER, RC_KIND = 0, 5, 6, 7, 8, 9
-KINDS = ('static', 'live', 'data-citing', 'data-derived', 'frozen-delivery')
-
-RE_OUT = re.compile(r'outputs/[A-Za-z0-9_\u4e00-\u9fff\-./]+')
-# 版本号只认"像版本"的: v1.0 / V2.1 / 版本1.2 / 客户版|外发版|正式版|评审版。
-# ★别写成 \d{8} 之类 —— 那会把页面里的机组编号(00019164)、时间戳当成版本号(2026-09-17 实测踩到)。
-RE_VER = re.compile(r'[vV]\d+(?:\.\d+)+|版本\s*[vV]?\d+(?:\.\d+)*|客户版|外发版|正式版|评审版')
-RE_DATE = re.compile(r'20\d\d-\d\d-\d\d|20\d{6}')
-RE_GEN = re.compile(r'生成[^0-9]{0,8}(20\d\d-\d\d-\d\d[ T]?\d?\d?:?\d?\d?)')
-
-
-def load(path=None):
-    import yaml
-    reg_p = pathlib.Path(path) if path else REG
-    if not reg_p.is_file():
-        raise SystemExit(f'缺登记表 {P.rel(reg_p)}')
-    reg = yaml.safe_load(reg_p.read_text(encoding='utf-8'))
-    reg['_file'] = str(reg_p)
-    return reg
-
-
-def sha256(p: pathlib.Path) -> str | None:
-    try:
-        return hashlib.sha256(p.read_bytes()).hexdigest()
-    except OSError:
-        return None
-
-
-def expand(entry) -> list[pathlib.Path]:
-    """一个登记项的 `file` 字段 → 实际文件列表 (支持 * 通配)。"""
-    pat = P.ROOT / entry['file']
-    s = str(pat)
-    if any(c in s for c in '*?['):
-        return sorted(pathlib.Path(x) for x in globmod.glob(s, recursive=True))
-    return [pat] if pat.is_file() else []
-
-
-def texts(files):
-    for f in files:
-        try:
-            yield f, f.read_text(encoding='utf-8', errors='replace')
-        except OSError:
-            continue
-
-
-def read_embedded_json(path: pathlib.Path, var: str):
-    """抓页面里内嵌的 window.<var> = {…} —— 模板里是 HTML 转义过的, 解析前先反转义。"""
-    try:
-        t = path.read_text(encoding='utf-8', errors='replace')
-    except OSError:
-        return None
-    m = re.search(r'window\.' + re.escape(var) + r'\s*=\s*(\{)', t)
-    if not m:
-        return None
-    raw, depth = '', 0
-    for ch in t[m.start(1):]:
-        raw += ch
-        if ch == '{':
-            depth += 1
-        elif ch == '}':
-            depth -= 1
-            if depth == 0:
-                break
-    for cand in (raw, htmllib.unescape(raw)):
-        try:
-            return json.loads(cand)
-        except Exception:
-            continue
-    return None
-
-
-def check_entry(e, results):
-    """按 kind 查一个登记项 → 往 results 里加 (level, 项, 说明, 退出码)。
-
-    level: OK 通过 / X 失败 / ! 要处理 / ? 只能人工看 / i 已知缺口(登记了 known_gap 并写了理由)。
-    ★ known_gap 的存在是**如实**的产物: 交付包里确实没有那件产物(例如振动线六层链的 cleaned/gearbox_life),
-      页面里却引用了它。我们不改交付件正文(客户手里那一版是冻结的), 但必须把"引用悬空"这件事记在明处,
-      所以它报 `i` 而不是装作通过, 也不当成新问题反复报警。
-    """
-    eid = e.get('id') or e.get('name') or '?'
-    kind = e.get('kind')
-    gap = e.get('known_gap') or {}
-    if kind not in KINDS:
-        results.append(('X', eid, f'kind 非法: {kind} (允许 {"/".join(KINDS)})', RC_KIND))
-        return
-    files = expand(e)
-    if not files and not e.get('members_in_zip'):
-        results.append(('X', eid, f'登记的 file 不存在或没匹配到: {e.get("file")}', RC_MISS))
-        return
-
-    # ① 引用的产物必须在位 (data-citing / data-derived; 登记了 cites 的也查)
-    for f, t in texts(files):
-        cited = set(e.get('cites') or [])
-        if kind in ('data-derived', 'data-citing'):
-            cited |= set(RE_OUT.findall(t))
-        for c in sorted(cited):
-            if (P.ROOT / c).exists():
-                continue
-            if gap and c in (gap.get('missing') or []):
-                results.append(('i', eid, f'引用悬空(已知): {c} —— {gap.get("why", "见登记表")}', RC_OK))
-            else:
-                results.append(('X', eid, f'{P.rel(f)} 引用的产物不在位: {c} (悬空引用)', RC_MISS))
-
-    # ② 分类规则
-    if kind == 'static':
-        for f, t in texts(files):
-            if RE_OUT.search(t):
-                results.append(('!', eid, f'{P.rel(f)} 正文里出现产物路径 —— 可能该改成 data-citing/data-derived',
-                                RC_KIND))
-
-    elif kind == 'live':
-        svc = {}
-        try:
-            svc = json.loads(P.SERVE_JSON.read_text(encoding='utf-8-sig'))
-        except Exception:
-            pass
-        ports = {int(v) for k, v in svc.items() if isinstance(v, int)}
-        tgt = e.get('live_target') or ''
-        m = re.search(r':(\d{2,5})', tgt)
-        if m and ports and int(m.group(1)) not in ports:
-            results.append(('X', eid, f'live_target 端口 {m.group(1)} 不在 configs/serve.json 的服务端口里', RC_MISS))
-
-    elif kind == 'data-derived':
-        src = e.get('source')
-        var = e.get('embedded_json_var')
-        skey, shkey = e.get('source_key'), e.get('source_sha_key')
-        found_sha, found_src, found_gen = None, src, None
-        if var and skey:
-            for f in files:
-                obj = read_embedded_json(f, var)
-                if isinstance(obj, dict):
-                    found_src = obj.get(skey) or found_src
-                    found_sha = obj.get(shkey) if shkey else None
-                    break
-        for f, t in texts(files):
-            if found_gen is None:
-                g = RE_GEN.search(t)
-                if g:
-                    found_gen = g.group(1)
-        if not found_src:
-            if gap:
-                results.append(('i', eid, f'溯源缺失(已知): {gap.get("why", "")}', RC_OK))
-            else:
-                results.append(('!', eid, 'data-derived 但没登记 source, 页面里也没有内嵌来源 ⇒ 溯源缺失', RC_PROV))
-        if not (found_sha or found_gen or e.get('generated_at_key')):
-            if not gap:
-                results.append(('!', eid, 'data-derived 但既无 source_sha256 也无生成时间 ⇒ 无法做陈旧检测', RC_PROV))
-        if found_src:
-            sp = P.ROOT / found_src
-            if not sp.exists():
-                if gap and found_src in (gap.get('missing') or []):
-                    results.append(('i', eid, f'内嵌来源不在位(已知): {found_src} —— {gap.get("why", "")}', RC_OK))
-                else:
-                    results.append(('!', eid, f'内嵌来源 {found_src} 不在位 (页面里烘的是别处/历史数据)', RC_MISS))
-            elif found_sha:
-                cur = sha256(sp)
-                if cur and cur.lower() != str(found_sha).lower():
-                    results.append(('X', eid,
-                                    f'**页面陈旧**: 内嵌快照 source_sha256={str(found_sha)[:16]} 与当前 '
-                                    f'{found_src} 的 sha256={cur[:16]} 不同 ⇒ 数据变了, 这份页面没跟着重生成',
-                                    RC_STALE))
-                else:
-                    results.append(('OK', eid, f'快照与当前产物一致 ({found_src} sha256 {str(found_sha)[:16]})', RC_OK))
-            else:
-                results.append(('?', eid, f'有来源 {found_src} 但无指纹, 只能按生成时间判断 '
-                                          f'(记录: {found_gen or "无"})', RC_PROV))
-
-    elif kind == 'frozen-delivery':
-        ok_ver = ok_date = False
-        for f in files:
-            name = f.name
-            ok_ver |= bool(RE_VER.search(name))
-            ok_date |= bool(RE_DATE.search(name))
-            if not (ok_ver and ok_date):
-                t = f.read_text(encoding='utf-8', errors='replace')      # 版本号/日期可能在正文深处
-                ok_ver |= bool(RE_VER.search(t))
-                ok_date |= bool(RE_DATE.search(t))
-        if e.get('version'):
-            ok_ver = True
-        if not ok_ver:
-            results.append(('!', eid, '冻结交付件但没有版本号 (文件名/正文都没有 v*/日期) ⇒ 客户拿到手分不清是哪一版',
-                            RC_VER))
-        if not ok_date and not e.get('version'):
-            results.append(('?', eid, '冻结交付件没有日期标记', RC_VER))
-
-    # ③ 子项递归
-    for c in e.get('children') or []:
-        check_entry(c, results)
-
-
-def declared_patterns(reg) -> list[str]:
-    """登记表里所有 file 字段 (含通配) —— 覆盖检查用。"""
-    pats = []
-
-    def walk(e):
-        f = e.get('file')
-        if f:
-            pats.append(str(f))
-        for c in e.get('children') or []:
-            walk(c)
-    for e in reg.get('pages') or []:
-        walk(e)
-    return pats
-
-
-def coverage_gaps(reg) -> list[str]:
-    """登记表里没写、但门户里真实存在的内嵌页 —— 漏登记就等于没管。"""
-    pats = declared_patterns(reg)
-    gaps = []
-    tdir = P.RELEASE / 'portal_src' / 'templates'
-    for f in sorted(tdir.glob('*.html')):
-        rel = f.relative_to(P.ROOT).as_posix()
-        if not any(fnmatch.fnmatch(rel, p) or rel == p for p in pats):
-            gaps.append(rel)
-    return gaps
-
-
-def render(reg, results) -> str:
-    lines = ['| 页面/子页 | kind | 随输入数据变? | 依据(为什么这么判) |', '|---|---|---|---|']
-    def walk(e, depth=0):
-        eid = e.get('id') or e.get('name') or '?'
-        title = e.get('title') or ''
-        pre = '&nbsp;&nbsp;└ ' * depth
-        lines.append(f"| {pre}`{eid}`{' ' + title if title and e.get('id') else ''} | `{e.get('kind')}` | "
-                     f"{'**是**' if e.get('changes_with_data') else '否'} | {e.get('evidence', '')} |")
-        for c in e.get('children') or []:
-            walk(c, depth + 1)
-    for e in reg.get('pages') or []:
-        walk(e)
-    return '\n'.join(lines)
-
-
-def audit(registry=None):
-    """→ (rc, results, unregistered)。
-
-    供 `guanlan.py check` / `scripts/rebuild_all.py` 直接调用 —— 走函数而不是子进程:
-    子进程的 stdout 是中文, Windows 控制台默认 cp936 会把 UTF-8 输出解成乱码, 解析结论就不可靠了。
-    """
-    reg = load(registry)
-    results: list[tuple[str, str, str, int]] = []
-    for e in reg.get('pages') or []:
-        check_entry(e, results)
-    rc_map = {r[3] for r in results if r[0] not in ('OK', 'i') and r[3]}
-    return (max(rc_map) if rc_map else RC_OK), results, coverage_gaps(reg)
-
-
-def main() -> int:
-    ap = argparse.ArgumentParser()
-    ap.add_argument('--list', action='store_true', help='只打表')
-    ap.add_argument('--check', action='store_true', help='只检查')
-    ap.add_argument('--write-doc', action='store_true', help='把表写进 docs/系统设计说明.md')
-    ap.add_argument('--json', action='store_true', help='机器可读输出')
-    ap.add_argument('--registry', default=None, help='登记表路径 (默认 configs/portal_pages.yaml; 测试用)')
-    a = ap.parse_args()
-    reg = load(a.registry)
-    rc, results, gaps = audit(a.registry)
-
-    if a.json:
-        print(json.dumps(dict(rc=rc,
-                              results=[dict(level=r[0], item=r[1], note=r[2], rc=r[3]) for r in results],
-                              unregistered=gaps), ensure_ascii=False, indent=1))
-        return rc
-
-    if a.list or not a.check:
-        print('== 门户页面归口 (configs/portal_pages.yaml) ==')
-        for e in reg.get('pages') or []:
-            def walk(x, d=0):
-                eid = x.get('id') or x.get('name')
-                print(f'   {"  " * d}{eid:34s} {x.get("kind"):16s} '
-                      f'{"随数据变" if x.get("changes_with_data") else "不随数据变"}')
-                for c in x.get('children') or []:
-                    walk(c, d + 1)
-            walk(e)
-
-    if not a.check or True:
-        lvl = {}
-        for level, item, note, r in results:
-            lvl[level] = lvl.get(level, 0) + 1
-        print(f'\n== 检查: {lvl.get("OK", 0)} 项一致, {lvl.get("i", 0)} 项已知缺口, '
-              f'{sum(v for k, v in lvl.items() if k not in ("OK", "i"))} 项要处理 ==')
-        for level, item, note, r in results:
-            if level not in ('OK', 'i'):
-                print(f'   [{level}] {item}: {note}')
-        if lvl.get('i'):
-            print(f'   [i] 已知缺口 {lvl["i"]} 条 (引用悬空/溯源缺失, 已在登记表里写明理由, 不重复刷屏):')
-            seen = set()
-            for level, item, note, r in results:
-                if level == 'i' and item not in seen:
-                    seen.add(item)
-                    print(f'        · {item}')
-        if gaps:
-            print(f'   [!] 门户里有 {len(gaps)} 个内嵌页/子页**未登记**(漏登记=没管):')
-            for g in gaps[:8]:
-                print(f'        · {g}')
-        print(f'结论: {"全部一致" if rc == RC_OK else "见上"}; 退出码 {rc}')
-        if a.write_doc:
-            write_doc(reg)
-        return rc
-    return rc
-
-
-def write_doc(reg) -> None:
-    if not DOC.is_file():
-        print(f'   (缺 {P.rel(DOC)}, 跳过写文档)')
-        return
-    t = DOC.read_text(encoding='utf-8')
-    block = (f'{DOC_BEGIN}\n### 10.1 门户页面归口(自动生成,勿手改)\n\n'
-             f'判据与说明见 `configs/portal_pages.yaml` 头部;检查器 `scripts/pages_audit.py`。\n\n'
-             f'{render(reg, None)}\n\n{DOC_END}')
-    if DOC_BEGIN in t and DOC_END in t:
-        t = re.sub(re.escape(DOC_BEGIN) + r'.*?' + re.escape(DOC_END), lambda m: block, t, flags=re.S)
-    else:
-        t = t.rstrip() + '\n\n---\n\n## 10. 页面归口(哪些页面是产物、该不该随数据变)\n\n' + block + '\n'
-    DOC.write_text(t, encoding='utf-8')
-    print(f'   已写入 {P.rel(DOC)} (§10 页面归口)')
-
-
-if __name__ == '__main__':
-    from src import console
-    console.soft()
-    sys.exit(main())
+#!/usr/bin/env python3
+# -*- coding: utf-8 -*-
+r"""页面归口审计 —— 把"这一页该不该随输入数据变 / 算不算产物"变成机器每天能查的事 (2026-09-17, 用户令 1)。
+
+## 背景(用户的问题)
+
+    http://127.0.0.1:28084/#findings  #sim  #documents 及它们包含的子页,
+    "应否随着输入数据的变化而变化"、"是否也属于产物"、"如是产物应纳入产物管理"。
+
+答案不是一句话能给的: 三页里既有**受管静态叙述**(方法论/架构), 也有**冻结交付件**(治理清单/报告),
+还有**真·数据派生快照**(把 outputs 里的数据烘进页面) —— 后者才是"属于产物、必须纳入产物管理"的那批。
+本模块把这份判断落成 `configs/portal_pages.yaml` 登记表 + 五条机器可查的规则, 于是:
+
+  · static            —— 正文不得引用产物(setting 变了就该改分类);
+  · live              —— iframe/链接指向的端口必须是 `configs/serve.json` 里的已知服务;
+  · data-citing       —— 引用的产物路径**必须存在**(否则是悬空引用, 挂羊头卖狗肉);
+  · data-derived      —— **必须有 source + (source_sha256 或 generated_at)**; 有 sha 的还要跟当前产物
+                         逐字节比对 ⇒ **陈旧检测**(数据换了、页面没换 = 报出来);
+  · frozen-delivery   —— 文件名或正文里必须能读到版本号与日期(客户拿到手才知道是哪一版)。
+
+## 为什么要有"陈旧检测"
+
+页面把产物烘进去, 就**脱离**了"产物即时进页面"的机制(§5.1 的指纹热重载只对实时读产物的服务生效)。
+数据重算后, 这种页面**不会自己变**; 没有检测, 它就一直挂着旧数字 —— 这正是用户担心的情形。
+
+## 退出码 (给 rebuild_all / check 用)
+
+    0 全部一致 · 5 登记缺项/文件缺失/未登记页 · 6 溯源缺失 · 7 页面陈旧(数据已变, 页面没变) · 8 交付件缺版本号 · 9 分类错误
+
+## 用法
+
+    python scripts/pages_audit.py                 # 表 + 检查(默认)
+    python scripts/pages_audit.py --list          # 只打表
+    python scripts/pages_audit.py --check         # 只检查(安静模式, 只打结论)
+    python scripts/pages_audit.py --write-doc     # 把表写进 docs/系统设计说明.md (标记区间内)
+    python scripts/pages_audit.py --json          # 机器可读结果
+"""
+from __future__ import annotations
+
+import argparse
+import fnmatch
+import glob as globmod
+import hashlib
+import html as htmllib
+import json
+import pathlib
+import re
+import sys
+
+ROOT = pathlib.Path(__file__).resolve().parents[1]
+sys.path.insert(0, str(ROOT))
+from src import paths as P                                          # noqa: E402
+
+REG = P.CONFIGS / 'portal_pages.yaml'
+DOC = P.ROOT / 'docs' / '系统设计说明.md'
+
+
+def dt_from(ts: float | None) -> str:
+    """mtime → 'YYYY-mm-dd HH:MM' (报告里给人看的时间)。"""
+    if not ts:
+        return '?'
+    import datetime as _dt
+    return _dt.datetime.fromtimestamp(ts).strftime('%Y-%m-%d %H:%M')
+DOC_BEGIN, DOC_END = '<!-- PAGES:BEGIN -->', '<!-- PAGES:END -->'
+
+RC_OK, RC_MISS, RC_PROV, RC_STALE, RC_VER, RC_KIND = 0, 5, 6, 7, 8, 9
+KINDS = ('static', 'live', 'data-citing', 'data-derived', 'frozen-delivery')
+
+RE_OUT = re.compile(r'outputs/[A-Za-z0-9_\u4e00-\u9fff\-./]+')
+# 版本号只认"像版本"的: v1.0 / V2.1 / 版本1.2 / 客户版|外发版|正式版|评审版。
+# ★别写成 \d{8} 之类 —— 那会把页面里的机组编号(00019164)、时间戳当成版本号(2026-09-17 实测踩到)。
+RE_VER = re.compile(r'[vV]\d+(?:\.\d+)+|版本\s*[vV]?\d+(?:\.\d+)*|客户版|外发版|正式版|评审版')
+RE_DATE = re.compile(r'20\d\d-\d\d-\d\d|20\d{6}')
+RE_GEN = re.compile(r'生成[^0-9]{0,8}(20\d\d-\d\d-\d\d[ T]?\d?\d?:?\d?\d?)')
+
+
+def load(path=None):
+    import yaml
+    reg_p = pathlib.Path(path) if path else REG
+    if not reg_p.is_file():
+        raise SystemExit(f'缺登记表 {P.rel(reg_p)}')
+    reg = yaml.safe_load(reg_p.read_text(encoding='utf-8'))
+    reg['_file'] = str(reg_p)
+    return reg
+
+
+def sha256(p: pathlib.Path) -> str | None:
+    try:
+        return hashlib.sha256(p.read_bytes()).hexdigest()
+    except OSError:
+        return None
+
+
+def expand(entry) -> list[pathlib.Path]:
+    """一个登记项的 `file` 字段 → 实际文件列表 (支持 * 通配)。"""
+    pat = P.ROOT / entry['file']
+    s = str(pat)
+    if any(c in s for c in '*?['):
+        return sorted(pathlib.Path(x) for x in globmod.glob(s, recursive=True))
+    return [pat] if pat.is_file() else []
+
+
+def texts(files):
+    for f in files:
+        try:
+            yield f, f.read_text(encoding='utf-8', errors='replace')
+        except OSError:
+            continue
+
+
+def read_embedded_json(path: pathlib.Path, var: str):
+    """抓页面里内嵌的 window.<var> = {…} —— 模板里是 HTML 转义过的, 解析前先反转义。"""
+    try:
+        t = path.read_text(encoding='utf-8', errors='replace')
+    except OSError:
+        return None
+    m = re.search(r'window\.' + re.escape(var) + r'\s*=\s*(\{)', t)
+    if not m:
+        return None
+    raw, depth = '', 0
+    for ch in t[m.start(1):]:
+        raw += ch
+        if ch == '{':
+            depth += 1
+        elif ch == '}':
+            depth -= 1
+            if depth == 0:
+                break
+    for cand in (raw, htmllib.unescape(raw)):
+        try:
+            return json.loads(cand)
+        except Exception:
+            continue
+    return None
+
+
+def check_entry(e, results):
+    """按 kind 查一个登记项 → 往 results 里加 (level, 项, 说明, 退出码)。
+
+    level: OK 通过 / X 失败 / ! 要处理 / ? 只能人工看 / i 已知缺口(登记了 known_gap 并写了理由)。
+    ★ known_gap 的存在是**如实**的产物: 交付包里确实没有那件产物(例如振动线六层链的 cleaned/gearbox_life),
+      页面里却引用了它。我们不改交付件正文(客户手里那一版是冻结的), 但必须把"引用悬空"这件事记在明处,
+      所以它报 `i` 而不是装作通过, 也不当成新问题反复报警。
+    """
+    eid = e.get('id') or e.get('name') or '?'
+    kind = e.get('kind')
+    gap = e.get('known_gap') or {}
+    if kind not in KINDS:
+        results.append(('X', eid, f'kind 非法: {kind} (允许 {"/".join(KINDS)})', RC_KIND))
+        return
+    files = expand(e)
+    if not files and not e.get('members_in_zip'):
+        results.append(('X', eid, f'登记的 file 不存在或没匹配到: {e.get("file")}', RC_MISS))
+        return
+
+    # ① 引用的产物必须在位 (data-citing / data-derived; 登记了 cites 的也查)
+    for f, t in texts(files):
+        cited = set(e.get('cites') or [])
+        if kind in ('data-derived', 'data-citing'):
+            cited |= set(RE_OUT.findall(t))
+        for c in sorted(cited):
+            if (P.ROOT / c).exists():
+                continue
+            if gap and c in (gap.get('missing') or []):
+                results.append(('i', eid, f'引用悬空(已知): {c} —— {gap.get("why", "见登记表")}', RC_OK))
+            else:
+                results.append(('X', eid, f'{P.rel(f)} 引用的产物不在位: {c} (悬空引用)', RC_MISS))
+
+    # ② 分类规则
+    if kind == 'static':
+        for f, t in texts(files):
+            if RE_OUT.search(t):
+                results.append(('!', eid, f'{P.rel(f)} 正文里出现产物路径 —— 可能该改成 data-citing/data-derived',
+                                RC_KIND))
+
+    elif kind == 'live':
+        svc = {}
+        try:
+            svc = json.loads(P.SERVE_JSON.read_text(encoding='utf-8-sig'))
+        except Exception:
+            pass
+        ports = {int(v) for k, v in svc.items() if isinstance(v, int)}
+        tgt = e.get('live_target') or ''
+        m = re.search(r':(\d{2,5})', tgt)
+        if m and ports and int(m.group(1)) not in ports:
+            results.append(('X', eid, f'live_target 端口 {m.group(1)} 不在 configs/serve.json 的服务端口里', RC_MISS))
+
+    elif kind == 'data-derived':
+        src = e.get('source')
+        var = e.get('embedded_json_var')
+        skey, shkey = e.get('source_key'), e.get('source_sha_key')
+        found_sha, found_src, found_gen = None, src, None
+        if var and skey:
+            for f in files:
+                obj = read_embedded_json(f, var)
+                if isinstance(obj, dict):
+                    found_src = obj.get(skey) or found_src
+                    found_sha = obj.get(shkey) if shkey else None
+                    break
+        for f, t in texts(files):
+            if found_gen is None:
+                g = RE_GEN.search(t)
+                if g:
+                    found_gen = g.group(1)
+        if not found_src:
+            if gap:
+                results.append(('i', eid, f'溯源缺失(已知): {gap.get("why", "")}', RC_OK))
+            else:
+                results.append(('!', eid, 'data-derived 但没登记 source, 页面里也没有内嵌来源 ⇒ 溯源缺失', RC_PROV))
+        if not (found_sha or found_gen or e.get('generated_at_key')):
+            if e.get('generated_at_source') == 'mtime':
+                # 页面件自己就是生成物: 用"它自己的 mtime"当生成时间, 与 source 的 mtime 比 ⇒ 可做陈旧代理判据
+                src = e.get('source')
+                mt = None
+                for f in files:
+                    try:
+                        mt = f.stat().st_mtime
+                        break
+                    except OSError:
+                        continue
+                sp = (P.ROOT / src) if isinstance(src, str) else None
+                if mt and sp is not None and sp.is_file() and (sp.stat().st_mtime - mt) > 86400:
+                    results.append(('X', eid, f'**页面件疑似陈旧**: {P.rel(files[0])} 写于 '
+                                              f'{dt_from(mt)}, 而它的源 {src} 更新于 {dt_from(sp.stat().st_mtime)} '
+                                              f'⇒ 源换了、这份页面件没重生成', RC_STALE))
+                else:
+                    results.append(('OK', eid, f'按 mtime 判: 页面件与源 {src or "(未登记)"} 的新旧关系正常 '
+                                                f'(页面件 {dt_from(mt) if mt else "?"})', RC_OK))
+                if not found_src:
+                    pass
+            elif not gap:
+                results.append(('!', eid, 'data-derived 但既无 source_sha256 也无生成时间 ⇒ 无法做陈旧检测', RC_PROV))
+        if found_src:
+            sp = P.ROOT / found_src
+            if not sp.exists():
+                if gap and found_src in (gap.get('missing') or []):
+                    results.append(('i', eid, f'内嵌来源不在位(已知): {found_src} —— {gap.get("why", "")}', RC_OK))
+                else:
+                    results.append(('!', eid, f'内嵌来源 {found_src} 不在位 (页面里烘的是别处/历史数据)', RC_MISS))
+            elif found_sha:
+                cur = sha256(sp)
+                if cur and cur.lower() != str(found_sha).lower():
+                    results.append(('X', eid,
+                                    f'**页面陈旧**: 内嵌快照 source_sha256={str(found_sha)[:16]} 与当前 '
+                                    f'{found_src} 的 sha256={cur[:16]} 不同 ⇒ 数据变了, 这份页面没跟着重生成',
+                                    RC_STALE))
+                else:
+                    results.append(('OK', eid, f'快照与当前产物一致 ({found_src} sha256 {str(found_sha)[:16]})', RC_OK))
+            elif e.get('generated_at_source') != 'mtime':
+                results.append(('?', eid, f'有来源 {found_src} 但无指纹, 只能按生成时间判断 '
+                                          f'(记录: {found_gen or "无"})', RC_PROV))
+
+    elif kind == 'frozen-delivery':
+        ok_ver = ok_date = False
+        for f in files:
+            name = f.name
+            ok_ver |= bool(RE_VER.search(name))
+            ok_date |= bool(RE_DATE.search(name))
+            if not (ok_ver and ok_date):
+                t = f.read_text(encoding='utf-8', errors='replace')      # 版本号/日期可能在正文深处
+                ok_ver |= bool(RE_VER.search(t))
+                ok_date |= bool(RE_DATE.search(t))
+        if e.get('version'):
+            ok_ver = True
+        if not ok_ver:
+            results.append(('!', eid, '冻结交付件但没有版本号 (文件名/正文都没有 v*/日期) ⇒ 客户拿到手分不清是哪一版',
+                            RC_VER))
+        if not ok_date and not e.get('version'):
+            results.append(('?', eid, '冻结交付件没有日期标记', RC_VER))
+
+    # ③ 子项递归
+    for c in e.get('children') or []:
+        check_entry(c, results)
+
+
+def declared_patterns(reg) -> list[str]:
+    """登记表里所有 file 字段 (含通配) —— 覆盖检查用。"""
+    pats = []
+
+    def walk(e):
+        f = e.get('file')
+        if f:
+            pats.append(str(f))
+        for c in e.get('children') or []:
+            walk(c)
+    for e in reg.get('pages') or []:
+        walk(e)
+    return pats
+
+
+def coverage_gaps(reg) -> list[str]:
+    """登记表里没写、但门户里真实存在的内嵌页 —— 漏登记就等于没管。"""
+    pats = declared_patterns(reg)
+    gaps = []
+    tdir = P.RELEASE / 'portal_src' / 'templates'
+    for f in sorted(tdir.glob('*.html')):
+        rel = f.relative_to(P.ROOT).as_posix()
+        if not any(fnmatch.fnmatch(rel, p) or rel == p for p in pats):
+            gaps.append(rel)
+    return gaps
+
+
+def render(reg, results) -> str:
+    lines = ['| 页面/子页 | kind | 随输入数据变? | 依据(为什么这么判) |', '|---|---|---|---|']
+    def walk(e, depth=0):
+        eid = e.get('id') or e.get('name') or '?'
+        title = e.get('title') or ''
+        pre = '&nbsp;&nbsp;└ ' * depth
+        lines.append(f"| {pre}`{eid}`{' ' + title if title and e.get('id') else ''} | `{e.get('kind')}` | "
+                     f"{'**是**' if e.get('changes_with_data') else '否'} | {e.get('evidence', '')} |")
+        for c in e.get('children') or []:
+            walk(c, depth + 1)
+    for e in reg.get('pages') or []:
+        walk(e)
+    return '\n'.join(lines)
+
+
+def audit(registry=None):
+    """→ (rc, results, unregistered)。
+
+    供 `guanlan.py check` / `scripts/rebuild_all.py` 直接调用 —— 走函数而不是子进程:
+    子进程的 stdout 是中文, Windows 控制台默认 cp936 会把 UTF-8 输出解成乱码, 解析结论就不可靠了。
+    """
+    reg = load(registry)
+    results: list[tuple[str, str, str, int]] = []
+    for e in reg.get('pages') or []:
+        check_entry(e, results)
+    results += ledger_coverage(reg)
+    rc_map = {r[3] for r in results if r[0] not in ('OK', 'i', '?') and r[3]}
+    return (max(rc_map) if rc_map else RC_OK), results, coverage_gaps(reg)
+
+
+def ledger_coverage(reg) -> list:
+    """页面侧产物的**台账覆盖**检查 (用户令 2026-09-17: 是产物就要纳入产物管理)。
+
+    凡登记表里指向 `outputs/<场>/…` 的路径 (`file` / `source`), 都必须出现在该场的
+    `_provenance.json::files` 里 —— 否则它虽然被页面用了, 却没进产物台账 (来源/类型无从查)。
+    """
+    out: list[tuple[str, str, str, int]] = []
+    prov_p = P.out_root() / '_provenance.json'
+    if not prov_p.is_file():
+        return out
+    try:
+        prov = set((json.loads(prov_p.read_text(encoding='utf-8')).get('files') or {}).keys())
+    except Exception:
+        return out
+    prefix = P.out_root().relative_to(P.ROOT).as_posix() + '/'      # outputs/<场>/
+
+    def walk(e):
+        eid = e.get('id') or e.get('name') or '?'
+        gap = e.get('known_gap') or {}
+        gap_missing = set(gap.get('missing') or [])
+        for key in ('file', 'source'):
+            v = e.get(key)
+            if not isinstance(v, str) or not v.startswith(prefix) or any(c in v for c in '*?['):
+                continue
+            if v in gap_missing or any(v.endswith(g.split('/')[-1]) for g in gap_missing):
+                out.append(('i', eid, f'页面侧产物 {v} 不在本包(已知缺口): {gap.get("why", "")[:60]}', RC_OK))
+                continue
+            rel = v[len(prefix):]
+            if rel not in prov:
+                out.append(('X', eid, f'页面侧产物 {v} 未进产物台账 (_provenance.json 里没有 {rel}) '
+                                      f'⇒ 来源/类型无从查, 不算"纳入产物管理"', RC_PROV))
+        for c in e.get('children') or []:
+            walk(c)
+    for e in reg.get('pages') or []:
+        walk(e)
+    return out
+
+
+def main() -> int:
+    ap = argparse.ArgumentParser()
+    ap.add_argument('--list', action='store_true', help='只打表')
+    ap.add_argument('--check', action='store_true', help='只检查')
+    ap.add_argument('--write-doc', action='store_true', help='把表写进 docs/系统设计说明.md')
+    ap.add_argument('--json', action='store_true', help='机器可读输出')
+    ap.add_argument('--registry', default=None, help='登记表路径 (默认 configs/portal_pages.yaml; 测试用)')
+    a = ap.parse_args()
+    reg = load(a.registry)
+    rc, results, gaps = audit(a.registry)
+
+    if a.json:
+        print(json.dumps(dict(rc=rc,
+                              results=[dict(level=r[0], item=r[1], note=r[2], rc=r[3]) for r in results],
+                              unregistered=gaps), ensure_ascii=False, indent=1))
+        return rc
+
+    if a.list or not a.check:
+        print('== 门户页面归口 (configs/portal_pages.yaml) ==')
+        for e in reg.get('pages') or []:
+            def walk(x, d=0):
+                eid = x.get('id') or x.get('name')
+                print(f'   {"  " * d}{eid:34s} {x.get("kind"):16s} '
+                      f'{"随数据变" if x.get("changes_with_data") else "不随数据变"}')
+                for c in x.get('children') or []:
+                    walk(c, d + 1)
+            walk(e)
+
+    if not a.check or True:
+        lvl = {}
+        for level, item, note, r in results:
+            lvl[level] = lvl.get(level, 0) + 1
+        print(f'\n== 检查: {lvl.get("OK", 0)} 项一致, {lvl.get("i", 0)} 项已知缺口, '
+              f'{sum(v for k, v in lvl.items() if k not in ("OK", "i"))} 项要处理 ==')
+        for level, item, note, r in results:
+            if level not in ('OK', 'i'):
+                print(f'   [{level}] {item}: {note}')
+        if lvl.get('i'):
+            print(f'   [i] 已知缺口 {lvl["i"]} 条 (引用悬空/溯源缺失, 已在登记表里写明理由, 不重复刷屏):')
+            seen = set()
+            for level, item, note, r in results:
+                if level == 'i' and item not in seen:
+                    seen.add(item)
+                    print(f'        · {item}')
+        if gaps:
+            print(f'   [!] 门户里有 {len(gaps)} 个内嵌页/子页**未登记**(漏登记=没管):')
+            for g in gaps[:8]:
+                print(f'        · {g}')
+        print(f'结论: {"全部一致" if rc == RC_OK else "见上"}; 退出码 {rc}')
+        if a.write_doc:
+            write_doc(reg)
+        return rc
+    return rc
+
+
+def write_doc(reg) -> None:
+    if not DOC.is_file():
+        print(f'   (缺 {P.rel(DOC)}, 跳过写文档)')
+        return
+    t = DOC.read_text(encoding='utf-8')
+    block = (f'{DOC_BEGIN}\n### 10.1 门户页面归口(自动生成,勿手改)\n\n'
+             f'判据与说明见 `configs/portal_pages.yaml` 头部;检查器 `scripts/pages_audit.py`。\n\n'
+             f'{render(reg, None)}\n\n{DOC_END}')
+    if DOC_BEGIN in t and DOC_END in t:
+        t = re.sub(re.escape(DOC_BEGIN) + r'.*?' + re.escape(DOC_END), lambda m: block, t, flags=re.S)
+    else:
+        t = t.rstrip() + '\n\n---\n\n## 10. 页面归口(哪些页面是产物、该不该随数据变)\n\n' + block + '\n'
+    DOC.write_text(t, encoding='utf-8')
+    print(f'   已写入 {P.rel(DOC)} (§10 页面归口)')
+
+
+if __name__ == '__main__':
+    from src import console
+    console.soft()
+    sys.exit(main())