Przeglądaj źródła

事实契约面落成源代码: 由重算台账生成 claim, 门户/问答/报告三处随重算刷新

用户令 2026-09-19 选定: 「要让观澜从重算台账生成 claim」+「pitch_daily 零位取停机段中位、满发段中位两者分列出」。

一、事实契约(本次落地)
- 新增 scripts/sop_findings_from_ledger.py: 按 guanlan_facts_contract.py 读的四份输入的**字段结构**生成
  sop/findings.json + paradigm_r1 的 E5/E3/E8_s0 三份底稿, 契约构建器一行未改。
  claim 六族全部由**重算产物**算出, 每条带 coverage(实算窗天数/覆盖台数)、named 台号、systems、
  mandatory_limitation(缺什么证据)、falsifiability(怎么证伪), 相对判据一律封顶「候选」:
    报警集中(alarms) · 重复检修(workorders) · 停机损失集中(loss_monthly) · 温度相对离群(temp_bins, n≥20 门)
    · 功率曲线偏离(powercurve_dev) · 振动面过闸线(model_run_l6, 机制未定⇒不命名部件)
  单位两处实逮修正: loss 列是 kWh(直接当 MWh 会大 1000 倍)、dev_w 是 MW(直接标 kW 会小 1000 倍);
  工单台号规范化(33→WTG33); 功率曲线改报 |偏差| 最大台 + 全场判别分布。
- 族表: 新增 ledger_claims(raw-derived, gen=sop_findings_from_ledger.py); guanlan_contract 从 shipped
  改为 raw-derived(gen=guanlan_facts_contract.py —— 它的输入现在由上游族自算, 两跳的生成端都在包内)。
- 链序: ④b → **⑤b 生成 claim → ⑤c 构建+渲染契约** → ⑤a 台账维护 → ⑤ 审计(新件进同一轮台账与审计)。
- 复检: /api/facts **200**(cards 6, 此前 503); 门户结论段由旧的人裁结论换成 RD-LEDGER-*(2.6KB, 6 条:
  4 候选/2 参考); 契约自检 PASS(sha/脱敏/枚举); 台账 1787 件全 raw-derived(shipped 0 / 人工件 0 / 未归类 0);
  反向呼应审计 rc=0。

二、变桨面(下一件, 口径已定)
- 消费者字段契约已核清(src/windscada/subsys/pitch.py): pitch_daily 需 date/turbine/n/hydlevel_timeon/
  hydfilt_timeon/hyd_band/hyd_over_relief/hyd_under_pump/hublub_timeon/pitchpum_timeon/pistA-C/n_op;
  pitch_zero_monthly 消费端读 zero_dev ⇒ 本次要求的两口径**分列**(zero_dev_stop / zero_dev_full), zero_dev 取停机段。
- 通道已勘察(10min 表头实测): din_wtc_HydLevel/HydrFilt_timeon、din_wtc_LubPistA/B/C_*、
  dot_wtc_HubLubPu_timeon、dot_wtc_PitchPum_timeon、prs_wtc_*/int_wtc_ValSuppV_*(待挑蓄能器压力);
  零位两口径取 **1min** 的 pitch_angle_blade_1/2/3(现已可由 scada_source 从 MDB 读)。
- 为什么这一件单独一轮: 它**直接驱动变桨面判级**(pitch.registry 分 D1-D4 维度给状态), 通道挑错=判级错,
  故按"先定口径→再落生成端→抽台对拍"的顺序做, 不猜通道。
- docs/源代码化清单_v0.1.md 已更新: §0 记已落成三件(融合面/总览页/事实契约), §2 写变桨面字段契约 +
  通道勘察 + 两口径取数办法, §4 修订执行顺序(releases → component_history/baseline_38 → tcm_replay → pitch → 低优先)。
zhouyang.xie 3 tygodni temu
rodzic
commit
33de3e9bea

+ 48 - 18
docs/源代码化清单_v0.1.md

@@ -11,14 +11,20 @@
 
 ---
 
-## 0. 今天已经补上源代码的两件(本令的第一批)
+## 0. 已经补上源代码的(本令的前几批)
 
 | 件 | 生成端(在包内) | 喂谁 | 复检 |
 |---|---|---|---|
 | `m5_cms_tcm/handoff_vibration_v2.json`(融合面) | **`scripts/rudong_fusion_handoff.py`**(观澜自算:报告状态等级 + `fusion_38.csv` + L6 过闸线 → 同结构件;`meta.from` 标"观澜自算",盘上若有振动线正本则不覆盖) | `/detail/v2` 的**「需要关注」+「全场状态」**(`/api/fleet` 的 `fus`) | 链盘 38 行 / 需要关注 **3 台**(WTG09 齿轮箱、WTG17/WTG27 发电机)/ 矩阵 38 行 ✔ |
-| `windscada/index.html`(总览「全场状态」页) | **`scripts/windscada_overview_build.py`**(用组件同一个快照烘焙器 `src/windscada/ui/snapshot.py::bake` 把 `/api/fleet`・`curves`・`vibcms`・本体响应烤进单文件页) | `/detail/v2` 的「全场状态总览」`<iframe src="/static/index.html">` | 3.78 MB 真页面(此前是 2 KB 的"产物不在位"说明页)✔ |
+| `windscada/index.html`(总览「全场状态」页) | **`scripts/windscada_overview_build.py`**(用组件同一个快照烘焙器 `src/windscada/ui/snapshot.py::bake`) | `/detail/v2` 的「全场状态总览」`<iframe src="/static/index.html">` | 3.78 MB 真页面(此前是 2 KB 的"产物不在位"说明页)✔ |
+| **事实契约面**:`sop/findings.json` + `paradigm_r1` 三份底稿 + `guanlan/**` | **`scripts/sop_findings_from_ledger.py`**(重算台账 → claim,按契约构建器要的字段结构写)+ **`scripts/guanlan_facts_contract.py build`**(原有生成端,输入换成自算件) | 门户「本场结论·契约生成」段、`/detail` 问答/报告面(`/api/facts`)、报告摘要 | `/api/facts` **200**(cards 6)· 门户结论段 2.6 KB 换成 `RD-LEDGER-*`(6 条: 4 候选 / 2 参考)· 契约自检 PASS ✔ |
 
-两件都已进重算链:④b 里 `model_run → fusion → report → **fusion_handoff** → kb`;⑦ 之后新增 **⑦b 全场状态总览页(自算烘焙)**。
+链上:④b(`model_run → fusion → report → fusion_handoff → kb`)→ **⑤b 由重算台账生成 claim** → **⑤c 构建+渲染契约** → ⑤a 台账维护 → ⑤ 审计 → … → ⑦b 总览页自算烘焙 → ⑧b 重装门户。
+
+**claim 六族**(全部由重算产物算,逐条带 coverage/缺证据/证伪判据,相对判据封顶「候选」):
+报警集中(`alarms`)· 重复检修(`workorders`)· 停机损失集中(`loss_monthly`,kWh→MWh 已换算)·
+温度相对离群(`temp_bins`,样本门 n≥20)· 功率曲线偏离(`powercurve_dev`,MW→kW 已换算 + 判别分布)·
+振动面过闸线(`model_run_l6`)。
 
 ---
 
@@ -30,7 +36,7 @@
 | 2 | `vib_handoff_and_scans` 其余件 | `m5_cms_tcm/component_history.json`、`baseline_38.json`、`oem_frequency_scan.parquet` 等扫描件 | 本体 `trend_ingest`(在升/闭环证据)、`vib_confidence`、`maintenance` 数据层页、报告附录 | 趋势/置信度面少一维;数据层页把这几件列成缺件 | `component_history`:由逐窗 L6+标量按部件聚合历史;`baseline_38`:由 `fleet_scalar_z`/标量算三层基线(自/机群/绝对) | **中**:结构简单,但"口径"要与消费端字段对齐(`trend_ingest.py` 读法已可参照) |
 | 3 | `tcm_replay` | `tcm_compatible_replay/model/mask_thresholds.tsv` 等 | CMS 报告/逐台页的**厂商黄红阈值线**;`data.load_masks` | 页面只画机群中位,阈值线缺失(现文案:"厂商未设阈值") | 由 `data/raw/如东/windcms` 导出件里的 `red_alarm/yellow_alarm/red_hys/…` 索引列反解出每 (测点,测量,工况) 的中位/黄/红标量 | **中偏低**:索引里确有这些列(54 列之一),口径需反推 + 抽台对拍 |
 | 4 | `ontology_releases` | `ontology/release_r1/*`、`release_r2/*` | 网关 `/release/*`(只读发布层)、门户"发布层"入口 | `/release/` 两件 503(如实报缺) | 由 `ontology/objects.json` + 层清单配置生成 r1/r2 层(`kb_ingest`/`populate` 之后跑一步) | **易**:纯搬运/投影,无判据 |
-| 5 | `guanlan_contract` + `sop_workspace` + `paradigm_r1` | `guanlan/facts_contract_v0.json` 及 `derived/*`;`sop/findings.json`;E3/E5/E8 底稿 | 门户「本场结论·契约生成」段、`/detail` 问答/报告页(`/api/facts` 503) | 门户结论段停在旧数;问答/报告面 503 | 新增"**从重算台账生成 claims**"的生成端(alarms/workorders/oil/temp/本体 → 结构化 claim:claim_id/裁决/时间窗/样本/证据引用),替代人裁底稿作为输入 | **难且涉及口径**:这一族原本是**人裁结论**(119 条),改成机器生成等于换口径 → 需你点头(见 §2) |
+| 5 | ~~`guanlan_contract` + `sop_workspace` + `paradigm_r1`~~ | — | 门户结论段 / `/api/facts` / 报告 | **已解决(见 §0)**:由重算台账生成 claim | — | ✅ 完成 |
 | 6 | `windcms_pages` 剩余件 | `windcms/` 下厂家报告转录、kb 之外的其他件 | CMS 首页/工作台的少数入口 | 个别入口空 | 多数已由 `windcms.py report/kb` 覆盖;剩余是"厂家报告转录"(输入是现场 PDF,属外部正本) | **易/不适用** |
 | 7 | `vib_figs` | `m5_cms_tcm/figs/*` | 振动报告里的插图 | 报告少图 | 由谱库渲染 PNG(可复用 `src/windcms` 的 SVG/绘图) | **低优先** |
 | 8 | `windscada_pages` 剩余件 | `windscada/turbines/*`、`review/*` | 静态副本(运行期逐台页由组件动态出,不依赖它) | 无实际影响 | 若要静态副本:复用 `windscada_overview_build.py` 的烘焙(逐台 slim 版) | **易** |
@@ -38,24 +44,48 @@
 
 ---
 
-## 2. 需要你定的两件事(都不动手就永远不会自己好)
+## 2. 变桨面(下一件,口径已定:**停机段中位 + 满发段中位两者分列出**)
+
+消费者字段契约(`src/windscada/subsys/pitch.py::load/registry`,改动需同步):
+
+```
+pitch/pitch_daily.parquet        date, turbine, n, hydlevel_timeon, hydfilt_timeon, hyd_band,
+                                 hyd_over_relief, hyd_under_pump, hublub_timeon, pitchpum_timeon,
+                                 pistA, pistB, pistC, n_op
+pitch/pitch_zero_monthly.parquet month, turbine, zero_dev            ← 消费端读这一列做并入判据
+                                 + zero_dev_stop, zero_dev_full      ← 本次要求的两口径**分列**
+```
+
+通道勘察(`data/raw/如东/scada_10min` 表头实测;`prs_wtc_*` 的压力列还需挑出"蓄能器压力"那一条):
+
+```
+din_wtc_HydLevel_timeon / din_wtc_HydrFilt_timeon      油位/滤网开关 on 秒 (=当日应测 n×600 − timeon → 报警秒)
+din_wtc_LubPistA/B/C_counts|timeon                     三叶柱塞 (pistA/B/C)
+dot_wtc_HubLubPu_timeon                                轮毂润滑泵 (hublub_timeon)
+dot_wtc_PitchPum_timeon                                变桨泵 (pitchpum_timeon)
+prs_wtc_* / int_wtc_ValSuppV_*                         压力族 → 挑蓄能器压力做 hyd_band/越线/跌泵
+```
+
+零位两口径的取数(**1min** 数据,现已可由 `scada_source` 从 MDB 读):
+`pitch_angle_blade_1/2/3`;
+· **停机段中位** = 该月"停机"状态下三叶桨距角中位 − 月内全场/该台基准;
+· **满发段中位** = 满发(高功率档)状态下三叶桨距角中位 − 同一基准。
+两个口径都按月×台输出;`zero_dev` 取**停机段**(零位=静止参考位),另一口径并列可见(页面依据行两个数都写)。
+
+## 3. 需要你定的(其余口径)
 
-1. **事实契约面(第 5 行)**:要不要让观澜**从重算台账生成 claims**(口径从"人裁结论"变成"机器判据 + 人复核")?
-   - 选"要":我按 `sop/findings.json` 的字段结构写生成端,门户结论段/问答/报告三处随之由重算驱动;
-   - 选"不要":门户那一段就按之前提的两条路处理(加"数据截至⋯、契约未随重算更新"横幅,或把该段摘掉)。
-2. **变桨面零位口径(第 1 行)**:`pitch_daily` 的"零位"用**停机段中位**还是**满发段中位**(或两者都出、分列)?
-   定了我就把生成端写进链(口径写进产物与台账)。
+(第 5 行的事实契约口径已按你的选择落地;第 1 行的变桨零位口径也已定为"两口径分列"。)
+目前**没有**待你裁的口径项了;剩余各族按 §4 顺序推进即可。
 
-## 3. 执行顺序(我建议)
+## 4. 执行顺序(修订)
 
 ```
-① ontology_releases 生成端(易,先把 /release 两件点亮)
-② component_history / baseline_38 生成端(中,补趋势与置信度面)
-③ tcm_replay 掩码阈值反解(中偏低,页面阈值线回来)
-④ pitch 面生成端(中,需你先定零位口径)
-⑤ 事实契约生成端(难,需你先定口径)
-⑥ vib_figs / 静态逐台副本(低优先)
+① ontology_releases 生成端(易,点亮 /release 两件)
+② component_history / baseline_38 生成端(中)
+③ tcm_replay 掩码阈值反解(中偏低)
+④ pitch 面生成端(中,口径已定;消费者字段契约见 §2)
+⑤ vib_figs / 静态逐台副本(低优先)
 ```
 
 每补一族都走同一套纪律:**生成端进包 + 进重算链 + 写进族表(gen/input) + 产物自登记 + 复检页面**,
-并在 `docs/系统设计说明.md` §13.6 与族表里把该族从"无生成端"改成"已在重算链上(附口径与验证依据)"。
+并把该族在族表与 `docs/系统设计说明.md` §13.6 里从"无生成端"改成"已在重算链上(附口径与验证依据)"。

+ 810 - 793
scripts/products_reverse_audit.py

@@ -1,793 +1,810 @@
-#!/usr/bin/env python3
-# -*- coding: utf-8 -*-
-r"""**反向**呼应审计:自输出 → 功能与算法 → 输入(用户令 2026-09-17)。
-
-## 与 `inventory_products.py` 的区别(两个方向,各管一件事)
-
-· `inventory_products.py` 是**正向**:从 `data/raw/<场>/` 的每一类输入出发,找它喂出来的产物,
-  比跨度/条数("输入到了 2026-07,产物只到 2026-04 → 未重算")。它只覆盖 5 类输入、十几件产物。
-· 本器是**反向**:从 `outputs/<场>/` 的**每一件产物**出发,问三个问题:
-    ① 它是谁算出来的?(功能与算法 = 生成端)
-    ② 它从哪份输入算出来的?(`data/raw/<场>/` 的哪一类)
-    ③ 这条"输出 ↔ 输入"的呼应关系**成立吗**?(生成端在位 ∧ 输入在位 ∧ 可机检的对应判据通过)
-  逐件判定,最后给出"成立 / 不成立(无生成端)/ 无法验证"三类账。
-
-## 为什么必须做反向
-
-正向只查"我关心的输入有没有被算成产物",查不出**"盘上这件产物到底有没有来路"**。
-2026-09-17 现场就是栽在这里:清了产物之后,页面缺的 `windscada/index.html`、
-`m5_cms_tcm/handoff_vibration_v2.json` 这类件**根本没有生成端**(全库 0 处写入方),
-正向那几条跨度判据一条都不会报 —— 它们压根不在 `INPUTS` 的 `feeds` 里。
-反向一对账就清楚了:**有来路的件**(生成端 + 输入 + 判据都过)与**没来路的件**(只能从交付包补)。
-
-## 判定口径(每条都写进结果里,不含糊)
-
-    ✓ 呼应成立          生成端在位 ∧ 输入在位 ∧ 判据通过(跨度 ⊆ 输入 / 键集 ⊆ 输入 / 计数一致 / 行数 > 0)
-    ~ 输入不在位         生成端知道,但 `data/raw/<场>/<类>` 不在 ⇒ **现在无法验证**(放数据后复跑本器)
-    ✗ 无生成端           全库 0 处写入方 ⇒ **输出↔输入的呼应在原理上不成立**:这件产物无法由输入推导出来,
-                        **无法由重算生出来**;按用户令 2026-09-17 #1 运行期也不从交付包补齐 ⇒ 只能由研发补生成端(本器 --feasibility 给出逐族可逆性)
-    ? 未归类             既不在族表里、台账里也没有 —— 需要人工认领(本器把它当缺口报出来)
-
-## 用法
-
-    python scripts/products_reverse_audit.py            # 全量反向审计(人看)
-    python scripts/products_reverse_audit.py --check    # 只出结论(有 ✗/? 时退出码 5)
-    python scripts/products_reverse_audit.py --write-doc # 把族表写进 docs/系统设计说明.md §13.6
-
-退出码: 0 全部成立(允许 ✗,只要它们都在台账里如实标了 shipped)· 5 有未归类件或判据失败
-"""
-from __future__ import annotations
-
-import argparse
-import fnmatch
-import json
-import pathlib
-import sys
-
-ROOT = pathlib.Path(__file__).resolve().parents[1]
-sys.path.insert(0, str(ROOT))
-sys.path.insert(0, str(ROOT / 'scripts'))
-from src import paths as P                                              # noqa: E402
-import inventory_products as INV                                        # noqa: E402
-
-DOC_BEGIN, DOC_END = '<!-- REVERSE-AUDIT:BEGIN -->', '<!-- REVERSE-AUDIT:END -->'
-
-# ── 族表: 输出(glob) → 功能 / 算法(生成端) / 输入类 / 判据 ────────────────────────────
-#   pred: span=产物时间跨度必须落在输入跨度内; turbines=产物机组集 ⊆ 输入机组集;
-#         rows=行数>0; count=件数/计数与说明一致; manifest=与清单自记一致
-FAMILIES: list[dict] = [
-    dict(id='vib_window_index', glob='m5_cms_tcm/windows/*/index.parquet', kind='raw-derived',
-         func='振动摄入 · 窗索引', algo='scripts/rudong_tcm_index.py(逐 decode JSON 解析 54 列:turbine/'
-              'sensor_name/meas_name/trigger_time/rpm/condition_key/alarm_type)',
-         gen='scripts/rudong_tcm_index.py', input='windcms', pred=('span', 'turbines', 'rows'),
-         time_col='trigger_time'),
-    dict(id='vib_spectra', glob=['m5_cms_tcm/windows/*/spectra/**', 'm5_cms_tcm/windows/*/spectra_meta.parquet'],
-         kind='raw-derived',
-         func='振动摄入 · 谱库', algo='scripts/rudong_tcm_spectra.py(FFT_ 测量 → npz 幅值数组, 分片存 `spectra/p<NN>/`)',
-         gen='scripts/rudong_tcm_spectra.py', input='windcms', pred=('npz',)),
-    dict(id='vib_raw_manifest', glob='m5_cms_tcm/vib_raw_manifest.json', kind='raw-derived',
-         func='振动摄入 · 清单', algo='scripts/vib_raw_build.py(汇总窗/谱件数、时间跨度、missing_chain 缺口)',
-         gen='scripts/vib_raw_build.py', input='windcms', pred=('manifest',)),
-    dict(id='scada_alarms', glob='windscada/alarms.parquet', kind='raw-derived',
-         func='三门台账 · 报警', algo='scripts/windscada_alarms_ingest.py(SpreadsheetML *.xls → 事件表)',
-         gen='scripts/windscada_alarms_ingest.py', input='故障报警', pred=('span', 'rows'), time_col='t_on'),
-    dict(id='scada_workorders', glob='windscada/workorders.parquet', kind='raw-derived',
-         func='三门台账 · 工单', algo='scripts/windscada_workorder_ingest.py(检修台账 → 工单表)',
-         gen='scripts/windscada_workorder_ingest.py', input='风机故障记录', pred=('span', 'rows'),
-         time_col='t_report'),
-    dict(id='scada_oil', glob='windscada/oil_samples_index.parquet', kind='raw-derived',
-         func='三门台账 · 油样', algo='scripts/windscada_watch_channels_build.py(油样 PDF → 索引)',
-         gen='scripts/windscada_watch_channels_build.py', input='油样报告', pred=('span', 'rows'),
-         time_col='date'),
-    dict(id='scada_monthly', glob='windscada/temp_monthly.parquet', kind='raw-derived',
-         func='月度派生件', algo='scripts/windscada_monthly_build.py(逐值对齐随包件 19,494/19,494)',
-         gen='scripts/windscada_monthly_build.py', input='scada_10min', pred=('span', 'rows'), time_col='month'),
-    dict(id='scada_derived', glob='windscada/*.parquet', kind='raw-derived',
-         func='SCADA 派生分析(功率曲线/损失/曲线/控制/停机/温度/偏航/液压/热链/系统辅助)',
-         algo='src/windscada/perf/{powercurve,availability,curves,control,faults} + subsys/{temp_nbm,yaw,'
-              'hydraulic,thermal_chain} + taxonomy.py',
-         gen='src/windscada/perf/powercurve.py', input='scada_10min', pred=('rows',)),
-    dict(id='ontology_core', glob='ontology/{objects.json,retrieval_index.json,turbine_params.parquet}',
-         kind='raw-derived',
-         func='本体层 · 码表/对象库/检索索引/实机参数',
-         algo='python -m src.ontology.kb_ingest → populate → chain_ingest → trend_ingest;'
-              'src.ontology.retrieval.build;src.ontology.maintenance.refresh_params',
-         gen='src/ontology/kb_ingest.py', input='西门子4.0技术资料', pred=('rows',)),
-
-    dict(id='ontology_aux_build', glob=['ontology/_id_alias.json', 'ontology/retrieval_vec.npy'],
-         kind='raw-derived',
-         func='本体层 · 别名表与向量索引缓存', algo='src.ontology.kb_ingest(别名)/ '
-              'src.ontology.retrieval.build(use_vec=True)(向量)',
-         gen='src/ontology/kb_ingest.py', input='西门子4.0技术资料', pred=('rows',)),
-    dict(id='ontology_domain', glob=['ontology/oil_2026H2_huabiao.json', 'ontology/parts_cost_public.json',
-                                     'ontology/scenario_29_repair.json'],
-         kind='shipped',
-         func='本体层 · 领域数据件(油样台账/备件价格/场景修复)', algo='(无生成端, 领域正本出件)',
-         gen=None, input=None, pred=(), why='领域正本随包发来, 不由 data/raw 推导'),
-
-    # ── 以下族: 全库 0 处写入方 ⇒ 反向呼应**在原理上不成立**(如实记账, 不假装成立)────
-    #    ★2026-09-18 用户令补充: 这些件**运行期一律不从交付包补**(用户令 2026-09-17 #1 的延伸),
-    #      页面如实为空/降级; 要出件只有两条路 —— ① 研发补生成端; ② 现场正本随原始件进 data/raw。
-    #      故本表的 why 里不再写"从交付包补齐"那种话(那是**离线人工补救**, 见 products_restore_missing.py)。
-    dict(id='vib_model_run_l6', glob=['m5_cms_tcm/model_run_l6.parquet', 'm5_cms_tcm/model_run_summary.json'],
-         kind='raw-derived',
-         func='六层链 model_run 步 · L6 过闸谱线表 + 层小结',
-         algo='scripts/rudong_model_run.py(候选线=oem_scan_plan 的 11 部件×{BPFI,BPFO,BSF}; 过闸='
-              'src/sop/discriminators.py::spectral_line_gates G1–G8; 定级=vib_verdict_and_writeback)',
-         gen='scripts/rudong_model_run.py', input='m5_cms_tcm', pred=('rows',),
-         why='★2026-09-19 按口径重建(**无标准答案**: 这两件从未随包, 无法逐值对拍); 峰值拾取仍是本器自定'
-             '(目标频率 ±2 bin 取最大), 不得当"复现"用 —— 见 docs/振动六层链_接口规格与缺口_v0.1.md §3'),
-    dict(id='vib_fusion_38', glob='m5_cms_tcm/fusion_38.csv', kind='raw-derived',
-         func='六层链 fusion 步 · 逐台融合表(报告"融合级"列的来源)',
-         algo='scripts/rudong_fusion_run.py::fusion_table(模型侧=model_run_l6 过闸线; CMS 侧=窗索引 '
-              'RedMask/YellowMask; 裁决=src/sop/fusion_diag.py::fuse)',
-         gen='scripts/rudong_fusion_run.py', input='m5_cms_tcm', pred=('rows',),
-         why='★按口径重建(无标准答案): 台号用 WTGnn 形态(消费端 registry 的键), 与 l6 的 f"{n}#" 不同'),
-    dict(id='vib_fleet_scalar_z', glob='m5_cms_tcm/fleet_scalar_z.parquet', kind='raw-derived',
-         func='振动线 · 全场标量 z 值 (fusion 步出件)',
-         algo='scripts/rudong_fusion_run.py(逆向工程实现: 窗索引 scalar_value → 同 (传感器,测量,工况) 族中位 '
-              '→ 稳健 z=(val−med)/(1.4826·MAD); 与随包样件 8,887 行逐值对拍 val 列 100% 一致)',
-         # input 写**原始件目录名**(族表口径: data/raw/<场>/<input>); 真实取数经窗索引一跳,
-         # 而窗索引本身是 raw-derived(vib_window_index 族) —— 写 windows/_*.parquet 会让审计判"输入不在位"。
-         gen='scripts/rudong_fusion_run.py', input='m5_cms_tcm', pred=('rows',),
-         why='2026-09-18 上链: 由 vib_raw_build.py 的 --with-report 分支按当次窗名调用, '
-             '使 /api/fleet 的振动标量 z 面不再空白'),
-    dict(id='vib_fusion_handoff', glob='m5_cms_tcm/handoff_vibration_v2.json', kind='raw-derived',
-         func='融合面 handoff (观澜自算, 同结构于振动线手交件)',
-         algo='scripts/rudong_fusion_handoff.py(报告状态等级 + fusion_38 融合裁决 + L6 过闸线 → 同结构, '
-              'meta.from 标明"观澜自算")',
-         gen='scripts/rudong_fusion_handoff.py', input='m5_cms_tcm', pred=('rows',),
-         why='★2026-09-19 用户令「所有的计算均要形成观澜的源代码」: 这件此前全库 0 处写入方(振动线人写), '
-             '新机器装完 `/detail/v2` 的「需要关注/全场状态」必然空白。现由观澜自算出同结构件; '
-             '盘上若有**振动线正本**, 生成端不覆盖(正本含人工裁决, 优先)'),
-    dict(id='vib_handoff_and_scans', glob='m5_cms_tcm/*.{json,parquet,csv,md,txt}',
-         kind='shipped',
-         func='振动线出件与专项扫描(handoff / 历史 / 基线 / 各类 freq scan / 判级台账)', algo='(无生成端)',
-         gen=None, input=None, pred=(), why='振动六层链四步脚本未随包(oem_frequency_scan 等)⇒ '
-                                            '运行期不从交付包补(用户令 2026-09-17); 要出件须研发补生成端, '
-                                            '或由现场正本随原始件进 data/raw 后重算'),
-    dict(id='vib_figs', glob='m5_cms_tcm/figs/*', kind='shipped',
-         func='振动线出图(谱图/趋势图)', algo='(无生成端)', gen=None, input=None, pred=(),
-         why='随包快照(振动线出件)'),
-    # ★2026-09-19 拆族 (用户令「所有的计算均要形成观澜的源代码」): `windscada/index.html` 是
-    #   `/detail/v2` 的「全场状态总览」iframe 目标, 此前整族记为"无生成端" ⇒ 新机器上那块必然显示
-    #   "产物不在位"。现在它由 `scripts/windscada_overview_build.py` 用**组件同一个快照烘焙器**
-    #   (src/windscada/ui/snapshot.py::bake, /v2/snapshot 用的就是它)从重算产物生成 ⇒ 单列一族。
-    dict(id='windscada_overview', glob='windscada/index.html', kind='raw-derived',
-         func='总览「全场状态」页 (单文件: 接口响应烤进页面, 离线可开)',
-         algo='scripts/windscada_overview_build.py(ctx 取数函数与在线 /api/* 同一实现; 经 '
-              'src/windscada/ui/snapshot.py::bake 烘焙)',
-         gen='scripts/windscada_overview_build.py', input='scada_10min', pred=('rows',),
-         why='2026-09-19 用户令后上链: 该页此前是随包件(全库 0 处写入方), 清过产物的机器上 iframe 只能'
-             '显示"产物不在位"; 现在由观澜自算, 每次重算随数据更新'),
-    dict(id='windscada_pages', glob='windscada/{turbines/*,review/*}', kind='shipped',
-         func='逐台静态页 + 评审记录(当前由组件按请求动态出页, 静态件仍属随包快照)', algo='(无生成端)',
-         gen=None, input=None, pred=(), why='逐台页在运行期由 /turbine/<台> 动态生成, 静态副本无写入方; '
-                                             '运行期不从交付包补(用户令 2026-09-17)'),
-    # ★2026-09-18 补: `windcms/cache/*` 是**读侧缓存**(标量按内容哈希落盘), 它不是随包快照 ——
-    #   原先 `windcms_pages`(shipped) 先认领了它, 于是重算后新写的缓存被记成"包内无生成端"。
-    #   写方在 src/windcms/data.py::load_scalars (窗索引 → 标量), 故单列一族并放在 windcms_pages 之前。
-    dict(id='windcms_cache', glob='windcms/cache/*', kind='raw-derived',
-         func='CMS 报告侧 · 标量缓存', algo='src/windcms/data.py::load_scalars(窗索引 → 标量数组, '
-              '按内容哈希命名落 cache/)', gen='src/windcms/data.py', input='windcms', pred=('rows',)),
-    # ★2026-09-18 补 (用户令"页面只能基于输入数据重算"之后): `windcms/**` 整族原先一律按
-    #   "无生成端"记账 —— 但**报告两步的生成端就在包内**(scripts/windcms.py report → src/windcms/
-    #   report.py + report_std.py), 只要 data/raw/<场>/windcms 里有 decode 导出就出得来。
-    #   实测(2026-09-18): 从重算出的窗 w0316 一键跑出标准模版报告 + 38 台逐台页 + 总览, 判级 38 台
-    #   (报警 3 / 优秀 35), 缺的只有 model_run/fusion 两步的融合级列(如实写 —)。故拆成两族按 raw-derived 记。
-    dict(id='windcms_report_std', glob=['windcms/报告_CMS振动状态评估报告_*.md',
-                                        'windcms/报告_CMS振动状态评估报告_*.docx',
-                                        'windcms/登记_三轴_38台.csv'], kind='raw-derived',
-         func='CMS 标准模版评估报告 (三轴: 状态等级 × 证据状态 × 行动等级建议)',
-         algo='src/windcms/report_std.py::build(结构化判级 → 报告 md + docx + 登记 csv; '
-              '由 scripts/windcms.py report 调用)',
-         gen='src/windcms/report_std.py', input='windcms', pred=('rows',),
-         why='★同名件有两种来源: 观澜自算(本族: 从 data/raw/<场>/windcms 的 *_decode.json 算) 与 '
-             '厂家报告转录(scripts/vib_reports_build.py, 输入是现场振动分析报告 PDF)。按名字认领不了'
-             '这层区别 ⇒ 以"是否由本步生成"为准, 判据看 m5_cms_tcm/vib_raw_manifest.json 的 steps'),
-    dict(id='windcms_report_overview', glob=['windcms/report.md', 'windcms/index.html',
-                                             'windcms/index_eng.html', 'windcms/overview.html',
-                                             'windcms/turbines/*'], kind='raw-derived',
-         func='CMS 诊断总览页 + 逐台页 + 结构化层报告',
-         algo='src/windcms/report.py::build(逐台 turbines/<台>.html + 总览 overview/index.html + '
-              'report.md; 融合级列要 model_run/fusion 出件, 缺则如实写"—")',
-         gen='src/windcms/report.py', input='windcms', pred=('rows',)),
-    dict(id='windcms_kb', glob='windcms/kb.json', kind='raw-derived',
-         func='CMS 知识库索引 (问答检索用)',
-         algo='src/windcms/knowledge.py::build(由 scripts/windcms.py kb 调用; 离线无模型时 0 chunks 也算跑成)',
-         gen='src/windcms/knowledge.py', input='windcms', pred=('rows',)),
-    dict(id='windcms_pages', glob='windcms/**', kind='shipped',
-         func='windcms 里其余件 (厂家报告转录 / 知识库出件等)',
-         algo='(无随包生成端)', gen=None, input=None, pred=(),
-         why='上述两族与 cache 之外剩下的 windcms 件: 运行期不从交付包补(用户令 2026-09-17), '
-             '要出件须研发补生成端'),
-    dict(id='sop_workspace', glob='sop/**', kind='shipped',
-         func='SOP 中间件/评审/事实契约底稿', algo='(无生成端)', gen=None, input=None, pred=(),
-         why='评审过程件, 属人工作业留痕'),
-    dict(id='tcm_replay', glob='tcm_compatible_replay/**', kind='shipped',
-         func='TCM 兼容链回放资产(模型表/掩码阈值/裁决记录)', algo='(无生成端)', gen=None, input=None,
-         pred=(), why='随包快照'),
-    dict(id='paradigm_r1', glob='paradigm_r1/**', kind='shipped',
-         func='范式实验件(E3/E5/E8 底稿)', algo='(无生成端)', gen=None, input=None, pred=(),
-         why='实验底稿, 事实契约的输入'),
-    dict(id='pitch_shipped', glob='pitch/**', kind='shipped',
-         func='变桨侧派生件(零位/日粒度)', algo='(无生成端;rebuild_from_raw --scada 只覆盖其中一部分)',
-         gen=None, input=None, pred=(), why='无生成端 ⇒ 运行期不从交付包补(用户令 2026-09-17), '
-                                            '页面降级; 要出件须研发补生成端'),
-    dict(id='ontology_releases', glob='ontology/release_*/*', kind='shipped',
-         func='本体发布层 r1/r2(只读暴露给网关 /release/)', algo='(无生成端)', gen=None, input=None,
-         pred=(), why='发布层清单, 由研发出件'),
-    dict(id='guanlan_contract', glob='guanlan/**', kind='shipped',
-         func='事实契约与对外派生(门户结论段/取数)', algo='scripts/guanlan_facts_contract.py(读者在包内, '
-              '输入是 sop/findings.json + paradigm_r1 三份**人裁底稿**)', gen=None, input=None, pred=(),
-         why='派生链的中间层: 它的"输入"本身是产物而非 data/raw ⇒ 与 raw 之间隔着不止一跳; '
-             '运行期不从交付包补(用户令 2026-09-17) ⇒ `/api/facts` 如实回 503 缺件'),
-    dict(id='windscada_release_manifest', glob='windscada/release-manifest.json', kind='source-derived',
-         func='发布清单快照(`/healthz` 与版本信息读它)',
-         algo='scripts/guanlan_baseline_manifest.py(从 src/windscada/ui 源重建 zh 页 → 摘要 '
-              '{css_tokens, nav, dom, sha256, source blobs})',
-         gen='scripts/guanlan_baseline_manifest.py', input=None, pred=(),
-         why='★ 2026-09-17 修正: 它一度被我按"交证件"错分成非产物; 实际**包内有生成端**'
-             '(guanlan_baseline_manifest.py 就写这个路径), 且 guanlan_gateway.py 运行期读它 ⇒ 是产物'),
-    dict(id='human_deliverables',
-         glob=['windscada/mblub_alarm_cross.json', 'windscada/review_presented.json',
-               'windscada/给振动线_温度轴回复_20260826.json', 'windscada/给振动线_温度轴回复_20260826.md',
-               'windscada/送审_系统结论与判定_20260828.md'],
-         kind='human',
-         func='**人工正本/交证件**:与振动线的跨线通知、送审结论、发布评审快照',
-         algo='(人写的, 不是算出来的)', gen=None, input=None, pred=(),
-         why='这几件是"人对人的交证件/回执", 内容由分析结论抄写与人判断坐实而成; '
-             'mblub_alarm_cross.json 还被 src/ontology/populate.py 当作**证据来源**引用 ⇒ '
-             '必须随包留在原位, 但**不参与呼应校验**(它本来就没有生成端, 也不该有)'),
-    dict(id='scratch_in_products', glob=['**/_*.js', '**/_*audit*.md', '**/*.out', '**/*.err',
-                                        '**/*_test_*.parquet', '**/*_test_*.json'],
-         kind='not-product',
-         func='**非产物**:自审脚本/记录与模型测试输出(躺在产物仓里)', algo='(工具脚本与运行痕迹, 不是产物)',
-         gen=None, input=None, pred=(),
-         why='2026-09-17: 17 件已移出到 reference/非产物留痕_20260917/ 并留了 _来源说明.json ⇒ '
-             '本族现在应为空; 一旦再出现(新脚本往产物仓里写 .out/_audit.js 之类), 本器当场报出来'),
-]
-
-
-def _files_of(glob) -> list[str]:
-    """族 glob(str 或 list)→ 台账里的相对键(支持 `{a,b}` 花括号与 `*`)。
-
-    ★ 深度必须**显式对齐**(2026-09-17 第一版踩过): `fnmatch` 的 `*` 是**跨 `/`** 的, 于是
-      `windscada/*.parquet` 会把 `windscada/turbines/xx.parquet`、`windscada/_pre_rebuild_*/xx.parquet`
-      一起吞进来 —— 反向审计最怕"族把不该管的件认领了", 那样账就假了。
-      规则: 模式里没有 `**` 时, 要求的 `/` 个数必须与键的 `/` 个数**相等**(逐条展开后各自判)。
-    """
-    pats = []
-    for g in ([glob] if isinstance(glob, str) else list(glob)):
-        if '{' in g:
-            head, tail = g.split('{', 1)
-            opts, rest = tail.split('}', 1)
-            pats += [head + o + rest for o in opts.split(',')]
-        else:
-            pats.append(g)
-    out = []
-    for rel in LEDGER:
-        for pt in pats:
-            if '**' not in pt and rel.count('/') != pt.count('/'):
-                continue
-            if fnmatch.fnmatch(rel, pt) or rel == pt:
-                out.append(rel)
-                break
-    return sorted(out)
-
-
-LEDGER: dict[str, dict] = {}
-
-
-def load_ledger(farm: str) -> dict[str, dict]:
-    """台账 = `_provenance.json`(逐件来源) + `_derived_manifest.json`(生成端自登记)。"""
-    out: dict[str, dict] = {}
-    root = P.out_root(farm)
-    for fn in ('_provenance.json', '_derived_manifest.json'):
-        f = root / fn
-        if not f.is_file():
-            continue
-        try:
-            d = json.loads(f.read_text(encoding='utf-8'))
-        except Exception:
-            continue
-        for rel, v in (d.get('files') or {}).items():
-            if isinstance(v, dict):
-                out.setdefault(rel, dict(source=v.get('source') or 'raw-derived',
-                                         builder=v.get('builder') or v.get('by') or '',
-                                         why=v.get('why') or '',))
-    return out
-
-
-def _input_home(station: pathlib.Path, sub: str) -> pathlib.Path | None:
-    """输入类目录在哪 —— 优先场站目录下, 其次 raw 根下。
-
-    ★ 机理层资料按 A2 约定**不在场站目录下**(`data/raw/西门子4.0技术资料/`, 见 src/windscada/config.py)。
-      只查场站目录会把本体层判成"输入不在位"(2026-09-17 第一版就这么误报过 3 件)。
-    """
-    for base in (station, station.parent):
-        d = base / sub
-        if d.is_dir():
-            return base
-    return None
-
-
-def _parquet_span(p: pathlib.Path, col: str | None):
-    import pandas as pd
-    try:
-        d = pd.read_parquet(p)
-    except Exception:
-        return None, None, 0
-    n = len(d)
-    if col and col in d.columns:
-        s = pd.to_datetime(d[col], errors='coerce')
-        if s.notna().any():
-            return str(s.min())[:10], str(s.max())[:10], n
-    return None, None, n
-
-
-def _npz_ok(p: pathlib.Path) -> bool:
-    """振动谱件: npz 能读出来且非空(列数不是"行", 不能拿 parquet 的行数判它)。"""
-    try:
-        import numpy as np
-        with np.load(p, allow_pickle=False) as z:
-            return any(z[k].size > 0 for k in z.files)
-    except Exception:
-        return False
-
-
-# 已知"含人工裁决/校准"的随包快照(按用户令 2026-09-17 #3 显式登记为人工件)。依据是文档里已经写明的性质:
-# docs §7「handoff_vibration_v2.json / component_history.json 是随包快照, 该件含人工裁决/校准更新,
-# 不是测量数据的函数」; baseline_38.json 的三层基线 self/absolute 两层由人工坐实; findings.json 是判级发现台账。
-HUMAN_SNAPSHOTS = {
-    'm5_cms_tcm/component_history.json',
-    'm5_cms_tcm/baseline_38.json',
-    'm5_cms_tcm/handoff_vibration_v1.json',
-    'm5_cms_tcm/findings.json',
-}
-
-# ── 逆向工程可行性 (用户令 2026-09-17 #2) ────────────────────────────────────────────
-# 问的是: "这批没有生成端的产物, 能不能从原始件/在包生成端**推**出来?"
-# 判定不靠印象, 靠两条机器证据:
-#   ① 生成端在不在包内(在 → 只是没接上, 跑一次就有);
-#   ② 产物正文里有没有**人工判断**的痕迹("裁决/审核/校准/评审/经验"这类字段)——
-#      含人工判断的件不是任何输入的纯函数, 逆向工程推不出来, 只能把那一步人工工作重做。
-import re as _re                                                         # noqa: E402
-JUDGE_PAT = _re.compile(r'人工|裁决|审核|复核|校准|评审|经验值|专家|judg|review|manual|calibrat|verdict|sign[_ ]?off',
-                        _re.I)
-# ★ 只认"**字段**叫这个名字"(`"裁决": …` / `裁决,`),不认正文里顺嘴提一句 ——
-#   2026-09-17 第一版拿整篇子串匹配, 结果把**页面**里的展示标签(`index.html` 里的"评审/人工")当成
-#   人工判断, 于是把"页面"整族误判成不可逆。这属于"判据太糙 → 结论假"。
-JUDGE_KEY_PAT = _re.compile(r'["\']?[^"\',:]{0,24}(人工|裁决|审核|复核|校准|评审|经验值|专家|'
-                            r'judg|review|manual|calibrat|verdict|sign[_ ]?off)[^"\',:]{0,24}["\']?\s*[:=]',
-                            _re.I)
-
-
-def judgement_evidence(farm: str, rels: list[str], cap: int = 8) -> tuple[int, list[str], int]:
-    """→ (含人工判断字段的抽样比例%, 样例, 被检查件数)。
-
-    只查**数据/文本件**(.json/.csv/.md/.txt); **页面(.html/.js)不算** —— 页面是产物的渲染,
-    里面出现"评审/人工"是展示标签, 不代表这份件内含判断。
-    """
-    hit, samples, checked, pages = 0, [], 0, 0
-    for rel in rels[:cap]:
-        p = P.out_root(farm) / rel
-        suf = p.suffix.lower()
-        if suf in ('.html', '.js', '.css', '.htm'):
-            pages += 1
-            continue
-        try:
-            if suf == '.json':
-                txt = p.read_text(encoding='utf-8', errors='replace')[:200000]
-                pat = JUDGE_KEY_PAT
-            elif suf in ('.csv', '.txt'):
-                txt = p.read_text(encoding='utf-8', errors='replace')[:50000]
-                pat = JUDGE_KEY_PAT
-            elif suf == '.md':
-                txt = p.read_text(encoding='utf-8', errors='replace')[:50000]
-                pat = JUDGE_PAT
-            else:
-                continue
-        except Exception:
-            continue
-        checked += 1
-        if pat.search(txt):
-            hit += 1
-            samples.append(rel)
-    return (100 * hit // checked if checked else 0), samples, pages
-
-
-def feasibility(fam: dict, rels: list[str], gen_ok: bool, in_ok: bool,
-                judge_pct: int, judge_samples: list[str], pages: int = 0) -> tuple[str, str]:
-    """→ (可逆性判定, 依据)。四类: 可逆(直接) / 可逆(需反推口径) / 不可逆(人工判断) / 应移出产物仓。"""
-    if fam['kind'] == 'not-product':
-        return '应移出产物仓', '它本就不是产物(工具脚本/交证件/测试输出) ⇒ 该从产物仓移走, 而不是"补生成端"'
-    if gen_ok and in_ok:
-        return '可逆(直接)', f"生成端在包内({fam['gen']})且输入在位 ⇒ 接上/重跑即可"
-    if pages and pages >= max(3, len(rels[:8]) // 2):
-        return '可逆(页面可再生)', ('这一族主要是**页面**(.html): 页面是产物的渲染, 不是独立数据 ⇒ '
-                                    '接上在包的页面生成端重出即可(未必逐字节同, 能力与数据一致)')
-    m_n = sum(1 for r in rels if r.endswith(('.parquet', '.csv', '.npz')))
-    t_n = len(rels) - m_n
-    split = f'({m_n} 件测量形态 / {t_n} 件文本·判断件)' if (m_n and t_n) else ''
-    # 两类都占相当比重时, 结论必须是"混合"而不是挑一类代表全族 —— 否则一句判定就把另一半骗过去了
-    if m_n and t_n and min(m_n, t_n) * 4 >= len(rels):
-        return '混合(测量件可反推 / 文本件需人定)', split + \
-               '测量形态的件可由 data/raw 反推口径 + 逐值对拍; 文本/判断件(评审、裁决、回复)推不出来'
-    if judge_pct >= 50 and judge_samples:
-        return '不可逆(人工判断)', (f"抽样 {len(judge_samples)} 件里都写着人工判断字段(如 {judge_samples[0]}) "
-                                    f"⇒ 不是输入的纯函数, 逆向工程推不出来")
-    if m_n and m_n * 2 >= len(rels):
-        return '可逆(需反推口径)', split + '测量形态的件是测量数据的函数 ⇒ 可从 data/raw 反推口径 + 逐值对拍' \
-                                        '(做法同 temp_monthly: 反推 → 逐值一致才敢用)'
-    if gen_ok:
-        return '半可逆(生成端在包, 输入不足)', '生成端在包内, 缺的是上游输入 ⇒ 上游补齐后即可重出'
-    if judge_pct > 0:
-        return '半可逆(需人工裁定)', f'抽样里 {judge_pct}% 的件含人工判断字段 ⇒ 机器部分可推、判断部分要人定'
-    return '需人工裁定', '既无生成端、形态也不是测量函数 ⇒ 需研发给口径或领域正本'
-
-
-def _turbines_of(p: pathlib.Path, col: str = 'turbine', cap: int = 200000):
-    import pandas as pd
-    try:
-        d = pd.read_parquet(p, columns=[col]) if col else pd.read_parquet(p)
-    except Exception:
-        return None
-    try:
-        return {str(x).upper() for x in d[col].dropna().unique()} if col in d.columns else None
-    except Exception:
-        return None
-
-
-def _input_turbines(station: pathlib.Path, sub: str):
-    """输入侧机组集合: 逐台 CSV 文件名 / 目录名里的 WTG 号。"""
-    d = station / sub
-    if not d.is_dir():
-        return None
-    out = set()
-    for p in d.rglob('*'):
-        if p.is_file():
-            for m in __import__('re').finditer(r'(WTG\s?\d{1,2})', p.name.upper()):
-                out.add(m.group(1).replace(' ', ''))
-    return out or None
-
-
-def classify_human(farm: str, rels: list[str], cap: int = 600) -> set[str]:
-    """逐件判"是不是人工件"(用户令 2026-09-17 #3:人工件不再按自动产物参与呼应校验)。
-
-    判据:**非测量形态**的件(.md/.json/.csv/.txt)里出现人工判断字段(按字段名匹配)。
-    测量形态件(.parquet/.csv 数值、.npz)一律不算人工件 —— 它们是数据的函数,属"可反推"那一路。
-    """
-    human: set[str] = set()
-    for rel in rels[:cap]:
-        if rel in HUMAN_SNAPSHOTS:            # ① 文档里已写明含人工裁决/校准的随包快照
-            human.add(rel)
-            continue
-        suf = pathlib.Path(rel).suffix.lower()
-        if suf == '.md':                      # ② 产物仓里的 .md 一律是报告/交证/评审记录(人写的)
-            human.add(rel)
-            continue
-        if suf == '.txt':
-            pat = JUDGE_PAT
-        elif suf in ('.json', '.csv'):
-            pat = JUDGE_KEY_PAT
-        else:
-            continue
-        p = P.out_root(farm) / rel
-        try:
-            txt = p.read_text(encoding='utf-8', errors='replace')[:120000]
-        except Exception:
-            continue
-        if pat.search(txt):                   # ③ 正文里有"人工判断字段"
-            human.add(rel)
-    return human
-
-
-def write_human_manifest(farm: str | None = None) -> int:
-    """把人工件逐件落成 `outputs/<场>/_human_artifacts.json`(机器台账,供自检与审计引用)。"""
-    farm = farm or P.farm()
-    global LEDGER
-    LEDGER = load_ledger(farm)
-    entries: dict[str, dict] = {}
-    for fam in FAMILIES:
-        if fam['kind'] not in ('shipped', 'human', 'not-product'):
-            continue
-        rels = [r for r in _files_of(fam['glob']) if r not in entries]
-        if not rels:
-            continue
-        if fam['kind'] == 'human':
-            for r in rels:
-                entries[r] = dict(kind='human', family=fam['id'], evidence='族内显式登记为人工正本/交证件')
-            continue
-        for r in classify_human(farm, rels):
-            if r not in entries:
-                entries[r] = dict(kind='human', family=fam['id'],
-                                  evidence='件内含人工判断字段(按字段名匹配),不是输入的纯函数')
-    out = P.out_root(farm) / '_human_artifacts.json'
-    out.write_text(json.dumps(dict(
-        at=__import__('time').strftime('%Y-%m-%d %H:%M:%S'),
-        note='人工件台账(用户令 2026-09-17 #3):这些件由人写成/坐实,不是任何输入的纯函数 ⇒ '
-             '不参与"输出↔输入呼应"校验;随包留在原位,重算不会也不该重写它们。'
-             '本文件由 scripts/products_reverse_audit.py --write-human-manifest 生成。',
-        count=len(entries), files=entries), ensure_ascii=False, indent=1), encoding='utf-8')
-    by_fam: dict[str, int] = {}
-    for v in entries.values():
-        by_fam[v['family']] = by_fam.get(v['family'], 0) + 1
-    print(f'人工件台账: {len(entries)} 件 → {P.rel(out)}')
-    for k, v in sorted(by_fam.items(), key=lambda kv: -kv[1]):
-        print(f'   {k:24s} {v:4d} 件')
-    return 0
-
-
-def audit(farm: str | None = None, verbose: bool = True):
-    farm = farm or P.farm()
-    global LEDGER
-    LEDGER = load_ledger(farm)
-    # ★ 场站原始件目录用唯一取用口 (场站名可与 farm 键不同: 键 = rudong, 目录 = data/raw/如东)
-    from src.windscada.config import raw_station_dir
-    station = pathlib.Path(raw_station_dir(farm))
-    verdict: dict[str, dict] = {}
-    fam_rows = []
-    fails: list[str] = []
-
-    for fam in FAMILIES:
-        # ★ 族按**声明顺序**认领, 先声明的先拿 (否则 `m5_cms_tcm/*` 这种宽通配会把 windows/ 下
-        #   1704 件也吞进来 —— fnmatch 的 `*` 是跨 `/` 的, 2026-09-17 第一版就这么误判过)
-        rels = [r for r in _files_of(fam['glob']) if r not in verdict]
-        rels = [r for r in rels if not r.lower().endswith(('.log', '.jsonl'))]
-        if not rels:
-            continue
-        row = dict(id=fam['id'], n=len(rels), func=fam['func'], algo=fam['algo'], input=fam['input'],
-                   pred='+'.join(fam['pred']) or '—', verdict='', note='')
-        if fam['kind'] == 'not-product':
-            # ★ 不是产物, 却躺在产物仓里: 反向审计对它的结论是"它压根不该按产物管" ——
-            #   既不需要生成端, 也谈不上与输入呼应(这一类最该被人看见, 故单独一类, 不当失败计)。
-            jp, js, pg = judgement_evidence(farm, rels)
-            fv, fw = feasibility(fam, rels, False, False, jp, js, pg)
-            row['verdict'] = '✗ 非产物'
-            row['note'] = fam.get('why', '')
-            row['rev'], row['rev_why'] = fv, fw
-            for r in rels:
-                verdict[r] = dict(fam=fam['id'], verdict='✗',
-                                  why='非产物(过程留痕/工具脚本躺在产物仓里): ' + fam.get('why', ''))
-        elif fam['kind'] == 'human':
-            # ★ 人工正本/交证件 (用户令 2026-09-17 #3): 由人写、被运行期当证据引用,
-            #   **不参与呼应校验**——它本来就没有、也不该有生成端; 单独计数, 不混进"无生成端"的失败堆里。
-            row['verdict'] = '◆ 人工件'
-            row['note'] = fam.get('why', '')
-            row['rev'], row['rev_why'] = '人工件(正本/交证)', '人写的, 不参与呼应校验; 随包留在原位'
-            for r in rels:
-                verdict[r] = dict(fam=fam['id'], verdict='◆',
-                                  why='人工件(人工正本/交证件): ' + fam.get('why', ''))
-        elif fam['kind'] == 'shipped':
-            jp, js, pg = judgement_evidence(farm, rels)
-            gen_file = ROOT / fam['gen'] if fam.get('gen') else None
-            gen_ok0 = bool(gen_file and gen_file.is_file())
-            # ★ 逐件分人工件 (用户令 2026-09-17 #3): 含人工判断字段的件**不按自动产物参与呼应校验**,
-            #   单独计 ◆; 余下才是真正的"无生成端"。
-            human = classify_human(farm, rels)
-            rest = [r for r in rels if r not in human]
-            fv, fw = feasibility(fam, rest or rels, gen_ok0, False, jp, js, pg)
-            row['verdict'] = '✗ 无生成端' if not human else f'✗ 无生成端 + ◆ 人工件 {len(human)}'
-            row['n'] = len(rest)
-            row['note'] = fam.get('why', '')
-            row['rev'], row['rev_why'] = fv, fw
-            if jp:
-                row['rev_why'] += f';人工判断字段抽样命中 {jp}%' + (f'(如 {js[0]})' if js else '')
-            for r in rest:
-                verdict[r] = dict(fam=fam['id'], verdict='✗', why='无生成端(全库 0 处写入方): ' + fam['why'])
-            for r in human:
-                verdict[r] = dict(fam=fam['id'], verdict='◆',
-                                  why='人工件(含人工判断字段): 不参与呼应校验, 随包留在原位')
-        else:
-            # ① 生成端在位
-            gen = ROOT / fam['gen'] if fam.get('gen') else None
-            gen_ok = bool(gen and gen.is_file())
-            # ② 输入在位(场站目录优先, 机理层资料在 raw 根下 —— 见 _input_home)
-            spec = INV.INPUTS.get(fam['input'] or '')
-            i_s = i_e = None
-            i_n = 0
-            i_gran = '年'
-            home = _input_home(station, fam['input']) if fam['input'] else None
-            if home and spec:
-                i_s, i_e, i_n, _note, i_gran = INV.input_span(home, fam['input'], spec)
-            elif home:
-                # 机理层资料这类"没有跨度判据"的输入: 按件数清点即算在位(INV.INPUTS 里没有它们的
-                # 跨度口径 —— 技术资料是文档而不是时序数据, 拿跨度比毫无意义)
-                i_n = sum(1 for p in (home / fam['input']).rglob('*') if p.is_file())
-                i_gran = '年'
-            elif fam['kind'] == 'source-derived':
-                # ★ 输入是**包内源码**而不是 data/raw(如发布清单快照: 从 src/windscada/ui 重建页再摘要)。
-                #   生成端在位 = 输入在位, 这类产物的"呼应"是对源码的, 不是对原始件的。
-                i_n = 1
-                i_s = i_e = '包内源码'
-            in_ok = i_n > 0
-            passed, notes = [], []
-            if fam['pred'] and in_ok:
-                if 'span' in fam['pred']:
-                    p_s, p_e, n = _parquet_span(P.out_root(farm) / rels[0], fam.get('time_col'))
-                    if p_s and i_s and (p_s[:7] < i_s[:7]):
-                        # ★ 年粒度输入(按文件名年份推的)不判失败: 与正向检查同一条教训 ——
-                        #   台账 xls 里可能含比文件名年份更早的历史记录, 拿年粒度当"输入起点"会造假缺口
-                        #   (2026-09-17 实测: 工单产物起点 2020-01 而输入文件名最早 2021 ⇒ 不是数据串了)。
-                        msg = f'产物起点 {p_s} 早于输入起点 {i_s}'
-                        if i_gran in ('日', '月'):
-                            fails.append(f"{fam['id']}: {msg}(跨窗混入 / 键错)")
-                            notes.append('⚠ ' + msg)
-                        else:
-                            notes.append(f'{msg}(输入为年粒度, 只作参考: 台账可能含更早的历史行)')
-                    passed.append(f'跨度 {p_s}~{p_e} ⊆ 输入 {i_s}~{i_e}')
-                if 'turbines' in fam['pred']:
-                    pt = _turbines_of(P.out_root(farm) / rels[0])
-                    it = _input_turbines(home, fam['input']) if home else None
-                    if pt and it:
-                        extra = pt - it
-                        if extra:
-                            fails.append(f"{fam['id']}: 产物含输入里没有的机组 {sorted(extra)[:5]}")
-                        passed.append(f'机组 {len(pt)} 台 ⊆ 输入 {len(it)} 台')
-                if 'rows' in fam['pred']:
-                    tot = 0
-                    for r in rels[:400]:
-                        _s, _e, n = _parquet_span(P.out_root(farm) / r, None) if r.endswith('.parquet') \
-                            else (None, None, 1)
-                        tot += n
-                    if tot <= 0:
-                        fails.append(f"{fam['id']}: 产物行数合计为 0")
-                    passed.append(f'件内数据行合计 {tot:,}' + ('(抽样 400 件)' if len(rels) > 400 else ''))
-                if 'npz' in fam['pred']:
-                    sample = rels[:5]
-                    bad = [r for r in sample if not _npz_ok(P.out_root(farm) / r)]
-                    if bad:
-                        fails.append(f"{fam['id']}: 抽样 {len(bad)}/{len(sample)} 件 npz 读不出或为空")
-                    passed.append(f'{len(rels)} 件谱(抽样 {len(sample)} 件 npz 均可读非空)· 输入源 {i_n} 件')
-                if 'manifest' in fam['pred']:
-                    try:
-                        mf = json.loads((P.out_root(farm) / rels[0]).read_text(encoding='utf-8'))
-                        passed.append('清单自记: ' + ', '.join(f'{k}={v}' for k, v in list(mf.items())[:3]))
-                    except Exception as e:
-                        fails.append(f"{fam['id']}: 清单读不出来 ({type(e).__name__})")
-            if not gen_ok:
-                row['verdict'] = '✗ 生成端不在位'
-                row['note'] = f"缺 {fam['gen']}"
-                fails.append(f"{fam['id']}: 生成端不在位 {fam['gen']}")
-            elif not in_ok:
-                row['verdict'] = '~ 输入不在位'
-                row['note'] = (f"缺 {(home / fam['input']) if home else station / str(fam['input'])}"
-                               ' ⇒ 现在无法验证(放数据后复跑本器)')
-            else:
-                row['verdict'] = '✓ 呼应成立'
-                row['note'] = ';'.join(notes + passed)
-                row['rev'], row['rev_why'] = '已在重算链上', '生成端在位、输入在位、判据通过 —— 无需逆向工程'
-            for r in rels:
-                verdict[r] = dict(fam=fam['id'], verdict=row['verdict'][0],
-                                  why=row['note'] if row['verdict'][0] != '✓' else row['note'][:120])
-        fam_rows.append(row)
-
-    # 未归类件 (台账里有, 但没有任何族接住)
-    unclassified = []
-    for rel, v in LEDGER.items():
-        if rel.lower().endswith(('.log', '.jsonl')):
-            continue
-        if rel in verdict:
-            continue
-        # 产物走通配没接住的, 也按台账来源判定
-        unclassified.append(rel)
-        verdict[rel] = dict(fam='(未归类)', verdict='?', why='既不在族表、也无法从台账判定来路')
-
-    counts = {}
-    for v in verdict.values():
-        counts[v['verdict']] = counts.get(v['verdict'], 0) + 1
-    if verbose:
-        print(f'== 反向呼应审计 · 场站 {farm} · 台账 {len(LEDGER)} 件 · 判定 {len(verdict)} 件 ==')
-        print(f'   输入根: {P.rel(station)}')
-        print()
-        print(f'{"族":26s} {"件数":>6s} {"判定":14s} {"输入类":12s} {"功能 / 算法 / 依据"}')
-        for r in fam_rows:
-            print(f'  {r["id"]:24s} {r["n"]:6d} {r["verdict"]:14s} {str(r["input"] or "—"):12s} '
-                  f'{r["func"][:34]}')
-            if r['note']:
-                print(f'      └─ {r["note"][:150]}')
-        if unclassified:
-            print(f'\n  【未归类 {len(unclassified)} 件】')
-            for r in unclassified[:20]:
-                print(f'      ? {r}')
-        print('\n  判定汇总: ' + ' · '.join(f'{k}={v}' for k, v in sorted(counts.items())))
-        print(f'  判据失败 {len(fails)} 条' + (':' + ';'.join(fails[:3]) if fails else ''))
-        ok = counts.get('✓', 0)
-        print(f'  结论: 输出↔输入呼应**成立** {ok} 件 / **不成立(无生成端)** {counts.get("✗", 0)} 件 / '
-              f'**人工件** {counts.get("◆", 0)} 件 / 无法验证 {counts.get("~", 0)} 件 / 未归类 {counts.get("?", 0)} 件')
-        # ── 逆向工程可行性 (用户令 #2): 没有呼应关系的那些, 能不能推出来 ──
-        import collections as _c
-        by_rev = _c.Counter()
-        for r in fam_rows:
-            by_rev[r.get('rev', '?')] += r['n']
-        print()
-        print('== 反向「可逆性」矩阵(用户令 2026-09-17 #2: 无生成端的件能否从原件推出来)==')
-        print(f'{"族":26s} {"件数":>6s} {"可逆性":22s} 依据')
-        for r in fam_rows:
-            if r['verdict'][0] == '✓':
-                continue
-            print(f'  {r["id"]:24s} {r["n"]:6d} {r.get("rev", "?"):22s} {r.get("rev_why", "")[:96]}')
-        print('  按件数汇总: ' + ' · '.join(f'{k}={v}' for k, v in by_rev.most_common()))
-    rc = 5 if (fails or unclassified) else 0
-    return fam_rows, verdict, counts, fails, unclassified, rc
-
-
-def doc_block(farm: str) -> str:
-    fam_rows, verdict, counts, fails, unclassified, _rc = audit(farm, verbose=False)
-    L = [DOC_BEGIN, '### 13.6 反向呼应审计:自输出 → 功能与算法 → 输入(自动生成,勿手改)', '',
-         f'台账 {len(verdict)} 件产物逐件回溯:**呼应成立 {counts.get("✓", 0)} 件** · '
-         f'**不成立(无生成端){counts.get("✗", 0)} 件** · 无法验证 {counts.get("~", 0)} 件 · '
-         f'未归类 {counts.get("?", 0)} 件。'
-         '判据:① 生成端在位 ② 输入在位 ③ 跨度 ⊆ 输入 / 机组集 ⊆ 输入 / 行数 > 0 / 计数与清单一致。', '',
-         '| 输出族(产物 glob) | 件数 | 反向判定 | 功能 | 算法 / 生成端 | 输入(data/raw/<场>/) | 判据 | 说明 |',
-         '|---|---:|---|---|---|---|---|---|']
-    for r in fam_rows:
-        fam = next(f for f in FAMILIES if f['id'] == r['id'])
-        L.append(f'| `{fam["glob"]}` | {r["n"]} | {r["verdict"]} | {r["func"]} | {r["algo"]} | '
-                 f'{r["input"] or "—(无生成端)"} | {r["pred"]} | {r["note"][:160]} |')
-    if unclassified:
-        L += ['', f'**未归类 {len(unclassified)} 件**(需人工认领):' + '、'.join(f'`{x}`' for x in unclassified[:12])]
-    L += ['', '#### 反向「可逆性」:没有呼应关系的那批,能否从原始件推出来(用户令 2026-09-17 #2)', '',
-          '判据两条机器证据:① 生成端在不在包内 ② 产物正文里有没有**人工判断**字段'
-          '(`人工/裁决/审核/校准/评审/经验/judge/review/calibrat/verdict`,抽样读件统计命中率)。'
-          '含人工判断的件不是任何输入的纯函数 ⇒ 逆向工程推不出来,只能把那一步人工工作重做。', '',
-          '| 输出族 | 件数 | 可逆性 | 依据 |', '|---|---:|---|---|']
-    for r in fam_rows:
-        if r['verdict'][0] == '✓':
-            continue
-        L.append(f'| `{r["id"]}` | {r["n"]} | {r.get("rev", "?")} | {r.get("rev_why", "")} |')
-    import collections as _c
-    by_rev = _c.Counter()
-    for r in fam_rows:
-        by_rev[r.get('rev', '?')] += r['n']
-    L += ['', '**按件数汇总**:' + ' · '.join(f'{k} = {v} 件' for k, v in by_rev.most_common()), '',
-          '**人工件**:含人工判断字段的件已逐件登记在 outputs/&lt;场&gt;/_human_artifacts.json(用户令 2026-09-17 #3),'
-          '它们**不参与**呼应校验(人写成/坐实的东西不是任何输入的纯函数);生成方式:'
-          'python scripts/products_reverse_audit.py --write-human-manifest。', '',
-          '> 结论口径:`可逆(直接)` = 包内已有生成端,接上即可;`可逆(需反推口径)` = 内容是测量数据的函数,'
-          '照 `temp_monthly` 的办法反推 + 逐值对拍;`不可逆(人工判断)` = 逆向工程推不出来,'
-          '要么由研发补生成端、要么承认它是人工件(不该按产物管)。', DOC_END]
-    return '\n'.join(L)
-
-
-def write_doc(farm: str) -> int:
-    doc = ROOT / 'docs' / '系统设计说明.md'
-    txt = doc.read_text(encoding='utf-8')
-    blk = doc_block(farm)
-    i, j = txt.find(DOC_BEGIN), txt.find(DOC_END)
-    if i >= 0 and j > i:
-        txt = txt[:i] + blk + txt[j + len(DOC_END):]
-    else:                                             # 首次: 追加到文末
-        txt = txt.rstrip('\n') + '\n\n' + blk + '\n'
-    doc.write_text(txt, encoding='utf-8', newline='\n')
-    print(f'已写入 {P.rel(doc)} (§13.6 反向呼应审计块)')
-    return 0
-
-
-def main() -> int:
-    ap = argparse.ArgumentParser(description='反向呼应审计: 输出 → 功能与算法 → 输入')
-    ap.add_argument('--farm', default=None)
-    ap.add_argument('--check', action='store_true', help='只出结论 (有未归类/判据失败 → rc=5)')
-    ap.add_argument('--write-doc', action='store_true', help='把族表写进 docs/系统设计说明.md §13.6')
-    ap.add_argument('--write-human-manifest', action='store_true',
-                    help='生成人工件台账 outputs/<场>/_human_artifacts.json (用户令 2026-09-17 #3)')
-    a = ap.parse_args()
-    if a.write_human_manifest:
-        return write_human_manifest(a.farm)
-    if a.write_doc:
-        return write_doc(a.farm or P.farm())
-    _rows, _v, counts, fails, uncl, rc = audit(a.farm, verbose=True)
-    if a.check:
-        print(f'[{"OK" if rc == 0 else "X"}] 反向呼应审计 rc={rc}'
-              + (f'(未归类 {len(uncl)} 件 / 判据失败 {len(fails)} 条)' if rc else '(未归类 0 件、判据无失败)'))
-    return rc
-
-
-if __name__ == '__main__':
-    for _s in (sys.stdout, sys.stderr):
-        try:
-            _s.reconfigure(errors='replace')
-        except Exception:
-            pass
-    sys.exit(main())
+#!/usr/bin/env python3
+# -*- coding: utf-8 -*-
+r"""**反向**呼应审计:自输出 → 功能与算法 → 输入(用户令 2026-09-17)。
+
+## 与 `inventory_products.py` 的区别(两个方向,各管一件事)
+
+· `inventory_products.py` 是**正向**:从 `data/raw/<场>/` 的每一类输入出发,找它喂出来的产物,
+  比跨度/条数("输入到了 2026-07,产物只到 2026-04 → 未重算")。它只覆盖 5 类输入、十几件产物。
+· 本器是**反向**:从 `outputs/<场>/` 的**每一件产物**出发,问三个问题:
+    ① 它是谁算出来的?(功能与算法 = 生成端)
+    ② 它从哪份输入算出来的?(`data/raw/<场>/` 的哪一类)
+    ③ 这条"输出 ↔ 输入"的呼应关系**成立吗**?(生成端在位 ∧ 输入在位 ∧ 可机检的对应判据通过)
+  逐件判定,最后给出"成立 / 不成立(无生成端)/ 无法验证"三类账。
+
+## 为什么必须做反向
+
+正向只查"我关心的输入有没有被算成产物",查不出**"盘上这件产物到底有没有来路"**。
+2026-09-17 现场就是栽在这里:清了产物之后,页面缺的 `windscada/index.html`、
+`m5_cms_tcm/handoff_vibration_v2.json` 这类件**根本没有生成端**(全库 0 处写入方),
+正向那几条跨度判据一条都不会报 —— 它们压根不在 `INPUTS` 的 `feeds` 里。
+反向一对账就清楚了:**有来路的件**(生成端 + 输入 + 判据都过)与**没来路的件**(只能从交付包补)。
+
+## 判定口径(每条都写进结果里,不含糊)
+
+    ✓ 呼应成立          生成端在位 ∧ 输入在位 ∧ 判据通过(跨度 ⊆ 输入 / 键集 ⊆ 输入 / 计数一致 / 行数 > 0)
+    ~ 输入不在位         生成端知道,但 `data/raw/<场>/<类>` 不在 ⇒ **现在无法验证**(放数据后复跑本器)
+    ✗ 无生成端           全库 0 处写入方 ⇒ **输出↔输入的呼应在原理上不成立**:这件产物无法由输入推导出来,
+                        **无法由重算生出来**;按用户令 2026-09-17 #1 运行期也不从交付包补齐 ⇒ 只能由研发补生成端(本器 --feasibility 给出逐族可逆性)
+    ? 未归类             既不在族表里、台账里也没有 —— 需要人工认领(本器把它当缺口报出来)
+
+## 用法
+
+    python scripts/products_reverse_audit.py            # 全量反向审计(人看)
+    python scripts/products_reverse_audit.py --check    # 只出结论(有 ✗/? 时退出码 5)
+    python scripts/products_reverse_audit.py --write-doc # 把族表写进 docs/系统设计说明.md §13.6
+
+退出码: 0 全部成立(允许 ✗,只要它们都在台账里如实标了 shipped)· 5 有未归类件或判据失败
+"""
+from __future__ import annotations
+
+import argparse
+import fnmatch
+import json
+import pathlib
+import sys
+
+ROOT = pathlib.Path(__file__).resolve().parents[1]
+sys.path.insert(0, str(ROOT))
+sys.path.insert(0, str(ROOT / 'scripts'))
+from src import paths as P                                              # noqa: E402
+import inventory_products as INV                                        # noqa: E402
+
+DOC_BEGIN, DOC_END = '<!-- REVERSE-AUDIT:BEGIN -->', '<!-- REVERSE-AUDIT:END -->'
+
+# ── 族表: 输出(glob) → 功能 / 算法(生成端) / 输入类 / 判据 ────────────────────────────
+#   pred: span=产物时间跨度必须落在输入跨度内; turbines=产物机组集 ⊆ 输入机组集;
+#         rows=行数>0; count=件数/计数与说明一致; manifest=与清单自记一致
+FAMILIES: list[dict] = [
+    dict(id='vib_window_index', glob='m5_cms_tcm/windows/*/index.parquet', kind='raw-derived',
+         func='振动摄入 · 窗索引', algo='scripts/rudong_tcm_index.py(逐 decode JSON 解析 54 列:turbine/'
+              'sensor_name/meas_name/trigger_time/rpm/condition_key/alarm_type)',
+         gen='scripts/rudong_tcm_index.py', input='windcms', pred=('span', 'turbines', 'rows'),
+         time_col='trigger_time'),
+    dict(id='vib_spectra', glob=['m5_cms_tcm/windows/*/spectra/**', 'm5_cms_tcm/windows/*/spectra_meta.parquet'],
+         kind='raw-derived',
+         func='振动摄入 · 谱库', algo='scripts/rudong_tcm_spectra.py(FFT_ 测量 → npz 幅值数组, 分片存 `spectra/p<NN>/`)',
+         gen='scripts/rudong_tcm_spectra.py', input='windcms', pred=('npz',)),
+    dict(id='vib_raw_manifest', glob='m5_cms_tcm/vib_raw_manifest.json', kind='raw-derived',
+         func='振动摄入 · 清单', algo='scripts/vib_raw_build.py(汇总窗/谱件数、时间跨度、missing_chain 缺口)',
+         gen='scripts/vib_raw_build.py', input='windcms', pred=('manifest',)),
+    dict(id='scada_alarms', glob='windscada/alarms.parquet', kind='raw-derived',
+         func='三门台账 · 报警', algo='scripts/windscada_alarms_ingest.py(SpreadsheetML *.xls → 事件表)',
+         gen='scripts/windscada_alarms_ingest.py', input='故障报警', pred=('span', 'rows'), time_col='t_on'),
+    dict(id='scada_workorders', glob='windscada/workorders.parquet', kind='raw-derived',
+         func='三门台账 · 工单', algo='scripts/windscada_workorder_ingest.py(检修台账 → 工单表)',
+         gen='scripts/windscada_workorder_ingest.py', input='风机故障记录', pred=('span', 'rows'),
+         time_col='t_report'),
+    dict(id='scada_oil', glob='windscada/oil_samples_index.parquet', kind='raw-derived',
+         func='三门台账 · 油样', algo='scripts/windscada_watch_channels_build.py(油样 PDF → 索引)',
+         gen='scripts/windscada_watch_channels_build.py', input='油样报告', pred=('span', 'rows'),
+         time_col='date'),
+    dict(id='scada_monthly', glob='windscada/temp_monthly.parquet', kind='raw-derived',
+         func='月度派生件', algo='scripts/windscada_monthly_build.py(逐值对齐随包件 19,494/19,494)',
+         gen='scripts/windscada_monthly_build.py', input='scada_10min', pred=('span', 'rows'), time_col='month'),
+    dict(id='scada_derived', glob='windscada/*.parquet', kind='raw-derived',
+         func='SCADA 派生分析(功率曲线/损失/曲线/控制/停机/温度/偏航/液压/热链/系统辅助)',
+         algo='src/windscada/perf/{powercurve,availability,curves,control,faults} + subsys/{temp_nbm,yaw,'
+              'hydraulic,thermal_chain} + taxonomy.py',
+         gen='src/windscada/perf/powercurve.py', input='scada_10min', pred=('rows',)),
+    dict(id='ontology_core', glob='ontology/{objects.json,retrieval_index.json,turbine_params.parquet}',
+         kind='raw-derived',
+         func='本体层 · 码表/对象库/检索索引/实机参数',
+         algo='python -m src.ontology.kb_ingest → populate → chain_ingest → trend_ingest;'
+              'src.ontology.retrieval.build;src.ontology.maintenance.refresh_params',
+         gen='src/ontology/kb_ingest.py', input='西门子4.0技术资料', pred=('rows',)),
+
+    dict(id='ontology_aux_build', glob=['ontology/_id_alias.json', 'ontology/retrieval_vec.npy'],
+         kind='raw-derived',
+         func='本体层 · 别名表与向量索引缓存', algo='src.ontology.kb_ingest(别名)/ '
+              'src.ontology.retrieval.build(use_vec=True)(向量)',
+         gen='src/ontology/kb_ingest.py', input='西门子4.0技术资料', pred=('rows',)),
+    dict(id='ontology_domain', glob=['ontology/oil_2026H2_huabiao.json', 'ontology/parts_cost_public.json',
+                                     'ontology/scenario_29_repair.json'],
+         kind='shipped',
+         func='本体层 · 领域数据件(油样台账/备件价格/场景修复)', algo='(无生成端, 领域正本出件)',
+         gen=None, input=None, pred=(), why='领域正本随包发来, 不由 data/raw 推导'),
+
+    # ── 以下族: 全库 0 处写入方 ⇒ 反向呼应**在原理上不成立**(如实记账, 不假装成立)────
+    #    ★2026-09-18 用户令补充: 这些件**运行期一律不从交付包补**(用户令 2026-09-17 #1 的延伸),
+    #      页面如实为空/降级; 要出件只有两条路 —— ① 研发补生成端; ② 现场正本随原始件进 data/raw。
+    #      故本表的 why 里不再写"从交付包补齐"那种话(那是**离线人工补救**, 见 products_restore_missing.py)。
+    dict(id='vib_model_run_l6', glob=['m5_cms_tcm/model_run_l6.parquet', 'm5_cms_tcm/model_run_summary.json'],
+         kind='raw-derived',
+         func='六层链 model_run 步 · L6 过闸谱线表 + 层小结',
+         algo='scripts/rudong_model_run.py(候选线=oem_scan_plan 的 11 部件×{BPFI,BPFO,BSF}; 过闸='
+              'src/sop/discriminators.py::spectral_line_gates G1–G8; 定级=vib_verdict_and_writeback)',
+         gen='scripts/rudong_model_run.py', input='m5_cms_tcm', pred=('rows',),
+         why='★2026-09-19 按口径重建(**无标准答案**: 这两件从未随包, 无法逐值对拍); 峰值拾取仍是本器自定'
+             '(目标频率 ±2 bin 取最大), 不得当"复现"用 —— 见 docs/振动六层链_接口规格与缺口_v0.1.md §3'),
+    dict(id='vib_fusion_38', glob='m5_cms_tcm/fusion_38.csv', kind='raw-derived',
+         func='六层链 fusion 步 · 逐台融合表(报告"融合级"列的来源)',
+         algo='scripts/rudong_fusion_run.py::fusion_table(模型侧=model_run_l6 过闸线; CMS 侧=窗索引 '
+              'RedMask/YellowMask; 裁决=src/sop/fusion_diag.py::fuse)',
+         gen='scripts/rudong_fusion_run.py', input='m5_cms_tcm', pred=('rows',),
+         why='★按口径重建(无标准答案): 台号用 WTGnn 形态(消费端 registry 的键), 与 l6 的 f"{n}#" 不同'),
+    dict(id='vib_fleet_scalar_z', glob='m5_cms_tcm/fleet_scalar_z.parquet', kind='raw-derived',
+         func='振动线 · 全场标量 z 值 (fusion 步出件)',
+         algo='scripts/rudong_fusion_run.py(逆向工程实现: 窗索引 scalar_value → 同 (传感器,测量,工况) 族中位 '
+              '→ 稳健 z=(val−med)/(1.4826·MAD); 与随包样件 8,887 行逐值对拍 val 列 100% 一致)',
+         # input 写**原始件目录名**(族表口径: data/raw/<场>/<input>); 真实取数经窗索引一跳,
+         # 而窗索引本身是 raw-derived(vib_window_index 族) —— 写 windows/_*.parquet 会让审计判"输入不在位"。
+         gen='scripts/rudong_fusion_run.py', input='m5_cms_tcm', pred=('rows',),
+         why='2026-09-18 上链: 由 vib_raw_build.py 的 --with-report 分支按当次窗名调用, '
+             '使 /api/fleet 的振动标量 z 面不再空白'),
+    dict(id='vib_fusion_handoff', glob='m5_cms_tcm/handoff_vibration_v2.json', kind='raw-derived',
+         func='融合面 handoff (观澜自算, 同结构于振动线手交件)',
+         algo='scripts/rudong_fusion_handoff.py(报告状态等级 + fusion_38 融合裁决 + L6 过闸线 → 同结构, '
+              'meta.from 标明"观澜自算")',
+         gen='scripts/rudong_fusion_handoff.py', input='m5_cms_tcm', pred=('rows',),
+         why='★2026-09-19 用户令「所有的计算均要形成观澜的源代码」: 这件此前全库 0 处写入方(振动线人写), '
+             '新机器装完 `/detail/v2` 的「需要关注/全场状态」必然空白。现由观澜自算出同结构件; '
+             '盘上若有**振动线正本**, 生成端不覆盖(正本含人工裁决, 优先)'),
+    dict(id='vib_handoff_and_scans', glob='m5_cms_tcm/*.{json,parquet,csv,md,txt}',
+         kind='shipped',
+         func='振动线出件与专项扫描(handoff / 历史 / 基线 / 各类 freq scan / 判级台账)', algo='(无生成端)',
+         gen=None, input=None, pred=(), why='振动六层链四步脚本未随包(oem_frequency_scan 等)⇒ '
+                                            '运行期不从交付包补(用户令 2026-09-17); 要出件须研发补生成端, '
+                                            '或由现场正本随原始件进 data/raw 后重算'),
+    dict(id='vib_figs', glob='m5_cms_tcm/figs/*', kind='shipped',
+         func='振动线出图(谱图/趋势图)', algo='(无生成端)', gen=None, input=None, pred=(),
+         why='随包快照(振动线出件)'),
+    # ★2026-09-19 拆族 (用户令「所有的计算均要形成观澜的源代码」): `windscada/index.html` 是
+    #   `/detail/v2` 的「全场状态总览」iframe 目标, 此前整族记为"无生成端" ⇒ 新机器上那块必然显示
+    #   "产物不在位"。现在它由 `scripts/windscada_overview_build.py` 用**组件同一个快照烘焙器**
+    #   (src/windscada/ui/snapshot.py::bake, /v2/snapshot 用的就是它)从重算产物生成 ⇒ 单列一族。
+    dict(id='windscada_overview', glob='windscada/index.html', kind='raw-derived',
+         func='总览「全场状态」页 (单文件: 接口响应烤进页面, 离线可开)',
+         algo='scripts/windscada_overview_build.py(ctx 取数函数与在线 /api/* 同一实现; 经 '
+              'src/windscada/ui/snapshot.py::bake 烘焙)',
+         gen='scripts/windscada_overview_build.py', input='scada_10min', pred=('rows',),
+         why='2026-09-19 用户令后上链: 该页此前是随包件(全库 0 处写入方), 清过产物的机器上 iframe 只能'
+             '显示"产物不在位"; 现在由观澜自算, 每次重算随数据更新'),
+    dict(id='windscada_pages', glob='windscada/{turbines/*,review/*}', kind='shipped',
+         func='逐台静态页 + 评审记录(当前由组件按请求动态出页, 静态件仍属随包快照)', algo='(无生成端)',
+         gen=None, input=None, pred=(), why='逐台页在运行期由 /turbine/<台> 动态生成, 静态副本无写入方; '
+                                             '运行期不从交付包补(用户令 2026-09-17)'),
+    # ★2026-09-18 补: `windcms/cache/*` 是**读侧缓存**(标量按内容哈希落盘), 它不是随包快照 ——
+    #   原先 `windcms_pages`(shipped) 先认领了它, 于是重算后新写的缓存被记成"包内无生成端"。
+    #   写方在 src/windcms/data.py::load_scalars (窗索引 → 标量), 故单列一族并放在 windcms_pages 之前。
+    dict(id='windcms_cache', glob='windcms/cache/*', kind='raw-derived',
+         func='CMS 报告侧 · 标量缓存', algo='src/windcms/data.py::load_scalars(窗索引 → 标量数组, '
+              '按内容哈希命名落 cache/)', gen='src/windcms/data.py', input='windcms', pred=('rows',)),
+    # ★2026-09-18 补 (用户令"页面只能基于输入数据重算"之后): `windcms/**` 整族原先一律按
+    #   "无生成端"记账 —— 但**报告两步的生成端就在包内**(scripts/windcms.py report → src/windcms/
+    #   report.py + report_std.py), 只要 data/raw/<场>/windcms 里有 decode 导出就出得来。
+    #   实测(2026-09-18): 从重算出的窗 w0316 一键跑出标准模版报告 + 38 台逐台页 + 总览, 判级 38 台
+    #   (报警 3 / 优秀 35), 缺的只有 model_run/fusion 两步的融合级列(如实写 —)。故拆成两族按 raw-derived 记。
+    dict(id='windcms_report_std', glob=['windcms/报告_CMS振动状态评估报告_*.md',
+                                        'windcms/报告_CMS振动状态评估报告_*.docx',
+                                        'windcms/登记_三轴_38台.csv'], kind='raw-derived',
+         func='CMS 标准模版评估报告 (三轴: 状态等级 × 证据状态 × 行动等级建议)',
+         algo='src/windcms/report_std.py::build(结构化判级 → 报告 md + docx + 登记 csv; '
+              '由 scripts/windcms.py report 调用)',
+         gen='src/windcms/report_std.py', input='windcms', pred=('rows',),
+         why='★同名件有两种来源: 观澜自算(本族: 从 data/raw/<场>/windcms 的 *_decode.json 算) 与 '
+             '厂家报告转录(scripts/vib_reports_build.py, 输入是现场振动分析报告 PDF)。按名字认领不了'
+             '这层区别 ⇒ 以"是否由本步生成"为准, 判据看 m5_cms_tcm/vib_raw_manifest.json 的 steps'),
+    dict(id='windcms_report_overview', glob=['windcms/report.md', 'windcms/index.html',
+                                             'windcms/index_eng.html', 'windcms/overview.html',
+                                             'windcms/turbines/*'], kind='raw-derived',
+         func='CMS 诊断总览页 + 逐台页 + 结构化层报告',
+         algo='src/windcms/report.py::build(逐台 turbines/<台>.html + 总览 overview/index.html + '
+              'report.md; 融合级列要 model_run/fusion 出件, 缺则如实写"—")',
+         gen='src/windcms/report.py', input='windcms', pred=('rows',)),
+    dict(id='windcms_kb', glob='windcms/kb.json', kind='raw-derived',
+         func='CMS 知识库索引 (问答检索用)',
+         algo='src/windcms/knowledge.py::build(由 scripts/windcms.py kb 调用; 离线无模型时 0 chunks 也算跑成)',
+         gen='src/windcms/knowledge.py', input='windcms', pred=('rows',)),
+    dict(id='windcms_pages', glob='windcms/**', kind='shipped',
+         func='windcms 里其余件 (厂家报告转录 / 知识库出件等)',
+         algo='(无随包生成端)', gen=None, input=None, pred=(),
+         why='上述两族与 cache 之外剩下的 windcms 件: 运行期不从交付包补(用户令 2026-09-17), '
+             '要出件须研发补生成端'),
+    # ★2026-09-19 用户令「要让观澜从重算台账生成 claim」: findings 与三份底稿改由**观澜自算**
+    #   (scripts/sop_findings_from_ledger.py 从 alarms/workorders/oil/temp_bins/powercurve_dev/loss_monthly/L6
+    #   算 claim, 按契约构建器要的字段结构写) ⇒ 单列一族; 契约四件随之由 scripts/guanlan_facts_contract.py 生成。
+    dict(id='ledger_claims', glob=['sop/findings.json', 'paradigm_r1/experiments/*/底稿*.json'],
+         kind='raw-derived',
+         func='事实契约的 claim 本体与底稿 (观澜自算, 替代人裁底稿)',
+         algo='scripts/sop_findings_from_ledger.py(六族 claim: 报警集中/重复检修/停机损失/温度离群/功率曲线偏离/振动过闸线; '
+              '每条带 coverage·缺证据·证伪判据, 相对判据封顶「候选」)',
+         gen='scripts/sop_findings_from_ledger.py', input='scada_10min', pred=('rows',),
+         why='输入是**重算产物**(windscada/*.parquet + m5_cms_tcm/*): 其 raw 祖先是 故障报警/风机故障记录/scada_10min; '
+             '本族单列以区别于旧的人裁底稿(同路径旧件无生成端)'),
+    dict(id='sop_workspace', glob='sop/**', kind='shipped',
+         func='SOP 中间件/评审/事实契约底稿', algo='(无生成端)', gen=None, input=None, pred=(),
+         why='评审过程件, 属人工作业留痕'),
+    dict(id='tcm_replay', glob='tcm_compatible_replay/**', kind='shipped',
+         func='TCM 兼容链回放资产(模型表/掩码阈值/裁决记录)', algo='(无生成端)', gen=None, input=None,
+         pred=(), why='随包快照'),
+    dict(id='paradigm_r1', glob='paradigm_r1/**', kind='shipped',
+         func='范式实验件(E3/E5/E8 底稿)', algo='(无生成端)', gen=None, input=None, pred=(),
+         why='实验底稿, 事实契约的输入'),
+    dict(id='pitch_shipped', glob='pitch/**', kind='shipped',
+         func='变桨侧派生件(零位/日粒度)', algo='(无生成端;rebuild_from_raw --scada 只覆盖其中一部分)',
+         gen=None, input=None, pred=(), why='无生成端 ⇒ 运行期不从交付包补(用户令 2026-09-17), '
+                                            '页面降级; 要出件须研发补生成端'),
+    dict(id='ontology_releases', glob='ontology/release_*/*', kind='shipped',
+         func='本体发布层 r1/r2(只读暴露给网关 /release/)', algo='(无生成端)', gen=None, input=None,
+         pred=(), why='发布层清单, 由研发出件'),
+    # ★2026-09-19 用户令「所有的计算均要形成观澜的源代码」+「从重算台账生成 claim」:
+    #   这一族从"无生成端"改成 raw-derived —— 生成端 guanlan_facts_contract.py 本来就在包内,
+    #   原来缺的是**它的输入**(人裁底稿); 现在输入由上游族 ledger_claims 自算
+    #   (scripts/sop_findings_from_ledger.py) ⇒ 新机器上重算即得, /api/facts 不再 503。
+    dict(id='guanlan_contract', glob='guanlan/**', kind='raw-derived',
+         func='事实契约与对外派生(门户结论段/取数)',
+         algo='scripts/guanlan_facts_contract.py(build→render: 契约 + portal_claims/detail_cards/qa_refs/'
+              'report_summary; 输入=上游族 ledger_claims 自算的 sop/findings.json + 三份底稿)',
+         gen='scripts/guanlan_facts_contract.py', input='scada_10min', pred=('rows',),
+         why='输入与 data/raw 隔着两跳(重算台账 → claims → 契约), 但**两跳的生成端都在包内**; '
+             'claim 的判据/缺证据/证伪逐条写在件里, 相对判据封顶「候选」'),
+    dict(id='windscada_release_manifest', glob='windscada/release-manifest.json', kind='source-derived',
+         func='发布清单快照(`/healthz` 与版本信息读它)',
+         algo='scripts/guanlan_baseline_manifest.py(从 src/windscada/ui 源重建 zh 页 → 摘要 '
+              '{css_tokens, nav, dom, sha256, source blobs})',
+         gen='scripts/guanlan_baseline_manifest.py', input=None, pred=(),
+         why='★ 2026-09-17 修正: 它一度被我按"交证件"错分成非产物; 实际**包内有生成端**'
+             '(guanlan_baseline_manifest.py 就写这个路径), 且 guanlan_gateway.py 运行期读它 ⇒ 是产物'),
+    dict(id='human_deliverables',
+         glob=['windscada/mblub_alarm_cross.json', 'windscada/review_presented.json',
+               'windscada/给振动线_温度轴回复_20260826.json', 'windscada/给振动线_温度轴回复_20260826.md',
+               'windscada/送审_系统结论与判定_20260828.md'],
+         kind='human',
+         func='**人工正本/交证件**:与振动线的跨线通知、送审结论、发布评审快照',
+         algo='(人写的, 不是算出来的)', gen=None, input=None, pred=(),
+         why='这几件是"人对人的交证件/回执", 内容由分析结论抄写与人判断坐实而成; '
+             'mblub_alarm_cross.json 还被 src/ontology/populate.py 当作**证据来源**引用 ⇒ '
+             '必须随包留在原位, 但**不参与呼应校验**(它本来就没有生成端, 也不该有)'),
+    dict(id='scratch_in_products', glob=['**/_*.js', '**/_*audit*.md', '**/*.out', '**/*.err',
+                                        '**/*_test_*.parquet', '**/*_test_*.json'],
+         kind='not-product',
+         func='**非产物**:自审脚本/记录与模型测试输出(躺在产物仓里)', algo='(工具脚本与运行痕迹, 不是产物)',
+         gen=None, input=None, pred=(),
+         why='2026-09-17: 17 件已移出到 reference/非产物留痕_20260917/ 并留了 _来源说明.json ⇒ '
+             '本族现在应为空; 一旦再出现(新脚本往产物仓里写 .out/_audit.js 之类), 本器当场报出来'),
+]
+
+
+def _files_of(glob) -> list[str]:
+    """族 glob(str 或 list)→ 台账里的相对键(支持 `{a,b}` 花括号与 `*`)。
+
+    ★ 深度必须**显式对齐**(2026-09-17 第一版踩过): `fnmatch` 的 `*` 是**跨 `/`** 的, 于是
+      `windscada/*.parquet` 会把 `windscada/turbines/xx.parquet`、`windscada/_pre_rebuild_*/xx.parquet`
+      一起吞进来 —— 反向审计最怕"族把不该管的件认领了", 那样账就假了。
+      规则: 模式里没有 `**` 时, 要求的 `/` 个数必须与键的 `/` 个数**相等**(逐条展开后各自判)。
+    """
+    pats = []
+    for g in ([glob] if isinstance(glob, str) else list(glob)):
+        if '{' in g:
+            head, tail = g.split('{', 1)
+            opts, rest = tail.split('}', 1)
+            pats += [head + o + rest for o in opts.split(',')]
+        else:
+            pats.append(g)
+    out = []
+    for rel in LEDGER:
+        for pt in pats:
+            if '**' not in pt and rel.count('/') != pt.count('/'):
+                continue
+            if fnmatch.fnmatch(rel, pt) or rel == pt:
+                out.append(rel)
+                break
+    return sorted(out)
+
+
+LEDGER: dict[str, dict] = {}
+
+
+def load_ledger(farm: str) -> dict[str, dict]:
+    """台账 = `_provenance.json`(逐件来源) + `_derived_manifest.json`(生成端自登记)。"""
+    out: dict[str, dict] = {}
+    root = P.out_root(farm)
+    for fn in ('_provenance.json', '_derived_manifest.json'):
+        f = root / fn
+        if not f.is_file():
+            continue
+        try:
+            d = json.loads(f.read_text(encoding='utf-8'))
+        except Exception:
+            continue
+        for rel, v in (d.get('files') or {}).items():
+            if isinstance(v, dict):
+                out.setdefault(rel, dict(source=v.get('source') or 'raw-derived',
+                                         builder=v.get('builder') or v.get('by') or '',
+                                         why=v.get('why') or '',))
+    return out
+
+
+def _input_home(station: pathlib.Path, sub: str) -> pathlib.Path | None:
+    """输入类目录在哪 —— 优先场站目录下, 其次 raw 根下。
+
+    ★ 机理层资料按 A2 约定**不在场站目录下**(`data/raw/西门子4.0技术资料/`, 见 src/windscada/config.py)。
+      只查场站目录会把本体层判成"输入不在位"(2026-09-17 第一版就这么误报过 3 件)。
+    """
+    for base in (station, station.parent):
+        d = base / sub
+        if d.is_dir():
+            return base
+    return None
+
+
+def _parquet_span(p: pathlib.Path, col: str | None):
+    import pandas as pd
+    try:
+        d = pd.read_parquet(p)
+    except Exception:
+        return None, None, 0
+    n = len(d)
+    if col and col in d.columns:
+        s = pd.to_datetime(d[col], errors='coerce')
+        if s.notna().any():
+            return str(s.min())[:10], str(s.max())[:10], n
+    return None, None, n
+
+
+def _npz_ok(p: pathlib.Path) -> bool:
+    """振动谱件: npz 能读出来且非空(列数不是"行", 不能拿 parquet 的行数判它)。"""
+    try:
+        import numpy as np
+        with np.load(p, allow_pickle=False) as z:
+            return any(z[k].size > 0 for k in z.files)
+    except Exception:
+        return False
+
+
+# 已知"含人工裁决/校准"的随包快照(按用户令 2026-09-17 #3 显式登记为人工件)。依据是文档里已经写明的性质:
+# docs §7「handoff_vibration_v2.json / component_history.json 是随包快照, 该件含人工裁决/校准更新,
+# 不是测量数据的函数」; baseline_38.json 的三层基线 self/absolute 两层由人工坐实; findings.json 是判级发现台账。
+HUMAN_SNAPSHOTS = {
+    'm5_cms_tcm/component_history.json',
+    'm5_cms_tcm/baseline_38.json',
+    'm5_cms_tcm/handoff_vibration_v1.json',
+    'm5_cms_tcm/findings.json',
+}
+
+# ── 逆向工程可行性 (用户令 2026-09-17 #2) ────────────────────────────────────────────
+# 问的是: "这批没有生成端的产物, 能不能从原始件/在包生成端**推**出来?"
+# 判定不靠印象, 靠两条机器证据:
+#   ① 生成端在不在包内(在 → 只是没接上, 跑一次就有);
+#   ② 产物正文里有没有**人工判断**的痕迹("裁决/审核/校准/评审/经验"这类字段)——
+#      含人工判断的件不是任何输入的纯函数, 逆向工程推不出来, 只能把那一步人工工作重做。
+import re as _re                                                         # noqa: E402
+JUDGE_PAT = _re.compile(r'人工|裁决|审核|复核|校准|评审|经验值|专家|judg|review|manual|calibrat|verdict|sign[_ ]?off',
+                        _re.I)
+# ★ 只认"**字段**叫这个名字"(`"裁决": …` / `裁决,`),不认正文里顺嘴提一句 ——
+#   2026-09-17 第一版拿整篇子串匹配, 结果把**页面**里的展示标签(`index.html` 里的"评审/人工")当成
+#   人工判断, 于是把"页面"整族误判成不可逆。这属于"判据太糙 → 结论假"。
+JUDGE_KEY_PAT = _re.compile(r'["\']?[^"\',:]{0,24}(人工|裁决|审核|复核|校准|评审|经验值|专家|'
+                            r'judg|review|manual|calibrat|verdict|sign[_ ]?off)[^"\',:]{0,24}["\']?\s*[:=]',
+                            _re.I)
+
+
+def judgement_evidence(farm: str, rels: list[str], cap: int = 8) -> tuple[int, list[str], int]:
+    """→ (含人工判断字段的抽样比例%, 样例, 被检查件数)。
+
+    只查**数据/文本件**(.json/.csv/.md/.txt); **页面(.html/.js)不算** —— 页面是产物的渲染,
+    里面出现"评审/人工"是展示标签, 不代表这份件内含判断。
+    """
+    hit, samples, checked, pages = 0, [], 0, 0
+    for rel in rels[:cap]:
+        p = P.out_root(farm) / rel
+        suf = p.suffix.lower()
+        if suf in ('.html', '.js', '.css', '.htm'):
+            pages += 1
+            continue
+        try:
+            if suf == '.json':
+                txt = p.read_text(encoding='utf-8', errors='replace')[:200000]
+                pat = JUDGE_KEY_PAT
+            elif suf in ('.csv', '.txt'):
+                txt = p.read_text(encoding='utf-8', errors='replace')[:50000]
+                pat = JUDGE_KEY_PAT
+            elif suf == '.md':
+                txt = p.read_text(encoding='utf-8', errors='replace')[:50000]
+                pat = JUDGE_PAT
+            else:
+                continue
+        except Exception:
+            continue
+        checked += 1
+        if pat.search(txt):
+            hit += 1
+            samples.append(rel)
+    return (100 * hit // checked if checked else 0), samples, pages
+
+
+def feasibility(fam: dict, rels: list[str], gen_ok: bool, in_ok: bool,
+                judge_pct: int, judge_samples: list[str], pages: int = 0) -> tuple[str, str]:
+    """→ (可逆性判定, 依据)。四类: 可逆(直接) / 可逆(需反推口径) / 不可逆(人工判断) / 应移出产物仓。"""
+    if fam['kind'] == 'not-product':
+        return '应移出产物仓', '它本就不是产物(工具脚本/交证件/测试输出) ⇒ 该从产物仓移走, 而不是"补生成端"'
+    if gen_ok and in_ok:
+        return '可逆(直接)', f"生成端在包内({fam['gen']})且输入在位 ⇒ 接上/重跑即可"
+    if pages and pages >= max(3, len(rels[:8]) // 2):
+        return '可逆(页面可再生)', ('这一族主要是**页面**(.html): 页面是产物的渲染, 不是独立数据 ⇒ '
+                                    '接上在包的页面生成端重出即可(未必逐字节同, 能力与数据一致)')
+    m_n = sum(1 for r in rels if r.endswith(('.parquet', '.csv', '.npz')))
+    t_n = len(rels) - m_n
+    split = f'({m_n} 件测量形态 / {t_n} 件文本·判断件)' if (m_n and t_n) else ''
+    # 两类都占相当比重时, 结论必须是"混合"而不是挑一类代表全族 —— 否则一句判定就把另一半骗过去了
+    if m_n and t_n and min(m_n, t_n) * 4 >= len(rels):
+        return '混合(测量件可反推 / 文本件需人定)', split + \
+               '测量形态的件可由 data/raw 反推口径 + 逐值对拍; 文本/判断件(评审、裁决、回复)推不出来'
+    if judge_pct >= 50 and judge_samples:
+        return '不可逆(人工判断)', (f"抽样 {len(judge_samples)} 件里都写着人工判断字段(如 {judge_samples[0]}) "
+                                    f"⇒ 不是输入的纯函数, 逆向工程推不出来")
+    if m_n and m_n * 2 >= len(rels):
+        return '可逆(需反推口径)', split + '测量形态的件是测量数据的函数 ⇒ 可从 data/raw 反推口径 + 逐值对拍' \
+                                        '(做法同 temp_monthly: 反推 → 逐值一致才敢用)'
+    if gen_ok:
+        return '半可逆(生成端在包, 输入不足)', '生成端在包内, 缺的是上游输入 ⇒ 上游补齐后即可重出'
+    if judge_pct > 0:
+        return '半可逆(需人工裁定)', f'抽样里 {judge_pct}% 的件含人工判断字段 ⇒ 机器部分可推、判断部分要人定'
+    return '需人工裁定', '既无生成端、形态也不是测量函数 ⇒ 需研发给口径或领域正本'
+
+
+def _turbines_of(p: pathlib.Path, col: str = 'turbine', cap: int = 200000):
+    import pandas as pd
+    try:
+        d = pd.read_parquet(p, columns=[col]) if col else pd.read_parquet(p)
+    except Exception:
+        return None
+    try:
+        return {str(x).upper() for x in d[col].dropna().unique()} if col in d.columns else None
+    except Exception:
+        return None
+
+
+def _input_turbines(station: pathlib.Path, sub: str):
+    """输入侧机组集合: 逐台 CSV 文件名 / 目录名里的 WTG 号。"""
+    d = station / sub
+    if not d.is_dir():
+        return None
+    out = set()
+    for p in d.rglob('*'):
+        if p.is_file():
+            for m in __import__('re').finditer(r'(WTG\s?\d{1,2})', p.name.upper()):
+                out.add(m.group(1).replace(' ', ''))
+    return out or None
+
+
+def classify_human(farm: str, rels: list[str], cap: int = 600) -> set[str]:
+    """逐件判"是不是人工件"(用户令 2026-09-17 #3:人工件不再按自动产物参与呼应校验)。
+
+    判据:**非测量形态**的件(.md/.json/.csv/.txt)里出现人工判断字段(按字段名匹配)。
+    测量形态件(.parquet/.csv 数值、.npz)一律不算人工件 —— 它们是数据的函数,属"可反推"那一路。
+    """
+    human: set[str] = set()
+    for rel in rels[:cap]:
+        if rel in HUMAN_SNAPSHOTS:            # ① 文档里已写明含人工裁决/校准的随包快照
+            human.add(rel)
+            continue
+        suf = pathlib.Path(rel).suffix.lower()
+        if suf == '.md':                      # ② 产物仓里的 .md 一律是报告/交证/评审记录(人写的)
+            human.add(rel)
+            continue
+        if suf == '.txt':
+            pat = JUDGE_PAT
+        elif suf in ('.json', '.csv'):
+            pat = JUDGE_KEY_PAT
+        else:
+            continue
+        p = P.out_root(farm) / rel
+        try:
+            txt = p.read_text(encoding='utf-8', errors='replace')[:120000]
+        except Exception:
+            continue
+        if pat.search(txt):                   # ③ 正文里有"人工判断字段"
+            human.add(rel)
+    return human
+
+
+def write_human_manifest(farm: str | None = None) -> int:
+    """把人工件逐件落成 `outputs/<场>/_human_artifacts.json`(机器台账,供自检与审计引用)。"""
+    farm = farm or P.farm()
+    global LEDGER
+    LEDGER = load_ledger(farm)
+    entries: dict[str, dict] = {}
+    for fam in FAMILIES:
+        if fam['kind'] not in ('shipped', 'human', 'not-product'):
+            continue
+        rels = [r for r in _files_of(fam['glob']) if r not in entries]
+        if not rels:
+            continue
+        if fam['kind'] == 'human':
+            for r in rels:
+                entries[r] = dict(kind='human', family=fam['id'], evidence='族内显式登记为人工正本/交证件')
+            continue
+        for r in classify_human(farm, rels):
+            if r not in entries:
+                entries[r] = dict(kind='human', family=fam['id'],
+                                  evidence='件内含人工判断字段(按字段名匹配),不是输入的纯函数')
+    out = P.out_root(farm) / '_human_artifacts.json'
+    out.write_text(json.dumps(dict(
+        at=__import__('time').strftime('%Y-%m-%d %H:%M:%S'),
+        note='人工件台账(用户令 2026-09-17 #3):这些件由人写成/坐实,不是任何输入的纯函数 ⇒ '
+             '不参与"输出↔输入呼应"校验;随包留在原位,重算不会也不该重写它们。'
+             '本文件由 scripts/products_reverse_audit.py --write-human-manifest 生成。',
+        count=len(entries), files=entries), ensure_ascii=False, indent=1), encoding='utf-8')
+    by_fam: dict[str, int] = {}
+    for v in entries.values():
+        by_fam[v['family']] = by_fam.get(v['family'], 0) + 1
+    print(f'人工件台账: {len(entries)} 件 → {P.rel(out)}')
+    for k, v in sorted(by_fam.items(), key=lambda kv: -kv[1]):
+        print(f'   {k:24s} {v:4d} 件')
+    return 0
+
+
+def audit(farm: str | None = None, verbose: bool = True):
+    farm = farm or P.farm()
+    global LEDGER
+    LEDGER = load_ledger(farm)
+    # ★ 场站原始件目录用唯一取用口 (场站名可与 farm 键不同: 键 = rudong, 目录 = data/raw/如东)
+    from src.windscada.config import raw_station_dir
+    station = pathlib.Path(raw_station_dir(farm))
+    verdict: dict[str, dict] = {}
+    fam_rows = []
+    fails: list[str] = []
+
+    for fam in FAMILIES:
+        # ★ 族按**声明顺序**认领, 先声明的先拿 (否则 `m5_cms_tcm/*` 这种宽通配会把 windows/ 下
+        #   1704 件也吞进来 —— fnmatch 的 `*` 是跨 `/` 的, 2026-09-17 第一版就这么误判过)
+        rels = [r for r in _files_of(fam['glob']) if r not in verdict]
+        rels = [r for r in rels if not r.lower().endswith(('.log', '.jsonl'))]
+        if not rels:
+            continue
+        row = dict(id=fam['id'], n=len(rels), func=fam['func'], algo=fam['algo'], input=fam['input'],
+                   pred='+'.join(fam['pred']) or '—', verdict='', note='')
+        if fam['kind'] == 'not-product':
+            # ★ 不是产物, 却躺在产物仓里: 反向审计对它的结论是"它压根不该按产物管" ——
+            #   既不需要生成端, 也谈不上与输入呼应(这一类最该被人看见, 故单独一类, 不当失败计)。
+            jp, js, pg = judgement_evidence(farm, rels)
+            fv, fw = feasibility(fam, rels, False, False, jp, js, pg)
+            row['verdict'] = '✗ 非产物'
+            row['note'] = fam.get('why', '')
+            row['rev'], row['rev_why'] = fv, fw
+            for r in rels:
+                verdict[r] = dict(fam=fam['id'], verdict='✗',
+                                  why='非产物(过程留痕/工具脚本躺在产物仓里): ' + fam.get('why', ''))
+        elif fam['kind'] == 'human':
+            # ★ 人工正本/交证件 (用户令 2026-09-17 #3): 由人写、被运行期当证据引用,
+            #   **不参与呼应校验**——它本来就没有、也不该有生成端; 单独计数, 不混进"无生成端"的失败堆里。
+            row['verdict'] = '◆ 人工件'
+            row['note'] = fam.get('why', '')
+            row['rev'], row['rev_why'] = '人工件(正本/交证)', '人写的, 不参与呼应校验; 随包留在原位'
+            for r in rels:
+                verdict[r] = dict(fam=fam['id'], verdict='◆',
+                                  why='人工件(人工正本/交证件): ' + fam.get('why', ''))
+        elif fam['kind'] == 'shipped':
+            jp, js, pg = judgement_evidence(farm, rels)
+            gen_file = ROOT / fam['gen'] if fam.get('gen') else None
+            gen_ok0 = bool(gen_file and gen_file.is_file())
+            # ★ 逐件分人工件 (用户令 2026-09-17 #3): 含人工判断字段的件**不按自动产物参与呼应校验**,
+            #   单独计 ◆; 余下才是真正的"无生成端"。
+            human = classify_human(farm, rels)
+            rest = [r for r in rels if r not in human]
+            fv, fw = feasibility(fam, rest or rels, gen_ok0, False, jp, js, pg)
+            row['verdict'] = '✗ 无生成端' if not human else f'✗ 无生成端 + ◆ 人工件 {len(human)}'
+            row['n'] = len(rest)
+            row['note'] = fam.get('why', '')
+            row['rev'], row['rev_why'] = fv, fw
+            if jp:
+                row['rev_why'] += f';人工判断字段抽样命中 {jp}%' + (f'(如 {js[0]})' if js else '')
+            for r in rest:
+                verdict[r] = dict(fam=fam['id'], verdict='✗', why='无生成端(全库 0 处写入方): ' + fam['why'])
+            for r in human:
+                verdict[r] = dict(fam=fam['id'], verdict='◆',
+                                  why='人工件(含人工判断字段): 不参与呼应校验, 随包留在原位')
+        else:
+            # ① 生成端在位
+            gen = ROOT / fam['gen'] if fam.get('gen') else None
+            gen_ok = bool(gen and gen.is_file())
+            # ② 输入在位(场站目录优先, 机理层资料在 raw 根下 —— 见 _input_home)
+            spec = INV.INPUTS.get(fam['input'] or '')
+            i_s = i_e = None
+            i_n = 0
+            i_gran = '年'
+            home = _input_home(station, fam['input']) if fam['input'] else None
+            if home and spec:
+                i_s, i_e, i_n, _note, i_gran = INV.input_span(home, fam['input'], spec)
+            elif home:
+                # 机理层资料这类"没有跨度判据"的输入: 按件数清点即算在位(INV.INPUTS 里没有它们的
+                # 跨度口径 —— 技术资料是文档而不是时序数据, 拿跨度比毫无意义)
+                i_n = sum(1 for p in (home / fam['input']).rglob('*') if p.is_file())
+                i_gran = '年'
+            elif fam['kind'] == 'source-derived':
+                # ★ 输入是**包内源码**而不是 data/raw(如发布清单快照: 从 src/windscada/ui 重建页再摘要)。
+                #   生成端在位 = 输入在位, 这类产物的"呼应"是对源码的, 不是对原始件的。
+                i_n = 1
+                i_s = i_e = '包内源码'
+            in_ok = i_n > 0
+            passed, notes = [], []
+            if fam['pred'] and in_ok:
+                if 'span' in fam['pred']:
+                    p_s, p_e, n = _parquet_span(P.out_root(farm) / rels[0], fam.get('time_col'))
+                    if p_s and i_s and (p_s[:7] < i_s[:7]):
+                        # ★ 年粒度输入(按文件名年份推的)不判失败: 与正向检查同一条教训 ——
+                        #   台账 xls 里可能含比文件名年份更早的历史记录, 拿年粒度当"输入起点"会造假缺口
+                        #   (2026-09-17 实测: 工单产物起点 2020-01 而输入文件名最早 2021 ⇒ 不是数据串了)。
+                        msg = f'产物起点 {p_s} 早于输入起点 {i_s}'
+                        if i_gran in ('日', '月'):
+                            fails.append(f"{fam['id']}: {msg}(跨窗混入 / 键错)")
+                            notes.append('⚠ ' + msg)
+                        else:
+                            notes.append(f'{msg}(输入为年粒度, 只作参考: 台账可能含更早的历史行)')
+                    passed.append(f'跨度 {p_s}~{p_e} ⊆ 输入 {i_s}~{i_e}')
+                if 'turbines' in fam['pred']:
+                    pt = _turbines_of(P.out_root(farm) / rels[0])
+                    it = _input_turbines(home, fam['input']) if home else None
+                    if pt and it:
+                        extra = pt - it
+                        if extra:
+                            fails.append(f"{fam['id']}: 产物含输入里没有的机组 {sorted(extra)[:5]}")
+                        passed.append(f'机组 {len(pt)} 台 ⊆ 输入 {len(it)} 台')
+                if 'rows' in fam['pred']:
+                    tot = 0
+                    for r in rels[:400]:
+                        _s, _e, n = _parquet_span(P.out_root(farm) / r, None) if r.endswith('.parquet') \
+                            else (None, None, 1)
+                        tot += n
+                    if tot <= 0:
+                        fails.append(f"{fam['id']}: 产物行数合计为 0")
+                    passed.append(f'件内数据行合计 {tot:,}' + ('(抽样 400 件)' if len(rels) > 400 else ''))
+                if 'npz' in fam['pred']:
+                    sample = rels[:5]
+                    bad = [r for r in sample if not _npz_ok(P.out_root(farm) / r)]
+                    if bad:
+                        fails.append(f"{fam['id']}: 抽样 {len(bad)}/{len(sample)} 件 npz 读不出或为空")
+                    passed.append(f'{len(rels)} 件谱(抽样 {len(sample)} 件 npz 均可读非空)· 输入源 {i_n} 件')
+                if 'manifest' in fam['pred']:
+                    try:
+                        mf = json.loads((P.out_root(farm) / rels[0]).read_text(encoding='utf-8'))
+                        passed.append('清单自记: ' + ', '.join(f'{k}={v}' for k, v in list(mf.items())[:3]))
+                    except Exception as e:
+                        fails.append(f"{fam['id']}: 清单读不出来 ({type(e).__name__})")
+            if not gen_ok:
+                row['verdict'] = '✗ 生成端不在位'
+                row['note'] = f"缺 {fam['gen']}"
+                fails.append(f"{fam['id']}: 生成端不在位 {fam['gen']}")
+            elif not in_ok:
+                row['verdict'] = '~ 输入不在位'
+                row['note'] = (f"缺 {(home / fam['input']) if home else station / str(fam['input'])}"
+                               ' ⇒ 现在无法验证(放数据后复跑本器)')
+            else:
+                row['verdict'] = '✓ 呼应成立'
+                row['note'] = ';'.join(notes + passed)
+                row['rev'], row['rev_why'] = '已在重算链上', '生成端在位、输入在位、判据通过 —— 无需逆向工程'
+            for r in rels:
+                verdict[r] = dict(fam=fam['id'], verdict=row['verdict'][0],
+                                  why=row['note'] if row['verdict'][0] != '✓' else row['note'][:120])
+        fam_rows.append(row)
+
+    # 未归类件 (台账里有, 但没有任何族接住)
+    unclassified = []
+    for rel, v in LEDGER.items():
+        if rel.lower().endswith(('.log', '.jsonl')):
+            continue
+        if rel in verdict:
+            continue
+        # 产物走通配没接住的, 也按台账来源判定
+        unclassified.append(rel)
+        verdict[rel] = dict(fam='(未归类)', verdict='?', why='既不在族表、也无法从台账判定来路')
+
+    counts = {}
+    for v in verdict.values():
+        counts[v['verdict']] = counts.get(v['verdict'], 0) + 1
+    if verbose:
+        print(f'== 反向呼应审计 · 场站 {farm} · 台账 {len(LEDGER)} 件 · 判定 {len(verdict)} 件 ==')
+        print(f'   输入根: {P.rel(station)}')
+        print()
+        print(f'{"族":26s} {"件数":>6s} {"判定":14s} {"输入类":12s} {"功能 / 算法 / 依据"}')
+        for r in fam_rows:
+            print(f'  {r["id"]:24s} {r["n"]:6d} {r["verdict"]:14s} {str(r["input"] or "—"):12s} '
+                  f'{r["func"][:34]}')
+            if r['note']:
+                print(f'      └─ {r["note"][:150]}')
+        if unclassified:
+            print(f'\n  【未归类 {len(unclassified)} 件】')
+            for r in unclassified[:20]:
+                print(f'      ? {r}')
+        print('\n  判定汇总: ' + ' · '.join(f'{k}={v}' for k, v in sorted(counts.items())))
+        print(f'  判据失败 {len(fails)} 条' + (':' + ';'.join(fails[:3]) if fails else ''))
+        ok = counts.get('✓', 0)
+        print(f'  结论: 输出↔输入呼应**成立** {ok} 件 / **不成立(无生成端)** {counts.get("✗", 0)} 件 / '
+              f'**人工件** {counts.get("◆", 0)} 件 / 无法验证 {counts.get("~", 0)} 件 / 未归类 {counts.get("?", 0)} 件')
+        # ── 逆向工程可行性 (用户令 #2): 没有呼应关系的那些, 能不能推出来 ──
+        import collections as _c
+        by_rev = _c.Counter()
+        for r in fam_rows:
+            by_rev[r.get('rev', '?')] += r['n']
+        print()
+        print('== 反向「可逆性」矩阵(用户令 2026-09-17 #2: 无生成端的件能否从原件推出来)==')
+        print(f'{"族":26s} {"件数":>6s} {"可逆性":22s} 依据')
+        for r in fam_rows:
+            if r['verdict'][0] == '✓':
+                continue
+            print(f'  {r["id"]:24s} {r["n"]:6d} {r.get("rev", "?"):22s} {r.get("rev_why", "")[:96]}')
+        print('  按件数汇总: ' + ' · '.join(f'{k}={v}' for k, v in by_rev.most_common()))
+    rc = 5 if (fails or unclassified) else 0
+    return fam_rows, verdict, counts, fails, unclassified, rc
+
+
+def doc_block(farm: str) -> str:
+    fam_rows, verdict, counts, fails, unclassified, _rc = audit(farm, verbose=False)
+    L = [DOC_BEGIN, '### 13.6 反向呼应审计:自输出 → 功能与算法 → 输入(自动生成,勿手改)', '',
+         f'台账 {len(verdict)} 件产物逐件回溯:**呼应成立 {counts.get("✓", 0)} 件** · '
+         f'**不成立(无生成端){counts.get("✗", 0)} 件** · 无法验证 {counts.get("~", 0)} 件 · '
+         f'未归类 {counts.get("?", 0)} 件。'
+         '判据:① 生成端在位 ② 输入在位 ③ 跨度 ⊆ 输入 / 机组集 ⊆ 输入 / 行数 > 0 / 计数与清单一致。', '',
+         '| 输出族(产物 glob) | 件数 | 反向判定 | 功能 | 算法 / 生成端 | 输入(data/raw/<场>/) | 判据 | 说明 |',
+         '|---|---:|---|---|---|---|---|---|']
+    for r in fam_rows:
+        fam = next(f for f in FAMILIES if f['id'] == r['id'])
+        L.append(f'| `{fam["glob"]}` | {r["n"]} | {r["verdict"]} | {r["func"]} | {r["algo"]} | '
+                 f'{r["input"] or "—(无生成端)"} | {r["pred"]} | {r["note"][:160]} |')
+    if unclassified:
+        L += ['', f'**未归类 {len(unclassified)} 件**(需人工认领):' + '、'.join(f'`{x}`' for x in unclassified[:12])]
+    L += ['', '#### 反向「可逆性」:没有呼应关系的那批,能否从原始件推出来(用户令 2026-09-17 #2)', '',
+          '判据两条机器证据:① 生成端在不在包内 ② 产物正文里有没有**人工判断**字段'
+          '(`人工/裁决/审核/校准/评审/经验/judge/review/calibrat/verdict`,抽样读件统计命中率)。'
+          '含人工判断的件不是任何输入的纯函数 ⇒ 逆向工程推不出来,只能把那一步人工工作重做。', '',
+          '| 输出族 | 件数 | 可逆性 | 依据 |', '|---|---:|---|---|']
+    for r in fam_rows:
+        if r['verdict'][0] == '✓':
+            continue
+        L.append(f'| `{r["id"]}` | {r["n"]} | {r.get("rev", "?")} | {r.get("rev_why", "")} |')
+    import collections as _c
+    by_rev = _c.Counter()
+    for r in fam_rows:
+        by_rev[r.get('rev', '?')] += r['n']
+    L += ['', '**按件数汇总**:' + ' · '.join(f'{k} = {v} 件' for k, v in by_rev.most_common()), '',
+          '**人工件**:含人工判断字段的件已逐件登记在 outputs/&lt;场&gt;/_human_artifacts.json(用户令 2026-09-17 #3),'
+          '它们**不参与**呼应校验(人写成/坐实的东西不是任何输入的纯函数);生成方式:'
+          'python scripts/products_reverse_audit.py --write-human-manifest。', '',
+          '> 结论口径:`可逆(直接)` = 包内已有生成端,接上即可;`可逆(需反推口径)` = 内容是测量数据的函数,'
+          '照 `temp_monthly` 的办法反推 + 逐值对拍;`不可逆(人工判断)` = 逆向工程推不出来,'
+          '要么由研发补生成端、要么承认它是人工件(不该按产物管)。', DOC_END]
+    return '\n'.join(L)
+
+
+def write_doc(farm: str) -> int:
+    doc = ROOT / 'docs' / '系统设计说明.md'
+    txt = doc.read_text(encoding='utf-8')
+    blk = doc_block(farm)
+    i, j = txt.find(DOC_BEGIN), txt.find(DOC_END)
+    if i >= 0 and j > i:
+        txt = txt[:i] + blk + txt[j + len(DOC_END):]
+    else:                                             # 首次: 追加到文末
+        txt = txt.rstrip('\n') + '\n\n' + blk + '\n'
+    doc.write_text(txt, encoding='utf-8', newline='\n')
+    print(f'已写入 {P.rel(doc)} (§13.6 反向呼应审计块)')
+    return 0
+
+
+def main() -> int:
+    ap = argparse.ArgumentParser(description='反向呼应审计: 输出 → 功能与算法 → 输入')
+    ap.add_argument('--farm', default=None)
+    ap.add_argument('--check', action='store_true', help='只出结论 (有未归类/判据失败 → rc=5)')
+    ap.add_argument('--write-doc', action='store_true', help='把族表写进 docs/系统设计说明.md §13.6')
+    ap.add_argument('--write-human-manifest', action='store_true',
+                    help='生成人工件台账 outputs/<场>/_human_artifacts.json (用户令 2026-09-17 #3)')
+    a = ap.parse_args()
+    if a.write_human_manifest:
+        return write_human_manifest(a.farm)
+    if a.write_doc:
+        return write_doc(a.farm or P.farm())
+    _rows, _v, counts, fails, uncl, rc = audit(a.farm, verbose=True)
+    if a.check:
+        print(f'[{"OK" if rc == 0 else "X"}] 反向呼应审计 rc={rc}'
+              + (f'(未归类 {len(uncl)} 件 / 判据失败 {len(fails)} 条)' if rc else '(未归类 0 件、判据无失败)'))
+    return rc
+
+
+if __name__ == '__main__':
+    for _s in (sys.stdout, sys.stderr):
+        try:
+            _s.reconfigure(errors='replace')
+        except Exception:
+            pass
+    sys.exit(main())

+ 9 - 0
scripts/rebuild_all.py

@@ -78,6 +78,15 @@ def build_plan(a) -> list:
         plan.append(step_cmd('④b 振动侧摄入 + CMS 报告 + 标量 z (原始导出 → 窗索引/谱/逐台页/融合 z)',
                              [PY, 'scripts/vib_raw_build.py'] + ([] if a.vib_report else ['--no-report']),
                              note='没有振动原始件时空跑属正常; 若解析出错会以非零退出 (不静默)'))
+    # ⑤b 事实契约的 claim 生成 (用户令 2026-09-19「要让观澜从重算台账生成 claim」):
+    #     由重算产物(alarms/workorders/oil/temp_bins/powercurve_dev/loss_monthly/L6/fusion)算 claim,
+    #     写成 sop/findings.json + paradigm_r1 三份底稿(与契约构建器同结构) ⇒ 契约/门户/问答三处随重算刷新。
+    #     ★必须排在 ⑤a 台账维护之前: 这两步产出的件要进同一轮台账与审计, 否则审计会按"溯源缺失"报 rc=6。
+    plan.append(step_cmd('⑤b 事实契约: 由重算台账生成 claim', [PY, 'scripts/sop_findings_from_ledger.py'],
+                         note='从重算产物生成; 判据与缺证据逐条写在 claim 里(相对判据封顶「候选」)'))
+    plan.append(step_cmd('⑤c 事实契约: 构建 + 渲染消费者', [PY, 'scripts/guanlan_facts_contract.py', 'build'],
+                         tolerate=(2,),
+                         note='rc=2 = 契约自检不过(sha/脱敏/枚举), 需要看输出; 门户结论段由 ⑧b 重装时刷新'))
     # ⑤a 台账维护 (用户令 2026-09-18: 页面只认"有来路的件" ⇒ 台账必须跟着重算一起更新)。
     #    ★实逮: 2026-09-18 用户「清除产物」后从 raw 全量重算, `outputs/<场>/_provenance.json`
     #      **整本丢了** —— 因为写它的唯一入口 (旧的第⑤步 products_restore_missing.py) 已按用户令

+ 354 - 0
scripts/sop_findings_from_ledger.py

@@ -0,0 +1,354 @@
+#!/usr/bin/env python3
+# -*- coding: utf-8 -*-
+r"""事实契约的**重算台账生成端**(用户令 2026-09-19:「要让观澜从重算台账生成 claim」)。
+
+## 背景(为什么换掉人裁底稿)
+
+门户的「本场结论·契约生成」段、`/detail` 的问答/报告面(`/api/facts`)都读事实契约
+(`outputs/<场>/guanlan/facts_contract_v0.json` + `derived/{detail_cards,portal_claims,qa_refs}.json`)。
+契约的输入此前是**人裁底稿**:`sop/findings.json`(119 条结论)+ `paradigm_r1` 的 E3/E5/E8 三份底稿
+—— 清过产物、只从 `data/raw` 重算的机器上它们必然不在位 ⇒ 契约面永远 503、门户结论段停在旧数。
+用户令选定:**由观澜从重算台账生成 claim**,这三处随重算刷新。
+
+## 本器做什么
+
+按 `scripts/guanlan_facts_contract.py` 读的四份输入的**字段结构**产出同结构件(契约构建器一行不改):
+
+    outputs/<场>/sop/findings.json                                  结论本体(claim 列表)
+    outputs/<场>/paradigm_r1/experiments/E5_coverage_table/底稿.json  finding_rows: 模块映射
+    outputs/<场>/paradigm_r1/experiments/E3_candidate_closure/底稿.json rows: 归宿提议
+    outputs/<场>/paradigm_r1/experiments/E8_rudong_rebuild/底稿_s0.json findings_rows: 系统/台号
+
+claim 全部由**重算产物**算出(每条都记 `source_refs` 的件与 sha16):
+
+  · 报警集中台        windscada/alarms.parquet(逐台计数 + 同族码集中度)
+  · 重复检修件        windscada/workorders.parquet(同台同部件重复更换)
+  · 停机损失集中      windscada/loss_monthly.parquet(停机态损失占比)
+  · 温度离群台        windscada/temp_bins.parquet(同工况同通道 dev 离群)
+  · 功率曲线偏离台    windscada/powercurve_dev.parquet(同风速残差)
+  · 振动面报警台      m5_cms_tcm/fusion_38.csv(融合级 定论/准定论·预警)
+
+## 判据纪律(不造数)
+
+· 每条 claim 的 `verdict` 只取契约六枚举,且**只能用台账能支撑的那一档**:全部相对判据 + 无实物锚 ⇒ 最高
+  「候选」(振动面来自观澜自算融合, 同样封顶候选);口径性/样本性问题用「参考」或「INSUFFICIENT」。
+· 逐条写 `mandatory_limitation`(缺什么证据)与 `falsifiability`(怎么证伪),而不是只给结论。
+· `coverage` 写**实算**的窗天数与覆盖台数; 点名台号写进 `sample_level.named`。
+· 与振动线/厂家正本无关: 本件是**观澜自算结论**, 不冒充人工裁决(generate 侧写在文件头与 meta)。
+
+用法:
+    python scripts/sop_findings_from_ledger.py            # 生成四份输入件并自登记
+    python scripts/sop_findings_from_ledger.py --dry-run   # 只报会生成几条 claim
+    python scripts/sop_findings_from_ledger.py --status
+"""
+from __future__ import annotations
+
+import argparse
+import hashlib
+import json
+import pathlib
+import sys
+import time
+
+import pandas as pd
+
+ROOT = pathlib.Path(__file__).resolve().parents[1]
+sys.path.insert(0, str(ROOT))
+sys.path.insert(0, str(ROOT / 'scripts'))
+from src import paths as P                                              # noqa: E402
+
+SIX = ('定论', '准定论·预警', '候选', '参考', 'INSUFFICIENT', '撤回')
+OUT_FILES = ('sop/findings.json',
+             'paradigm_r1/experiments/E5_coverage_table/底稿.json',
+             'paradigm_r1/experiments/E3_candidate_closure/底稿.json',
+             'paradigm_r1/experiments/E8_rudong_rebuild/底稿_s0.json')
+
+
+def _sha16(p: pathlib.Path) -> str:
+    return hashlib.sha256(p.read_bytes()).hexdigest()[:16]
+
+
+def _win(t, days: int) -> tuple:
+    """(起, 止) 与窗天数 —— 从数据实算, 不写死。"""
+    end = pd.to_datetime(t.max())
+    start = end - pd.Timedelta(days=days)
+    return start, end, int((end - start).days)
+
+
+def claims(store: pathlib.Path, m5: pathlib.Path) -> list[dict]:
+    out: list[dict] = []
+
+    def add(cid, title, verdict, refs, *, named=(), systems=(), module='', window_days=None,
+            n_machines=None, limitation='', falsify='', nature='', claim_class='相对判据',
+            coverage_note=''):
+        out.append(dict(id=cid, title=title, verdict=verdict, source_refs=refs, named=list(named),
+                        systems=list(systems), module=module, window_days=window_days,
+                        n_machines=n_machines, mandatory_limitation=limitation, falsifiability=falsify,
+                        problem_nature=nature, claim_class=claim_class, coverage_note=coverage_note))
+
+    # ── ① 报警集中台 ────────────────────────────────────────────────────────────
+    ap = store / 'alarms.parquet'
+    if ap.is_file():
+        al = pd.read_parquet(ap)
+        al['t_on'] = pd.to_datetime(al['t_on'], errors='coerce')
+        s, e, days = _win(al['t_on'], 365)
+        w = al[(al.t_on >= s) & (al.t_on <= e)]
+        cnt = w.groupby('turbine').size().sort_values(ascending=False)
+        if len(cnt) > 2:
+            med = float(cnt.median())
+            top = cnt.head(3)
+            named = [f'{t}({int(n)}条)' for t, n in top.items()]
+            x = float(top.iloc[0]) / med if med else float('nan')
+            refs = [dict(file=P.rel(ap), sha16=_sha16(ap))]
+            code_top = (w[w.turbine == top.index[0]].groupby('code').size().sort_values(ascending=False)
+                        .head(3).to_dict()) if 'code' in w else {}
+            add('RD-LEDGER-ALM-01',
+                f'报警集中: {top.index[0]} {int(top.iloc[0])} 条(机群中位 {med:.0f} 条的 {x:.1f} 倍)',
+                '候选', refs, named=named, systems=['报警面'], module='台账·报警',
+                window_days=days, n_machines=int(w.turbine.nunique()),
+                limitation='报警条数不等于缺陷严重度: 单码抖动(同一码反复报)会抬高条数; 未做同码基线归一。',
+                falsify=f'若该台的高发码是全场共因(同一码在多台同量出现)或与限电/调度同向, 则该"集中"是工况不是缺陷 —— 判据: 逐码机群基率>{1 / max(len(cnt), 1):.0%}。',
+                nature='口径性(计数面)', coverage_note=f'窗 {s:%Y-%m-%d}~{e:%Y-%m-%d}, 全员 {int(w.turbine.nunique())} 台',
+                claim_class=f'计数面相对判据; 头部码 {code_top}')
+
+    # ── ② 重复检修件 ────────────────────────────────────────────────────────────
+    wo = store / 'workorders.parquet'
+    if wo.is_file():
+        d = pd.read_parquet(wo)
+        comp_col = next((c for c in ('故障位置二级', '故障位置一级', '故障名称') if c in d.columns), None)
+        tm = next((c for c in ('故障报出时间', '复位运行时间') if c in d.columns), None)
+        unit = next((c for c in ('机组编号',) if c in d.columns), None)
+        if comp_col and unit:
+            d['_c'] = d[comp_col].astype(str).str.strip()
+            g = d[d['_c'] != ''].groupby([unit, '_c']).size().sort_values(ascending=False)
+            if len(g):
+                (t0, c0), n0 = g.index[0], int(g.iloc[0])
+                tm_ = pd.to_datetime(d[tm], errors='coerce') if tm else None
+                s, e, days = _win(tm_, 365) if tm_ is not None else ('', '', None)
+                _norm = lambda x: (str(x) if str(x).upper().startswith('WTG') else f'WTG{int(str(x)):02d}' if str(x).isdigit() else str(x))
+                named = [f'{_norm(t)}·{c}({int(n)}次)' for (t, c), n in g.head(3).items()]
+                add('RD-LEDGER-WO-01',
+                    f'重复检修: {_norm(t0)} 的「{c0}」累计 {n0} 次(窗内)',
+                    '候选', [dict(file=P.rel(wo), sha16=_sha16(wo))], named=named, systems=['传动链'],
+                    module='台账·工单', window_days=days, n_machines=int(d[unit].nunique()),
+                    limitation='检修台账记录的是"做过什么", 不等于"坏过几次": 同一部件的例行检查也会计次; '
+                               '未做计划性/故障性停机区分(缺 work_type 字段)。',
+                    falsify='若这些次记录多为计划性检查/预防性更换, 则该"重复"无非计划性含义 —— 判据: 记录里标注为故障停机的比例。',
+                    nature='趋势性(重复发生)', coverage_note=f'窗 {s}~{e}' if s else '窗未标(无时间列)',
+                    claim_class='计数面相对判据; 部件名取台账原文')
+
+    # ── ③ 停机损失集中 ─────────────────────────────────────────────────────────
+    lm = store / 'loss_monthly.parquet'
+    if lm.is_file():
+        d = pd.read_parquet(lm)
+        st_col = 'state' if 'state' in d.columns else None
+        if st_col and 'loss' in d.columns and 'turbine' in d.columns:
+            stop = d[d[st_col].astype(str).str.contains('停机', na=False)]
+            # ★单位: loss 列是 **kWh**(实测全场合计 4,428 万、停机态 2,273 万)—— 直接当 MWh 报会
+            #   把量级说大 1000 倍。这里统一换算成 MWh 再报, 分子分母同口径。
+            tot_stop = float(stop['loss'].sum()) / 1000.0
+            per = (stop.groupby('turbine')['loss'].sum() / 1000.0).sort_values(ascending=False)
+            if len(per) and tot_stop > 0:
+                sh = float(per.iloc[0]) / tot_stop
+                named = [f'{t}({v:,.0f}MWh)' for t, v in per.head(3).items()]
+                add('RD-LEDGER-LOSS-01',
+                    f'停机损失集中: {per.index[0]} 占全场停机损失 {sh:.0%}({per.iloc[0]:,.0f} MWh, 全场停机 {tot_stop:,.0f} MWh)',
+                    '参考', [dict(file=P.rel(lm), sha16=_sha16(lm))], named=named, systems=['发电量'],
+                    module='台账·损失', window_days=int(d['month'].nunique()) * 30 if 'month' in d else None,
+                    n_machines=int(d['turbine'].nunique()),
+                    limitation='损失是"损失了多少电", 不是"为什么损失"; 含调度令停机(非缺陷)与限电, 需与故障类分开看。',
+                    falsify='若该台损失主要来自调度令/限电(状态列已区分)而非故障停机, 则与设备无关。',
+                    nature='口径性(电量面)', coverage_note='按月汇总, 月份数取自产物')
+
+    # ── ④ 温度离群台 ────────────────────────────────────────────────────────────
+    tb = store / 'temp_bins.parquet'
+    if tb.is_file() and (m5 / 'fusion_38.csv').is_file():
+        d = pd.read_parquet(tb)
+        if {'turbine', 'channel', 'pbin', 'med', 'n'}.issubset(d.columns):
+            d = d[d['n'] >= 20]                                     # 样本量门 (n 太少的中位不稳)
+            if len(d):
+                piv = d.pivot_table(index=['channel', 'pbin', 'win'], columns='turbine', values='med')
+                dev = piv.sub(piv.median(axis=1), axis=0).abs()
+                worst = dev.max(axis=1).sort_values(ascending=False)
+                if len(worst):
+                    (ch, pb, win), val = worst.index[0], float(worst.iloc[0])
+                    t0 = str(dev.loc[worst.index[0]].idxmax())
+                    add('RD-LEDGER-TEMP-01',
+                        f'温度相对离群: {t0} 的 {ch} 在 {pb} 功率档偏离机群中位 {val:.1f}K(工况窗 {win})',
+                        '候选', [dict(file=P.rel(tb), sha16=_sha16(tb))],
+                        named=[f'{t0}·{ch} +{val:.1f}K'], systems=['温度面'], module='台账·温度 NBM',
+                        window_days=None, n_machines=int(d['turbine'].nunique()),
+                        limitation='同工况中位比对的是"机群", 若全机群共因升温和(环境/冷却水温)则判据失效; '
+                                   '未接绝对限值(无 ISO/厂商温度锚)。',
+                        falsify='若该通道全场中位同期同向上升(共模), 则该台"离群"不成立 —— 判据: 逐月机群中位趋势。',
+                        nature='趋势性(个体偏离)', coverage_note=f'{win} 工况段, 样本门 n≥20',
+                        claim_class='同工况机群残差; 功率档分箱')
+
+    # ── ⑤ 功率曲线偏离台 ───────────────────────────────────────────────────────
+    pc = store / 'powercurve_dev.parquet'
+    if pc.is_file():
+        d = pd.read_parquet(pc)
+        if 'dev_w' in d.columns:
+            d2 = d.assign(_a=pd.to_numeric(d['dev_w'], errors='coerce').abs()).sort_values('_a', ascending=False)
+            r0 = d2.iloc[0]
+            # ★单位: dev_w 是 **MW**(实测 -0.147 → -147 kW)—— 直接标 kW 会把量级说小 1000 倍。
+            _dv = float(pd.to_numeric(pd.Series([r0.get('dev_w')]), errors='coerce').iloc[0]) * 1000.0
+            _cnt = (d['判别'].astype(str).value_counts().to_dict() if '判别' in d.columns else {})
+            _bad = {k: v for k, v in _cnt.items() if k not in ('—', 'nan', '')}
+            add('RD-LEDGER-PC-01',
+                f'功率曲线偏离: {r0.get("turbine")} 同风速下功率偏差 {_dv:+,.0f} kW(判别: {r0.get("判别", "—")})'
+                + (f'; 全场判别分布 {_bad}' if _bad else ''),
+                '参考', [dict(file=P.rel(pc), sha16=_sha16(pc))],
+                named=[f'{r0.get("turbine")} {float(r0.get("dev_w")):,.0f}kW'],
+                systems=['发电量', '功率曲线'], module='台账·功率曲线', window_days=None,
+                n_machines=int(len(d)),
+                limitation='同风速残差受机位风资源(尾流/扇区)影响; 未做机位修正, 故只作"待核"不作缺陷结论。',
+                falsify='若该台所处扇区长期受尾流(邻机偏置)影响, 残差可由机位解释 ⇒ 非机组问题。',
+                nature='口径性(机位/风资源)', coverage_note='逐台一行(残差+判别列)')
+
+    # ── ⑥ 振动面过闸线(观澜自算六层链) ───────────────────────────────────────
+    f38 = m5 / 'fusion_38.csv'
+    l6p = m5 / 'model_run_l6.parquet'
+    cand = pd.DataFrame()
+    if l6p.is_file():
+        l6 = pd.read_parquet(l6p)
+        if '定级' in l6.columns and len(l6):
+            cand = l6[l6['定级'].astype(str).str.contains('候选|准定论|定论', na=False)]
+            if len(cand):
+                named = [f"{r['台']}·{str(r['线'])[:22]} {r['定级']}" for _, r in cand.head(5).iterrows()]
+                add('RD-LEDGER-VIB-01',
+                    f"振动面过闸线: {len(cand)} 条特征线达「候选」及以上({', '.join(n.split(' ')[0] for n in named[:3])})",
+                    '候选', [dict(file=P.rel(l6p), sha16=_sha16(l6p))], named=named, systems=['振动面'],
+                    module='振动·六层链 L6 过闸线', window_days=None, n_machines=int(cand['台'].nunique()),
+                    limitation='机制未定 ⇒ **不命名部件**(滚道/圈侧须实物或换件闭环); 峰值拾取为本器口径'
+                               '(目标频率 ±2 bin 取最大), 无正样本锚 ⇒ 封顶「候选」, 不出「定论」。',
+                    falsify='若复测同一判据回落(或拆检未见对应损伤), 则该候选撤回 —— 判据: 下一窗同线定级 + 实物证据。',
+                    nature='趋势性(证据面)', coverage_note='L6 过闸线表; 闸分布见 m5_cms_tcm/model_run_summary.json',
+                    claim_class='观澜自算(非振动线正裁)')
+    if not len(cand) and f38.is_file():
+        d = pd.read_csv(f38, dtype=str)
+        if '融合' in d.columns:
+            bad = d[d['融合'].astype(str).str.contains('候选|定论|准定论', na=False)]
+            if len(bad):
+                named = [f"{r['台']}·{r['融合']}" for _, r in bad.head(5).iterrows()]
+                add('RD-LEDGER-VIB-01',
+                    f"振动面融合级: {len(bad)} 台达「候选」及以上({', '.join(named[:3])})",
+                    '候选', [dict(file=P.rel(f38), sha16=_sha16(f38))], named=named, systems=['振动面'],
+                    module='振动·六层链融合', window_days=None, n_machines=int(len(d)),
+                    limitation='机制未定 ⇒ 不命名部件; 无正样本锚 ⇒ 封顶「候选」。',
+                    falsify='复测同一判据回落或拆检未见损伤即撤回。',
+                    nature='趋势性(证据面)', coverage_note='融合表逐台一行',
+                    claim_class='观澜自算融合(非振动线正裁)')
+
+    return out
+
+
+def build(farm: str | None = None, dry=False) -> int:
+    root = P.out_root(farm)
+    store, m5 = P.store(farm), P.m5(farm)
+    cl = claims(store, m5)
+    if not cl:
+        print('[X] 一条 claim 也没算出来 —— 台账产物不在位? 先跑重算链(②三门台账 / ③SCADA 侧)')
+        return 2
+    if dry:
+        for c in cl:
+            print(f'   [{c["verdict"]:6}] {c["id"]:20} {c["title"][:70]}')
+        print(f'   (dry-run, 共 {len(cl)} 条;未写文件)')
+        return 0
+    t0 = time.strftime('%Y-%m-%d %H:%M:%S')
+    findings = dict(
+        schema='guanlan-sop-findings/v1',
+        generated_by='scripts/sop_findings_from_ledger.py (用户令 2026-09-19: 从重算台账生成 claim)',
+        built=t0, n=len(cl),
+        note='本件由**观澜自算**(重算台账 → claim), 不含人工裁决; 逐条 source_refs 记件与 sha16, '
+             'mandatory_limitation/falsifiability 如实写明缺证据与证伪判据。',
+        findings=[dict(
+            id=c['id'], title=c['title'], verdict=c['verdict'],
+            coverage=dict(window_days=c['window_days'], window_days_note=c['coverage_note'],
+                          per_day_evidence=None, n_machines=c['n_machines']),
+            claim_class=c['claim_class'], module=c['module'],
+            mandatory_limitation=c['mandatory_limitation'], falsifiability=c['falsifiability'],
+            problem_nature=c['problem_nature'], source_chapter='重算台账',
+            evidence=[c['title']], named=c['named'], systems=c['systems'],
+            source_refs=c['source_refs']) for c in cl])
+    p1 = root / 'sop' / 'findings.json'
+    p1.parent.mkdir(parents=True, exist_ok=True)
+    p1.write_text(json.dumps(findings, ensure_ascii=False, indent=1), encoding='utf-8')
+
+    base = root / 'paradigm_r1' / 'experiments'
+    # E5 底稿: 模块映射(契约的 scope.module / module_map_how)
+    (base / 'E5_coverage_table').mkdir(parents=True, exist_ok=True)
+    (base / 'E5_coverage_table' / '底稿.json').write_text(json.dumps(dict(
+        schema='guanlan-paradigm-e5/v1', generated_by='scripts/sop_findings_from_ledger.py (自算)',
+        note='模块映射: 每条 claim 归到哪个分析模块(观澜自算口径)',
+        finding_rows=[dict(i=i, module=c['module'], map_how='ledger:' + c['module']) for i, c in enumerate(cl)]),
+        ensure_ascii=False, indent=1), encoding='utf-8')
+    # E3 底稿: 归宿提议(契约的 closure.machine_proposed / human=空)
+    (base / 'E3_candidate_closure').mkdir(parents=True, exist_ok=True)
+    (base / 'E3_candidate_closure' / '底稿.json').write_text(json.dumps(dict(
+        schema='guanlan-paradigm-e3/v1', generated_by='scripts/sop_findings_from_ledger.py (自算)',
+        note='归宿提议: 机器只提"下一步该做什么", 人裁字段留空(不在本器职责内)',
+        rows=[dict(i=i, human=None,
+                   proposed=('复测同一判据并核实物证据' if c['verdict'] == '候选' else '例行跟踪, 不单独排程'))
+                  for i, c in enumerate(cl)]), ensure_ascii=False, indent=1), encoding='utf-8')
+    # E8 底稿_s0: 系统/台号(契约的 scope.systems/units 与 sample_level.named)
+    (base / 'E8_rudong_rebuild').mkdir(parents=True, exist_ok=True)
+    (base / 'E8_rudong_rebuild' / '底稿_s0.json').write_text(json.dumps(dict(
+        schema='guanlan-paradigm-e8-s0/v1', generated_by='scripts/sop_findings_from_ledger.py (自算)',
+        note='系统/台号: 由 claim 的点名台与系统列直出',
+        findings_rows=[dict(i=i, n_tid=len(c['named']), tids=[n.split('(')[0].split('·')[0] for n in c['named']],
+                            systems=c['systems'], units=c['named']) for i, c in enumerate(cl)]),
+        ensure_ascii=False, indent=1), encoding='utf-8')
+
+    print(f'已写 {P.rel(p1)}({len(cl)} 条 claim)+ 3 份底稿(E5/E3/E8_s0)')
+    for c in cl:
+        print(f'   [{c["verdict"]:6}] {c["id"]:20} {c["title"][:64]}')
+    try:
+        from src import derived_manifest as DM
+        rels = {f: 'scripts/sop_findings_from_ledger.py (重算台账 → claim; 观澜自算)'
+                for f in OUT_FILES}
+        DM.record(root, rels, by='sop_findings_from_ledger')
+        print('   已自登记 → _derived_manifest.json')
+    except Exception as e:
+        print(f'   [i] 自登记跳过: {type(e).__name__}: {e}')
+    return 0
+
+
+def status(farm: str | None = None) -> int:
+    root = P.out_root(farm)
+    p = root / 'sop' / 'findings.json'
+    if not p.is_file():
+        print(json.dumps(dict(path=P.rel(p), exists=False,
+                              note='不在位 —— 门户结论段/问答/报告面会 503 或停在旧数; 跑本器可由重算台账生成'),
+                         ensure_ascii=False))
+        return 0
+    try:
+        d = json.loads(p.read_text(encoding='utf-8'))
+    except Exception as e:
+        print(json.dumps(dict(path=P.rel(p), exists=True, err=str(e)[:120])))
+        return 1
+    print(json.dumps(dict(path=P.rel(p), exists=True, n=d.get('n'), built=d.get('built'),
+                          generated_by=d.get('generated_by'),
+                          kind=('观澜自算' if 'sop_findings_from_ledger' in str(d.get('generated_by')) else '人裁底稿')),
+                     ensure_ascii=False))
+    return 0
+
+
+def main() -> int:
+    ap = argparse.ArgumentParser(description='事实契约的 claim 生成端(重算台账 → findings/底稿)')
+    ap.add_argument('--farm', default=None)
+    ap.add_argument('--dry-run', action='store_true')
+    ap.add_argument('--status', action='store_true')
+    a = ap.parse_args()
+    if a.status:
+        return status(a.farm)
+    return build(a.farm, a.dry_run)
+
+
+if __name__ == '__main__':
+    for _s in (sys.stdout, sys.stderr):
+        try:
+            _s.reconfigure(errors='replace')
+        except Exception:
+            pass
+    sys.exit(main())