Jelajahi Sumber

交付包 v0.4.0: 打包不含 输入数据/产物/日志 (用户令 2026-09-16)

zhouyang.xie 3 minggu lalu
induk
melakukan
d7e253bfc0
56 mengubah file dengan 3733 tambahan dan 1455 penghapusan
  1. 26 17
      README_先读我.txt
  2. 0 323
      _修复记录_20260911/README.md
  3. 0 249
      _修复记录_20260911/fix_guanlan_bom.py
  4. 0 140
      _修复记录_20260911/fix_guanlan_pythonpath.py
  5. 0 84
      _修复记录_20260911/fix_guanlan_startbat.py
  6. 0 67
      _修复记录_20260911/restore_release_rudong.py
  7. 0 50
      _修复记录_20260911/scan_portal_links.py
  8. 0 44
      _修复记录_20260911/verify_bug2.py
  9. 0 61
      _修复记录_20260911/verify_guanlan_simsys.py
  10. 238 0
      docs/振动数据接入_v0.1.md
  11. 69 9
      docs/数据目录结构与落位约定_v0.2.md
  12. 3 2
      docs/移植与独立运行_v0.1.md
  13. 294 0
      docs/系统设计说明.md
  14. 4 3
      docs/说明书_观澜如东样板v2_v0.2.md
  15. 48 41
      docs/重算操作手册_v0.1.md
  16. 24 9
      docs/重算缺口与补件清单_v0.1.md
  17. 47 3
      guanlan.py
  18. 1 1
      release/portal_src/README.md
  19. 5 4
      release/portal_src/manifest.json
  20. 4 4
      release/portal_src/shell.html
  21. 7 7
      scripts/_ops_launch.py
  22. 9 1
      scripts/_ops_run.py
  23. 4 1
      scripts/_ops_start_and_open.py
  24. 20 1
      scripts/check_transferable.py
  25. 8 3
      scripts/guanlan_facts_contract.py
  26. 14 4
      scripts/guanlan_gateway.py
  27. 86 63
      scripts/guanlan_ops.py
  28. 150 0
      scripts/guanlan_start_hidden.py
  29. 17 4
      scripts/ingest_ops_2025.py
  30. 349 0
      scripts/inventory_products.py
  31. 61 21
      scripts/pack_dist.py
  32. 133 33
      scripts/place_raw_data.py
  33. 37 0
      scripts/portal_build.py
  34. 78 13
      scripts/products_restore_missing.py
  35. 55 130
      scripts/products_state.py
  36. 33 2
      scripts/rebuild_all.py
  37. 383 0
      scripts/rudong_tcm_index.py
  38. 334 0
      scripts/rudong_tcm_spectra.py
  39. 271 0
      scripts/vib_raw_build.py
  40. 357 0
      scripts/vib_reports_build.py
  41. 113 27
      scripts/windscada_serve.py
  42. 82 0
      src/derived_manifest.py
  43. 21 16
      src/ontology/maintenance.py
  44. 18 0
      src/paths.py
  45. 156 0
      src/proc.py
  46. 29 2
      src/sop/wrapup.py
  47. 11 0
      src/windcms/config.py
  48. 34 4
      src/windcms/data.py
  49. 30 1
      src/windcms/pipeline.py
  50. 5 1
      src/windcms/plugins.py
  51. 8 3
      src/windscada/config.py
  52. 2 1
      src/windscada/subsys/fusion.py
  53. 8 4
      src/windscada/subsys/pitch.py
  54. 2 1
      src/windscada/taxonomy.py
  55. 45 1
      start.bat
  56. TEMPAT SAMPAH
      start_hidden.vbs

+ 26 - 17
README_先读我.txt

@@ -1,8 +1,12 @@
-观澜·如东样板 v2 0.2.0 — 单包全量离线应用 (Windows / Linux / macOS)
+观澜·如东样板 v2 0.4.0 — 单包全量离线应用 (Windows / Linux / macOS)
 
 
 <PY> = Windows 的 .venv\Scripts\python.exe, 或 Linux·macOS 的 .venv/bin/python;
 <PY> = Windows 的 .venv\Scripts\python.exe, 或 Linux·macOS 的 .venv/bin/python;
        命令里路径写 / 即可两平台通用; Windows 也可以只双击 .bat, 不必敲命令。
        命令里路径写 / 即可两平台通用; Windows 也可以只双击 .bat, 不必敲命令。
 
 
+★ 本包**不含 输入数据(data/) · 产物(outputs/) · 日志(logs/, run/)**(用户令 2026-09-16):
+  解压安装后页面会显示"无产物"; 把现场数据包放好后跑一次重算(第四节), 页面立刻有数。
+  随包内容 = 程序/配置/门户与三维资产/交付文档/离线依赖轮子/便携运行时 (没有数据与产物)。
+
 一、安装 (约 10 分钟; 安装目录不要有空格和中文)
 一、安装 (约 10 分钟; 安装目录不要有空格和中文)
   Windows : 解压到 D:\guanlan\app → 双击 install.bat        (自动选 Python → 建 .venv → 包内轮子离线装依赖)
   Windows : 解压到 D:\guanlan\app → 双击 install.bat        (自动选 Python → 建 .venv → 包内轮子离线装依赖)
   Linux   : 解压到 /opt/guanlan  → sh install.sh            (麒麟 V10 / 统信 UOS 20 / Ubuntu / CentOS / macOS;
   Linux   : 解压到 /opt/guanlan  → sh install.sh            (麒麟 V10 / 统信 UOS 20 / Ubuntu / CentOS / macOS;
@@ -17,26 +21,31 @@
   Linux   : <PY> guanlan.py check → <PY> guanlan.py serve           停止: <PY> guanlan.py stop
   Linux   : <PY> guanlan.py check → <PY> guanlan.py serve           停止: <PY> guanlan.py stop
   三个入口: 门户 http://127.0.0.1:28084/ · 工作台 /detail/ · 运维控制台 /ops
   三个入口: 门户 http://127.0.0.1:28084/ · 工作台 /detail/ · 运维控制台 /ops
 
 
-三、运维控制台 /ops — 一个按钮一件事: 停/启服务 · 执行重算 · 清除产物
-  按钮按真实状态启用(不能做的灰, 后端也拒绝, 直接打 API 返回 409); "启动服务"起来后自动打开门户;
-  "停止组件服务"保留控制台本身; 同时刻只允许一个动作, 页面上有进度/退出码/日志尾巴。
-  清除产物 = 挪到 _products_off(可恢复), 不动 data/raw、release、logs。(控制台与重算链只用标准库, 两平台同源)
+三、重算与产物 (门户菜单「数据重算」, 或直接开 http://127.0.0.1:28084/ops)
+  一个按钮一件事: 停/启服务 · 执行重算 · 清除产物。按钮按真实状态启用(不能做的灰, 后端也拒绝,
+  直接打 API 返回 409); "启动服务"起来后自动打开门户; "停止组件服务"保留控制台本身; 同时刻只允许
+  一个动作, 页面上有进度/退出码/实时日志尾巴。重算进行中按钮按设计禁用 —— 那不是坏了。
+  ★ 清除产物 = **直接删除, 不留备份**(用户令 2026-09-16; 旧的 --on 还原已移除), 不动 data/raw、release;
+    要补回"包内没有生成端"的随包件: <PY> scripts/products_restore_missing.py --stash <交付包.zip>
 
 
-四、换数据后重算 (一条命令; 控制台点"执行重算"等价)
+四、放数据后重算 (一条命令; 控制台点"执行重算"等价)
   <PY> scripts/place_raw_data.py --src <现场包目录> --scope full      放数据(同尺寸自动跳过, 可反复跑)
   <PY> scripts/place_raw_data.py --src <现场包目录> --scope full      放数据(同尺寸自动跳过, 可反复跑)
   <PY> scripts/rebuild_all.py [--skip-scada] [--src <现场包目录>]     默认含 SCADA 侧约 15 分钟;
   <PY> scripts/rebuild_all.py [--skip-scada] [--src <现场包目录>]     默认含 SCADA 侧约 15 分钟;
        只换台账类数据加 --skip-scada(约 2 分钟); --dry-run 只看计划; 完事自动重启组件服务。
        只换台账类数据加 --skip-scada(约 2 分钟); --dry-run 只看计划; 完事自动重启组件服务。
   算成功判据: /detail/ 左栏「系统维护」两屏, 或 <PY> -m src.ontology.maintenance
   算成功判据: /detail/ 左栏「系统维护」两屏, 或 <PY> -m src.ontology.maintenance
-  本包锚点: 报警 39211 · 工单 5876 · 油样 404 · temp_monthly 19494 · 本体 9702 对象(审计 0 问题)
+  重算后应达锚点: 报警 39211 · 工单 5876 · 油样 404 · temp_monthly 19494 · 本体 9702 对象(审计 0 问题)
 
 
 五、必读 (如实)
 五、必读 (如实)
-  · 页面的数来自 outputs/rudong 下的产物, 产物由 data/raw 下的原始件算出来(下一级目录名=场站名, 扫描辨识)。
-  · 随包里有一批产物**没有生成端**(pitch_daily、pc_monthly_bins、CMS/TCM 链等), 重算后由随包件补齐;
-    逐件来源见 outputs/rudong/_provenance.json (raw-derived 20 件 / shipped 568 件)。
-  · 某页没数据: 先看控制台「最近一次动作」的退出码与 logs/, 再看 docs/重算缺口与补件清单_v0.1.md。
-
-包内含: 程序与配置 / 如东分析产物(含 _provenance.json 来源台账) / 门户与仿真 / 三维资产 release/viewer /
-        治理清单交付件 release/如东(客户交付物勿外传) / 离线依赖轮子 / 便携运行时
-
-详细说明: docs/说明书_观澜如东样板v2_v0.2.md · docs/重算操作手册_v0.1.md
-          docs/数据目录结构与落位约定_v0.2.md · docs/重算缺口与补件清单_v0.1.md · _修复记录_20260911/README.md
+  · 页面的数来自 outputs/<场站> 下的产物, 产物由 data/raw 下的原始件算出来(下一级目录名=场站名, 扫描辨识)。
+    本包两样都没带: 没有 data/raw 时重算会"没有源件可算", 没有 outputs/ 时页面显示"无产物"。
+  · 有一批产物**没有生成端**(pitch_daily、pc_monthly_bins、CMS/TCM 链等): 从交付包按需补齐
+    (products_restore_missing.py --stash <交付包.zip>), 补齐后逐件来源见 outputs/<场站>/_provenance.json。
+  · 某页没数据: 先看控制台的退出码与 logs/, 再看 docs/重算缺口与补件清单_v0.1.md。
+
+包内含: 程序与配置 / 门户与仿真 / 三维资产 release/viewer / 治理清单交付件 release/如东(客户交付物勿外传) /
+        离线依赖轮子 wheels/ / 便携运行时 vendor/ / 交付文档 docs/
+包内不含: data/(输入数据) · outputs/(产物) · logs/ run/(日志) · .venv(安装时重建) · .git(版本库)
+
+详细说明: docs/系统设计说明.md · docs/重算操作手册_v0.1.md
+          docs/数据目录结构与落位约定_v0.2.md · docs/重算缺口与补件清单_v0.1.md
+          docs/振动数据接入_v0.1.md(振动侧 CMS/TCM 落位与摄入) · docs/说明书_观澜如东样板v2_v0.2.md

+ 0 - 323
_修复记录_20260911/README.md

@@ -1,323 +0,0 @@
-# 观澜·如东样板 v2 (v0.2.0) 启动报错修复记录 — 2026-09-11
-
-本目录**不属于原始交付包**, 是现场修完之后留下的记录: 每个脚本都可重复运行 (幂等),
-用来复现 "改了什么" 或在别的机器/另一份拷贝上重放同一批修复。
-
-## 现场现象
-
-```
-powershell -ExecutionPolicy Bypass -File install.ps1     -> 装完最后一步崩:
-    json.decoder.JSONDecodeError: Unexpected UTF-8 BOM (decode using utf-8-sig)
-start.bat                                                -> 同样崩; 随后
-    [X] Server did not start (reason above). Not opening the browser
-```
-
-## 根因 (5 条, 前 3 条是同一个坑的不同后果)
-
-| # | 位置 | 问题 |
-|---|------|------|
-| 1 | `install.ps1` 第 4 步 `Set-Content configs\serve.json -Encoding UTF8` | **Windows PowerShell 5.1 的 `-Encoding UTF8` 写出带 BOM 的 UTF-8**; `guanlan.py` 用裸 `utf-8` 读 -> `JSONDecodeError`。启动器在解析配置时就死了, 所以 `check`/`serve` 全跑不起来。 |
-| 2 | 同一行的 `Get-Content configs\serve.json -Raw` | PS 5.1 对**无 BOM** 的 UTF-8 按 ANSI 代码页 (中文机 = GBK) 解码, 于是把文件里的中文注释读成乱码, 再 `Set-Content` 写回 -> `_note` / `_note_raw` 两个字段永久损坏 (部分字节被替换成 `?`, 不可逆)。 |
-| 3 | 机器全局 `PYTHONPATH=D:\Program Files\Python\Lib\site-packages;` | ① pip 认为系统那套包 "已满足", **没把传递依赖装进 `.venv`** -> venv 缺 `urllib3`/`polars`/`pyyaml`/`jinja2`/`python-dotenv` 等, 换个没设 PYTHONPATH 的 shell 直接 `ModuleNotFoundError`; ② `PYTHONPATH` 排在 sys.path 前面, 把 pin 住的版本顶掉 (实测 `requests` 2.33.0 顶掉 2.34.2)。 |
-| 4 | `wheels\win_amd64\` | 缺 `colorama` 轮子 (`tqdm` 在 Windows 上的依赖)。之前被根因 3 掩盖: pip 看到系统已有 colorama 就不下载, 于是离线轮子集不全; 一旦 PYTHONPATH 摘干净, `pip install --no-index` 直接 `ERROR: No matching distribution found for colorama`。 |
-| 5 | `release/sim_sys_server.py` 第 36 行 `from scrub_rules import scrub, residual` | **`scrub_rules.py` 根本没随包发出** (全盘搜不到, zip 里也没有) -> `/sim/sys/ 仿真·四系统合页` 502, `logs/sim_sys.log` 里是 `ModuleNotFoundError`; 网关因此 `degraded`, `guanlan.py serve` 退出码 1。 |
-
-## 改了什么
-
-| 文件 | 改动 |
-|------|------|
-| `guanlan.py` | 新增 `jload()` = `utf-8-sig` 读 JSON (兼容有/无 BOM), 5 处配置读取全部改走它; 配置坏了只打印提示并退回内置默认端口, 不再抛栈。新增 `_hermetic()`: 启动时把 `PYTHONPATH` 从 `sys.path` 与环境里摘掉 (子进程继承干净环境)。 |
-| `install.ps1` | 第 4 步不再用 PS 的 cmdlet 碰配置, 改由 `.venv` 里的 Python 读写 (读 `utf-8-sig` / 写无 BOM 的 UTF-8), 并加注释说明为什么不能改回去; 顶部加 `Remove-Item Env:PYTHONPATH`, 让 pip 老老实实把包装进 `.venv`。文件仍保持 **UTF-8 带 BOM + CRLF** (PS 5.1 解码中文的前提)。 |
-| `install.sh` | 同样加 `unset PYTHONPATH` (保持 LF/无 BOM)。 |
-| `configs/serve.json` | 去掉 BOM; 还原 `_note` / `_note_raw` 两段中文注释 (依据残留可逆部分 + 说明书 §7 的措辞)。 |
-| `wheels\win_amd64\colorama-0.4.6-py2.py3-none-any.whl` | 补上缺失的轮子, 离线安装才完整。 |
-| `release/scrub_rules.py` | **恢复缺失文件**: 不新写任何脱敏规则, 只把包内唯一那份规则表 `src/windscada/deid_public.py` 转出 `scrub` / `residual`(= `audit`) / `scrub_or_die`。若拿到交付方原版, 直接覆盖。 |
-| `start.bat` | `guanlan.py serve` 的退出码 1 有两种含义 (网关没起来 / 网关起来了但有模块降级)。现在先探一次 `/healthz` 区分: 真没起来才报 `[X]` 并停下; 起来了但有降级则打 `[!]` 说明, 浏览器照常打开。 |
-
-## 备份与回滚
-
-原文件都留了副本, 直接改名覆盖即可回到改前状态:
-
-```
-guanlan.py.bak-bomfix      install.ps1.bak-bomfix     configs\serve.json.bak-bomfix
-guanlan.py.bak-pyfix       install.ps1.bak-pyfix      install.sh.bak-pyfix
-start.bat.bak-healthz
-```
-
-## 重放 / 验证
-
-```
-.venv\Scripts\python.exe fix_guanlan_bom.py          # 1 2 (BOM + 乱码)
-.venv\Scripts\python.exe fix_guanlan_pythonpath.py   # 3
-.venv\Scripts\python.exe fix_guanlan_startbat.py     # start.bat 判定
-.venv\Scripts\python.exe verify_guanlan_simsys.py    # 5 个仿真页 200 + 脱敏回扫干净
-```
-
-三个 `fix_*` 都是幂等的 (已改过会打印 "已修过")。`wheels\colorama*.whl` 与
-`release\scrub_rules.py` 属于新增文件, 没有对应的 `fix_*` 脚本。
-
-## bug-2 · `#documents` 页「如东液压系统治理清单 · 网页版」报 `{"err":"not found"}`
-
-**现象**: 门户 `http://127.0.0.1:28084/#documents` 里那个 iframe 显示
-`{"err": "not found", "path": "/如东/如东治理清单_交付_20260901/02_治理清单/如东_液压系统治理清单_v1.2_2026-09-01.html"}`。
-
-**根因**: v0.2.0 发行包里**没有 `release/如东/`** —— 0.2.0 的 zip 里 `release/` 下只有 `viewer/`,
-而门户那个 iframe 指的就是 `release/如东/.../_液压系统治理清单_v1.2_2026-09-01.html`;
-网关 `guanlan_gateway.py::_static()` 在 `release/` 里找不到文件就回那段 404 JSON
-(见 `_static()` 里 `return self._json(dict(err="not found", path=path), 404)`)。
-顺带核过: 门户里指向 `release/如东/` 的引用**只有这一条**, 所以不是路径写错, 是交付件没随 v0.2.0 发出来。
-
-**修复**: 从本机的 v0.1.0 全量包 `<临时目录>/guanlan-rudong-v2_0.1.0_all.zip`(这份交付件在里面是齐的)
-把整个 `release/如东/` 恢复到 v0.2.0 安装目录 —— 75 个文件 / 51.0 MB
-(`如东治理清单_交付_20260901/` 33 件、`如东取数单_2026-08-21/` 18 件、`观澜离线版.app/` 10 件、根下散件 14 件)。
-`release/如东/` 按 `.gitignore` 的既有约定不入库(客户交付物只在受保护工作区), 所以 git 里看不到这次恢复。
-
-**验收**: iframe URL 走网关取 → **[200] text/html, 103760 字节**, 标题 `如东液压系统治理清单`,
-正文 0 个外部相对引用(自包含); 同页另外两个引用 (`127.0.0.1:18020` CMS、`127.0.0.1:18033/v2`) 均 200;
-全门户扫一遍含「如东 / release」的静态引用 → 坏链 0。
-
-脚本: `restore_release_rudong.py` (恢复) · `verify_bug2.py` (验收) · `scan_portal_links.py` (同类坏链全量扫描)
-
-
-
-## 路径跨平台化(①+②, 2026-09-11)
-
-**用户令**: 后续部署到其它电脑; 兼容 Windows / Linux; **系统内目录路径必须使用相对路径**。
-
-**三条口径**(写进 `docs\数据目录结构与落位约定_v0.2.md` §7, 并由门禁自动检查):
-① 只写相对路径(相对**安装根**), 不写机器绝对路径;② 相对基准是安装根、**不是 cwd**, 一律经 `src\paths.py` 解析;
-③ 写进产物/清单/页面的用 POSIX 相对形式(`P.rel()`), 显示给人看时才用本机分隔符(`P.disp_dir()`)。
-
-**新增**
-- `src\paths.py` —— 唯一路径真源(`ROOT/RAW_ROOT` + `store()/ont()/cms()/m5()/sop()/guanlan()/pitch()/tcm_replay()`
-  + `RELEASE/PORTAL/VIEWER/SIM_DIR/LOGS/RUN` + `rel()/disp()/disp_dir()/resolve()/venv_python()`)。
-  `ROOT` = env `WINDSCADA_ROOT` → 否则按本文件位置回溯 ⇒ **整个安装目录可整拷到别的电脑/别的盘**。
-- `scripts\check_portability.py` —— 静态门禁: 机器绝对路径 / cwd 相对字面量 / `os.sep` 拼库存字符串 = ERROR;
-  例外必须具名登记理由。**当前通过**。
-- `scripts\page_fingerprint.py` —— 10 个端点的回归指纹(`status` + 归一化 sha256, 自动剔除 `checked/ms/repo_head`
-  这类设计上会变的字段、并把安装根前缀归一成 `<ROOT>`)。换机验收也用它: 基准机 `--save`, 新机 `--diff`。
-
-**清理的机器绝对路径**(换机必失配, 且是静默故障): windcms `MAIN_ROOT` 默认 mac 检出、windcms/ontology 写死的
-mac venv 解释器、`kb_ingest` 的 mac 技术资料路径、`ingest_ops_2025` 的 `/Volumes/BIG/…`、`place_raw_data` 的
-`<临时目录>/…`、界面占位符里的 `/Users/…`。**清理的 cwd 相对路径**: 20+ 处 `Path('outputs/rudong/…')`(ontology 14 个
-模块、windscada 4 个、windcms、serve/gateway、sop 8 个)全部归到 `paths`; `sop` 里写死的 `rudong` 改为跟场走。
-
-**配置/安装器**: `configs\serve.json` 的 `python` 改为相对(`.venv/Scripts/python.exe`); 启动器 `py()` 支持
-"相对→按安装根解析 / 绝对 / 留空自动探测 venv"; `install.ps1|sh` 写配置时写相对路径。
-
-**验证**
-- 门禁: 运行期 0 处机器路径、0 处 cwd 相对, `os.sep` 只留在显示函数里 ✔
-- 回归指纹: 改造前后 **10/10 端点逐字节一致** ⇒ 只换了路径解析, 页面与取数面零变化 ✔
-- 换 cwd 实测: 从 `C:\` 且不设 `PYTHONPATH` 跑 `scan_stations.py` / `farm()` → 场站辨识与 `store/src_10min` 均正确 ✔
-- 过程中我自己引入并修掉两处: 服务端 `from src import paths` 插在 `sys.path` 引导之前(导致 detail 起不来),
-  以及 `subsys` 子包导入深度写成两点(应为三点)。这两处都是**靠"重启+指纹比对"逮到**的 —— 印证了那把尺子的价值。
-
-
-
-## 场站扫描辨识 + 两个验收场景 (2026-09-11)
-
-**用户令**: 离线数据统一放 `data\raw\`, **下一级目录约定为「场站名称」**, 观澜系统**扫描辨识**。
-
-- `src/windscada/config.py`: 新增 `scan_stations()` / `station_scan()` / `station_report()`; `_expand()`
-  按扫描结果**派生** `src_*`(外场配置若显式给了 `src_*` 则不覆盖)。辨识规则: ①`raw_station` 全等
-  → ②与 `src_farm_names` 互为子串 → ③只有一个场站目录(单站部署)→ ④都不中 = 不猜, 报"未识别"。
-- `scripts/scan_stations.py`: 打印扫到的场站、辨识依据(`how`)、五个约定子目录的件数与存量。
-- 维护页新增一行「场站数据目录 (扫描辨识)」: 扫到 → `覆盖=如东/存在=True`; 扫不到 → `存在=False/覆盖=—`。
-- `scripts/windscada_serve.py`: `adhoc_query` 在原始件缺失时不再抛给 HTTP 处理器(前端只看到 500),
-  改为 **HTTP 200 + 结构化无数据**(`err=no_source` + "本机应在 …\scada_10min\WTG01.csv" 的提示)。
-
-**场景① 清空 `data\raw` + 挪开产物** → 扫描"0 个场站目录 / how=none"; 维护页五项全部 `条数=0 / 存在=False`;
-实时接口从 500 变 200+`no_source`。 ✅
-
-**场景② 按约定放回** → 扫描"1 个场站目录 如东(592 件), how=raw_station", 子目录 38/16/134/404;
-摄入后 报警 **0→39211**、工单 **0→5876**、油样 **0→404**、loss_monthly 3729(powercurve 31s + availability 88s);
-维护页五项覆盖区间全部回来, 实时接口 `months=7`。 ✅
-
-**产物等价**(从 raw 重算 vs 随包基线): `loss_monthly` 3729/3729、`powercurve_dev` 38/38、
-`powercurve_bins` 912/912、`alarms` 39211/39211 **逐值完全一致**; `workorders` 5876 行仅 1 条记录
-**纳秒尾差**(同一时刻 `.0000015` vs `.0000010`); `oil_samples_index` **506→404** —— 清空重建会丢
-**102 行华标 2026-07 批**(源件 `BG-2026-07-YP013 …pdf` 不在现场数据包里, 且一份覆盖多台), 已恢复完整件
-并写入文档 §4 作为缺口。
-
-文档: `docs\数据目录结构与落位约定_v0.2.md`(目录结构、落位约定、消费方式、可重建性边界、两个场景实测)。
-
-## 数据链补齐 · 让系统真正吃 data/raw/如东 (2026-09-11)
-
-**问题**: 说明书 §11 写着「不含从原始数据重生成产物的链 (P0)」, 而维护页四条「摄入命令」指向的文件
-**在包里一个都不存在**; 于是现场把新台账放进 `data/raw/如东/…` 也不会更新, 工单台账一直停在 2024-11-21。
-
-**新增 4 个脚本** (都在 `scripts/`; 维护页四条命令现在全部「命令可用=True」):
-
-| 脚本 | 源 → 产物 | 关键实现 |
-|---|---|---|
-| `windscada_alarms_ingest.py` | `故障报警/*.xls` → `alarms.parquet` | `.xls` 实为 SpreadsheetML(XML); 累计快照件自动跳过 (2025年全年 = Q1..Q4 精确并集 20155 行) |
-| `windscada_workorder_ingest.py` | `风机故障记录/**` → `workorders.parquet` | 按场名过滤 (集团表里混着民勤/来福/宝力格等十来个场); 归并跨表重复 1085 行; 日期按 `src/sop/ledger_dates.py` 同一纪律解析, 解析不了的整行保留 |
-| `windscada_watch_channels_build.py` | `油样报告/**/*.pdf` → `oil_samples_index.parquet` | 文件名解析 (日期/台号/部件/sample_id); 合并语义: 认不出的老行原样保留 |
-| `rebuild_from_raw.py` | 总入口 | `--verify` 等价验收; `--scada` 再跑包内 10 个 SCADA 侧构建器 |
-
-**等价验收** (`python scripts/rebuild_from_raw.py --verify`, 对 `_pre_rebuild_20260911/` 随包基线):
-
-```
-报警事件       39211 行 / 39211 行   逐值完全一致            ✔
-油液化验         506 行 /   506 行   逐值完全一致            ✔
-检修工单台账    1574 种内容复现 1566 种; 8 种差异逐条定位:
-   · 7 种 = 复位运行时间: 随包 NA、本链有值 — 旧链没认源列名「复位时间时间」这个错别字
-   · 1 种 = 随包把 Excel 序列号当纳秒解析成 1970-01-01, 本链解析为 2023-05-13
-   · 随包 78 行副本重复已归并; 另多出 69 行 = 旧链丢弃的日期不可解析行
-```
-
-**效果** (`/detail/v2#tab=system` 数据层, 已重启服务实测):
-
-| 行 | 改前 | 改后 |
-|---|---|---|
-| 检修工单台账 | 2020-01-03 ~ 2024-11-21, 1652 行 | **2020-01-03 ~ 2026-07-08, 5876 行** |
-| 四条「摄入命令」命令可用 | False | **True** |
-
-**依赖**: `requirements.txt` 加 `xlrd==2.0.2` (2023/2024 年总表是真 BIFF `.xls`, openpyxl 读不了),
-轮子已放进 `wheels/win_amd64/` (`wheels/` 按 .gitignore 不入库, 随包分发)。
-
-**SCADA 侧**: `python scripts/rebuild_from_raw.py --scada` 把 10 个构建器按依赖顺序跑一遍
-(powercurve → loss_monthly → curves/control/stop_events/温度/偏航/液压/热链/系统辅助)。
-抽验过等价性 (WTG01/WTG02 的 `temp_bins` 与随包件 1296/1296 行逐值相同、med 差 0.0),
-但**未整体重跑** —— 要跑建议先备份 `outputs/rudong/windscada/` 再逐项比对覆盖区间。
-
-## 门户拆包 + 闸门退出码 (2026-09-11 收尾)
-
-**起因**: 有人问 "`release/portal.html` 20 MB 里到底是什么、能不能不进 git"。查下来 99.3% 是内嵌的
-交付文档正文 (28 个 `<template>`: 治理清单分册、单机/整机报告、仿真台面板), 只有 138 KB 是门户自己的壳。
-原先整份文件入库, 每轮重算产物都往 git 塞 20 MB。
-
-**做法 (用户选 ②)**: 拆成"受管外壳 + 不入库内嵌件 + 可验证装配":
-
-| | 入库 | 说明 |
-|---|---|---|
-| `release\portal_src\shell.html` | ✅ | 138,544 B, 内嵌件处留 `<!--@TEMPLATE:id-->`, 资料索引处留 `<!--@GOVERNANCE_SOURCES-->` |
-| `release\portal_src\manifest.json` | ✅ | 各件 sha256 + 期望门户 sha256 + 行尾约定 |
-| `release\portal_src\README.md` | ✅ | 重建说明与两个坑 |
-| `release\portal_src\templates\` | ❌ 产物 | 28 件 20.04 MB (交付件正文) |
-| `release\portal_src\governance_sources.json` | ❌ 产物 | 1,512 B 脱敏资料索引 |
-| `release\portal.html` | ❌ 产物 | 装配结果, 服务/网关照读 |
-
-新增 `scripts\portal_build.py`: `--extract` 拆 / 默认装配 / `--verify` 逐字节比对 / `--check` 漂移检查。
-**实测 `--verify` 两行同为 `sha256 9b6aabeb6ca18d15`, 20,226,052 B —— 拆→装回到同一个文件**; `--check` 全件一致。
-
-**两个坑**:
-
-1. **行尾**: 门户通体 LF (0 处 CRLF)。Python 文本模式在 Windows 上写文件会把 `\n` 变 `\r\n` —— 20 MB
-   整体改写, `/api/version` 的 `portal_sha256` 与页面指纹全变。两个就地注入器
-   (`guanlan_portal_fix_anchors.py` / `guanlan_portal_inject_claims.py`) 补 `newline=""`,
-   `portal_build.py` 全程字节级读写并在装配后自检 CRLF。第一次拆出来的 `portal_src/` 就是这么废的。
-2. **闸门退出码**: 同一类坑让"校验通过"被报成失败 —— 中文控制台代码页 936, `print('… 一致 ✔')`
-   抛 `UnicodeEncodeError` 使进程以 1 退出。`portal_build --verify` 与 `page_fingerprint --diff`
-   各踩一次 (后者是回归闸门: 10 个端点全部回到基线, 却报失败)。新增 `src\console.py` 的 `soft()`
-   把编码错误降级为 `?`, 接入 5 个"以退出码讲话"的脚本; 复跑四个闸门均正确返回 0。
-
-**回归**: 拆包后 10 个端点全部与基线一致 (`portal.home 64b4b81158d09d9c`, `detail.v2 a0abff29c000f570` …)。
-其中一度看到 `/cms/` 掉线 —— 原因是最后那次 `guanlan.py serve` 发生在产物被挪走时, CMS 启动即因
-`src\windcms\data.py` "No objects to concatenate" 退出 (本该如此); 产物还原后重启即恢复
-(414,139 B, 与基线一致)。**教训: 产物开关与重启顺序有关, 挪/还产物后必须重启一遍。**
-
-## 从零重算 (2026-09-11 下午, 用户令: 不要随包产物, 基于 data/raw 重算)
-
-**做法**: 产物全部挪走 (`products_state.py --off`), 只留 `data/raw` 重算, 一个随包产物都不吃。
-
-```bat
-.venv\Scripts\python.exe scripts\products_state.py --off          :: 产物挪走
-.venv\Scripts\python.exe scripts\rebuild_from_raw.py              :: 三门台账
-.venv\Scripts\python.exe scripts\rebuild_from_raw.py --scada      :: SCADA 侧 10 个构建器 (逐台读 14.7 GB CSV)
-.venv\Scripts\python.exe scripts\rebuild_from_raw.py --verify     :: 等价验收 (需临时放回随包基线)
-```
-
-**结果**: 报警 0→39211 行 · 工单 0→5876 行 · 油样 0→404 行; SCADA 侧 10 个构建器全跑通,
-L0 仓 16 件 1.78 MB (随包 118 件 —— 差额就是"包内没有生成端"那些件, 见 docs §4);
-本体层 `objects.json` 2340 个对象 + 检索索引 + `turbine_params.parquet` 1706 条。
-等价验收: **alarms 39211/39211 逐值完全一致**; workorders 8 行差异全部可归类(旧链时间解析缺陷)
-+ 69 行本链新增覆盖; 唯一人工项是油样 506→404。
-
-**用户追问"缺的数据从 <现场包目录> 抽"** → `place_raw_data.py --scope mech` 落 355 件 15.5 GB:
-`data\raw\西门子4.0技术资料\`(317 件, 本体层的源件) 与 `<场站>\scada_1min\`(38 件 12.8 GB)。
-抽完本体层就从"包内没有源件⛔"变成**可重算**。**仍然缺的**: 油样那 102 行的源件
-(`BG-2026-07-YP013 … .pdf`) 不在包里(整个包只有 `2025年油样` 一批) → 那 102 行找不回来。
-
-### 这一轮逮到的三个真 bug (都不是参数问题, 是静默失效)
-
-| # | 症状 | 根因 | 影响面 |
-|---|---|---|---|
-| 1 | `python -m src.ontology.kb_ingest` 崩: `UnicodeEncodeError: 'gbk' codec can't encode '\u200b'` | `src/ontology/store.py` 的 `read_text()/write_text()` 没写 `encoding=` → Python 文本 I/O 用**系统 locale 编码**, 中文 Windows = cp936。读会乱码/解码失败, 写遇到 GBK 编不出的字符就崩 | **本体层在中文 Windows 上根本不可用**(开发机是 UTF-8 locale 所以从没暴露)。已修 `src/ontology/` 全层 10 处; 存量 12 处已进 `check_portability.py` 的新 WARN 项 6 |
-| 2 | `place_raw_data.py` 一跑就 `NameError: name 'os' is not defined` | 模块级 `DEFAULT_SRC` 用了 `os.environ` 却从没 `import os` | **该脚本自 `d90da05`(路径跨平台化) 起一直跑不通**; A3 那次落位不是它干的 → 说明"写成表的落位映射"从未被真正执行过 |
-| 3 | A2 规则 `…/大部件维修记录.20240619143912557.xlsx` 从没落过盘 | 包内前缀恰好等于条目本身时, 去前缀后 `rest` 为空, 旧代码把它当"目录条目"`continue` 掉 | 单文件形态的规则**全都静默丢件**: 受害两例 = 上面那个 25.8 MB 工单源 + 本轮新加的对译表。修后按前缀的文件名落位 |
-
-> 教训: 这三个都是"跑一次就暴露、但没人跑过"的类型。**从零重算本身就是最好的回归测试** ——
-> 它不依赖任何随包产物, 因而能照出所有"靠旧产物遮掩"的失效。
-
-**顺带发现**: 落位后工单源件 79→80 张, 但总行数不变(5876)——新件的行与既有来源内容重复,
-按"累计快照/副本必须归并"的纪律合并, 脚本把这类情况逐条打了出来(不是静默丢)。
-
-**要补什么、找谁补**: 那次重算照出来的缺口整理成了独立清单 ——
-`docs\重算缺口与补件清单_v0.1.md`(按"源件缺失(现场能补)"与"生成端缺失(研发补脚本)"两类,
-每项带证据路径、影响面、补齐判据与验证命令)。
-
-**另一件同类收尾**: 把剩下的 18 处"文本读写缺 `encoding=`"也清了(见下), 门禁 WARN 项 6 归零。
-
-## 文本 I/O 的 encoding (2026-09-11 收尾)
-
-上面 bug ① 只是冰山一角: 全库搜下来共 **28 处**文本 I/O 没写 `encoding=`(本体层 10 + 其余 18)。
-已全部补上 `encoding='utf-8'`, 其中写侧把 `json.dump(x, open(p,'w'))` 这类匿名句柄改成 `with` 块(正确关文件),
-`src\windcms\plugins.py` 里给子进程当 stdout 的那个句柄**不 with-close**(父进程持引用, 免得提前关掉)。
-
-`check_portability.py` 的 WARN 项 6 同时扩了扫描面: 除 `read_text/write_text` 外, 新增
-`open(..., 'w')` 与 `open(p)`(无 mode = 文本读)两种形态(二进制模式不算), 并与 ERROR 扫描共用 `ALLOW` 例外表
-(门禁自己的规则/文档里必须写出这些形态, 整份豁免 —— 否则会误报 10 处)。**实测命中 0。**
-
-冒烟: 11 个文件 `py_compile` 通过 · windcms 五个模块 + fusion/taxonomy/scenario_29 导入成功 ·
-重启服务 5/7 ok · 页面字节与改前逐项一致(`/` 20225828 · `/detail/` 229643 · `/sim/` 50998 ·
-`/sim/sys/` 123455 · `/viewer/` 4986) · 维护页数据层四个数字不变(报警 39211 · 工单 5876 · 油样 404 · 月表 3729)。
-
-## 探活页提速: /healthz 6.5 s → 冷探 1.5 s / 命中缓存 ~10 ms (2026-09-11)
-
-收尾测页面时发现 `/healthz` 要 **6.5 秒**才回来(而 `/` 那个 20 MB 门户只要 0.04 s)。拆开看是三笔叠加:
-
-| 项 | 原实现 | 代价 |
-|---|---|---|
-| 6 个上游探活 | **串行** `for` 循环, 单探针默认 3 s | 4.4 s(其中 CMS / Ollama 两个掉线模块各卡满超时) |
-| Ollama 模型查询 | 探活之后再串一次 | +2 s |
-| 门户 sha256 | **每次** `PORTAL.read_bytes()` 重读 20 MB 再哈希 | 0.1–数十秒(磁盘忙时) |
-
-改法(返回结构一字未改): ① 全部探针**并发**(`ThreadPoolExecutor`, 总耗时 = 最慢那个); ② 单探针 1.5 s;
-③ 门户 sha 按 `(mtime_ns, size)` 缓存(值不变就不重读); ④ 整个结果缓存 5 s —— 连续调用近乎零成本。
-
-实测: 冷探 **6.5 s → 1.53 s**, 命中缓存 **15 ms**; 重启后首探 8.8 ms(`guanlan.py serve` 自己探过一遍,
-缓存已热)。结构逐键比对无差异(顶层键/模块集合/每模块键/`local_ai` 键/`status`/门户 `file_sha256` 全同),
-`/api/version` 的 `portal_sha256` 与磁盘门户逐字节一致(`9b6aabeb…`)。
-
-**★一台机器上的实测结论**(写进代码注释了): 这台机器**连任何已关闭的 `127.0.0.1` 端口都不回 RST**,
-而是把 SYN 丢掉等超时 —— 对照用的随机闭端口 49996/49997 同样 1.5 s 超时, 不是我们端口或进程的问题,
-是系统防火墙/安全软件行为。所以"有模块掉线"时冷探下限就是 `PROBE_TIMEOUT`; 想更快只能靠缓存,
-**别为此把超时压到健康模块也可能被误判的程度**(实测最慢的健康探针 `sim_sys` 345 ms, 其余 <25 ms,
-1.5 s 留了 4 倍余量)。
-
-## 非法转义序列 + 自检加一行 (2026-09-11)
-
-全量编译 `src/` + `scripts/`(168 个文件)把 `SyntaxWarning` 当错误报出来, 逮到最后一处潜伏炸点:
-
-- `src\windcms\serve.py:12,68` —— JS 正则写在普通字符串里: `/\*\*([^*]+)\*\*/g` 与 `/\d+/g`。
-  Python 3.12 只警告, **未来版本会变成 SyntaxError**; 而且开发机上永远看不见, 换解释器才炸。
-- **不能简单改成 raw 字符串**: 同一串里混着 `\\n`(有意转义成 JS 的 `\n`) 与 `\d`(想写成 JS 的 `\d`),
-  改 raw 会把 `\\n` 变成"两个反斜杠 + n", 语义就变了。正确改法是**把非法转义补成合法**: `\*` → `\\*`。
-- **证明没改语义**: 改前/改后各把该文件 237 个字符串常量的求值内容算指纹 —— 都是 `70bd628f72f72101`,
-  JS 文本一字未变; 改后全库非法转义 **0 处**, `-W error::SyntaxWarning` 编译通过。
-
-`guanlan.py check` 顺带加了一行自检: **源码可编译 (168 个文件, 无非法转义/语法错)** —— 换机器/换解释器
-之前先跑它, 这类问题不必等运行到那一页才发现。
-
-> 自检现状(原始重算后): 3 项 FAIL —— `outputs/rudong/sop/findings.json` 与
-> `outputs/rudong/guanlan/facts_contract_v0.json` 不在(缺口清单 A3), 本机 Ollama 无模型(环境项);
-> L0/L1 产物、本体对象库、门户、仿真、viewer、原始件目录都是 OK。
-
-## 仍然存在 (不是代码问题)
-
-- `/local-ai/` 仍是 DOWN: 本机 Ollama 没运行, 且 `%USERPROFILE%\.ollama` 下**没有任何模型**
-  (`manifests` 都没有), 所以问答与本地审核页不可用。其余 6 个页面不受影响。
-  启用: 启动 Ollama, 按 `configs\models.json` 的 `pull_commands` 拉模型 (需联网/大流量),
-  或把有网机器的 `%USERPROFILE%\.ollama\models` 整个目录拷过来。
-- `release\如东` (治理清单交付件) 本来就是可选项, 不影响页面。

+ 0 - 249
_修复记录_20260911/fix_guanlan_bom.py

@@ -1,249 +0,0 @@
-#!/usr/bin/env python3
-# -*- coding: utf-8 -*-
-"""一次性修复: 观澜·如东样板 v2 启动器的 UTF-8 BOM 崩溃 + serve.json 中文乱码。
-
-报错: guanlan.py cfg() -> json.loads(p.read_text(encoding="utf-8"))
-      json.decoder.JSONDecodeError: Unexpected UTF-8 BOM (decode using utf-8-sig)
-
-根因 (install.ps1 第 4 步 "写配置", Windows PowerShell 5.1):
-  1) `Set-Content configs\\serve.json -Encoding UTF8` 写出的是 **带 BOM** 的 UTF-8;
-     guanlan.py 用裸 utf-8 读 -> 直接 JSONDecodeError, 启动器连 check 都跑不了。
-  2) 紧随其前的 `Get-Content` 对无 BOM 的 UTF-8 按 ANSI 代码页 (中文机 = GBK) 解码,
-     把 serve.json 里的中文注释读成乱码再写回 -> 注释字段损坏 (部分字节已被替换成 '?', 不可逆)。
-
-修复:
-  A. guanlan.py     : 所有 JSON 配置改走 jload() = utf-8-sig (兼容有/无 BOM); 配置损坏时打印提示并退回
-                      内置默认端口, 不再抛栈 (hand-edit 配置写坏 BOM 也不该让产品起不来)。
-  B. install.ps1    : 第 4 步改由 venv 里的 Python 读写配置 (读 utf-8-sig / 写无 BOM 的 UTF-8),
-                      不再用 PS 的 cmdlet 碰 serve.json; 保持文件 UTF-8 带 BOM + CRLF 的既有约定。
-  C. serve.json     : 去掉 BOM, 还原两个 _note 中文注释 (依据原文件残留可逆部分 + 说明书 §7)。
-
-幂等: 重复运行只报 "已修过"。原文件备份为 *.bak-bomfix。
-"""
-from __future__ import annotations
-
-import json
-import pathlib
-import shutil
-import sys
-
-ROOT = pathlib.Path(r"F:\temp\guanlan-rudong-v2_0.2.0")
-GUANLAN = ROOT / "guanlan.py"
-INSTALL_PS1 = ROOT / "install.ps1"
-SERVE_JSON = ROOT / "configs" / "serve.json"
-BAK = ".bak-bomfix"
-
-report: list[str] = []
-
-
-def fail(msg: str) -> None:
-    print("  [X] " + msg)
-    sys.exit(1)
-
-
-def backup(p: pathlib.Path) -> None:
-    b = p.with_name(p.name + BAK)
-    if not b.exists():
-        shutil.copy2(p, b)
-        print(f"  备份 {p.name} -> {b.name}")
-
-
-def read_bytes(p: pathlib.Path) -> bytes:
-    if not p.exists():
-        fail(f"文件不存在: {p}")
-    return p.read_bytes()
-
-
-def has_bom(b: bytes) -> bool:
-    return b.startswith(b"\xef\xbb\xbf")
-
-
-# ---------------------------------------------------------------- A. guanlan.py
-def fix_guanlan() -> None:
-    raw = read_bytes(GUANLAN)
-    if has_bom(raw):
-        fail("guanlan.py 竟然带 BOM, 先人工确认再修")
-    src = raw.decode("utf-8")  # 该文件是 UTF-8 无 BOM + LF
-
-    if "def jload(" in src:
-        print("  [=] guanlan.py 已修过 (存在 jload), 跳过")
-        return
-    backup(GUANLAN)
-
-    anchor_cfg = (
-        'def cfg():\n'
-        '    c = dict(DEFAULT); p = ROOT / "configs/serve.json"\n'
-        '    if p.exists(): c.update(json.loads(p.read_text(encoding="utf-8")))\n'
-        '    return c\n'
-    )
-    new_cfg = (
-        'def jload(p: Path):\n'
-        '    """读 JSON 配置。必须用 utf-8-sig: 记事本和 PowerShell 5.1 的 `Set-Content -Encoding UTF8`\n'
-        '    写出的都是 **带 BOM** 的 UTF-8, 裸 utf-8 读会 JSONDecodeError (Unexpected UTF-8 BOM)。"""\n'
-        '    return json.loads(Path(p).read_text(encoding="utf-8-sig"))\n'
-        '\n'
-        '\n'
-        'def cfg():\n'
-        '    c = dict(DEFAULT); p = ROOT / "configs/serve.json"\n'
-        '    if p.exists():\n'
-        '        try: c.update(jload(p))\n'
-        '        except Exception as ex:  # 配置写坏不该让整个产品起不来: 退回内置默认端口, 但要说清楚\n'
-        '            print(f"  [X] {p} 读不了 ({ex.__class__.__name__}: {str(ex)[:100]}); 本次用内置默认端口。"\n'
-        '                  f" 重跑 install 脚本可重写该文件")\n'
-        '    return c\n'
-    )
-
-    # (旧串, 新串, 期望出现次数)
-    edits = [
-        (anchor_cfg, new_cfg, 1),
-        ('mcfg = json.loads((ROOT / "configs/models.json").read_text(encoding="utf-8"))',
-         'mcfg = jload(ROOT / "configs/models.json")', 1),
-        ('pids = json.loads(PIDS.read_text(encoding="utf-8")) if PIDS.exists() else {}',
-         'pids = jload(PIDS) if PIDS.exists() else {}', 2),
-        ('pids = json.loads(PIDS.read_text(encoding="utf-8"))', 'pids = jload(PIDS)', 1),
-    ]
-    for old, new, want in edits:
-        got = src.count(old)
-        if got != want:
-            fail(f"guanlan.py 锚点命中 {got} 次 (期望 {want}), 未改动: {old[:60]!r}")
-        src = src.replace(old, new)
-
-    left = src.count('read_text(encoding="utf-8")')  # 修完不该再有裸 utf-8 的 JSON 读
-    if left:
-        fail(f"guanlan.py 仍有 {left} 处裸 utf-8 读 JSON")
-    compile(src, str(GUANLAN), "exec")  # 语法自检, 不过就不落盘
-    GUANLAN.write_bytes(src.encode("utf-8"))
-    report.append("guanlan.py: JSON 配置改走 utf-8-sig; 配置损坏时给提示不抛栈")
-    print("  [OK] guanlan.py 已修")
-
-
-# --------------------------------------------------------------- B. install.ps1
-NEW_STEP4 = '''Say "== 4/5 写配置"
-# 别改回 Get-Content / Set-Content -Encoding UTF8 (2026-09-10 实机踩过):
-#   PS 5.1 的 Get-Content 把无 BOM 的 UTF-8 当 ANSI(GBK) 解码 -> 中文注释变乱码;
-#   Set-Content -Encoding UTF8 又写出 BOM -> Python 侧用裸 utf-8 json.loads 直接 JSONDecodeError。
-#   交给 venv 里的 Python 读写: 读 utf-8-sig (有 BOM 也认), 写无 BOM 的 UTF-8。
-#   下面 Python 代码故意只用单引号, 避开 PS 5.1 传参时的引号转义坑。
-$cfgFix = @'
-import json, pathlib, sys
-p = pathlib.Path('configs/serve.json')
-c = json.loads(p.read_text(encoding='utf-8-sig'))
-c['python'] = sys.executable
-p.write_text(json.dumps(c, ensure_ascii=False, indent=2) + '\\n', encoding='utf-8')
-print('   configs/serve.json -> ' + sys.executable)
-'@
-& $vpy -c $cfgFix
-if ($LASTEXITCODE -ne 0) { Say "   X 写 configs\\serve.json 失败"; exit 4 }
-'''
-
-OLD_STEP4 = (
-    'Say "== 4/5 写配置"\n'
-    '$cfg = Get-Content configs\\serve.json -Raw | ConvertFrom-Json\n'
-    '$cfg.python = $vpy\n'
-    '$cfg | ConvertTo-Json -Depth 5 | Set-Content configs\\serve.json -Encoding UTF8\n'
-)
-
-
-def fix_install_ps1() -> None:
-    raw = read_bytes(INSTALL_PS1)
-    if not has_bom(raw):
-        fail("install.ps1 丢了 BOM —— 文件头注释要求 UTF-8 带 BOM + CRLF, 先人工确认")
-    text = raw.decode("utf-8-sig").replace("\r\n", "\n")
-    if "cfgFix" in text:
-        print("  [=] install.ps1 已修过 (存在 cfgFix), 跳过")
-        return
-    backup(INSTALL_PS1)
-
-    if text.count(OLD_STEP4) != 1:
-        fail(f"install.ps1 第 4 步锚点命中 {text.count(OLD_STEP4)} 次, 未改动")
-    text = text.replace(OLD_STEP4, NEW_STEP4.replace("\n", "\n"))
-
-    # 恢复该文件的编码约定: UTF-8 带 BOM + CRLF
-    out = text.replace("\r\n", "\n").replace("\n", "\r\n")
-    INSTALL_PS1.write_bytes(out.encode("utf-8-sig"))
-    report.append("install.ps1: 第 4 步改由 venv Python 写配置 (utf-8-sig 读 / 无 BOM 写)")
-    print("  [OK] install.ps1 已修 (保持 UTF-8 BOM + CRLF)")
-
-
-# ------------------------------------------------------ C. configs/serve.json
-NOTE = ("端口只在这里改 (网关路由表 scripts/guanlan_gateway.py 里的上游端口须同步; "
-        "v0.1 仍是硬编码, 见说明书 §7)")
-NOTE_RAW = (r"原始 SCADA/技术资料目录 (相对本目录或绝对路径, 如 D:\guanlan\data\raw); "
-            r"启动器导出为 WINDSCADA_RUDONG_SRC; 原始件不随包分发")
-
-
-def fix_serve_json() -> None:
-    raw = read_bytes(SERVE_JSON)
-    was_bom = has_bom(raw)
-    try:
-        cfg = json.loads(raw.decode("utf-8-sig"))
-    except Exception as ex:
-        fail(f"serve.json 内容已经是坏 JSON ({ex}); 从 configs\\serve.json{BAK} 或重跑 install 恢复")
-    if not isinstance(cfg, dict):
-        fail("serve.json 不是对象")
-
-    if not was_bom and cfg.get("_note") == NOTE and cfg.get("_note_raw") == NOTE_RAW:
-        print("  [=] serve.json 已修过 (无 BOM 且注释正常), 跳过")
-        return
-    backup(SERVE_JSON)
-
-    cfg["_note"] = NOTE
-    cfg["_note_raw"] = NOTE_RAW
-    SERVE_JSON.write_text(json.dumps(cfg, ensure_ascii=False, indent=2) + "\n", encoding="utf-8")
-    report.append("serve.json: 去掉 BOM, 还原 _note/_note_raw 中文注释")
-    print(f"  [OK] serve.json 已修 (原 BOM={was_bom}, python={cfg.get('python')})")
-
-
-def verify() -> None:
-    print("\n== 校验")
-    # 1) serve.json: 无 BOM, 且裸 utf-8 也能解析 (即 guanlan.py 老写法也不会再炸)
-    b = SERVE_JSON.read_bytes()
-    print(f"  serve.json BOM={has_bom(b)} bytes={len(b)}")
-    if has_bom(b):
-        fail("serve.json 仍带 BOM")
-    cfg = json.loads(b.decode("utf-8"))
-    for k in ("_note", "_note_raw"):
-        if "\ufffd" in cfg[k] or "?" in cfg[k]:
-            fail(f"{k} 仍有乱码: {cfg[k]!r}")
-    if not cfg.get("python", "").lower().endswith(r".venv\scripts\python.exe"):
-        fail(f"python 路径不对: {cfg.get('python')!r}")
-    print(f"  python -> {cfg['python']}")
-    print(f"  _note     : {cfg['_note']}")
-    print(f"  _note_raw : {cfg['_note_raw']}")
-
-    # 2) guanlan.py: 编译 + 现场跑一遍 cfg()
-    src = GUANLAN.read_bytes().decode("utf-8")
-    compile(src, str(GUANLAN), "exec")
-    ns: dict = {"__name__": "not_main", "__file__": str(GUANLAN)}
-    exec(compile(src, str(GUANLAN), "exec"), ns)
-    c = ns["cfg"]()
-    print(f"  guanlan.cfg() -> gateway={c['gateway']} python={c['python']} (无异常)")
-
-    # 3) install.ps1: 字节约定
-    pb = INSTALL_PS1.read_bytes()
-    txt = pb.decode("utf-8-sig")
-    crlf, lf = txt.count("\r\n"), txt.count("\n")
-    print(f"  install.ps1 BOM={has_bom(pb)} CRLF={crlf} LF-only={lf - crlf}")
-    if not has_bom(pb) or crlf != lf:
-        fail("install.ps1 不再是 UTF-8 BOM + 全 CRLF")
-
-
-def main() -> int:
-    print(f"目标: {ROOT}")
-    for p in (GUANLAN, INSTALL_PS1, SERVE_JSON):
-        if not p.exists():
-            fail(f"缺文件 {p}")
-    print("\n== A guanlan.py"); fix_guanlan()
-    print("\n== B install.ps1"); fix_install_ps1()
-    print("\n== C configs/serve.json"); fix_serve_json()
-    verify()
-    print("\n完成:")
-    for r in report:
-        print("  - " + r)
-    if not report:
-        print("  (无需改动, 之前已修)")
-    return 0
-
-
-if __name__ == "__main__":
-    sys.exit(main())

+ 0 - 140
_修复记录_20260911/fix_guanlan_pythonpath.py

@@ -1,140 +0,0 @@
-#!/usr/bin/env python3
-# -*- coding: utf-8 -*-
-"""第二处修复: 机器全局 PYTHONPATH 污染 .venv。
-
-实测 (本机 PYTHONPATH="D:\\Program Files\\Python\\Lib\\site-packages;") :
-  pip 在 venv 里看到系统 site-packages 的包, 判定 "Requirement already satisfied ... outside environment",
-  于是没有把传递依赖装进 .venv: venv 里缺 urllib3 / polars / pyyaml / jinja2 / python-dotenv 等,
-  并且 sys.path 里 PYTHONPATH 排在 venv site-packages **前面**, requests 被系统 2.33.0 顶掉了 pin 住的 2.34.2。
-  后果: 一旦换个没设 PYTHONPATH 的 shell (或别的机器用户), import requests/polars/yaml 直接 ModuleNotFoundError。
-
-修复 (都不改变功能, 只要求 "跑的就是 .venv 里那一套"):
-  D. install.ps1 / install.sh : 装依赖前摘掉 PYTHONPATH, 让 pip 把 requirements.txt 真正装进 .venv。
-  E. guanlan.py                : 启动时把 PYTHONPATH 从 sys.path/b环境里摘掉 (check/serve/stop 与子进程都干净)。
-
-幂等: 重复运行报 "已修过"。
-"""
-from __future__ import annotations
-
-import pathlib
-import shutil
-import sys
-
-ROOT = pathlib.Path(r"F:\temp\guanlan-rudong-v2_0.2.0")
-GUANLAN = ROOT / "guanlan.py"
-INSTALL_PS1 = ROOT / "install.ps1"
-INSTALL_SH = ROOT / "install.sh"
-BAK = ".bak-pyfix"
-
-changed: list[str] = []
-
-
-def fail(msg: str) -> None:
-    print("  [X] " + msg)
-    sys.exit(1)
-
-
-def backup(p: pathlib.Path) -> None:
-    b = p.with_name(p.name + BAK)
-    if not b.exists():
-        shutil.copy2(p, b)
-        print(f"  备份 {p.name} -> {b.name}")
-
-
-def sub1(text: str, old: str, new: str, what: str) -> str:
-    n = text.count(old)
-    if n != 1:
-        fail(f"{what}: 锚点命中 {n} 次 (期望 1)")
-    return text.replace(old, new)
-
-
-# ------------------------------------------------------------------ D1. install.ps1
-PS1_ANCHOR = '$ErrorActionPreference = "Stop"; $env:PYTHONUTF8 = "1"\n'
-PS1_INSERT = PS1_ANCHOR + (
-    '# 机器上若全局设了 PYTHONPATH (例如指向 D:\\Program Files\\Python\\Lib\\site-packages), pip 会把那里的包\n'
-    '# 当成 "已满足", 于是不装进 .venv —— venv 里就缺 urllib3 / polars / pyyaml / jinja2 等传递依赖,\n'
-    '# 换个没设 PYTHONPATH 的 shell 立刻 import 失败; 且 PYTHONPATH 排在 venv 之前会顶掉 pin 住的版本。\n'
-    '# 先把环境变量摘掉, 让下面 pip 老老实实按 requirements.txt 装进 .venv。\n'
-    'Remove-Item Env:PYTHONPATH -ErrorAction SilentlyContinue\n'
-)
-
-
-def fix_install_ps1() -> None:
-    raw = INSTALL_PS1.read_bytes()
-    text = raw.decode("utf-8-sig").replace("\r\n", "\n")
-    if "Remove-Item Env:PYTHONPATH" in text:
-        print("  [=] install.ps1 已修过, 跳过")
-        return
-    backup(INSTALL_PS1)
-    text = sub1(text, PS1_ANCHOR, PS1_INSERT, "install.ps1")
-    INSTALL_PS1.write_bytes(text.replace("\r\n", "\n").replace("\n", "\r\n").encode("utf-8-sig"))
-    changed.append("install.ps1: 装依赖前 Remove-Item Env:PYTHONPATH")
-    print("  [OK] install.ps1 已修 (保持 UTF-8 BOM + CRLF)")
-
-
-# ------------------------------------------------------------------ D2. install.sh
-SH_ANCHOR = 'set -e\n'
-SH_INSERT = SH_ANCHOR + (
-    '# 同理: 全局 PYTHONPATH 会让 pip 误判依赖已满足而不装进 .venv (见 install.ps1 注释)\n'
-    'unset PYTHONPATH\n'
-)
-
-
-def fix_install_sh() -> None:
-    raw = INSTALL_SH.read_bytes()
-    if raw.startswith(b"\xef\xbb\xbf"):
-        fail("install.sh 不该有 BOM")
-    text = raw.decode("utf-8")
-    if "unset PYTHONPATH" in text:
-        print("  [=] install.sh 已修过, 跳过")
-        return
-    backup(INSTALL_SH)
-    text = sub1(text, SH_ANCHOR, SH_INSERT, "install.sh")
-    INSTALL_SH.write_bytes(text.encode("utf-8"))  # 保持 LF / 无 BOM
-    changed.append("install.sh: 装依赖前 unset PYTHONPATH")
-    print("  [OK] install.sh 已修 (保持 LF)")
-
-
-# ------------------------------------------------------------------ E. guanlan.py
-G_ANCHOR = 'WIN = os.name == "nt"\n'
-G_INSERT = G_ANCHOR + '''
-
-def _hermetic():
-    """摘掉机器全局的 PYTHONPATH: 本项目依赖都装在 .venv 里, 而 PYTHONPATH 会被插到 sys.path 前面,
-    顶掉 venv 里 pin 住的版本 (实测 requests 2.33.0 顶掉 2.34.2), 还会让 pip 少装传递依赖。
-    子进程从 env() 继承的是摘干净之后的环境。"""
-    for d in [x for x in os.environ.pop("PYTHONPATH", "").split(os.pathsep) if x]:
-        while d in sys.path: sys.path.remove(d)
-
-
-_hermetic()
-'''
-
-
-def fix_guanlan() -> None:
-    raw = GUANLAN.read_bytes()
-    text = raw.decode("utf-8")
-    if "_hermetic()" in text:
-        print("  [=] guanlan.py 已修过, 跳过")
-        return
-    backup(GUANLAN)
-    text = sub1(text, G_ANCHOR, G_INSERT, "guanlan.py")
-    compile(text, str(GUANLAN), "exec")
-    GUANLAN.write_bytes(text.encode("utf-8"))
-    changed.append("guanlan.py: 启动时摘掉 PYTHONPATH (自身与子进程都跑 .venv 的那套)")
-    print("  [OK] guanlan.py 已修")
-
-
-def main() -> int:
-    print(f"目标: {ROOT}")
-    print("\n== D1 install.ps1"); fix_install_ps1()
-    print("\n== D2 install.sh"); fix_install_sh()
-    print("\n== E guanlan.py"); fix_guanlan()
-    print("\n完成:")
-    for c in changed or ["(无需改动, 之前已修)"]:
-        print("  - " + c)
-    return 0
-
-
-if __name__ == "__main__":
-    sys.exit(main())

+ 0 - 84
_修复记录_20260911/fix_guanlan_startbat.py

@@ -1,84 +0,0 @@
-#!/usr/bin/env python3
-# -*- coding: utf-8 -*-
-"""第三处修复: start.bat 把 "网关已起但有模块降级" 误报成 "没起来"。
-
-现状: guanlan.py serve 的退出码 1 有**两种**含义 —— ① 网关 healthz 压根没起来;
-② 网关起来了, 但有模块 [DOWN] (本机 Ollama 没跑 / 没拉模型时必然如此)。
-start.bat 只看 errorlevel, 于是 ② 也走 "[X] Server did not start ... Not opening the browser",
-用户看到的是 "服务没起", 实际网关在跑、6/7 页面都正常 —— 与原始报障里那条
-"[X] Server did not start (reason above)" 同源。
-
-修复: 退出码非 0 时先探一次 /healthz 区分两种情形; 真没起来才报 [X] 并停;
-起来了但有降级则打 [!] 说明 (哪一类页面受影响), 浏览器照常打开。
-保持 .bat 的 CRLF / 无 BOM 约定。
-"""
-from __future__ import annotations
-
-import pathlib
-import shutil
-import sys
-
-ROOT = pathlib.Path(r"F:\temp\guanlan-rudong-v2_0.2.0")
-START = ROOT / "start.bat"
-
-OLD = (
-    '".venv\\Scripts\\python.exe" guanlan.py serve\n'
-    'if errorlevel 1 (\n'
-    '  echo.\n'
-    '  echo [X] Server did not start ^(reason above^). Not opening the browser -\n'
-    '  echo     a "site cannot be reached" page would look like a network problem.\n'
-    '  echo     See logs\\gateway.log\n'
-    '  pause\n'
-    '  exit /b 1\n'
-    ')\n'
-    'start "" http://127.0.0.1:28084/\n'
-)
-
-NEW = '''".venv\\Scripts\\python.exe" guanlan.py serve
-if errorlevel 1 (
-  rem serve exits 1 in two different cases: the gateway never came up, or the gateway is
-  rem up with at least one degraded module, e.g. Ollama not running. Probe healthz to tell
-  rem them apart, otherwise a working server gets reported as "did not start".
-  ".venv\\Scripts\\python.exe" -c "import urllib.request,sys; sys.exit(0 if urllib.request.urlopen('http://127.0.0.1:28084/healthz',timeout=6).status==200 else 1)" 2>nul
-  if errorlevel 1 (
-    echo.
-    echo [X] Gateway did not come up ^(reason above^). Not opening the browser -
-    echo     a "site cannot be reached" page would look like a network problem.
-    echo     See logs\\gateway.log
-    pause
-    exit /b 1
-  )
-  echo.
-  echo [!] Gateway is up, but at least one module is degraded ^(the [DOWN] lines above^).
-  echo     Usual cause: local Ollama is not running or has no models pulled, which only
-  echo     disables the Q^&A / local-review pages. Every other page works.
-  echo     Details: logs\\ for the degraded module.
-)
-start "" http://127.0.0.1:28084/
-'''
-
-
-def main() -> int:
-    raw = START.read_bytes()
-    if raw.startswith(b"\xef\xbb\xbf"):
-        print("  [X] start.bat 不该有 BOM")
-        return 1
-    text = raw.decode("utf-8").replace("\r\n", "\n")   # .bat 是 CRLF, 先归一化再匹配
-    if "Gateway did not come up" in text:
-        print("  [=] start.bat 已修过, 跳过")
-        return 0
-    if text.count(OLD) != 1:
-        print(f"  [X] start.bat 锚点命中 {text.count(OLD)} 次, 未改动")
-        return 1
-    b = START.with_name(START.name + ".bak-healthz")
-    if not b.exists():
-        shutil.copy2(START, b)
-        print(f"  备份 start.bat -> {b.name}")
-    out = text.replace(OLD, NEW).replace("\r\n", "\n").replace("\n", "\r\n")
-    START.write_bytes(out.encode("utf-8"))     # UTF-8 无 BOM + CRLF
-    print("  [OK] start.bat 已修 (CRLF / 无 BOM 保持不变)")
-    return 0
-
-
-if __name__ == "__main__":
-    sys.exit(main())

+ 0 - 67
_修复记录_20260911/restore_release_rudong.py

@@ -1,67 +0,0 @@
-#!/usr/bin/env python3
-# -*- coding: utf-8 -*-
-"""bug-2 修复: 从 v0.1.0 发行包把缺失的 release/如东/ 交付件恢复到 v0.2.0 安装目录。
-
-根因: v0.2.0 包**根本没带** release/如东/ (0.2.0 zip 里 release/ 下只有 viewer/),
-而 release/portal.html 的 #documents 区块用 iframe 指向
-  http://127.0.0.1:28084/如东/如东治理清单_交付_20260901/02_治理清单/如东_液压系统治理清单_v1.2_2026-09-01.html
-该路径由网关 _static() 从 release/ 目录取文件, 文件不在 → 404 {"err":"not found", "path": …},
-iframe 里就直接显示那段 JSON。v0.1.0 包里这份交付件是齐的 (76 项), 直接恢复。
-
-(release/如东/ 按 .gitignore 约定不入库 —— 客户交付物只在受保护工作区。)
-"""
-from __future__ import annotations
-
-import pathlib
-import shutil
-import sys
-import zipfile
-
-ZIP = pathlib.Path(r"F:\temp\guanlan-rudong-v2_0.1.0_all.zip")
-REL = pathlib.Path(r"F:\temp\guanlan-rudong-v2_0.2.0\release")
-PREFIX = "release/如东/"
-
-
-def nm(info):
-    raw = info.filename
-    for enc in ("gbk", "utf-8"):
-        try:
-            return raw.encode("cp437").decode(enc)
-        except Exception:
-            continue
-    return raw
-
-
-def main() -> int:
-    if not ZIP.is_file():
-        print(f"  [X] 缺源包 {ZIP}")
-        return 1
-    print(f"源: {ZIP}")
-    print(f"目标: {REL}\\如东\\\n")
-    n_files = 0
-    n_bytes = 0
-    with zipfile.ZipFile(ZIP) as zf:
-        items = [(nm(i), i) for i in zf.infolist()]
-        rel = [(n.replace("\\", "/"), i) for n, i in items if n.replace("\\", "/").startswith(PREFIX)]
-        if not rel:
-            print("  [X] 源包里没有 release/如东/")
-            return 1
-        for name, info in rel:
-            target = REL / name[len("release/"):]
-            if info.is_dir():
-                target.mkdir(parents=True, exist_ok=True)
-                continue
-            target.parent.mkdir(parents=True, exist_ok=True)
-            with zf.open(info) as fsrc, open(target, "wb") as fdst:
-                shutil.copyfileobj(fsrc, fdst, 1024 * 1024)
-            if target.stat().st_size != info.file_size:
-                print(f"  [X] 大小不符: {target}")
-                return 1
-            n_files += 1
-            n_bytes += info.file_size
-    print(f"恢复 {n_files} 个文件, {n_bytes/1024/1024:.1f} MB")
-    return 0
-
-
-if __name__ == "__main__":
-    sys.exit(main())

+ 0 - 50
_修复记录_20260911/scan_portal_links.py

@@ -1,50 +0,0 @@
-#!/usr/bin/env python3
-# -*- coding: utf-8 -*-
-"""全量核查 portal.html 里指向本机静态件的引用 (含相对写法), 找出同类"交付件没随包"的坏链。"""
-import pathlib
-import re
-import urllib.error
-import urllib.parse
-import urllib.request
-
-REL = pathlib.Path(r"F:\temp\guanlan-rudong-v2_0.2.0\release")
-html = (REL / "portal.html").read_text(encoding="utf-8", errors="replace")
-
-# 所有 href/src 里含 如东 或 release 的
-cand = set()
-for m in re.finditer(r'(?:href|src)="([^"]+)"', html):
-    u = m.group(1)
-    if ("如东" in u or "/release" in u or "release/" in u) and not u.startswith(("mailto:", "javascript:")):
-        cand.add(u)
-print(f"候选引用 {len(cand)} 个\n")
-
-
-def check(u: str):
-    if u.startswith("http"):
-        full = u
-    else:
-        base = "http://127.0.0.1:28084/" if u.startswith("/") else "http://127.0.0.1:28084/"
-        full = base + u.lstrip("/")
-    p = urllib.parse.urlsplit(full)
-    full = urllib.parse.urlunsplit((p.scheme, p.netloc, urllib.parse.quote(p.path), p.query, ""))
-    try:
-        with urllib.request.urlopen(full, timeout=30) as r:
-            r.read(1)
-            return r.status
-    except urllib.error.HTTPError as e:
-        return e.code
-    except Exception as e:
-        return f"{e.__class__.__name__}"
-
-
-bad = []
-for u in sorted(cand):
-    s = check(u)
-    tag = "ok  " if s == 200 else "BAD "
-    if s != 200:
-        bad.append(u)
-    print(f"  [{s}] {tag} {u}")
-
-print(f"\n坏链 {len(bad)} 个")
-for u in bad:
-    print("   - " + u)

+ 0 - 44
_修复记录_20260911/verify_bug2.py

@@ -1,44 +0,0 @@
-#!/usr/bin/env python3
-# -*- coding: utf-8 -*-
-"""bug-2 验收: 走网关取那个 iframe URL, 并把它自身的相对引用也逐条取一遍。"""
-import re
-import urllib.parse
-import urllib.request
-
-URL = ("http://127.0.0.1:28084/如东/如东治理清单_交付_20260901/"
-       "02_治理清单/如东_液压系统治理清单_v1.2_2026-09-01.html")
-
-
-def get(u):
-    req = urllib.request.Request(u, headers={'User-Agent': 'verify'})
-    try:
-        with urllib.request.urlopen(req, timeout=30) as r:
-            return r.status, r.headers.get('Content-Type'), r.read()
-    except urllib.error.HTTPError as e:
-        return e.code, e.headers.get('Content-Type'), e.read()
-
-
-def q(u):
-    """URL 里中文需编码, 网关会 unquote。"""
-    p = urllib.parse.urlsplit(u)
-    return urllib.parse.urlunsplit((p.scheme, p.netloc,
-                                    urllib.parse.quote(p.path), p.query, p.fragment))
-
-
-st, ct, body = get(q(URL))
-print(f"iframe URL : [{st}] {ct} {len(body)} 字节")
-txt = body.decode('utf-8', 'replace')
-print(f"首行       : {txt.splitlines()[0][:120] if txt.strip() else '(空)'}")
-print(f"含 not found: {'not found' in txt[:400] and st != 200}")
-
-refs = sorted({m for m in re.findall(r'(?:href|src)="([^"]+)"', txt)
-               if not m.startswith(('http', '#', 'data:', 'mailto:', 'javascript:'))})
-print(f"\n文中相对引用 {len(refs)} 个, 逐条取:")
-bad = 0
-for r in refs:
-    u = urllib.parse.urljoin(URL, r)
-    s, c, b = get(q(u))
-    if s != 200:
-        bad += 1
-    print(f"  [{s}] {len(b):8d} 字节  {r}")
-print(f"\n结论: {'全部 200' if not bad else f'{bad} 个引用取不到'}")

+ 0 - 61
_修复记录_20260911/verify_guanlan_simsys.py

@@ -1,61 +0,0 @@
-#!/usr/bin/env python3
-# -*- coding: utf-8 -*-
-"""验证 /sim/sys/ (release/sim_sys_server.py) 恢复后: 页面能出 + 脱敏回扫干净。"""
-import sys
-import urllib.request
-import zipfile
-import pathlib
-
-ROOT = pathlib.Path(r"F:\temp\guanlan-rudong-v2_0.2.0")
-sys.path.insert(0, str(ROOT))
-from src.windscada import deid_public as DP
-
-
-def readable(name):
-    try:
-        return name.encode('cp437').decode('utf-8')
-    except UnicodeError:
-        return name
-
-
-zip_path = ROOT / 'release' / '如东SWT40_控制律仿真台_20260906.zip'
-with zipfile.ZipFile(zip_path) as z:
-    pages = [readable(n.rsplit('/', 1)[-1]) for n in z.namelist()
-             if n.endswith('.html') and readable(n.rsplit('/', 1)[-1])[:1] in {'0', '1', '2', '3', '4'}]
-
-print(f"页面 {len(pages)} 个")
-bad = 0
-for name in sorted(pages):
-    url = 'http://127.0.0.1:18792/' + urllib.parse.quote(name)
-    try:
-        with urllib.request.urlopen(url, timeout=20) as r:
-            body = r.read().decode('utf-8')
-            code = r.status
-    except Exception as ex:
-        print(f"  [X] {name}: {ex.__class__.__name__}: {ex}")
-        bad += 1
-        continue
-    leaks = DP.audit(body)
-    leak_txt = '干净' if not leaks else '; '.join(f'{n}×{c} {s}' for n, c, s in leaks)
-    print(f"  [{code}] {name:28s} {len(body):7d} 字节  回扫: {leak_txt}")
-    bad += bool(leaks)
-
-print("\n网关路由:")
-for path in ('/healthz', '/sim/sys/'):
-    try:
-        with urllib.request.urlopen('http://127.0.0.1:28084' + path, timeout=20) as r:
-            txt = r.read().decode('utf-8', 'replace')
-            if path == '/healthz':
-                import json
-                d = json.loads(txt)
-                print(f"  [{r.status}] {path}  status={d['status']} {d['mode']}")
-                for m in d['modules']:
-                    print(f"        {'ok  ' if m['ok'] else 'DOWN'} {m['path']} {m['name']}")
-            else:
-                print(f"  [{r.status}] {path}  {len(txt)} 字节, 含返回链接: {'guanlan-top' in txt}")
-    except Exception as ex:
-        print(f"  [X] {path}: {ex.__class__.__name__}: {ex}")
-        bad += 1
-
-print("\n结论:", "全部通过" if not bad else f"有 {bad} 项需要注意")
-sys.exit(0 if not bad else 1)

+ 238 - 0
docs/振动数据接入_v0.1.md

@@ -0,0 +1,238 @@
+# 振动数据接入(CMS / TCM)· v0.1
+
+> 日期: 2026-09-12 · 适用: `<安装目录>`(安装目录)
+> 本文回答: ①振动侧的现场件长什么样、放哪;②谁把它变成系统能用的东西;③改了哪些代码、怎么验收;
+> ④哪些还做不到(如实列,不假装)。
+> 相关: `docs\数据目录结构与落位约定_v0.2.md`(§2b 落位速查)· `docs\重算操作手册_v0.1.md`(一键重算)
+
+---
+
+## 1. 用户令与现场件盘点
+
+**用户令原文**(路径按本包文档约定写成占位符): 「数据层里的「CMS 振动评估报告」应遵循
+`<安装目录>\data\raw\如东\windcms`、「振动线 handoff」应遵循 `<安装目录>\data\raw\如东\m5_cms_tcm`
+存放; 要求 1. 修改 `<安装目录>` 下的观澜系统, 支持振动数据参与系统运行、重算等;
+2. 自 `<现场包目录>` 提取相应振动数据存放至上述目录」;随后用户把 `CMS_RuDong_CGN_202603-04.zip`
+放进现场包 —— **那正是此前缺的原始件**。
+
+现场包(`<现场包目录>`)里的振动相关件共三类:
+
+| # | 件 | 形态 | 能不能算 |
+|---|---|---|---|
+| ① | `CMS_RuDong_CGN_202603-04.zip`(38.8 GB 压缩 / 150 GB 解压 / 25,679 件) | Brande TCM Enterprise 导出: `measurement\<年>\<月>\<WTGxx>\<WTGxx>_<uuid>_decode.json`;每件是一次 API 响应 `{body:{body:{"<时间戳>":[{"Record":{…}}]}}}` | **能**(索引/谱/报告全部由它算) |
+| ② | 12 份月度「振动分析报告(用印版)」PDF(2025-06…2026-05) | **纯扫描件** | **不能** |
+| ③ | 上海电气 2026年07月报告 docx(4.87 MB)· 大生科技传动链报告 docx(20.81 MB) | Word, 有文本层 | 能(逐台判级转录) |
+
+**②为什么不能算(实测, 不是推测)**: 用 pypdf 打开 2026年5月那份 → `len(pages)=12`,
+每页 `len(images)=1`, `page.extract_text()` 长度 **0**(12 页全 0)。字节面也印证: `/Font` 0 处、
+`FlateDecode` 0 处、`/Image` 24 处、`DCTDecode` 12 处 = 每页一张 JPEG。
+⇒ 没有 OCR 就取不出任何数值。**处理方式: 只登记归档**(`_扫描件清单.json` 记 sha256/大小/无文本层),
+不参与判级、不假装读过。要它们的数值只有两条路: 现场给电子件(docx/xlsx),或上 OCR(本包不装)。
+
+**顺带核过的一条旧判断**: v0.2.0 初版把「`(8)…\振动分析报告\`」列在"故意不落"里,理由是
+"成品牌报告而非可再加工的测量数据, 且 windcms 要的测点索引包里没有"。**该判断在 2026-09-12 被推翻**
+(①到货后索引源件具备),故改为落位。现场包里的**同件副本**(根目录 `…2026年5月…(2).pdf`、
+嵌套包 `如东海上振动报告11份.zip`)不重复落 —— 落三遍会得到 3 份同名报告。
+
+---
+
+## 2. 落位约定
+
+```
+data\raw\如东\
+├─ windcms\                                   ← 数据层「CMS 振动评估报告」
+│    ├─ CMS_RuDong_CGN_202603-04\
+│    │    └─ measurement\2026\{03,04}\WTGxx\*_decode.json     25,679 件 / 150 GB
+│    └─ 厂家报告\上海电气_月度\
+│         ├─ 中广核如东海上风电场2025年06月…2026年05月振动分析报告用印版.pdf    12 件(扫描件)
+│         ├─ 中广核如东海上风电场2026年07月振动分析报告_上海电气.docx
+│         └─ _扫描件清单.json
+└─ m5_cms_tcm\                                ← 数据层「振动线 handoff」
+     └─ 厂家报告\
+          └─ 中广核如东海上风电场传动链振动分析报告_大生科技_20260311(1).docx
+```
+
+- 两个目录名已写进 `src\windscada\config.py` 的 `STATION_SUBDIRS`(扫描/落位/维护页从此认它们;
+  改目录名 = 换接口,要同步维护页与本文)。
+- 原始导出**保留包内 `measurement\` 这一层**(来源可追溯);摄入按 `rglob` 找 `*_decode.json`,
+  套不套这层都能吃。
+- 若现场给出 `handoff_vibration_v2.json` / `component_history.json` **正本**,放 `m5_cms_tcm\` 下即被优先采用
+  (包内那两份目前仍是随包快照 —— 振动线分支的产物没随包,见 §7)。
+- 落位命令(映射表可复核、可重跑,同尺寸文件自动跳过):
+  ```
+  .venv\Scripts\python.exe scripts\place_raw_data.py --src <现场包目录> --scope vib --dry-run
+  .venv\Scripts\python.exe scripts\place_raw_data.py --src <现场包目录> --scope vib
+  ```
+
+---
+
+## 3. 摄入链(谁把原始件变成系统能用的东西)
+
+```
+data/raw/如东/windcms/…/*_decode.json
+   │  scripts/rudong_tcm_index.py      → outputs/rudong/m5_cms_tcm/windows/<窗>/index.parquet   (54 列)
+   │  scripts/rudong_tcm_spectra.py    → outputs/rudong/m5_cms_tcm/windows/<窗>/spectra/*.npz
+   │                                     + <窗>/spectra_meta.parquet(谱库目录内另存一份, 兼顾两种读取路径)
+   │  [可选] scripts/windcms.py report / kb  → outputs/rudong/windcms/*(★默认不跑, 见 §3b)
+   ▼
+data/raw/如东/{windcms,m5_cms_tcm}/…/*.docx
+   │  scripts/vib_reports_build.py     → outputs/rudong/windcms/厂家报告提取_<报告期>.json
+   │                                     + 报告_CMS振动状态评估报告_<报告期>.md(与自产报告同构)
+   │                                     + outputs/rudong/m5_cms_tcm/报告_TCM传动链振动分析_<报告期>.md
+```
+
+**一键**: `scripts/vib_raw_build.py`(索引→谱;`--skip-spectra` 可关;`--with-report` 才跑报告/知识库)。
+它也写 `outputs/rudong/m5_cms_tcm/vib_raw_manifest.json`(源件路径、窗名、行数、谱数、时间窗、
+各步耗时/rc,以及 `missing_chain`)。
+
+### 3b. ★ 为什么"重生成 CMS 报告"默认不跑(2026-09-12 实测)
+
+跑一次 `scripts/windcms.py report` 在本包**会让产物变差**:
+
+| 件 | 随包快照 | 本包重生成后 | 变化 |
+|---|---|---|---|
+| `windcms/report.md` | 16,857 B(含逐台融合级表 + L4 过闸谱线) | **467 B**(融合级表空、L4 写"无") | **−97.2%** |
+| `windcms/overview.html` | 646,218 B | 548,466 B | −15.1% |
+| `windcms/index_eng.html` | 443,841 B | 408,112 B | −8.0% |
+| `windcms/index.html` | 389,997 B | 382,544 B | −1.9%(逐台页仍 38 个) |
+
+根因不在数据,在**缺件**:`src/windcms/data.py::load_model()` 要读 `m5/model_run_l6.parquet` 与
+`m5/fusion_38.csv`,而这两件属六层链的 `model_run` / `fusion` 两步 —— 那四步脚本没随包(§7),产物也不在。
+于是"重生成"= 用残缺输入覆盖完整快照。
+
+**处置**: 摄入默认只做索引/谱(加性、不覆盖任何随包件);报告/知识库改为 `--with-report` 显式开启,
+且要求先备份 `outputs/<场>/windcms/`。本次实测后已把 `outputs/rudong/windcms` **逐字节还原**为随包快照
+(56 件,哈希比对 0 差异),只保留厂家报告转录那两件。六层链补齐后这个默认值应当翻过来。
+
+**来源登记**: 摄入与转录产出的每一件都由构建脚本自登记进 `outputs/<场>/_derived_manifest.json`
+(`src/derived_manifest.py`),`_provenance.json` 生成时据此把它们记成 `raw-derived` ——
+避免"新造的件因不在随包快照里而整条不进台账"。
+
+**窗名口径**: `wMMDD` = **数据起始日**(与既有 `w0127`/`w0707`/`w0811` 同口径)。
+`src\windcms\pipeline.py::window_year_gate` 会拦"窗名无年份 + 数据过老"的误用(w1226 事故的护栏)。
+重复摄入同一批数据时,已存在的窗自动改名为 `<窗>_reimport_<时分>`,而 `data.EXCLUDE_DEFAULT`
+把带 `_reimport` 的窗排除在生产集外 —— **重复摄入不会污染分析集**。
+
+**消费者(无需改一行代码,窗是自动发现的)**:
+
+| 消费者 | 拿什么 |
+|---|---|
+| `src\windcms\data.py::windows()` | 扫 `m5\windows\w????\index.parquet` → 新窗进分析集 |
+| `src\windcms\data.py::load_scalars()` | 各窗标量(`ds_size==1` 行)拼接 → CMS 报告/页面 |
+| `src\windcms\data.py::spectrum()` | `spectra_meta.parquet` + `npz['values'][shard_row]` → **谱图能取到** |
+| `src\windscada\taxonomy.py` / `subsys\fusion.py::windcms_grades()` | 最新 `报告_CMS振动状态评估报告_*.md` 的 `## 附录 A` 表 → 设备状态转录 |
+
+> ★ `spectra` 这一环此前是**死的**: 包内 `m5\spectra\`、`m5\windows\` 两个目录整个不在,
+> `data.spectrum()` 只能返回 None。本次摄入让谱图第一次有数据可画。
+
+---
+
+## 4. 格式契约与对拍(改这块必看)
+
+摄入产物与**包内既有产物**是同构关系,基准就是包内那份 `outputs\rudong\m5_cms_tcm\tcm_index.parquet`
+(330,308 行 × 54 列,2026-01-27~02-03 窗,同一摄取逻辑的产物):
+
+| 检查 | 结果 |
+|---|---|
+| 列名与列序 | **完全一致**(54 列,逐字抄自基准,含列序) |
+| dtype | 差异 3 列: `rec_i`(float64/int64)、`ds_dim`(float64/str)、`parse_error`(str/object)—— 消费者按列名取数,不受影响 |
+| 取值(FFT 行) | `x_offset=0.0` `x_delta=0.9375` `x_unit='Hz'` `y_unit='m/s²'` 与基准同行**逐格一致** |
+| 内部一致性 | `ds_size == lines+1` 比例 **1.000** |
+
+**TCM 导出里踩过的两个坑(写在这里防复发)**:
+
+1. **X/Y 轴四件在 `Measurement.DataSets` 层,不在 `DataSet` 里**。`Size`/`Dimension`/`Values` 在
+   `DataSets.DataSet`,而 `X-axisOffset`/`X-axisDelta`/`X-axisUnit`/`Y-axisUnit` 在**上一层**。
+   第一版两层取错 → 四列整列为空,内部一致性检查算出 0.000(本该 1.000)才暴露。**先看一致性数再信列**。
+2. `DataSets.DataSet.Values` 是**空格分隔的数值字符串**(不是数组)。谱: `Size=Lines+1`;
+   `X-axisDelta = 带宽/Lines`(例: 6000/6400 = 0.9375)。标量: `Size==1` 时 `Values` 本身就是标量值。
+   另有少数测量(`FFT_250/2000/10000_Tr`, `Lines=400`)的 `x_delta = 带宽/800` = 半带宽口径 ——
+   **源数据自己的约定**,摄入原样转录,不做"修正"。
+
+---
+
+## 5. 厂商报告的转录纪律
+
+`scripts\vib_reports_build.py` 只做**转录**,不做判级:
+
+- **合并单元格必须按 `tc` 恒等还原**:python-docx 对合并区会把同一个 `tc` 在相邻列/行重复给出,
+  照抄会把 `优秀 | 优秀 | 优秀` 压成一格(第一版手工 dump 就吃了这个亏: 表里 `WTG-01` 行显示成
+  `WTG-01 | 优秀 | 无`,实际是 主轴承=优秀 齿轮箱=优秀 发电机=优秀 结论=无)。规则: 同行 `tc` 与左邻
+  相同 = 横向合并(沿用左值);与上行同列相同 = 纵向合并(沿用上值)。
+- **用厂家自己的总览句当校验和**:2026年07月报告原文"…优秀机组34台,良好3台,预警0台,无数据1台,
+  测点异常1台"。转录结果 `{优秀34, 良好3, 不可判1}` —— **逐项一致**("无数据"记 `不可判`: 没测到 ≠ 正常)。
+- **厂家没给的不猜**:报告按「主轴承/齿轮箱/发电机」给级(**不分前后**),转录表里主轴承前/后填同一值并注明;
+  「综合」= 三项**取严**(危险>报警>良好>优秀>不可判)的聚合规则结果,必须写明"这是我们的规则不是厂家判级";
+  「融合级/CMS 红黄/行动等级建议」厂家未给 → 写 `—`。
+- **两条轴不混**:厂商报告转录的 md 与观澜自产报告同构(消费端不必改代码),报告期用**文件名月份取月末**,
+  于是它不会顶掉更新的自产报告;随包自产报告缺失时它就是最新版(从零重算的机器上正好用得上)。
+  大生科技那份只写 `报告_TCM传动链振动分析_*.md`,**不冒充 handoff 判级**。
+- **口径不同别直接比**:厂家四级(优秀/良好/预警/报警)+「数据异常」≠ 观澜五级。实测两者对 38 台
+  有 33 台结论不同(厂家 2026-07 vs 观澜 2026-08)—— 来源、日期、判据都不同,属预期,不是谁错。
+
+---
+
+## 6. 验收(命令 + 实测数字)
+
+```
+.venv\Scripts\python.exe scripts\place_raw_data.py --src <现场包目录> --scope vib --dry-run   # 先看计划
+.venv\Scripts\python.exe scripts\place_raw_data.py --src <现场包目录> --scope vib             # 落位 (150 GB)
+.venv\Scripts\python.exe scripts\vib_raw_build.py --jobs 10                                   # 摄入 (索引+谱)
+.venv\Scripts\python.exe scripts\vib_reports_build.py --dump                                  # 只打印转录, 人工核对
+.venv\Scripts\python.exe scripts\vib_reports_build.py                                         # 写产物
+.venv\Scripts\python.exe scripts\scan_stations.py                                             # 两个新目录是否被认到
+```
+
+**落位**(实测): `windcms\CMS_RuDong_CGN_202603-04` 25,679 件 150.1 GB +
+`windcms\厂家报告\上海电气_月度` 13 件 51.4 MB + `m5_cms_tcm\厂家报告` 1 件 20.8 MB = **25,693 件 / 150.2 GB**
+(`rc=0`;余量要求: `F:` 需 ≥ 165 GB)。重跑幂等(同尺寸跳过;实测第二次跳过 2 件已存在的 docx)。
+
+**摄入**(实测, 10 并行): 索引步 **392.8 s** → `windows\w0316\index.parquet` **2,066,686 行 × 54 列**
+(标量行 861,917 · 谱行 1,204,769;38 台;时间 2026-03-16 17:27:22 → 2026-04-21 10:18:33);
+谱步 **477.2 s** → **420,742 条谱** / **1,712 个 npz 分片**(只转 `FFT_`,故小于"谱行"数;
+`Size>1` 行里另有 `Time_*` 波形 78 万行未转)。窗体合计 **3.27 GB**。全链 **875 s(15 分钟)**。
+
+**消费端**(实测, 改代码前后都是同一套消费者):
+
+| 环节 | 实测结果 |
+|---|---|
+| `data.windows()` | 发现 `['w0127', 'w0316']` —— 新窗自动进分析集 |
+| `data.load_scalars()` | 193,902 行 × 11 列(w0316 贡献 168,880;7 个关键标量齐) |
+| `data.spectra_meta()` | 420,742 行 × 18 列,8 个测点全 |
+| `data.spectrum()` | **✅ 取到真实谱**: `WTG01 / Gear_IMS / FFT_6000_Tr` → 6,401 点,x 0→6000 Hz,y 单位 m/s²(此前 `m5/spectra` 缺失 ⇒ 只能返回 None) |
+| 索引列契约 | `data._read_index()` 要的 10 列齐;无缺列 |
+
+**厂商报告转录**: 上海电气 2026-07 → 38 台,取严 `{优秀34, 良好3, 不可判1}` 与报告原文总览句
+`优秀34/良好3/预警0/无数据1/测点异常1` 逐项一致;大生科技 2026-03-11 → 38 台
+`{优秀29, 良好7, 预警1, 数据异常1}`(与其正文"WTG02 未取到数据"一致);
+消费端解析器(取 `## 附录 A` 后 `| WTG` 行第 6 列)读到 38 台不缺台。
+
+**来源台账**(`outputs/rudong/_provenance.json`): 改造前 **17 raw-derived / 572 shipped**;
+本次摄入+转录后 **1,740 / 571**(`m5_cms_tcm` 由 0/90 变 1,718/90、`windcms` 由 0/56 变 2/54)。
+逐件登记由 `src/derived_manifest.py` 承载(谁算的谁登记),`products_restore_missing.py` 读它;
+已存在的窗若漏登记,用 `python scripts/vib_raw_build.py --register-only` 幂等补登记。
+
+**页面侧**(`http://127.0.0.1:28084/detail/v2#tab=system` 的数据层表,接口 `/detail/api/maint_survey`):
+两行位置已改为 `<安装目录>\data\raw\如东\windcms\` 与 `<安装目录>\data\raw\如东\m5_cms_tcm\`。
+★ 组件服务**启动时会把这张表烘进缓存**,改完 `src/ontology/maintenance.py` 必须重启组件才生效
+(本次用运营台同一套脚本: `_ops_stop_keep_gateway.py` + `_ops_start_and_open.py --no-open`)。
+
+**其余闸门**: 全库 183 个 py 在 `-W error::SyntaxWarning` 下编译 0 失败;本体审计 `rc=0`;
+`check_transferable.py` 命中数 350 → **344**(`data\raw` 已排除出文本扫描;我的产出 0 处,
+manifest 里的源件路径已改为相对安装根)。
+
+## 7. 仍然做不到的(如实列)
+
+- **六层链四步未随包**: `rudong_tcm_oem_scan.py` / `rudong_line_energy_share.py` / `rudong_model_run.py` /
+  `rudong_fusion_run.py`(原在振动线分支 `claude/vibration-data-diagnosis-32b69e`)。
+  因此扫描线/能量占比/模型层/融合层那几类 parquet(`oem_frequency_scan`、`gear_freq_scan`、
+  `blade_1p_*`、`model_run_l6`、`fusion_38.csv` 等)**仍走包内随包快照**,不由 `data\raw` 重算;
+  连带的后果是 **CMS 报告/页面不能重生成**(§3b 实测 −97%)。`vib_raw_manifest.json` 的
+  `missing_chain` 字段如实列着这四项;缺口台账里对应 **B5**。
+- **扫描件 PDF 的数值**取不出(无文本层, 本包不带 OCR)。
+- `handoff_vibration_v2.json` / `component_history.json` 目前是**随包快照**:它们含大量人工裁决、
+  校准更新与开放项,不是能从测量数据直接算出来的东西 —— 现场给正本才改由现场件驱动。
+- 谱库的磁盘代价:整窗全转(`--meas ALL`)会显著吃盘(含 `Time_*` 波形, 单点 65,536/200,000);
+  默认只转 `FFT_` 谱(本轮 42 万条谱 = 3.2 GB)。
+- 一个可选的后续:把纯 Python 的 `pypdf` 轮子并入 `wheels\`,让 `vib_reports_build.py` 能**离线**
+  逐件核验扫描件页数/有无文本层(本轮是靠临时装的 pypdf 抽检一件得出的结论,清单里记了这一点)。

+ 69 - 9
docs/数据目录结构与落位约定_v0.2.md

@@ -16,7 +16,9 @@
 │         ├─ scada_1min\              (可选) 1min 导出
 │         ├─ scada_1min\              (可选) 1min 导出
 │         ├─ 故障报警\                 报警事件导出 (SpreadsheetML *.xls)
 │         ├─ 故障报警\                 报警事件导出 (SpreadsheetML *.xls)
 │         ├─ 风机故障记录\             检修工单台账 (*.xls/xlsx, 内按 {年}年故障记录\ 分年)
 │         ├─ 风机故障记录\             检修工单台账 (*.xls/xlsx, 内按 {年}年故障记录\ 分年)
-│         └─ 油样报告\                 油液化验报告 (*.pdf)
+│         ├─ 油样报告\                 油液化验报告 (*.pdf)
+│         ├─ windcms\                 ★振动侧: CMS 原始测量导出 + 厂商月度评估报告 (2026-09-12 新增)
+│         └─ m5_cms_tcm\              ★振动侧: TCM 侧深度分析报告 + 现场给的 handoff 正本
 ├─ outputs\rudong\                   ★系统取数用的**产物仓** (页面 99% 读这里, 不直接读 raw)
 ├─ outputs\rudong\                   ★系统取数用的**产物仓** (页面 99% 读这里, 不直接读 raw)
 │    ├─ windscada\                   L0 标准仓: 37 个 parquet + 索引/日志
 │    ├─ windscada\                   L0 标准仓: 37 个 parquet + 索引/日志
 │    ├─ ontology\                    本体对象库 objects.json + 检索索引 + release_r1/r2
 │    ├─ ontology\                    本体对象库 objects.json + 检索索引 + release_r1/r2
@@ -38,9 +40,8 @@
 ├─ reference\rudong\                 契约与语言资产 (windscada_contract.yaml / 英文语言库 / 报警码表 …)
 ├─ reference\rudong\                 契约与语言资产 (windscada_contract.yaml / 英文语言库 / 报警码表 …)
 ├─ wheels\win_amd64\                 离线轮子 (随包分发, 不入 git)
 ├─ wheels\win_amd64\                 离线轮子 (随包分发, 不入 git)
 ├─ vendor\                           便携 Python / Ollama 离线包 (不入 git)
 ├─ vendor\                           便携 Python / Ollama 离线包 (不入 git)
-├─ docs\                             说明书与本文
-├─ logs\ run\                        运行日志 / pids.json (不入 git)
-└─ _修复记录_20260911\                现场修复记录与可重放脚本
+├─ docs\                             说明书、落位约定、振动接入与本文
+└─ (不随包: `_修复记录_*`/临时目录 —— 本包已纳入 git, 变更历史由版本库承载)
 ```
 ```
 
 
 ---
 ---
@@ -58,7 +59,7 @@
 | ③ | `data\raw\` 下**只有这一个**场站目录(单站部署) | 采用它, 但报告里标"凭单站唯一性", 不假装精确匹配 |
 | ③ | `data\raw\` 下**只有这一个**场站目录(单站部署) | 采用它, 但报告里标"凭单站唯一性", 不假装精确匹配 |
 | ④ | 多目录且都不匹配 | **不猜**: 报"未识别", 页面显示无数据并列出扫到的目录 |
 | ④ | 多目录且都不匹配 | **不猜**: 报"未识别", 页面显示无数据并列出扫到的目录 |
 
 
-约定子目录(这五个名字是摄入接口, 改名等于换接口):
+约定子目录(这些名字是摄入接口, 改名等于换接口; 前五个是 SCADA/台账侧, 后两个是振动侧, 见 §2b):
 
 
 | 子目录 | 放什么 | 文件形态要求 |
 | 子目录 | 放什么 | 文件形态要求 |
 |---|---|---|
 |---|---|---|
@@ -67,6 +68,8 @@
 | `故障报警\` | 报警事件导出 | `.xls` 实为 SpreadsheetML(XML), `<row>` 需带 `TimeOn` + `Alarmcode`;「全年/年至今」累计快照件会自动跳过 |
 | `故障报警\` | 报警事件导出 | `.xls` 实为 SpreadsheetML(XML), `<row>` 需带 `TimeOn` + `Alarmcode`;「全年/年至今」累计快照件会自动跳过 |
 | `风机故障记录\` | 检修工单台账 | 模板表(表头含 `机组编号`+`故障名称`+`故障代码`);表里 `风场名称` 必须属本场(集团导出件混着十来个场);按年分组不限层级 |
 | `风机故障记录\` | 检修工单台账 | 模板表(表头含 `机组编号`+`故障名称`+`故障代码`);表里 `风场名称` 必须属本场(集团导出件混着十来个场);按年分组不限层级 |
 | `油样报告\` | 油液化验报告 | `.pdf`, 文件名须含 `日期_台号_部件`(如 `17072025_10303681_…_1#_主轴后.pdf`) |
 | `油样报告\` | 油液化验报告 | `.pdf`, 文件名须含 `日期_台号_部件`(如 `17072025_10303681_…_1#_主轴后.pdf`) |
+| `windcms\` | CMS 原始测量导出 + 厂商月度评估报告 | 导出为 `测量根\<年>\<月>\<WTGxx>\*_decode.json`(Brande TCM 导出, 目录层级不限); 报告为 `.docx`(可解析) 或 `.pdf`(扫描件只归档) |
+| `m5_cms_tcm\` | TCM 侧深度分析报告 + handoff 正本 | `.docx` 报告; 若现场给出 `handoff_vibration_v2.json` / `component_history.json` 正本, 放这里即被优先采用 |
 
 
 机理层(厂商资料)按 A2 仍放 `data\raw\西门子4.0技术资料\`, **不在场站目录下**。
 机理层(厂商资料)按 A2 仍放 `data\raw\西门子4.0技术资料\`, **不在场站目录下**。
 
 
@@ -82,13 +85,61 @@
 |---|---|---|
 |---|---|---|
 | `--scope a2`(默认) | 上表四类(`scada_10min`/`故障报警`/`风机故障记录`/`油样报告`) | A2 约定的四项数据层 |
 | `--scope a2`(默认) | 上表四类(`scada_10min`/`故障报警`/`风机故障记录`/`油样报告`) | A2 约定的四项数据层 |
 | `--scope mech` | `data\raw\西门子4.0技术资料\`(317 件 2.7 GB,含那份**对译表**)+ `<场站>\scada_1min\`(38 件 12.8 GB) | 本体层构建器 `kb_ingest.py` 的 `TECH` 与 `rd()` 四个文件名;`config.STATION_SUBDIRS` 声明的 `scada_1min` |
 | `--scope mech` | `data\raw\西门子4.0技术资料\`(317 件 2.7 GB,含那份**对译表**)+ `<场站>\scada_1min\`(38 件 12.8 GB) | 本体层构建器 `kb_ingest.py` 的 `TECH` 与 `rd()` 四个文件名;`config.STATION_SUBDIRS` 声明的 `scada_1min` |
-| `--scope full` | 两组一起 | 从零重算要用的全集 |
+| `--scope vib` | `<场站>\windcms\`(CMS 原始测量导出 25,679 件 150 GB + 上海电气月度报告)+ `<场站>\m5_cms_tcm\`(大生科技 TCM 报告) | §2b |
+| `--scope full` | 三组一起 | 从零重算要用的全集 |
 
 
 > 为什么 `scada_1min` 也属"该落的": 它写在 `config.STATION_SUBDIRS` 里, `scan_stations.py` 会把它报成
 > 为什么 `scada_1min` 也属"该落的": 它写在 `config.STATION_SUBDIRS` 里, `scan_stations.py` 会把它报成
 > "缺"。但**包内没有消费者** —— 落它是补数据层完整性, 不改任何页面数值。同理, 现场包里的
 > "缺"。但**包内没有消费者** —— 落它是补数据层完整性, 不改任何页面数值。同理, 现场包里的
-> `scada数据(如东)\`(19 个月原始通道导出)、`fastlog数据\`、`(8)…\振动分析报告\` 三种**故意不落**:
-> 前者是 `scada_10min` 的上游且无人读, 中间那种全库只有 2 处注释提到, 后一种是振动线的成品牌报告
-> 而非可再加工的测量数据 —— 理由都逐条写在脚本的 `SKIPPED` 表里。
+> `scada数据(如东)\`(19 个月原始通道导出)、`fastlog数据\` 两种**故意不落**:
+> 前者是 `scada_10min` 的上游且无人读, 后者全库只有 2 处注释提到 —— 理由都逐条写在脚本的 `SKIPPED` 表里。
+>
+> ★ 关于 `(8)…\振动分析报告\`(12 份月度用印版 PDF): v0.2.0 初版把它列在"故意不落"里, 理由是
+> "成品牌报告而非可再加工的测量数据, 且 windcms 要的测点索引包里没有"。**那条判断在 2026-09-12 被推翻**:
+> 用户把 CMS 原始测量导出 (`CMS_RuDong_CGN_202603-04.zip`) 补进了现场包, 索引源件已具备, 于是
+> 报告作为数据层证据一并落位(见 §2b)。脚本的 `SKIPPED` 表里保留了这条记录并标注了推翻原因。
+
+---
+
+## 2b. 振动侧落位与摄入链(2026-09-12 新增)
+
+**用户令**: 「数据层里的「CMS 振动评估报告」应遵循 `<安装目录>\data\raw\如东\windcms`、
+「振动线 handoff」应遵循 `<安装目录>\data\raw\如东\m5_cms_tcm` 存放; 修改系统支持振动数据参与运行、重算;
+自现场包提取相应振动数据存放至上述目录」。
+
+落位后的实际形态(`--scope vib`):
+
+```
+data\raw\如东\
+├─ windcms\
+│    ├─ CMS_RuDong_CGN_202603-04\          ← 原始测量导出原样保留包内层级 (可追溯"哪来的")
+│    │    └─ measurement\2026\{03,04}\WTGxx\WTGxx_<uuid>_decode.json     25,679 件 / 150 GB
+│    └─ 厂家报告\上海电气_月度\
+│         ├─ 中广核如东海上风电场2025年06月…2026年05月振动分析报告用印版.pdf   12 件 (扫描件, 只归档)
+│         ├─ 中广核如东海上风电场2026年07月振动分析报告_上海电气.docx          (有文本层, 摄入)
+│         └─ _扫描件清单.json                                                (登记: sha256/大小/无文本层)
+└─ m5_cms_tcm\厂家报告\
+     └─ 中广核如东海上风电场传动链振动分析报告_大生科技_20260311(1).docx      (TCM M-system, 摄入)
+```
+
+**摄入链(两条, 分工不同, 别混)**:
+
+| 源 | 摄入命令 | 产出 | 谁消费 |
+|---|---|---|---|
+| CMS 原始测量导出(`*_decode.json`) | `python scripts\vib_raw_build.py`(= `rudong_tcm_index.py` → `rudong_tcm_spectra.py`;★**不**跑 `windcms.py report` —— 缺六层链产物时重生成会掉内容, 实测 `report.md` −97%, 见 `docs\振动数据接入_v0.1.md` §3b) | `outputs\rudong\m5_cms_tcm\windows\<窗>\index.parquet`(54 列)+ `…\<窗>\spectra\*.npz` + `spectra_meta.parquet` | `src\windcms\data.py`(窗/标量/谱图自动收录)、`scripts\windcms.py report/serve`、`src\windscada\taxonomy.py`(设备状态转录) |
+| 厂商 docx 评估报告 | `python scripts\vib_reports_build.py` | `outputs\rudong\windcms\厂家报告提取_<报告期>.json` + `报告_CMS振动状态评估报告_<报告期>.md`;`outputs\rudong\m5_cms_tcm\报告_TCM传动链振动分析_<报告期>.md` | 同上(报告 md 与自产报告**同构**, 消费端不必改代码) |
+
+窗名口径: `wMMDD` = **数据起始日**(与既有 `w0127`/`w0707`/`w0811` 同口径, `pipeline.window_year_gate`
+会拦"窗名无年份"的历史误用)。同一批数据重复摄入时, 已存在的窗会改名为 `<窗>_reimport_<时分>`,
+而 `data.EXCLUDE_DEFAULT` 会把带 `_reimport` 的窗排除在生产集外 —— 重复摄入不会污染分析集(2026-08-24 的教训)。
+
+**这一步之后"振动数据就真的参与了"的证据链**: `data.windows()` 自动发现新窗 → `load_scalars()` 把
+新窗标量并进 CMS 报告 → `data.spectrum()` 能取到谱(此前 `m5\spectra` 整个目录缺失, 谱图是死的)→
+`fusion.windcms_grades()` / `taxonomy.system_matrix()` 读最新报告 md 做设备状态转录。
+
+**仍然缺的(如实标注, 不冒充)**: 六层链的四步脚本 `rudong_tcm_oem_scan.py` / `rudong_line_energy_share.py` /
+`rudong_model_run.py` / `rudong_fusion_run.py` **未随包**(原在振动线分支 `claude/vibration-data-diagnosis-32b69e`)。
+`vib_raw_build.py` 产出的 `outputs\rudong\m5_cms_tcm\vib_raw_manifest.json` 里 `missing_chain` 字段列明这四项。
+因此扫描线/能量占比/模型层/融合层那几类 parquet 仍走"包内 shipped 快照", 只有**索引/谱/报告/知识库**是可重算的。
 
 
 ---
 ---
 
 
@@ -102,6 +153,8 @@
 | `故障报警\` | 否(只有摄入脚本读) | `scripts\windscada_alarms_ingest.py` | 数据层「报警事件」、故障分析、停机事件、限电绑定 |
 | `故障报警\` | 否(只有摄入脚本读) | `scripts\windscada_alarms_ingest.py` | 数据层「报警事件」、故障分析、停机事件、限电绑定 |
 | `风机故障记录\` | 否 | `scripts\windscada_workorder_ingest.py` | 数据层「检修工单台账」、检修面、闭环验证 |
 | `风机故障记录\` | 否 | `scripts\windscada_workorder_ingest.py` | 数据层「检修工单台账」、检修面、闭环验证 |
 | `油样报告\` | 否 | `scripts\windscada_watch_channels_build.py` | 数据层「油液化验」、油液时效胶囊、融合面油样轴 |
 | `油样报告\` | 否 | `scripts\windscada_watch_channels_build.py` | 数据层「油液化验」、油液时效胶囊、融合面油样轴 |
+| `windcms\`(CMS 原始导出) | 否(窗是**摄入时**落盘的) | `scripts\vib_raw_build.py`(索引→谱→报告/知识库) | CMS 系统的谱图/标量与报告(`windcms.py serve`);`taxonomy` 的设备状态转录 |
+| `windcms\`/`m5_cms_tcm\`(厂商报告) | 否 | `scripts\vib_reports_build.py` | 数据层的厂商评估报告转录(与自产报告同构的 md) |
 
 
 命令(在 `<安装目录>` 下跑):
 命令(在 `<安装目录>` 下跑):
 
 
@@ -111,6 +164,7 @@
 .venv\Scripts\python.exe scripts\rebuild_from_raw.py --scada # 再加 SCADA 侧 10 个构建器 (慢)
 .venv\Scripts\python.exe scripts\rebuild_from_raw.py --scada # 再加 SCADA 侧 10 个构建器 (慢)
 .venv\Scripts\python.exe scripts\rebuild_from_raw.py --verify# 与随包基线逐值等价验收
 .venv\Scripts\python.exe scripts\rebuild_from_raw.py --verify# 与随包基线逐值等价验收
 .venv\Scripts\python.exe scripts\place_raw_data.py           # 现场压缩包 → 约定子目录 (映射表)
 .venv\Scripts\python.exe scripts\place_raw_data.py           # 现场压缩包 → 约定子目录 (映射表)
+.venv\Scripts\python.exe scripts\vib_raw_build.py            # 振动侧: CMS 原始导出 → 窗索引/谱/报告
 ```
 ```
 
 
 ---
 ---
@@ -129,6 +183,9 @@
 | `powercurve_dev/bins` · `loss_monthly` · `curve_lenses/liveness` · `control_profile/schedule` · `stop_events` · `temp_bins` · `yaw_daily` · `hydraulic_accum` · `thermal_chain` · `system_aux` | scada_10min + alarms | `rebuild_from_raw.py --scada`(实测: `loss_monthly` 3729/3729、`powercurve_dev` 38/38、`powercurve_bins` 912/912 与随包基线**逐值完全一致**) |
 | `powercurve_dev/bins` · `loss_monthly` · `curve_lenses/liveness` · `control_profile/schedule` · `stop_events` · `temp_bins` · `yaw_daily` · `hydraulic_accum` · `thermal_chain` · `system_aux` | scada_10min + alarms | `rebuild_from_raw.py --scada`(实测: `loss_monthly` 3729/3729、`powercurve_dev` 38/38、`powercurve_bins` 912/912 与随包基线**逐值完全一致**) |
 | `objects.json`(本体 2340 个对象)· `retrieval_index.json` | **厂商技术资料** `data\raw\西门子4.0技术资料`(不在场站目录下) | `python -m src.ontology.kb_ingest`;检索索引 `python -c "from src.ontology import retrieval as R; R.build(use_vec=False)"`(BM25 词法索引; 向量那半要 Ollama 的 `bge-m3`, 本机没装模型时会跳过并保持纯词法可用) |
 | `objects.json`(本体 2340 个对象)· `retrieval_index.json` | **厂商技术资料** `data\raw\西门子4.0技术资料`(不在场站目录下) | `python -m src.ontology.kb_ingest`;检索索引 `python -c "from src.ontology import retrieval as R; R.build(use_vec=False)"`(BM25 词法索引; 向量那半要 Ollama 的 `bge-m3`, 本机没装模型时会跳过并保持纯词法可用) |
 | `turbine_params.parquet`(1706 条整定值) | 同上, `Turbine+Parameters.*.xlsx` | `python -c "from src.ontology.maintenance import refresh_params as f; f()"` |
 | `turbine_params.parquet`(1706 条整定值) | 同上, `Turbine+Parameters.*.xlsx` | `python -c "from src.ontology.maintenance import refresh_params as f; f()"` |
+| `m5_cms_tcm\windows\<窗>\index.parquet`(54 列标量索引) | **CMS 原始测量导出** `data\raw\如东\windcms\**\*_decode.json` | `scripts\rudong_tcm_index.py`(旧包缺这个脚本 → 2026-09-12 补齐; 列名列序 dtype 与包内 `tcm_index.parquet` 逐列对齐, 实测 FFT 行取值逐格一致) |
+| `m5_cms_tcm\windows\<窗>\spectra\*.npz` + `spectra_meta.parquet` | 同上 | `scripts\rudong_tcm_spectra.py`(`DataSets.DataSet.Values` 是空格分隔字符串: `Size=Lines+1`, `X-axisDelta=带宽/Lines`) |
+| `windcms\报告_CMS振动状态评估报告_<报告期>.md`(厂商报告转录) | 厂商 docx 评估报告 | `scripts\vib_reports_build.py`(合并单元格还原 + 厂家总览句作**校验和**: 实测 34优秀/3良好/1无数据逐项一致) |
 
 
 > 本体层这三件的前提是**技术资料在盘上**(`place_raw_data.py --scope mech` 会落, 见 §2);技术资料不在时
 > 本体层这三件的前提是**技术资料在盘上**(`place_raw_data.py --scope mech` 会落, 见 §2);技术资料不在时
 > `kb_ingest` 会打印 `⚠ 源缺失` 并降级 —— 不静默。
 > `kb_ingest` 会打印 `⚠ 源缺失` 并降级 —— 不静默。
@@ -300,3 +357,6 @@ P.venv_python()   # 跨平台探测 .venv/Scripts/python.exe 或 .venv/bin/pytho
 | 2026-09-11 | 验证闸门退出码修复:控制台 GBK 编不出 `✔` 时降级为 `?`(新增 `src\console.py`),修 `page_fingerprint` / `check_portability` / `rebuild_from_raw` / `scan_stations` / `portal_build` —— 此前"全部一致"会因 print 抛异常而**退出码 1(通过被报成失败)** |
 | 2026-09-11 | 验证闸门退出码修复:控制台 GBK 编不出 `✔` 时降级为 `?`(新增 `src\console.py`),修 `page_fingerprint` / `check_portability` / `rebuild_from_raw` / `scan_stations` / `portal_build` —— 此前"全部一致"会因 print 抛异常而**退出码 1(通过被报成失败)** |
 | 2026-09-11 | **从零重算实测** + 落位扩范围:`place_raw_data.py` 新增 `--scope a2\|mech\|full`(机理层技术资料 + 对译表 + `scada_1min`)、修"单文件规则恒不落件"与缺 `import os`(自 d90da05 起脚本跑不通)、同尺寸文件跳过;`src\ontology\` 全层 10 处文本 I/O 补 `encoding='utf-8'`(**中文 Windows 上本体层原本根本不可用**:默认 cp936 写 objects.json 崩在 `\u200b`);`check_portability.py` 新增 WARN 项 6「文本读写缺 encoding=」;实测数字见 §4 末 |
 | 2026-09-11 | **从零重算实测** + 落位扩范围:`place_raw_data.py` 新增 `--scope a2\|mech\|full`(机理层技术资料 + 对译表 + `scada_1min`)、修"单文件规则恒不落件"与缺 `import os`(自 d90da05 起脚本跑不通)、同尺寸文件跳过;`src\ontology\` 全层 10 处文本 I/O 补 `encoding='utf-8'`(**中文 Windows 上本体层原本根本不可用**:默认 cp936 写 objects.json 崩在 `\u200b`);`check_portability.py` 新增 WARN 项 6「文本读写缺 encoding=」;实测数字见 §4 末 |
 | 2026-09-11 | 收尾这 18 处同类文本 I/O(`scripts\ontology_p2_verify.py`、`scripts\windscada_serve.py`、`src\ontology\scenario_29.py`、`src\windcms\{knowledge,pipeline,plugins}.py`、`src\windscada\subsys\fusion.py`、`src\windscada\taxonomy.py`),门禁 WARN 项 6 扫描面同时扩到 `open('w')` 与 `open(p)`(无 mode = 文本读);**WARN 归零** |
 | 2026-09-11 | 收尾这 18 处同类文本 I/O(`scripts\ontology_p2_verify.py`、`scripts\windscada_serve.py`、`src\ontology\scenario_29.py`、`src\windcms\{knowledge,pipeline,plugins}.py`、`src\windscada\subsys\fusion.py`、`src\windscada\taxonomy.py`),门禁 WARN 项 6 扫描面同时扩到 `open('w')` 与 `open(p)`(无 mode = 文本读);**WARN 归零** |
+| 2026-09-12 | **振动侧落位与摄入(本文 §2b)**:`data\raw\如东\` 新增 `windcms`(CMS 原始测量导出 25,679 件 150 GB + 上海电气月度报告) / `m5_cms_tcm`(大生科技 TCM 报告) 两个约定子目录(`config.STATION_SUBDIRS` 同步);补齐振动线分支缺失的摄入端 `scripts\rudong_tcm_index.py`(54 列窗索引, 与包内 `tcm_index.parquet` 同构)与 `scripts\rudong_tcm_spectra.py`(npz 谱库 + `spectra_meta.parquet`);新增一键 `scripts\vib_raw_build.py`(索引→谱→报告/知识库, 并行+窗体清单)与厂商报告摄入 `scripts\vib_reports_build.py`;`place_raw_data.py` 新增 `--scope vib`(含散装件规则与**每次开包一次**的写法修正: 旧写法对 2.5 万条目是平方复杂度);`rebuild_all.py` 新增 ④b 步;`check_transferable.py` 文本扫描排除 `data\raw\`(现场原始件不随包, 扫它只会把秒级闸门拖成小时级)。**细节见 `docs\振动数据接入_v0.1.md`** |
+| 2026-09-12 | **不再随包"修复记录/临时"类目录**(用户令: 本包已纳入 git 管理): `pack_dist.py` 的 `INCLUDE_DIRS` 去掉 `_修复记录_20260911`(其内容并入 `docs\振动数据接入_v0.1.md` 与本文 §2b), 相关文档/README 的引用一并改指正式文档 |
+| 2026-09-12 | 振动侧**实测复核**(见 `docs\振动数据接入_v0.1.md`): 落位 25,693 件/150.2 GB; 摄入出窗 `w0316`(2,066,686 行 × 54 列 + 420,742 条谱, 全链 875 s); `data.windows/load_scalars/spectrum` 三处消费实测通过(**谱图此前是死的**)。★ 同时逮到一条坑: 在本包跑 `windcms.py report` 会**用残缺输入覆盖完整快照**(`report.md` −97%)—— 因为六层链的 `model_run`/`fusion` 产物缺失, 故该步改为 `--with-report` 显式开启, `outputs\rudong\windcms` 已逐字节还原为随包快照(56 件/0 差异)。新增 `src\derived_manifest.py`(产物来源自登记)让台账认出新造件 |

+ 3 - 2
docs/移植与独立运行_v0.1.md

@@ -75,7 +75,7 @@ sh install.sh
 
 
 | 项 | 为什么 |
 | 项 | 为什么 |
 |---|---|
 |---|---|
-| `_products_off/` `_products_off_prev_*/` | 本机"清除产物"的暂存与存档(每份都是整份产物 ~220 MB); 目标机用不到 |
+| `_products_off/` `_products_off_prev_*/` | **旧设计**的"清除产物"暂存与存档(每份都是整份产物 ~220 MB)。2026-09-16 用户令"清除产物不留备份"后**不再产生**; 老机器上若还有, 确认不需要可手工删除 |
 | `logs/` `run/pids.json` | 本机运行日志与**旧 PID**; pids.json 到新机器上是无效引用(启动器会重写) |
 | `logs/` `run/pids.json` | 本机运行日志与**旧 PID**; pids.json 到新机器上是无效引用(启动器会重写) |
 | `.git/` | 55 MB 仓库; 交付不需要(自己留版除外) |
 | `.git/` | 55 MB 仓库; 交付不需要(自己留版除外) |
 | `data/raw/`(视交付约定) | 30.9 GB 现场原始件; 按约定"原始件不随包分发"时不必带(缺它只影响"重算原始数据", 不影响页面) |
 | `data/raw/`(视交付约定) | 30.9 GB 现场原始件; 按约定"原始件不随包分发"时不必带(缺它只影响"重算原始数据", 不影响页面) |
@@ -94,4 +94,5 @@ sh install.sh
 > 现状(2026-09-12 实测): 文本里仍有 350 处机器相关路径, **但都不在运行路径上** —— 分布是
 > 现状(2026-09-12 实测): 文本里仍有 350 处机器相关路径, **但都不在运行路径上** —— 分布是
 > `release/viewer/*`(三维单位的 `source_path` 出处注记, 186 处) · `outputs/**`(产物里记的历史路径/provenance, 95) ·
 > `release/viewer/*`(三维单位的 `source_path` 出处注记, 186 处) · `outputs/**`(产物里记的历史路径/provenance, 95) ·
 > `src/sop/farm_paths.py`(盘符映射表本体, 门禁具名登记) · `scripts/build_oem_lexicon.py`(开发机 OEM 资料盘) ·
 > `src/sop/farm_paths.py`(盘符映射表本体, 门禁具名登记) · `scripts/build_oem_lexicon.py`(开发机 OEM 资料盘) ·
-> `_修复记录_20260911/fix_*.py`(一次性修复脚本留档) · 文档里的反面示例。**都不影响在别的机器上运行**。
+> 文档里的反面示例。**都不影响在别的机器上运行**。(`_修复记录_*` 这类目录已按用户令不再随包;
+> 文本扫描同时排除 `data/raw/` —— 现场原始件不随包分发, 而单那一个目录就是 2.5 万个文件/150 GB。)

+ 294 - 0
docs/系统设计说明.md

@@ -0,0 +1,294 @@
+# 观澜 · 如东样板 v2 · 系统设计说明
+
+> 版本 v0.2.0 · 2026-09-16 · 适用 `<安装目录>`(安装目录)
+> 本文是**系统设计正本**: ① 产物全景(类型/功用/输出路径/生成端/消费端);② 路径与进程的统一约定;
+> ③ 输入 ↔ 产物呼应关系;④ 重算的两种运行状态与"产物即时进页面"的实现;⑤ 启动与门户。
+> 与本文配套、可复跑的机器校验: `python scripts/inventory_products.py`(清点 + 呼应校验,`--check` 出退出码)。
+
+---
+
+## 1. 系统是什么(一页)
+
+```
+现场数据 data/raw/<场站>/  ──摄入──▶  outputs/<场>/(产物仓)  ──现读──▶  页面/接口
+     ▲  7 类约定目录                        ▲  9 个产物仓                  ▲
+     │  scada_10min / scada_1min /           │  windscada  L0 标准仓        │  网关 :28084(门户/运维控制台)
+     │  故障报警 / 风机故障记录 /             │  ontology   本体对象库        │  detail :18033(分析工作台)
+     │  油样报告 / windcms / m5_cms_tcm      │  windcms   CMS 振动           │  cms    :18020(振动诊断)
+     └─ 西门子4.0技术资料(机理层,不在场站目录下)│  m5_cms_tcm 振动线出件/窗     │  sim :18791 · sim_sys :18792 · viewer :64292
+                                            │  tcm_compatible_replay      │
+                                            │  sop / guanlan / pitch / paradigm_r1
+```
+
+四条设计铁律(贯穿全部代码,改代码时先读这四条):
+
+1. **路径唯一真源** `src/paths.py` —— 代码/配置里只写 ROOT 相对路径,运行期由助手解析成绝对路径(基准是安装根,**不是 cwd**)。
+2. **产物只落在 `outputs/<场>/`**;原始件只读 `data/raw/<场站>/`;`release/` 是交付静态层。三者不互相写。
+3. **判断在代码、模型只转述**;每个产物都要能说清"谁生成、谁消费、来源是 raw 重算还是随包补齐"(`_provenance.json`)。
+4. **窗口/进程统一口径** `src/proc.py` —— 子进程一律 `CREATE_NO_WINDOW`,日志落 `logs/`,不弹命令窗口。
+
+---
+
+## 2. 产物全景(自动生成,勿手改)
+
+> 由 `python scripts/inventory_products.py --write-doc` 生成;数据源是**盘上的实际文件** +
+> `outputs/<场>/_provenance.json`(逐件来源台账)+ `outputs/<场>/_derived_manifest.json`(构建脚本自登记)。
+
+<!-- INVENTORY:BEGIN (由 scripts/inventory_products.py --write-doc 生成, 勿手改) -->
+*自动生成于 2026-09-16 16:26;数据源: `outputs/rudong/_provenance.json` + `_derived_manifest.json` + 实际文件*
+
+| 产物仓 | 输出路径 (相对安装根) | 功用 | 生成端 | 消费端 | 件数 | 大小 | 来源(raw 重算/随包) |
+|---|---|---|---|---|---|---|---|
+| `windscada` | `outputs/rudong/windscada/` | L0 标准仓: SCADA/台账/派生分析的全部 parquet (页面主取数处) | rebuild_from_raw.py (三门台账) + --scada (10 个构建器) + windscada_monthly_build.py | scripts/windscada_serve.py 各视图 · src/windscada/taxonomy · subsys/fusion | 115 | 4.5 MB | 17 / 98 |
+| `ontology` | `outputs/rudong/ontology/` | 本体对象库: 码表/手册/工单展开/失效树 + 检索索引 + 实机参数 | python -m src.ontology.kb_ingest → populate → chain_ingest → trend_ingest → retrieval.build | 脚本 windscada_serve.py 本体页/问答 · scripts/guanlan_facts_contract.py | 25 | 82.9 MB | 3 / 22 |
+| `windcms` | `outputs/rudong/windcms/` | CMS 振动诊断产物: 状态评估报告/逐台页/工作台页/知识库 + 厂家报告转录 | scripts/windcms.py report/kb · scripts/vib_reports_build.py | 自服务 :18020 (每请求现读) · taxonomy.system_matrix (转录设备状态) | 56 | 30.1 MB | 2 / 54 |
+| `m5_cms_tcm` | `outputs/rudong/m5_cms_tcm/` | 振动线出件与窗级分析: handoff 接口 + 窗索引/谱库 + TCM 兼容件 | scripts/vib_raw_build.py (窗索引/谱) · 振动线出件 (handoff, 随包快照) | src/windscada/subsys/fusion.py · src/windcms/data.py · scripts/windcms.py report | 1808 | 3329.7 MB | 1718 / 90 |
+| `tcm_compatible_replay` | `outputs/rudong/tcm_compatible_replay/` | TCM 兼容链回放资产 (模型表/掩码阈值/裁决记录) | 随包快照 (无生成端) | src/windcms/report*.py · config.mask_thresholds | 59 | 32.7 MB | 0 / 59 |
+| `sop` | `outputs/rudong/sop/` | SOP 中间件/评审/台账与事实契约底稿 | 随包快照 (无生成端) | scripts/guanlan_facts_contract.py · 门户结论段 | 210 | 15.8 MB | 0 / 210 |
+| `guanlan` | `outputs/rudong/guanlan/` | 事实契约与对外派生 (可上云面孔) | scripts/guanlan_facts_contract.py | 门户 #findings · /api/facts | 5 | 0.5 MB | 0 / 5 |
+| `pitch` | `outputs/rudong/pitch/` | 变桨侧派生件 (零位/日粒度) | 随包快照 + rebuild_from_raw --scada | 脚本 windscada_serve.py 变桨面 | 2 | 0.3 MB | 0 / 2 |
+| `paradigm_r1` | `outputs/rudong/paradigm_r1/` | 范式实验件 (E3/E5/E8 底稿, 事实契约输入) | 随包快照 (无生成端) | scripts/guanlan_facts_contract.py | 29 | 0.5 MB | 0 / 29 |
+
+合计 2309 件 / 3496.9 MB;其中 raw 重算 1740 件、随包补齐 571 件。
+<!-- INVENTORY:END -->
+
+### 2.1 逐件来源口径(`_provenance.json`)
+
+| 来源 | 含义 | 判据 |
+|---|---|---|
+| `raw-derived` | 由 `data/raw` 重算出来的 | 该件在 `RAW_DERIVED` 表里,或由**构建脚本自登记**(`src/derived_manifest.py`) |
+| `shipped` | 包内没有生成端 / 规则未复现,用随包件补齐 | 既不在上表、也无自登记,且件在随包快照里存在 |
+
+*自登记(2026-09-16 新增)*: 振动侧的产物名随"窗名/分片序号"变化,写不进精确路径表,而按名字通配会误伤
+同名旧件(`报告_CMS振动状态评估报告_*.md` 既有随包/自产的、也有厂家报告转录的)。改成**谁算的谁登记**:
+`vib_raw_build.py` / `vib_reports_build.py` 落盘后把相对路径写进 `_derived_manifest.json`,
+`products_restore_missing.py` 生成台账时据此记为 `raw-derived`。漏登记可用
+`python scripts/vib_raw_build.py --register-only` 幂等补登记。
+
+---
+
+## 3. 路径约定(唯一真源 `src/paths.py`)
+
+### 3.1 全部路径助手(写新代码时只用这些)
+
+| 助手 | 返回 | 用途 |
+|---|---|---|
+| `P.ROOT` | `<安装目录>` | 一切解析的基准(`WINDSCADA_ROOT` 可覆盖,冻结构建时由启动器设) |
+| `P.RAW_ROOT` | `data/raw` | 现场原始件根(`WINDSCADA_RUDONG_SRC` / `serve.json.raw_dir` 可覆盖) |
+| `P.station_dir(name)` | `data/raw/<场站名称>` | 兜底约定位置(权威值来自 `config.raw_station_dir()` 的扫描辨识) |
+| `P.out_root(name)` | `outputs/<场>` | 产物仓根 |
+| `P.store(name)` | `outputs/<场>/windscada` | L0 标准仓(页面主取数处) |
+| `P.ont(name)` | `outputs/<场>/ontology` | 本体对象库 |
+| `P.objects_json(name)` | `…/ontology/objects.json` | 对象库文件 |
+| `P.cms(name)` | `outputs/<场>/windcms` | CMS 振动诊断产物 |
+| `P.m5(name)` | `outputs/<场>/m5_cms_tcm` | 振动线出件 / 窗索引 / 谱库 |
+| `P.tcm_replay(name)` | `outputs/<场>/tcm_compatible_replay` | TCM 兼容链回放资产 |
+| `P.sop(name)` | `outputs/<场>/sop` | SOP 中间件与评审落盘 |
+| `P.guanlan(name)` | `outputs/<场>/guanlan` | 事实契约与对外派生 |
+| `P.pitch(name)` | `outputs/<场>/pitch` | 变桨侧派生件 |
+| `P.paradigm(name)` | `outputs/<场>/paradigm_r1` | 范式实验件(E3/E5/E8 底稿)**2026-09-16 新增** |
+| `P.report_dir(name)` | `outputs/<场>/report` | 报告交付件(交接单/现场单)**2026-09-16 新增** |
+| `P.cloud(name)` | `outputs/<场>/guanlan/cloud` | 可上云面孔(脱敏后的契约/派生/页面)**2026-09-16 新增** |
+| `P.contract(name)` | `reference/<场>/windscada_contract.yaml` | 机型判据契约(属 reference 侧,不在 outputs) |
+| `P.rel(p)` / `P.disp(p)` / `P.disp_dir(p)` | 相对 POSIX 串 / 显示串 / 目录显示串 | 写进产物用 `rel()`;给人看(页面「位置」列)用 `disp()` |
+| `P.resolve(p)` | 绝对路径 | 把"可能是相对"的值按 **ROOT**(不是 cwd)解析 |
+| `P.venv_python()` / `P.python_exe()` | 解释器路径 | 跨平台探测 `.venv/Scripts/python.exe` 或 `.venv/bin/python` |
+
+### 3.2 2026-09-16 统一掉的路径问题(源码已改)
+
+| # | 位置 | 原样 | 改成 | 为什么要改 |
+|---|---|---|---|---|
+| 1 | `scripts/products_restore_missing.py` | `P.STORE if hasattr(P,'STORE') else ROOT/'outputs'/'rudong'` | `P.out_root()` | `P.STORE` **不存在** ⇒ 恒落到写死的 `rudong`;多场部署会把 rudong 的随包件补进别的场(静默串场) |
+| 2 | `scripts/ingest_ops_2025.py` | `pathlib.Path('outputs/rudong/ontology/objects.json')` + 裸 `write_text` | `_P.objects_json()` + `Store(…).save()` | cwd 相对 + 写死场名 + **绕过 store 的全库校闸与原子写**;且该脚本缺 `import os` 一直 NameError,属"上膛但没响"的凶器 |
+| 3 | `scripts/ontology_p2_verify.py` | `ROOT/'outputs/rudong/ontology/objects.json'`、`…/guanlan/facts_contract_v0.json` | `P.objects_json()`、`P.guanlan()` | 写死场名 |
+| 4 | `scripts/guanlan_facts_contract.py` | 4 条 `ROOT/"outputs/rudong/…"` | `P.guanlan()/P.sop()/P.paradigm()` | 同上(事实契约的输入/输出全在这里) |
+| 5 | `src/ontology/maintenance.py` | 自写 `ROOT=parents[2]`、`_RAW_STR`、`disp()` | `P.ROOT` / `P.RAW_ROOT` / `P.disp` | 同一文件里三处"影子真源"(显示规则两个实现) |
+| 6 | `src/sop/wrapup.py` | `…/cleaned/turbine.parquet`(硬编码) | 读 `clean_gate.json` 的 `section` → 否则取目录内唯一 parquet → 都取不到则**响亮打印** | 写侧落的是 `cleaned/<section>.parquet`(section 可变)⇒ 读侧硬编码时静默 `None`,柱2 悄悄降级 INSUFFICIENT |
+| 7 | `src/windscada/subsys/pitch.py` | `P.store()/'alarms.parquet'` | `cfg['store']`(`registry(cfg)` 透传) | `P.store()` 跟环境变量,`cfg['store']` 跟显式选定的场 ⇒ 多场下跨场串数据且不报错 |
+| 8 | `scripts/check_transferable.py` | 门户期望字节数写死 `20225828` | 运行时从主实例现量 | 门户外壳一改常量即过期,核验会打印永远不成立的"与主包不同" |
+| 9 | `src/windscada/taxonomy.py`、`subsys/fusion.py`、`scripts/windscada_serve.py` | 读侧硬绑 `P.cms()` | 新增 `windcms.config.cms_out()`,读侧与写侧同一解析口 | `WINDCMS_OUT` 只有写侧认 ⇒ 设了它就"**写到旁路、页面读生产**",表现为"重算完了页面还是旧数"且**不报错**(冻结构建自检正是这个组合) |
+| 10 | `src/windcms/pipeline.py::ingest` | 任何 `.zip/.rar/.7z` 都调 `rudong_tcm_ingest_raw.py`(**未随包**) | 先判包内容:含 `*_decode.json` → 按"已解码导出"走 `rudong_tcm_index/spectra`(支持 zip 当 root);否则报明确错误并给两条可行路径 | 把 CMS 导出 zip 直接指过来时原本必崩 `FileNotFoundError`,而其实不需要那个脚本 |
+
+### 3.3 仍然存在的不统一(**如实列出,未擅自大改**)
+
+| 类别 | 具体 | 影响 | 处置建议 |
+|---|---|---|---|
+| 一次性/历史脚本自拼路径 | `scripts/guanlan_cloud_{face,page,qa}.py`、`guanlan_portal_inject_claims.py`、`guanlan_baseline_manifest.py`、`llm_known_answer.py` 里仍有 `outputs/rudong/…` 字面量 | 只影响这些**一次性派生/云端面孔**脚本;多场时需手改 | 换场前把这几处改走 `P.guanlan()/P.cloud()/P.store()`(`P.paradigm()/P.report_dir()/P.cloud()` 已备好) |
+| 历史随包留档 | `scripts/sim_hub/*`(含 `/Users/yuanying/...` mac 路径)、`outputs/rudong/sop/*.py` | **非运行期**(装配脚本留档) | 已在 `check_portability.py` 具名登记;随包只为"可重放",不参与运行 |
+| 多场机制未接通 | `src/windscada/config.py::available()` glob `configs/farms/*.json`,而磁盘上 11 个场配置是 `*.yaml` | 现在只有内置 `rudong` 可用;`cfg['store']` 恒等于 `P.store('rudong')` | 二选一:把 `available()` 同时认 `*.yaml`,或把场配置转成 json |
+| `store` seam 三种写法 | `cfg['store']`(14 处写点)/ `P.store()`(1 处,已修)/ `ROOT/'outputs'/…`(4 处,已修 3 处) | 多场前无实际差异,属"潜在缺陷" | 新代码一律 `cfg['store']` 或 `P.store(name)` |
+| 配置/依赖指向不存在的件 | `configs/scenario_registry.yaml`、`configs/<场>/value_assumptions.yaml`、`configs/analysis_lock*.yaml`、`scripts/sop_check.py`、`scripts/analysis_lock_check.py`、`docs/审核规则_经验固化_v1.md`、`docs/振动诊断模型_六层_v1.md`、`.claude/skills/…`、`m5_cms_tcm/model_run_l6.parquet`、`m5_cms_tcm/fusion_38.csv` | SOP 场景模块读注册表必抛;锁闸恒空转;`wrapup` 电价只能走"假定锚";CMS 报告**不能重生成**(会掉内容) | 见 §7 缺口表;这批是"随包没发全",要么补件要么显式降级(`wrapup` 已改为响亮说明) |
+| 被引用但未随包的脚本(10 个) | `rudong_tcm_ingest_raw.py`(已加前置判断,见 §3.2-10)、`rudong_tcm_oem_scan.py`、`rudong_line_energy_share.py`、`rudong_model_run.py`、`rudong_fusion_run.py`、`rudong_build_baseline.py`、`dsh_learn_page.py`、`intake_scan.py`、`guanlan_cloud_sims.py`、`deploy_gate_check.py` | 前五个属振动六层链(见 §7);其余是开发/部署辅助脚本的引用 | 逐个二选一:补件,或在调用处加"缺失即明确报错 + 替代路径"(`pipeline.ingest` 已按此改) |
+| 有读无写的"孤儿产物"(12 件) | `genbearing_monthly` / `mblub_monthly` / `yaw_dynamic_monthly` / `yaw1min_liveness` / `yaw_err_clean` / `sector_power` / `duty_monthly` / `pc_monthly_bins` / `thermal_monthly` / `control_monthly` / `structure.parquet` / `watch_channels_monthly` | 趋势件、热链、扇区、偏航/润滑/控制面 | 维持随包件;要"自己算"须研发给口径(照 `temp_monthly` 的办法反推 + 逐值验证) |
+| **有意**重复落盘(不是缺陷,但要知道) | ① `spectra_meta.parquet` 同时写 `<窗>/spectra/` 与 `<窗>/`(约 4.6 MB/窗)—— 为兼容 `data.spectra_meta()` 的两条读取约定;② `报告_CMS振动状态评估报告_<期>.md` 有两个来源(`report_std.py` 自算 vs `vib_reports_build.py` 厂家转录),靠**日期口径**区分(自算=当天、转录=报告期月末) | 磁盘占用翻倍(仅 meta,非谱数据);命名空间共享 | 已在此登记;`data.py` 优先读 `<窗>/spectra_meta.parquet`,两份内容逐字节一致 |
+| `P.VIEWER` 常量全仓 0 引用 | `src/paths.py` 定义了 `VIEWER`,实际取 `viewer_dir` 走 `configs/serve.json` | 无功能影响 | 下次统一时二选一(用起来或删掉),避免"看着像真源其实没人用" |
+
+---
+
+## 4. 进程与命令窗口(统一口径 `src/proc.py`)
+
+**问题**: Windows 上"**无控制台的父进程** + 裸 spawn 一个控制台程序" = 系统给子进程**新建一个可见控制台窗口**。
+本系统里无控制台的父进程很多(网关、运维动作进程),于是"点页面按钮"会闪窗、跑重算会挂一个 20 分钟的黑窗、
+`/ops` 每 2 s 轮询 `tasklist` 会反复闪窗。
+
+**统一写法**: `src/proc.py`
+
+| API | 用途 |
+|---|---|
+| `proc.NO_WINDOW` / `proc.flags()` | `CREATE_NO_WINDOW`(Windows)/ 0(POSIX)。★与 `DETACHED_PROCESS` **互斥**,只认这一种 |
+| `proc.spawn(cmd, log=…, env=…, cwd=…, **popen_kw)` | 后台起无窗口子进程;`log=` 落日志,也可自己传 `stdout=`/`stderr=` 句柄(**透传**,不覆盖) |
+| `proc.run(cmd, **kw)` / `proc.run_text(cmd)` | 前台等待的无窗口子进程(`tasklist` / `git` / 短命令);`run_text` 默认 `text=True, errors='replace'`(控制台输出是 GBK,严格解码会得到 `None`) |
+
+已收敛的调用点(**凡"父进程无可见控制台"的入口都必须走这里**): `guanlan.py`(组件 spawn、`tasklist`、`taskkill`)、
+`scripts/_ops_launch.py`、`scripts/_ops_run.py`(整套重算的宿主进程)、`scripts/guanlan_ops.py`(启动器调用 + 每 2 s 的 `tasklist` 轮询)、
+`scripts/guanlan_gateway.py`(`git`,每次 `/api/version`)、`scripts/_ops_start_and_open.py`(起 serve)、
+`src/windcms/plugins.py`(CMS 页面「分析」按钮起的全链)、`scripts/windscada_serve.py`(详情工作台云端档问答)。
+
+**判定口径**(新加子进程前先问一句): 我的父进程**有没有可见控制台**?有(用户从终端跑的 CLI)→ 让它继承,用户能看见进度;
+没有(服务/网关/运维动作进程)→ **必须**走 `src/proc.py`,否则 Windows 会给子进程新建一个可见窗口。
+
+### 4.1 三个"静默故障"坑(都实炸过,已写进 `guanlan.py check` 作回归守卫)
+
+| # | 坑 | 现场表现 | 处置 |
+|---|---|---|---|
+| 1 | `CREATE_NO_WINDOW` 新建的控制台**会把子进程标准句柄吸走** | 从页面点"执行重算":`_ops_run` 自己那行进了日志,`rebuild_all.py` 及之后**一个字都没有**,任务随后挂住(0 CPU / 0 I/O / 1 线程);页面只剩"运行中" | `src/proc.py::_inherit_stdio()` 把父进程当前 `sys.stdout/stderr` **显式**交给子进程(STARTF_USESTDHANDLES) |
+| 2 | `capture_output=True` 与显式 `stdout=` **互斥**,同给抛 `ValueError` | 异常被 `guanlan_ops.job_running()` 的 `except` 吞成 `alive=False` ⇒ **任何在跑的任务都被立刻改写成"被强杀"**,页面显示完成、重算按钮重新可点(可能并发起两个重算) | `_inherit_stdio` 先让路(有 `capture_output` 就什么都不加);自检加一条"run_text(capture_output) 可用" |
+| 3 | 内嵌版 `/ops/recalc` 裁掉了「服务」卡,但 JS 仍 `$('#b_start').onclick=…` | 元素为 null → 抛 `TypeError` → **后面所有按钮绑定与 `refresh()` 全不执行** ⇒ 用户看到的正是"清除产物/执行重算点了没反应" | 所有元素访问与绑定走 `set()`/`on()` 容错包装;并加 try/catch 把脚本错误**显示在页面上**(不再静默) |
+
+**服务端配套**: `/ops*` 响应现在带 `Cache-Control: no-store, must-revalidate` —— 否则页面 HTML/脚本更新后浏览器仍用旧版,
+表现同样是"改了没反应"(`guanlan_gateway.py::_bytes(no_store=True)`)。
+
+### 4.2 用户怎么启动(不弹命令窗口)
+
+| 入口 | 行为 |
+|---|---|
+| `start_hidden.vbs`(推荐,可建桌面快捷方式) | `WScript.Shell.Run(..., 0, False)` 以 `pythonw.exe` 跑启动器 —— **完全不出现命令窗口** |
+| `start.bat`(双击) | 默认走上面的隐藏路径(自己立刻退出);`start.bat console` = 旧的前台模式(排障看实时日志) |
+| `scripts/guanlan_start_hidden.py` | 隐藏启动器本体: 端口已在 → 只开浏览器(幂等);否则无窗口起 `guanlan.py serve`、等 `/healthz`、开浏览器 |
+| 运维控制台/门户「数据重算」 | 动作进程同样无窗口;进度与日志尾巴在页面里看 |
+
+**日志去哪了**: `logs/serve.log`(隐藏模式下 serve 的输出)、`logs/start_hidden.log`(启动器轨迹:就绪耗时/失败原因/是否已开浏览器)、
+各组件 `logs/<name>.log`、运维动作 `logs/ops_<动作>_<时间>.log`。失败时隐藏启动器还会弹一个**消息框**(无窗口模式下唯一能让人看见的通道)。
+
+**实测**(2026-09-16): 全停后可见窗口 0;`start.bat` 2.3 s 返回;启动后可见窗口 **0**,6 个服务端口全开,
+7 s 内 `/healthz` 就绪。
+
+---
+
+## 5. 页面取数与"产物即时进页面"
+
+### 5.1 产品数据的读取口径
+
+| 服务 | 读法 | 产物改了要不要重启 |
+|---|---|---|
+| detail `:18033`(`scripts/windscada_serve.py`) | 首次请求把产物读进内存 `_CACHE` | **不需要**:每次取数前比**产物指纹**(`products_stamp()` = 相关产物文件的 `(路径, mtime_ns, size)` 摘要,TTL 2 s),指纹变了自动重载 |
+| cms `:18020`(`src/windcms/serve.py`) | 报告/页面/文件**每请求现读** | 不需要 |
+| cms 的谱元数据(`src/windcms/data.py::spectra_meta`) | 进程内缓存,缓存键含**各窗 index/spectra_meta 指纹** | 不需要(新摄入/重算后自动失效) |
+| 门户(网关 `:28084`) | `portal.html` 按 `(mtime, size)` 缓存改写结果 | 不需要(重建门户即生效) |
+| 网关的 `/ops` 页面模块 | `guanlan_ops` 模块级缓存 | **需要重启网关**(改了这个文件才需要) |
+
+> 实测证据: 只改一个产物文件的 mtime,`logs/detail.log` 立刻出现
+> `[reload] 产物指纹变化 b15e2652… → f0b457f7…, 重载`,页面随即用新数(无需重启)。
+> 另有显式入口 `GET /detail/api/reload`(清缓存并回显新旧指纹)。
+
+### 5.2 门户菜单
+
+顶部菜单: `总览 | 系统架构 | 方法 | 经验发现 | 案例·液压 | 振动·CMS | 仿真与回放 | 交付文档 | 系统状态 | 数据重算 | 登录`
+(2026-09-16 用户令: 「数据重算」落在「交付文档」与「登录」之间)。
+
+- 门外壳正本 `release/portal_src/shell.html` → 装配 `release/portal.html`(`scripts/portal_build.py`)。
+  改外壳后流程: `--check`(看漂移)→ `--rebaseline`(**有意**改动后重设基准)→ 装配 → `--check` 应"全部一致 ✔"。
+- `#recalc` 段落用 iframe **懒加载** `http://127.0.0.1:28084/ops/recalc`(同源;切到本页才挂 src,
+  否则隐藏时它还会每 2 s 轮询 `/ops/api/state`)。
+- `/ops/recalc` 是 `/ops` 的**内嵌版**(只保留 重算 + 产物 + 动作进度;服务卡与页头由门户承担)。
+  走独立路由而不是 `?embed=`:网关转给 `guanlan_ops.handle()` 的是 `u.path`,查询串会被丢掉 —— 用查询串
+  会变成"看起来支持、实际不生效"的静默坑。原 `/ops` 整页保留不变。
+
+---
+
+## 6. 输入 ↔ 产物呼应 与 重算
+
+### 6.1 呼应关系(机器可校验)
+
+`python scripts/inventory_products.py --check` 逐类输入算跨度/条数,与它喂出的产物对拍(退出码 5 = 有问题)。
+2026-09-16 实测**呼应正常**:
+
+| 输入类 | 输入跨度 | 件数 | 对拍产物 | 产物跨度/条数 |
+|---|---|---|---|---|
+| `scada_10min` | 2025-01-01 ~ 2026-07-07 | 38 | `temp_monthly.parquet` / `loss_monthly.parquet` / `powercurve_bins.parquet` | 2025-01 ~ 2026-07 / 19494 · 3729 行 |
+| `故障报警` | 2025 ~ 2026(文件名年粒度) | 16 | `alarms.parquet` | 2025-01-01 ~ 2026-07-15 / 39211 行 |
+| `风机故障记录` | 2021 ~ 2026(年粒度) | 135 | `workorders.parquet` | 2020-01-03 ~ 2026-07-08 / 5876 行 |
+| `油样报告` | 2024-11-19 ~ 2025-08-29 | 404 | `oil_samples_index.parquet` | 同跨度 / 404 行 |
+| `windcms`(CMS 原始导出) | 2026-03 ~ 2026-04 | 25693 | `m5_cms_tcm/windows/w0316/index.parquet` | 2026-03-16 ~ 2026-04-21 / 2,066,686 行 |
+
+★ 年粒度输入(报警/工单的"年度汇总表")**不参与"产物落后"判定** —— 用文件名年份推出来的跨度天然是年粒度,
+拿它当"输入到 2026-12"会造出假缺口(本脚本第一版就这么误报过 2 条)。
+
+### 6.2 重算的两种运行状态(都实测过)
+
+| 状态 | 入口 | 说明 |
+|---|---|---|
+| **系统运行中** | 门户「数据重算」或 `/ops` 的「执行重算」按钮;等价命令行 `python scripts/rebuild_all.py` | 动作经 `_ops_launch.py` 二次启动(不挂网关的父子树,避免 `taskkill /T` 把自己杀掉);跑完**不需要重启**就能在页面看到新数(§5.1 指纹重载)。重算期间按钮全灰、并发动作被后端拒(HTTP 409) |
+| **系统未运行** | 同一套命令行(先 `guanlan.py stop` 或本就关机状态) | 全部构建器都是普通 CLI,不依赖服务;跑完再 `start.bat`/`start_hidden.vbs` 起来,页面直接读新产物。**实测**: 全停后跑 `rebuild_from_raw.py` rc=0、`inventory_products.py --check` rc=0 |
+
+一键顺序(`rebuild_all.py`,14 步): ① 放数据(`--src` 才跑)/② 三门台账/③ SCADA 侧 10 构建器/④ 月度派生件/
+**④b 振动侧摄入**/⑤ 补齐随包件(★2026-09-16 用户令"清除产物不留备份"之后, 随包件不再有 `_products_off/` 暂存区
+⇒ 该步固定返回 6 跳过并说明; 要补齐须显式给交付包 `products_restore_missing.py --stash <交付包.zip>`, **不再打断整条链**)/
+⑥ 重启组件服务/
+⑦ 本体六步/⑧ 本体审计(+可选等价验收)。
+
+---
+
+## 7. 缺口与边界(如实列)
+
+| 缺口 | 影响的产物 | 依据/证据 | 补齐判据 |
+|---|---|---|---|
+| 振动六层链四步脚本未随包(`rudong_tcm_oem_scan` / `rudong_line_energy_share` / `rudong_model_run` / `rudong_fusion_run`) | `oem_frequency_scan`、`gear_freq_scan`、`blade_1p_*`、`model_run_l6.parquet`、`fusion_38.csv` 等 | `outputs/<场>/m5_cms_tcm/vib_raw_manifest.json` 的 `missing_chain` | 给脚本或口径;源件已在 `data/raw/<场>/windcms/` |
+| 厂商月度报告 12 份 PDF 是**纯扫描件** | 「CMS 振动评估报告」的厂商侧 | pypdf 实测 12 页 12 图、`extract_text()` 长度 0 | 现场给电子件(docx/xlsx),或上 OCR(本包不装) |
+| `handoff_vibration_v2.json` / `component_history.json` 是随包快照 | 融合面判级、`/cms/` | 该件含人工裁决/校准更新,不是测量数据的函数 | 现场给正本,放 `data/raw/<场>/m5_cms_tcm/` |
+| 随包件"暂存区"**按设计不再存在**(2026-09-16 用户令"清除产物不留备份") | 第⑤步"补齐随包件"固定返回 6 跳过 | 旧设计把清掉的产物挪到 `_products_off/` 以便还原;现口径 = 真删除 | 要补回"包内没有生成端"的随包件: `python scripts/products_restore_missing.py --stash <交付包.zip>`(从交付件按需补齐, **不在安装目录里留备份**)。`products_state.py --off` 需显式 `--yes`;`--on` 已移除 |
+| 台账等价验收基线 `outputs/<场>/windscada/_pre_rebuild_20260911/` 曾缺失 | 第④步与 `rebuild_from_raw.py --verify` 的逐值比对(缺基线时 ④ 返回 5) | 该目录在打包时被"清除产物"挪进了暂存区, 而暂存区随后被清掉 ⇒ 机器上无标准答案 | **2026-09-16 已重建**: 从当天 dist 包里取 4 件(`alarms/workorders/oil_samples_index/temp_monthly`)+ `_来源说明.json` 标明来历。★链已加固: ④ 容忍 rc=5 —— **缺基线只跳过"等价验收", 不再打断整条重算链**。注意: 「清除产物」会把这个目录一并删掉(它在 `windscada/` 里), 届时需按同样办法重建 |
+| 油样 2026-07 批 102 行(华标合并报告) | 数据层「油液化验」 | 源件 `BG-2026-07-YP013 ….pdf` 不在现场包 | 补那份 PDF 后 `rebuild_from_raw.py --verify` 无人工项 |
+| 无生成端的组级产物(`pc_monthly_bins` / `duty_monthly` / `thermal_monthly` / `sector_power` / `yaw_*` / `genbearing_monthly` / `mblub_monthly` / `pitch_daily` 等) | 趋势件、热链、扇区、偏航/润滑面 | 全库只有读取方、0 处写入方 | 研发补口径(照 `temp_monthly` 的办法反推 + 逐值验证) |
+| `configs/farms/*.yaml` 无代码读者;`available()` 只认 `*.json` | 多场部署 | `src/windscada/config.py::available()` | 二选一(见 §3.3) |
+
+---
+
+## 8. 变更记录
+
+| 日期 | 变更 |
+|---|---|
+| 2026-09-16 | **本文建立**(用户令 1): 产物全景(自动生成清单)+ 路径唯一真源与本次统一的 8 处 + 未统一项如实列表 + 进程无窗口口径 `src/proc.py` + 产物指纹热重载 + 输入↔产物呼应 + 重算两种状态 + 门户「数据重算」 |
+| 2026-09-16 | 用户令 2/3/4/5 的落地: `inventory_products.py`(清点+呼应校验,`--check` 出码);⑤步不再打断重算链(`StashMissing` rc=6);`start_hidden.vbs` + `start.bat` 默认无窗口;门户菜单新增「数据重算」+ `/ops/recalc` 内嵌版;`portal_build.py --rebaseline` |
+| 2026-09-16 | 振动侧接入(详见 `docs/振动数据接入_v0.1.md`): `data/raw/<场>/{windcms,m5_cms_tcm}` 两类源件、`rudong_tcm_index.py`/`rudong_tcm_spectra.py`/`vib_raw_build.py`/`vib_reports_build.py`、窗 `w0316` |
+| 2026-09-16 | 用户令"清除产物不留备份": `products_state.py --off --yes` 改为**真删除**(不再产生 `_products_off*/`)、`--on` 与门户「恢复产物」按钮移除;随包件的唯一来源改为**交付包 zip**(`products_restore_missing.py --stash <交付包.zip>`);`derived_manifest.prune()` 清掉陈旧自登记(`raw-derived` 台账 3444 → 1740 件,回到真实) |
+| 2026-09-16 | 用户令"打包不含 输入数据/产物/日志" → 交付包 **v0.4.0**(见 §9): `pack_dist.py` 增 `--no-products`、`VERSION='0.4.0'`、`dist-manifest.json` 记 `no_data/no_products/no_logs` 与逐条排除理由;`guanlan.py check` 读该清单,产物缺失显示 `[--] 待重算` 而非 FAIL |
+
+---
+
+## 9. 交付包组成(v0.4.0,2026-09-16 用户令:不含 输入数据 / 产物 / 日志)
+
+打包器:`scripts/pack_dist.py`。默认文件名 `<父目录>/guanlan-v<VERSION>_dist_<日期>.zip`,可用 `--out` 指定。
+
+| 进包 | 内容 |
+|---|---|
+| `src/` `scripts/` `guanlan.py` | 程序与全部构建/运维脚本 |
+| `configs/` `resources/` `reference/` | 配置(含 `serve.json` 端口真源、场站/通道台账) |
+| `release/` | 门户、仿真页、三维资产 `viewer/`、治理清单交付件 `release/如东/`(客户交付物,勿外传) |
+| `docs/` `README_先读我.txt` `测试须知.txt` | 交付文档与说明书 |
+| `wheels/win_amd64/`(42 件)、`vendor/python/`(3 平台便携运行时) | 离线安装件:Windows 完全离线可装;Linux/macOS 走联网安装(要离线就把轮子放进 `wheels/linux_x86_64` / `wheels/macos_arm64`) |
+| `install.bat/.ps1/.sh`、`check/start/stop.bat`、`requirements.txt` | 安装与起停入口 |
+
+| 不进包 | 理由(同时写进包内 `dist-manifest.json` 的 `excluded`) |
+|---|---|
+| `data/` | 现场原始输入件;约定"原始件不随包分发"(要带用 `--with-data`) |
+| `outputs/` | **用户令**:产物不进包(`--no-products`)。目标机放数据后 `rebuild_all.py` 重算;或缺"无生成端"的随包件时用 `products_restore_missing.py --stash <交付包.zip>` |
+| `logs/` `run/` | **用户令**:日志不进包(`run/pids.json` 里的 PID 到新机器上是无效引用) |
+| `.venv/` `.git/` `.github/` | venv 换机必失效(安装时重建);版本库不随交付件分发 |
+| `_products_off*/` | 旧设计的"清除产物"暂存档(2026-09-16 起清除=真删,不再产生;老机器上若有可手工删) |
+| `__pycache__/` `*.pyc`、顶层 `*.zip` | 解释器缓存、旧的交付压缩包(避免包中包) |
+
+开箱验证(一条命令给出"能不能装、能不能跑"的证据):`python scripts/pack_dist.py --verify <zip>` —— 解压到
+临时目录 → 离线安装 → 起服务核验 `/`·`/detail/`·`/cms/`·`/ops`·`/healthz` → 收尾清理,要求 ≥4 个页面可用。
+★ 核验要求组件端口(18033/18020/18791/18792/64292)空闲,**验证前先 `guanlan.py stop`**,否则整段核验被跳过(打印 `[!]`)。
+

+ 4 - 3
docs/说明书_观澜如东样板v2_v0.2.md

@@ -74,11 +74,12 @@ ollama serve &
 - 门户每页首屏一句话; "经验发现" 与 "案例" 页的每条结论旁有三枚标签: **证据级** (定论 / 准定论·预警 / 候选 / 参考 / 数据不足 / 撤回)、**审级**、**处置措施** (谁、什么时候、做什么)。候选不进业主正文。
 - 门户每页首屏一句话; "经验发现" 与 "案例" 页的每条结论旁有三枚标签: **证据级** (定论 / 准定论·预警 / 候选 / 参考 / 数据不足 / 撤回)、**审级**、**处置措施** (谁、什么时候、做什么)。候选不进业主正文。
 - 工作台 `/detail/`: 系统矩阵 (38 台 × 九系统判级) → 点台号看证据窗; 检修助手 (检修四链); **问答区**: 选模型 (默认 8B; 复杂问题会自动升档到 27B 并在尾注标明), 回答后自动附"结构化事实契约"引文块; 引不到契约条时会明写"未被契约背书"。汇报纸可导出。
 - 工作台 `/detail/`: 系统矩阵 (38 台 × 九系统判级) → 点台号看证据窗; 检修助手 (检修四链); **问答区**: 选模型 (默认 8B; 复杂问题会自动升档到 27B 并在尾注标明), 回答后自动附"结构化事实契约"引文块; 引不到契约条时会明写"未被契约背书"。汇报纸可导出。
 - `/cms/`: 每台机组六层判读, 机制未定不点零件。
 - `/cms/`: 每台机组六层判读, 机制未定不点零件。
-- **运维控制台 `/ops`** (2026-09-12 新增): 停/启服务 · 执行重算 · 清除产物 —— 一个按钮一件事。
+- **运维控制台** (2026-09-12 新增; 2026-09-16 起菜单入口为门户「**数据重算**」): 停/启服务 · 执行重算 · 清除产物 —— 一个按钮一件事。
   按钮按**真实状态**启用 (不能做的动作是灰的; 后端同样校验并返回 409, 直接打 API 也绕不过);
   按钮按**真实状态**启用 (不能做的动作是灰的; 后端同样校验并返回 409, 直接打 API 也绕不过);
   "启动服务"起来后会自动打开门户; "停止组件服务"**保留控制台本身**(否则点完页面就没了);
   "启动服务"起来后会自动打开门户; "停止组件服务"**保留控制台本身**(否则点完页面就没了);
-  "清除产物"是把产物挪到 `_products_off\` (可恢复), 不动 `data\raw\`。页面同时显示各服务端口状态、
-  产物件数、逐件来源台账、验收锚点, 以及当前任务的进度/退出码/日志尾巴。
+  "清除产物"是**直接删除、不留备份**(2026-09-16 用户令; 要随包件用
+  `scripts\products_restore_missing.py --stash <交付包.zip>` 从交付包补齐), 不动 `data\raw\`。
+  页面同时显示各服务端口状态、产物件数、逐件来源台账、验收锚点, 以及当前任务的进度/退出码/日志尾巴。
 
 
 ### 4.2 问答的可信边界
 ### 4.2 问答的可信边界
 - 台号与数字必须能回抓到源记录, 回抓不到的句子被校闸拦下并提示; 两档模型都拦则**不出文并给原因** (不会硬编)。
 - 台号与数字必须能回抓到源记录, 回抓不到的句子被校闸拦下并提示; 两档模型都拦则**不出文并给原因** (不会硬编)。

+ 48 - 41
docs/重算操作手册_v0.1.md

@@ -39,8 +39,9 @@ cd /d <安装目录>
 | ① | 放数据(给了 `--src` 才跑) | `place_raw_data.py --scope full` |
 | ① | 放数据(给了 `--src` 才跑) | `place_raw_data.py --scope full` |
 | ② | 三门台账(报警/工单/油样) | `rebuild_from_raw.py` |
 | ② | 三门台账(报警/工单/油样) | `rebuild_from_raw.py` |
 | ③ | SCADA 侧 10 个构建器 | `rebuild_from_raw.py --scada` |
 | ③ | SCADA 侧 10 个构建器 | `rebuild_from_raw.py --scada` |
-| ④ | 月度派生件 | `windscada_monthly_build.py` |
-| ⑤ | 补齐"包内没有生成端"的产物 | `products_restore_missing.py` |
+| ④ | 月度派生件 | `windscada_monthly_build.py`(★无随包基线时返回 5 = "无法比对, 这不是通过" → 本步**容忍 5 并跳过等价验收**, 不再打断整条链; 要验收请先把随包件恢复成暂存区/基线) |
+| ④b | **振动侧摄入**(CMS 原始导出 → 窗索引/谱) | `vib_raw_build.py`(`--skip-vib` 可关;没振动原始件时空跑退出 0。★**不**重生成 CMS 报告/页面: 缺六层链产物时那会用残缺输入覆盖随包快照, 实测 `report.md` −97%, 见 `docs\振动数据接入_v0.1.md` §3b; 要跑用 `--vib-report`) |
+| ⑤ | 补齐"包内没有生成端"的产物 | `products_restore_missing.py`(★没有暂存区时返回 6 = "这一步没得做" → **跳过并说明**, 不打断链条; 想补齐先点"清除产物"生成暂存区) |
 | ⑥ | 重启服务(**必须**: ⑦ 要吃 `/api/fleet`) | `guanlan.py stop` / `serve` |
 | ⑥ | 重启服务(**必须**: ⑦ 要吃 `/api/fleet`) | `guanlan.py stop` / `serve` |
 | ⑦ | 本体层: 码表→铺开→决策链→趋势→检索→参数表 | `-m src.ontology.*` |
 | ⑦ | 本体层: 码表→铺开→决策链→趋势→检索→参数表 | `-m src.ontology.*` |
 | ⑧ | 本体审计 + (可选)等价验收 | `-m src.ontology.audit` 等 |
 | ⑧ | 本体审计 + (可选)等价验收 | `-m src.ontology.audit` 等 |
@@ -83,8 +84,7 @@ cd /d <安装目录>
 | **停止组件服务(保留控制台)** | 有组件服务在运行 | 只停 detail/cms/sim/sim_sys/viewer —— **保留网关**, 否则按钮点完页面就没了 |
 | **停止组件服务(保留控制台)** | 有组件服务在运行 | 只停 detail/cms/sim/sim_sys/viewer —— **保留网关**, 否则按钮点完页面就没了 |
 | **完整重启(含网关)** | 有组件服务在运行 | 分离进程先停再起; 页面断开约 15 s 后自动恢复 |
 | **完整重启(含网关)** | 有组件服务在运行 | 分离进程先停再起; 页面断开约 15 s 后自动恢复 |
 | **执行重算** | 没有任务在跑 | `rebuild_all.py`(默认跳过 SCADA; 可勾选含 SCADA / 末尾加等价验收) |
 | **执行重算** | 没有任务在跑 | `rebuild_all.py`(默认跳过 SCADA; 可勾选含 SCADA / 末尾加等价验收) |
-| **恢复产物** | 产物被清空过 | `products_state.py --on`(恢复后请点"启动服务") |
-| **清除产物(可恢复)** | 产物在位 且 没有任务在跑 | `products_state.py --off`, 暂存区有同名产物时自动加 `--archive-old` |
+| **清除产物(直接删除,不留备份)** | 产物在位 且 没有任务在跑 | `products_state.py --off --yes` —— ★2026-09-16 用户令: **不留备份、不可恢复**; 需要随包件时用 `products_restore_missing.py --stash <交付包.zip>` 从交付件补齐 |
 
 
 页面还实时显示: 各服务端口通不通 · 产物件数与分仓 · **来源台账**(raw 重算多少件 / 随包补齐多少件) ·
 页面还实时显示: 各服务端口通不通 · 产物件数与分仓 · **来源台账**(raw 重算多少件 / 随包补齐多少件) ·
 **验收锚点**(报警 39211 行 · 工单 5876 · temp_monthly 19494 · 本体 9702 对象) · 最近一次动作的状态、
 **验收锚点**(报警 39211 行 · 工单 5876 · temp_monthly 19494 · 本体 9702 对象) · 最近一次动作的状态、
@@ -238,26 +238,33 @@ cd /d <安装目录>
 `--verify` 拿**随包基线**逐键逐值比对重算件, 把差异分成"旧链已知缺陷 / 本链新增覆盖 / 无法归类"三类,
 `--verify` 拿**随包基线**逐键逐值比对重算件, 把差异分成"旧链已知缺陷 / 本链新增覆盖 / 无法归类"三类,
 只有第三类才算不通过(退出码 4)。
 只有第三类才算不通过(退出码 4)。
 
 
-基线**不参与运行**, 平时是单独存着的(`_products_off\_baseline_kept\`), 验收时临时放回:
+基线是**标准答案**, 不参与运行; 它就在 `outputs\rudong\windscada\_pre_rebuild_20260911\`(4 件)。
+若该目录不在(例如刚清过产物), 先从交付包把它取回来:
 
 
 ```bat
 ```bat
-:: ① 临时放回基线(3 件 1.1 MB)
-xcopy /E /I /Y "_products_off\_baseline_kept\_pre_rebuild_20260911" "outputs\rudong\windscada\_pre_rebuild_20260911"
+:: ① 基线不在时: 从交付包取回台账件(或直接从同事的安装目录拷这个目录)
+.venv\Scripts\python.exe scripts\products_restore_missing.py --stash <交付包.zip>
 
 
 :: ② 验收(只比不算)
 :: ② 验收(只比不算)
 .venv\Scripts\python.exe scripts\rebuild_from_raw.py --verify
 .venv\Scripts\python.exe scripts\rebuild_from_raw.py --verify
-
-:: ③ 验完撤走, 让产物仓只留重算出来的东西
-rmdir /S /Q "outputs\rudong\windscada\_pre_rebuild_20260911"
 ```
 ```
 
 
 **期望结论**: `alarms` 39211/39211 逐值完全一致; `workorders` 8 行差异全部可归类(旧链的时间解析缺陷)
 **期望结论**: `alarms` 39211/39211 逐值完全一致; `workorders` 8 行差异全部可归类(旧链的时间解析缺陷)
 + 69 行本链新增覆盖; **1 处"需人工看" = 油样 506→404**(那 102 行的源件不在包里)。
 + 69 行本链新增覆盖; **1 处"需人工看" = 油样 506→404**(那 102 行的源件不在包里)。
 
 
-## 5. 重启服务(**必做**)
+> ★ 基线目录位于 `windscada/` 内 ⇒ **「清除产物」会把它一并删掉**。要长期保留, 把它复制到包外
+> (例如交付包 zip 里那份), 或清完再用上面第 ① 步取回。
+
+## 5. 重启服务(通常**不必**)
+
+2026-09-16 起, 产物变了**不需要重启**: detail 服务每次取数前比"产物指纹", 变了自动重载
+(实测日志 `[reload] 产物指纹变化 …`); CMS 的报告/页面是**每请求现读**; 门户按 `(mtime,size)` 失效。
+所以重算完直接刷新页面即可。以下两种情况才需要重启:
 
 
-产物换了必须重启: CMS 等模块是**启动时**加载数据的, 不重启页面还是旧数; 反过来, 产物被挪走期间起的
-那次 CMS 会直接崩(表现为 `/cms/` 503)。
+- 改的是 `scripts/guanlan_ops.py`(运维控制台页面/接口) → **必须重启网关**(它被 `_ops_module()` 缓存);
+- 产物被清空期间起的 CMS 服务可能已崩(表现 `/cms/` 503) → 点一次"启动服务"。
+
+接口层的显式重载入口: `GET /detail/api/reload`(清缓存并回显新旧指纹)。
 
 
 ```bat
 ```bat
 .venv\Scripts\python.exe guanlan.py stop
 .venv\Scripts\python.exe guanlan.py stop
@@ -266,36 +273,33 @@ rmdir /S /Q "outputs\rudong\windscada\_pre_rebuild_20260911"
 
 
 期望: `状态 degraded offline  模块 5/7 ok`, 并列出 DOWN 的项。
 期望: `状态 degraded offline  模块 5/7 ok`, 并列出 DOWN 的项。
 
 
-## 5b. 清掉产物 / 再重算(完整操作, 2026-09-12 实测)
+## 5b. 清掉产物 / 再重算(完整操作, 2026-09-16 口径)
 
 
 ### 清掉产物(人工检查空状态用)
 ### 清掉产物(人工检查空状态用)
 
 
+> ★ 2026-09-16 用户令: **清除产物不留备份**。清掉就是**真删除**, 不再有 `_products_off/` 可 `--on` 还原;
+> 要随包件就从**交付包 zip** 按需补齐。演练实测(2026-09-16): 清 11 仓 / 2316 件 / 3499 MB 用时 <10 s;
+> 从交付包补回 589 件用时 <1 min。
+
 ```bat
 ```bat
-:: ① 清掉 (整体挪走, 不是删除 —— 一条命令能还原)
-.venv\Scripts\python.exe scripts\products_state.py --off --archive-old
-::    暂存区里已有上一代产物时**必须加 --archive-old**: 否则会被拦住(见下), 因为挪走的落点是
-::    _products_off\<场>\<目录>, 那里已有同名目录时 shutil.move 会嵌套成 windscada\windscada\,
-::    之后 --on 就把两代混在一起还原。加了它 = 自动把上一代存档到 _products_off_prev_<时间戳>\
-::    (随包基线 _baseline_kept\ 留在原地, rebuild_from_raw --verify 还要用)。
-
-:: ② 重启服务 (产物没了必须重启, CMS 等是启动时加载数据)
-.venv\Scripts\python.exe guanlan.py stop
-.venv\Scripts\python.exe guanlan.py serve
+:: ① 清掉 (真删除, 不可恢复; 必须显式加 --yes 防手滑)
+.venv\Scripts\python.exe scripts\products_state.py --off --yes
 
 
-:: ③ 看空状态: / 门户仍 200 (门户在 release\, 不是产物); /cms/ 变 503;
-::    /detail/api/fleet 回 err=no_products; 维护页「数据层」各项 条数=0 / 存在=False
+:: ② 看空状态: / 门户仍 200 (门户在 release\, 不是产物);
+::    /detail/api/fleet 回 err=no_products; /cms/ 变 503; 维护页「数据层」各项 条数=0 / 存在=False
+::    (服务不必重启: 产物指纹变了页面会自动重载; 见 §5c)
 
 
-:: ④ 还原
-.venv\Scripts\python.exe scripts\products_state.py --on
-.venv\Scripts\python.exe guanlan.py stop ; .venv\Scripts\python.exe guanlan.py serve
+:: ③ 恢复: 从**交付包 zip** 按需补回"包内没有生成端"的随包件
+.venv\Scripts\python.exe scripts\products_restore_missing.py --stash <交付包.zip>
+::    再把 raw 派生件重算出来 (或按需只跑某几环):
+.venv\Scripts\python.exe scripts\rebuild_all.py --skip-scada
 ```
 ```
 
 
-**不想用开关也行**(整目录搬走, 最直观):
+**不删产物、只想验证"没有产物会怎样"**: 把整个 `outputs\<场>` 目录改个名再改回来(手工), 效果等同清空且随时可逆:
 
 
 ```bat
 ```bat
-Move-Item outputs\rudong <暂存目录>       :: 清
-.venv\Scripts\python.exe guanlan.py stop ; .venv\Scripts\python.exe guanlan.py serve
-Move-Item <暂存目录> outputs\rudong       :: 还原
+Rename-Item outputs\rudong outputs\rudong_off      :: 相当于清空(页面立刻呈现无产物)
+Rename-Item outputs\rudong_off outputs\rudong      :: 还原
 ```
 ```
 
 
 > 清的是 `outputs\<场>\`(产物仓), **不动** `data\raw\`(现场数据)、`release\`(门户外壳与交付件)、
 > 清的是 `outputs\<场>\`(产物仓), **不动** `data\raw\`(现场数据)、`release\`(门户外壳与交付件)、
@@ -371,14 +375,17 @@ Move-Item <暂存目录> outputs\rudong       :: 还原
 | `populate` / `audit` 退出 1 | 卡在 `pitch\pitch_daily.parquet`(变桨产物无生成端) | 同上 |
 | `populate` / `audit` 退出 1 | 卡在 `pitch\pitch_daily.parquet`(变桨产物无生成端) | 同上 |
 | `/local-ai/` 503 | 本机 Ollama 未运行或没有模型(环境项, 与数据无关) | `ollama pull` 见 `configs\models.json` |
 | `/local-ai/` 503 | 本机 Ollama 未运行或没有模型(环境项, 与数据无关) | `ollama pull` 见 `configs\models.json` |
 
 
-## 8. 两个必须知道的坑
+## 8. 三个必须知道的坑
 
 
-1. **`products_state.py --on` 会删掉重算产物**。它按清单把随包产物挪回来, 且**目标目录存在就先 `rmtree`**
-   (`scripts\products_state.py` 第 110-112 行)。当前清单覆盖 `outputs\rudong\{ontology, windscada, …}` ——
-   也就是说: **想让"原始重算"的结果留着, 就不要跑 `--on`**。要拿随包产物就先备份自己的重算结果。
-   只想取验收基线 → 用 §4 的 xcopy, 别用 `--on`。
-2. **顺序永远是: 停服务 → 挪/放数据与产物 → 算 → 起服务**。挪产物期间起的服务会带着"没有数据"的
-   状态跑起来(CMS 直接崩), 表现为页面 503, 而数据其实是好的。
+1. **「清除产物」= 真删除, 没有备份**(2026-09-16 用户令)。`products_state.py --off` 需显式 `--yes`;
+   旧的 `--on` 已移除。清完要恢复:
+   `products_restore_missing.py --stash <交付包.zip>`(补随包件) + `rebuild_all.py`(重算 raw 派生件)。
+   清之前如果想留一份回退路, 就把整个 `outputs\<场>` 改名(§5b 的手工做法) —— 那是**你自己的临时目录**, 不是系统备份。
+2. **产物变了通常不用重启服务**(2026-09-16 起): detail 按产物指纹自动重载、CMS 每请求现读、门户按 mtime 失效。
+   只有改 `scripts/guanlan_ops.py`(控制台本身)才必须重启**网关**。
+3. **重算中的按钮是灰的, 这是设计**: 有任务在跑时「执行重算/清除产物/停服务」全禁用, 防止并发;
+   此时页面显示"正在执行 … "与实时日志尾巴(2 s 刷新)。**日志真的一直空白**才是异常(2026-09-16 修过一次:
+   `CREATE_NO_WINDOW` 会吸走子进程句柄)。
 
 
 ---
 ---
 
 
@@ -406,4 +413,4 @@ Move-Item <暂存目录> outputs\rudong       :: 还原
 ```
 ```
 
 
 相关文档: `docs\数据目录结构与落位约定_v0.2.md`(§2 落位 · §4 重算边界表) ·
 相关文档: `docs\数据目录结构与落位约定_v0.2.md`(§2 落位 · §4 重算边界表) ·
-`docs\重算缺口与补件清单_v0.1.md`(缺什么找谁补) · `_修复记录_20260911\README.md`(踩过的坑)
+`docs\重算缺口与补件清单_v0.1.md`(缺什么找谁补) · `docs\振动数据接入_v0.1.md`(振动侧 CMS/TCM 落位与摄入)

+ 24 - 9
docs/重算缺口与补件清单_v0.1.md

@@ -24,13 +24,14 @@
 | # | 缺什么 | 类型 | 卡住的页面/功能 | 找谁 | 补齐判据 |
 | # | 缺什么 | 类型 | 卡住的页面/功能 | 找谁 | 补齐判据 |
 |---|---|---|---|---|---|
 |---|---|---|---|---|---|
 | **A1** | 油样 2026-07 批的源件(1 份**合并报告**) | 源件 | 数据层「油液化验」**506→404**; 油液时效胶囊、油液判级 | 现场 / 化验机构 | 放进 `data\raw\如东\油样报告\` 重跑摄入后, `rebuild_from_raw.py --verify` **无人工项** |
 | **A1** | 油样 2026-07 批的源件(1 份**合并报告**) | 源件 | 数据层「油液化验」**506→404**; 油液时效胶囊、油液判级 | 现场 / 化验机构 | 放进 `data\raw\如东\油样报告\` 重跑摄入后, `rebuild_from_raw.py --verify` **无人工项** |
-| **A2** | 振动线 **CMS 测点数据(handoff)**, 随包 90 件 56 MB | 源件 | ~~`/cms/` 503~~ **已由随包件补齐(页面已恢复)**; 但要"自己算"仍需测点件 | 振动线 / 研发 | `/cms/` 由重算件驱动而非随包件 |
+| **A2** | 振动线 **CMS 测点数据(handoff)**, 随包 90 件 56 MB | 源件 | ~~`/cms/` 503~~ **已由随包件补齐(页面已恢复)**; **2026-09-12 源件已到** → 索引/谱/报告可重算(见 §1 A2) | 振动线 / 研发 | 剩余缺口见 B5 |
 | **A3** | SOP 底稿 + 范式实验件(`sop` 210 件 16.5 MB, `paradigm_r1` 29 件) | 源件(内含脚本) | ~~门户契约段、`/api/facts`~~ **已由随包件补齐** | 研发 | `/api/facts` 由重算件驱动 |
 | **A3** | SOP 底稿 + 范式实验件(`sop` 210 件 16.5 MB, `paradigm_r1` 29 件) | 源件(内含脚本) | ~~门户契约段、`/api/facts`~~ **已由随包件补齐** | 研发 | `/api/facts` 由重算件驱动 |
 | **A4** | `temp_monthly.parquet` | 生成端 | ~~工作台全标签页无数据~~ **已解决**: 已反推出口径并 100% 复现 | — | ✅ 已闭合(见上) |
 | **A4** | `temp_monthly.parquet` | 生成端 | ~~工作台全标签页无数据~~ **已解决**: 已反推出口径并 100% 复现 | — | ✅ 已闭合(见上) |
 | **B1** | 6 件组级/月度产物(见 §2) | 生成端 | 趋势件、热链、扇区、结构面 | 研发 | 提供口径后照 ① 反推+验证 |
 | **B1** | 6 件组级/月度产物(见 §2) | 生成端 | 趋势件、热链、扇区、结构面 | 研发 | 提供口径后照 ① 反推+验证 |
 | **B2** | 6 件偏航/温度/润滑产物(见 §2) | 生成端 | 偏航面、润滑面 | 研发 | 同上。**源件 `scada_1min` 已落位 12.8 GB** |
 | **B2** | 6 件偏航/温度/润滑产物(见 §2) | 生成端 | 偏航面、润滑面 | 研发 | 同上。**源件 `scada_1min` 已落位 12.8 GB** |
 | **B3** | `pitch\pitch_daily.parquet` | 生成端 | 变桨面; 本体 `populate`/`audit`(已因随包件补齐而可跑) | 研发 | 液压四列口径未知(试过日界位移/越限穿越都不对) |
 | **B3** | `pitch\pitch_daily.parquet` | 生成端 | 变桨面; 本体 `populate`/`audit`(已因随包件补齐而可跑) | 研发 | 液压四列口径未知(试过日界位移/越限穿越都不对) |
-| **B4** | CMS/TCM 兼容链(`windcms` 53 件 · `tcm_compatible_replay` 59 件) | 生成端 + 源件 | ~~`/cms/`、振动融合页~~ **已由随包件补齐** | 振动线 + 研发 | 由重算件驱动 |
+| **B4** | CMS/TCM 兼容链(`windcms` 53 件 · `tcm_compatible_replay` 59 件) | 生成端 + 源件 | ~~`/cms/`、振动融合页~~ **已由随包件补齐**; 其中**索引/谱/报告** 2026-09-12 起可重算 | 振动线 + 研发 | 由重算件驱动 |
+| **B5** | 振动六层链四步脚本(`rudong_tcm_oem_scan` / `rudong_line_energy_share` / `rudong_model_run` / `rudong_fusion_run`) | 生成端 | 扫描线/能量占比/模型层/融合层那批 parquet(`oem_frequency_scan`、`gear_freq_scan`、`blade_1p_*`、`model_run_l6`、`fusion_38.csv` 等) | 振动线 / 研发 | 给脚本或口径; 源件已在 `data\raw\如东\windcms\`(25,679 件 150 GB)。`vib_raw_manifest.json` 的 `missing_chain` 字段即此项 |
 
 
 ---
 ---
 
 
@@ -50,15 +51,29 @@
 
 
 ### A2 · 振动线 CMS 测点数据(handoff)
 ### A2 · 振动线 CMS 测点数据(handoff)
 
 
-- **缺的**: 随包里 `outputs\rudong\m5_cms_tcm\`(90 件 56 MB: handoff + TCM 兼容件 + 窗/谱/模型产物)
+> **2026-09-12 更新: 源件已到, 本条**部分闭合**。** 用户把 `CMS_RuDong_CGN_202603-04.zip`
+> (Brande TCM 导出 25,679 件 / 150 GB, 2026-03~04) 补进现场包, 落位到
+> `data\raw\如东\windcms\CMS_RuDong_CGN_202603-04\`, 并补齐了**摄入端**:
+> `scripts\rudong_tcm_index.py`(→ `windows\<窗>\index.parquet`, 54 列) +
+> `scripts\rudong_tcm_spectra.py`(→ `spectra\*.npz` + `spectra_meta.parquet`) +
+> 一键 `scripts\vib_raw_build.py`。**现在"索引/谱/报告/知识库"是可重算的了**
+> (此前连"含 `*_decode.json` 的目录"这条路径都跑不通 —— 两个摄入脚本随包缺失)。
+> **仍未闭合的**: 六层链四步(见新增 B5)、以及 handoff 接口件本身(它是振动线出件, 不是测量数据的函数)。
+
+- **缺的(原记录)**: 随包里 `outputs\rudong\m5_cms_tcm\`(90 件 56 MB: handoff + TCM 兼容件 + 窗/谱/模型产物)
   以及 `outputs\rudong\windcms\`(53 件 29 MB, 41 个 HTML 报告页)。
   以及 `outputs\rudong\windcms\`(53 件 29 MB, 41 个 HTML 报告页)。
 - **证据**: `src\windscada\subsys\fusion.py` 的 `load_handoff()` 直接抛
 - **证据**: `src\windscada\subsys\fusion.py` 的 `load_handoff()` 直接抛
   `FileNotFoundError: 振动 handoff 接口缺失 … 融合面不可静默降级`;
   `FileNotFoundError: 振动 handoff 接口缺失 … 融合面不可静默降级`;
   `src\windcms\data.py` 的 `load_scalars()` 读的是 CMS **测点索引**(不是报告); 现场包里只有
   `src\windcms\data.py` 的 `load_scalars()` 读的是 CMS **测点索引**(不是报告); 现场包里只有
-  `(8)…\振动分析报告\` 12 份**月度用印版 PDF** —— 那是成品牌报告, 不是可再加工的测量数据。
+  `(8)…\振动分析报告\` 12 份**月度用印版 PDF** —— 那是成品牌报告, 不是可再加工的测量数据,
+  **且实测为纯扫描件**(12 页 12 图, `extract_text()` 长度 0) ⇒ 连"从 PDF 反推"这条路也不通。
 - **影响面**: `/cms/` 503; 融合面判级; 台账里"振动级"这一轴; `trend_ingest`(在升/闭环证据)。
 - **影响面**: `/cms/` 503; 融合面判级; 台账里"振动级"这一轴; `trend_ingest`(在升/闭环证据)。
-- **怎么补**: 向振动线要 **CMS 测点导出/handoff 件**(按 `m5_cms_tcm` 的目录约定), 而不是 PDF 报告。
-- **判定**: `/cms/` 回 200 且能列出报告页。
+- **怎么补**: ~~向振动线要 CMS 测点导出/handoff 件~~ → **2026-09-12 已到**: CMS 测点导出落在
+  `data\raw\如东\windcms\`, TCM 侧报告落在 `data\raw\如东\m5_cms_tcm\`。
+  若现场还能给 `handoff_vibration_v2.json` **正本**, 放 `m5_cms_tcm\` 下即被优先采用。
+- **判定**: ① `/cms/` 回 200 且能列出报告页(已满足); ② `windcms` 的**谱图**能取到数据
+  (此前 `m5\spectra` 整个目录缺失, `data.spectrum()` 只能返回 None —— 摄入后可取);
+  ③ 新窗进入 `data.windows()` 分析集。细节与实测数字见 `docs\振动数据接入_v0.1.md`。
 
 
 ### A3 · SOP 底稿 + 范式实验件(事实契约的输入)
 ### A3 · SOP 底稿 + 范式实验件(事实契约的输入)
 
 
@@ -77,7 +92,7 @@
 
 
 - **缺的**: 随包 `outputs\rudong\windscada\` 里 118 件中的**非 16 件**(我们只重算得出 16 件)。
 - **缺的**: 随包 `outputs\rudong\windscada\` 里 118 件中的**非 16 件**(我们只重算得出 16 件)。
 - **证据**: 工作台取数面直接回结构化"无产物":
 - **证据**: 工作台取数面直接回结构化"无产物":
-  `{"err":"no_products","note":"temp_monthly.parquet 不存在 → outputs/rudong/windscada/temp_monthly.parquet; 先跑 scripts/rebuild_from_raw.py (SCADA 侧加 --scada) 生成产物, 或用 scripts/products_state.py --on 还原"}`。
+  `{"err":"no_products","note":"temp_monthly.parquet 不存在 → outputs/rudong/windscada/temp_monthly.parquet; 先跑 scripts/rebuild_from_raw.py (SCADA 侧加 --scada) 生成产物; "包内没有生成端"的件从交付包补齐: python scripts/products_restore_missing.py --stash <交付包.zip> (2026-09-16 起清除产物不留备份, 故不再有 --on 还原)"}`。
   注意那句提示在**这一件**上是误导的: `temp_monthly` 不在 `rebuild_from_raw.py` 的 10 个构建器里。
   注意那句提示在**这一件**上是误导的: `temp_monthly` 不在 `rebuild_from_raw.py` 的 10 个构建器里。
 - **怎么补**: 研发补 `temp_monthly` 的构建器(见 B1), 或明确它属于哪个上游产物。
 - **怎么补**: 研发补 `temp_monthly` 的构建器(见 B1), 或明确它属于哪个上游产物。
 
 
@@ -133,7 +148,7 @@
 ## 4. 补齐之后怎么验(照这个顺序)
 ## 4. 补齐之后怎么验(照这个顺序)
 
 
 > **手工操作的逐步手册见 `docs\重算操作手册_v0.1.md`** —— 含每步的期望输出、会踩的坑
 > **手工操作的逐步手册见 `docs\重算操作手册_v0.1.md`** —— 含每步的期望输出、会踩的坑
-> (尤其:`products_state.py --on` 会删掉重算产物)、以及"算不出来"对照表。下面是压缩版命令。
+> (尤其: 2026-09-16 起「清除产物」= **真删除、不留备份**, 旧的 `--on` 已移除)、以及"算不出来"对照表。下面是压缩版命令。
 
 
 ```bat
 ```bat
 :: ① 源件落位 (现场包 → 约定目录; 同尺寸自动跳过, 可反复跑)
 :: ① 源件落位 (现场包 → 约定目录; 同尺寸自动跳过, 可反复跑)
@@ -158,4 +173,4 @@
 - **累计快照必须归并**: 工单/报警目录里的"全年/年至今"件是季度件的累计快照, 摄入时自动跳过; 跨源内容重复自动归并 ——
 - **累计快照必须归并**: 工单/报警目录里的"全年/年至今"件是季度件的累计快照, 摄入时自动跳过; 跨源内容重复自动归并 ——
   新加源件不会让台账虚高(本轮 `大部件维修记录` 就是被归并的)。
   新加源件不会让台账虚高(本轮 `大部件维修记录` 就是被归并的)。
 - **本清单的维护**: 新增一项的条件是"能用一条命令重复验证它的缺失"; 每项都要写清证据路径与补齐判据。
 - **本清单的维护**: 新增一项的条件是"能用一条命令重复验证它的缺失"; 每项都要写清证据路径与补齐判据。
-- 相关文档: `docs\数据目录结构与落位约定_v0.2.md`(§4 重算边界表) · `_修复记录_20260911\README.md`(从零重算一节)
+- 相关文档: `docs\数据目录结构与落位约定_v0.2.md`(§4 重算边界表) · `docs\振动数据接入_v0.1.md`(振动侧接入与缺口)

+ 47 - 3
guanlan.py

@@ -87,7 +87,12 @@ def services(c):
 def spawn(cmd, log: Path, e):
 def spawn(cmd, log: Path, e):
     f = open(log, "ab")
     f = open(log, "ab")
     kw = dict(cwd=str(ROOT), env=e, stdout=f, stderr=subprocess.STDOUT, stdin=subprocess.DEVNULL)
     kw = dict(cwd=str(ROOT), env=e, stdout=f, stderr=subprocess.STDOUT, stdin=subprocess.DEVNULL)
-    if WIN: kw["creationflags"] = subprocess.CREATE_NEW_PROCESS_GROUP | getattr(subprocess, "DETACHED_PROCESS", 0)
+    if WIN:
+        # 2026-09-16: 统一到 CREATE_NO_WINDOW (src/proc.py) —— 原先用 DETACHED_PROCESS, 子进程没有控制台,
+        # 而 .venv 的 python.exe 是**转发器**, 它会再拉一个真解释器; 落地实测那一步会冒出可见黑窗。
+        # CREATE_NO_WINDOW 给的是"没有窗口的控制台", stdio 照旧重定向到日志, 且与 DETACHED 互斥故只用前者。
+        from src import proc as _proc
+        kw["creationflags"] = _proc.flags()
     else: kw["start_new_session"] = True
     else: kw["start_new_session"] = True
     return subprocess.Popen(cmd, **kw).pid
     return subprocess.Popen(cmd, **kw).pid
 
 
@@ -99,14 +104,18 @@ def alive(pid):
         # errors="replace" 必须有: tasklist 的输出是**控制台代码页**(中文 Windows = GBK), 而本进程
         # errors="replace" 必须有: tasklist 的输出是**控制台代码页**(中文 Windows = GBK), 而本进程
         # 可能是 PYTHONUTF8=1 起的(默认文本编码变 UTF-8) → 解码失败会让 r.stdout 变成 None,
         # 可能是 PYTHONUTF8=1 起的(默认文本编码变 UTF-8) → 解码失败会让 r.stdout 变成 None,
         # 下一行 `str(pid) in None` 直接 TypeError。2026-09-12 由运维控制台的动作进程实测逮到。
         # 下一行 `str(pid) in None` 直接 TypeError。2026-09-12 由运维控制台的动作进程实测逮到。
-        r = subprocess.run(["tasklist", "/FI", f"PID eq {pid}"], capture_output=True, text=True, errors="replace")
+        # ★ 走 src.proc.run_text: 本函数也会被**无控制台**的进程调用 (网关/运维任务), 裸 spawn tasklist
+        #   会给它新建可见控制台 → 反复闪窗 (2026-09-16)。
+        from src import proc as _proc
+        r = _proc.run_text(["tasklist", "/FI", f"PID eq {pid}"])
         return r.stdout is not None and str(pid) in r.stdout
         return r.stdout is not None and str(pid) in r.stdout
     try: os.kill(pid, 0); return True
     try: os.kill(pid, 0); return True
     except OSError: return False
     except OSError: return False
 
 
 
 
 def kill(pid):
 def kill(pid):
-    if WIN: subprocess.run(["taskkill", "/PID", str(pid), "/T", "/F"], capture_output=True)
+    from src import proc as _proc
+    if WIN: _proc.run(["taskkill", "/PID", str(pid), "/T", "/F"], capture_output=True)
     else:
     else:
         try: os.killpg(os.getpgid(pid), signal.SIGTERM)
         try: os.killpg(os.getpgid(pid), signal.SIGTERM)
         except OSError:
         except OSError:
@@ -132,13 +141,48 @@ def cmd_check(c):
         for it in w:
         for it in w:
             if issubclass(it.category, SyntaxWarning): bad.append(f"{Path(str(it.filename)).name}:{it.lineno} {it.message}")
             if issubclass(it.category, SyntaxWarning): bad.append(f"{Path(str(it.filename)).name}:{it.lineno} {it.message}")
     row(f"源码可编译 ({len(srcs)} 个文件, 无非法转义/语法错)", not bad, "; ".join(bad[:3]))
     row(f"源码可编译 ({len(srcs)} 个文件, 无非法转义/语法错)", not bad, "; ".join(bad[:3]))
+    # 子进程口径自检 (2026-09-16 加): 这三条是**回归守卫** —— 它们各自都真炸过一次, 而炸法都是"静默":
+    #   ① capture_output=True 与显式 stdout 互斥 → 异常被 except 吞成 alive=False ⇒ 在跑的任务被
+    #      误判成"被强杀", 页面显示完成、重算按钮可再点(可能并发重算);
+    #   ② CREATE_NO_WINDOW 新建的控制台会把子进程标准句柄吸走 ⇒ 重算日志一个字都没有、任务挂住;
+    #   ③ 不弹窗 (NO_WINDOW 位必须真的设上, 否则"点按钮闪黑窗"回来)。
+    try:
+        from src import proc as _proc
+        _r = _proc.run_text([sys.executable, "-c", "print('proc-ok')"])
+        row("子进程口径: run_text(capture_output) 可用", _r.returncode == 0 and "proc-ok" in (_r.stdout or ""),
+            "src/proc.py (capture_output 与显式 stdout 互斥, 冲突会让 job_running 误判任务已死)")
+        import tempfile as _tf
+        _lg = Path(_tf.gettempdir()) / "_guanlan_proc_check.log"
+        _p = _proc.spawn([sys.executable, "-c", "print('spawn-ok')"], log=_lg); _p.wait(timeout=60)
+        _txt = _lg.read_text(encoding="utf-8", errors="replace") if _lg.exists() else ""
+        try: _lg.unlink()
+        except OSError: pass
+        row("子进程口径: spawn(log=) 输出进日志", "spawn-ok" in _txt,
+            "CREATE_NO_WINDOW 会吸走句柄, 不显式继承则日志空白 (重算看起来卡住)")
+        row("子进程口径: 无窗口位已设 (CREATE_NO_WINDOW)", bool(_proc.flags() & 0x08000000) or os.name != "nt",
+            f"flags={hex(_proc.flags())}")
+    except Exception as _e:
+        row("子进程口径自检", False, f"{type(_e).__name__}: {_e}")
     for m in ("numpy", "pandas", "pyarrow", "polars", "yaml", "matplotlib", "plotly", "jinja2", "docx"):
     for m in ("numpy", "pandas", "pyarrow", "polars", "yaml", "matplotlib", "plotly", "jinja2", "docx"):
         try: __import__(m); row(f"依赖 {m}", True)
         try: __import__(m); row(f"依赖 {m}", True)
         except Exception as ex: row(f"依赖 {m}", False, f"未安装: {ex.__class__.__name__} (运行 install 脚本)")
         except Exception as ex: row(f"依赖 {m}", False, f"未安装: {ex.__class__.__name__} (运行 install 脚本)")
+    # ★2026-09-16 用户令: 打包"不含 输入数据/产物/日志"。产物缺失时, 若包内 dist-manifest.json 就声明了
+    #   no_products/no_data, 这里显示 `--`(待重算) 而**不是 FAIL** —— 否则开箱自检会把"本来就该重算"的
+    #   状态报成失败, install 脚本收尾的那次 check 也会失败 (把正常状态说成问题)。
+    _man = {}
+    try:
+        _mp = ROOT / "dist-manifest.json"
+        _man = jload(_mp) if _mp.is_file() else {}
+    except Exception:
+        _man = {}
+    _no_products = bool(_man.get("no_products"))
     for rel, n in (("outputs/rudong/windscada", "L0/L1 产物 (parquet)"), ("outputs/rudong/ontology/objects.json", "本体对象库"), ("outputs/rudong/sop/findings.json", "findings"), ("outputs/rudong/guanlan/facts_contract_v0.json", "事实契约"),
     for rel, n in (("outputs/rudong/windscada", "L0/L1 产物 (parquet)"), ("outputs/rudong/ontology/objects.json", "本体对象库"), ("outputs/rudong/sop/findings.json", "findings"), ("outputs/rudong/guanlan/facts_contract_v0.json", "事实契约"),
                    (c["release_dir"] + "/portal.html", "门户"), (c["release_dir"] + "/sim_sys_server.py", "仿真合页服务"), (c["release_dir"] + "/如东SWT40_控制律仿真台_20260906.zip", "仿真合页资料包"), (c["viewer_dir"], "三维 viewer 资产"), (c["sim_dir"], "仿真回放资产"), (c["release_dir"] + "/如东", "治理清单交付件 (可选)")):
                    (c["release_dir"] + "/portal.html", "门户"), (c["release_dir"] + "/sim_sys_server.py", "仿真合页服务"), (c["release_dir"] + "/如东SWT40_控制律仿真台_20260906.zip", "仿真合页资料包"), (c["viewer_dir"], "三维 viewer 资产"), (c["sim_dir"], "仿真回放资产"), (c["release_dir"] + "/如东", "治理清单交付件 (可选)")):
         p = ROOT / rel; exists = p.exists() and (any(p.iterdir()) if p.is_dir() else p.stat().st_size > 0)
         p = ROOT / rel; exists = p.exists() and (any(p.iterdir()) if p.is_dir() else p.stat().st_size > 0)
         if "可选" in n: print(f"  [{'OK' if exists else '--'}] {n}: {rel}")
         if "可选" in n: print(f"  [{'OK' if exists else '--'}] {n}: {rel}")
+        elif _no_products and rel.startswith("outputs/"):
+            print(f"  [--] {n}: {rel}  (本包按用户令**未随产物** —— 把现场包放好后跑 "
+                  f"scripts/place_raw_data.py --src <现场包目录> --scope full 再 scripts/rebuild_all.py)")
         else: row(n, exists, rel)
         else: row(n, exists, rel)
     raw = Path(c.get("raw_dir", "data/raw")); raw = raw if raw.is_absolute() else (ROOT / raw)
     raw = Path(c.get("raw_dir", "data/raw")); raw = raw if raw.is_absolute() else (ROOT / raw)
     print(f"  [{'OK' if raw.exists() else '--'}] 原始件目录 (补数据放这里): {raw}{'' if raw.exists() else ' (不存在; 原始件不随包分发, 只影响摄入命令, 不影响页面)'}")
     print(f"  [{'OK' if raw.exists() else '--'}] 原始件目录 (补数据放这里): {raw}{'' if raw.exists() else ' (不存在; 原始件不随包分发, 只影响摄入命令, 不影响页面)'}")

+ 1 - 1
release/portal_src/README.md

@@ -54,5 +54,5 @@ release/portal.html   20,226,052 B   sha256 9b6aabeb6ca18d15…
 ## 相关
 ## 相关
 
 
 - 契约与产物地图:`docs/数据目录结构与落位约定_v0.2.md` §6
 - 契约与产物地图:`docs/数据目录结构与落位约定_v0.2.md` §6
-- 产物进出 git 的约定与那次误提交的复盘:仓根 `.gitignore` 注释 + `_修复记录_20260911/README.md`
+- 产物进出 git 的约定与那次误提交的复盘:仓根 `.gitignore` 注释 + git 提交历史(`_修复记录_*` 类目录已按 2026-09-12 用户令不再随包/不再保留)
 - 人工检查空状态:`scripts/products_state.py --off/--on/--status`
 - 人工检查空状态:`scripts/products_state.py --off/--on/--status`

+ 5 - 4
release/portal_src/manifest.json

@@ -5,8 +5,8 @@
  },
  },
  "shell": {
  "shell": {
   "file": "shell.html",
   "file": "shell.html",
-  "bytes": 138544,
-  "sha256": "0a42ef684034be2f78cb8d38b16039d18a903d116d44154800d6abbaef8c3a71"
+  "bytes": 140103,
+  "sha256": "957e43d5e85ad7fe41a6e8c7e7c0e31921f4755dbba4bf7938825bd5b98f5329"
  },
  },
  "governance_sources": {
  "governance_sources": {
   "file": "governance_sources.json",
   "file": "governance_sources.json",
@@ -159,13 +159,14 @@
   }
   }
  },
  },
  "claims_section_stripped": 1,
  "claims_section_stripped": 1,
- "expected_portal_sha256": "9b6aabeb6ca18d15c3c0cd874dbb3dc1f820db050bdef7944e661d93d4ead843",
+ "expected_portal_sha256": "67ce6c881e5ffa09d6ae4e2099d8e78ce6f618349a43911164c288061d57e447",
  "line_ending": "LF",
  "line_ending": "LF",
  "notes": [
  "notes": [
   "shell.html = 门户外壳 (CSS/JS/导航/面板结构), 含 <!--@TEMPLATE:id--> 与 <!--@GOVERNANCE_SOURCES--> 标记;",
   "shell.html = 门户外壳 (CSS/JS/导航/面板结构), 含 <!--@TEMPLATE:id--> 与 <!--@GOVERNANCE_SOURCES--> 标记;",
   "templates/ 与 governance_sources.json 是**产物/交付件内容** (按\"产物不进 git\"的约定不入库); manifest 里留 sha256 以便漂移检测 (--check);",
   "templates/ 与 governance_sources.json 是**产物/交付件内容** (按\"产物不进 git\"的约定不入库); manifest 里留 sha256 以便漂移检测 (--check);",
   "contract-claims 段由 scripts/guanlan_portal_inject_claims.py 在装配最后注入 (内容来自 outputs/<场>/guanlan/…);",
   "contract-claims 段由 scripts/guanlan_portal_inject_claims.py 在装配最后注入 (内容来自 outputs/<场>/guanlan/…);",
   "锚点修复 scripts/guanlan_portal_fix_anchors.py 的输出是**代码**, 已固化在 shell.html 里;",
   "锚点修复 scripts/guanlan_portal_fix_anchors.py 的输出是**代码**, 已固化在 shell.html 里;",
-  "行尾必须是 LF: 三个脚本 (本器 + 两个注入器) 都已用字节级/newline=\"\" 写文件, 装配后有 CRLF 自检。"
+  "行尾必须是 LF: 三个脚本 (本器 + 两个注入器) 都已用字节级/newline=\"\" 写文件, 装配后有 CRLF 自检。",
+  "重基线 2026-09-16 16:20: shell 0a42ef684034be2f→957e43d5e85ad7fe, portal 9b6aabeb6ca18d15→67ce6c881e5ffa09 (有意改动门户外壳后重设漂移基准)"
  ]
  ]
 }
 }

+ 4 - 4
release/portal_src/shell.html

@@ -1,4 +1,4 @@
-<!DOCTYPE html><html lang="zh"><head><meta charset="utf-8"><meta name="viewport" content="width=device-width,initial-scale=1"><title>观澜中文系统 · 详细分析</title><style>:root{--ground:#070C14;--ground-2:#0B1220;--panel:#101A2B;--panel-2:#152238;--line:#1F2E44;--hair:#2A3B55;--ink:#E6EDF5;--ink-2:#C2CEDB;--muted:#8393A8;--accent:#3FC1D3;--accent-deep:#1F8FA0;--accent-soft:#0F2A35;--critical:#E4645C;--warning:#E0A43A;--good:#4CC39B;--excellent:#9AD8C4;--sans:"PingFang SC","Hiragino Sans GB","Microsoft YaHei",Inter,ui-sans-serif,system-ui,sans-serif;--mono:"SF Mono",Menlo,Consolas,monospace;--r:8px}
+<!DOCTYPE html><html lang="zh"><head><meta charset="utf-8"><meta name="viewport" content="width=device-width,initial-scale=1"><title>观澜中文系统 · 详细分析</title><style>:root{--ground:#070C14;--ground-2:#0B1220;--panel:#101A2B;--panel-2:#152238;--line:#1F2E44;--hair:#2A3B55;--ink:#E6EDF5;--ink-2:#C2CEDB;--muted:#8393A8;--accent:#3FC1D3;--accent-deep:#1F8FA0;--accent-soft:#0F2A35;--critical:#E4645C;--warning:#E0A43A;--good:#4CC39B;--excellent:#9AD8C4;--sans:"PingFang SC","Hiragino Sans GB","Microsoft YaHei",Inter,ui-sans-serif,system-ui,sans-serif;--mono:"SF Mono",Menlo,Consolas,monospace;--r:8px}
 *{box-sizing:border-box;margin:0} html{background:var(--ground)} body{background:var(--ground);color:var(--ink);font:14px/1.6 var(--sans)} a{color:var(--accent);text-decoration:none} a:hover{text-decoration:underline}
 *{box-sizing:border-box;margin:0} html{background:var(--ground)} body{background:var(--ground);color:var(--ink);font:14px/1.6 var(--sans)} a{color:var(--accent);text-decoration:none} a:hover{text-decoration:underline}
 .top{height:56px;display:flex;align-items:center;gap:18px;padding:0 24px;border-bottom:1px solid var(--line);background:#0A1120;position:sticky;top:0;z-index:5} .logo{width:30px;height:30px;border-radius:7px;background:linear-gradient(135deg,#1F8FA0,#0F2A35);display:grid;place-items:center;color:#fff;font-weight:700} .brand{font-weight:600} .brand small{display:block;color:var(--muted);font-size:11px;font-weight:400;letter-spacing:.08em}
 .top{height:56px;display:flex;align-items:center;gap:18px;padding:0 24px;border-bottom:1px solid var(--line);background:#0A1120;position:sticky;top:0;z-index:5} .logo{width:30px;height:30px;border-radius:7px;background:linear-gradient(135deg,#1F8FA0,#0F2A35);display:grid;place-items:center;color:#fff;font-weight:700} .brand{font-weight:600} .brand small{display:block;color:var(--muted);font-size:11px;font-weight:400;letter-spacing:.08em}
 .tabs{display:flex;gap:2px;margin-left:12px} .tabs a{padding:8px 12px;color:var(--ink-2);border-bottom:2px solid transparent} .tabs a.on{color:var(--ink);border-color:var(--accent)} .tabs a:hover{text-decoration:none;color:var(--ink)} .live{margin-left:auto;border:1px solid var(--accent-deep);color:var(--accent);border-radius:6px;padding:6px 12px;font-size:13px}
 .tabs{display:flex;gap:2px;margin-left:12px} .tabs a{padding:8px 12px;color:var(--ink-2);border-bottom:2px solid transparent} .tabs a.on{color:var(--ink);border-color:var(--accent)} .tabs a:hover{text-decoration:none;color:var(--ink)} .live{margin-left:auto;border:1px solid var(--accent-deep);color:var(--accent);border-radius:6px;padding:6px 12px;font-size:13px}
@@ -43,7 +43,7 @@ body{background:linear-gradient(120deg,#f1f7f9,#edf5f7);}.hero{background:transp
 @media(max-width:1050px) and (min-width:761px){#architecture>.hero>.wrap,#case_hydraulic>.hero>.wrap{padding-left:360px}#architecture>.hero>.wrap:after,#case_hydraulic>.hero>.wrap:after{width:290px}}
 @media(max-width:1050px) and (min-width:761px){#architecture>.hero>.wrap,#case_hydraulic>.hero>.wrap{padding-left:360px}#architecture>.hero>.wrap:after,#case_hydraulic>.hero>.wrap:after{width:290px}}
 @media(max-width:760px){#index>.hero>.wrap{padding:148px 20px 0}#index>.hero>.wrap:after{position:absolute;left:20px;right:20px;top:0;height:125px;margin:0;width:auto;background-size:auto 180px,cover}#index>.hero h1{font-size:29px}#architecture>.hero>.wrap,#case_hydraulic>.hero>.wrap,#method>.hero>.wrap,#admin>.hero>.wrap,#login>.hero>.wrap{padding:0 20px}#architecture>.hero>.wrap:after,#case_hydraulic>.hero>.wrap:after,#login>.hero>.wrap:after{position:relative;left:auto;right:auto;width:100%;height:160px;border-radius:12px}#method>.hero>.wrap:after{height:130px;background-size:auto 230px,cover}#cms>.hero>.wrap{padding:0 20px 150px}#cms>.hero>.wrap:after{position:absolute;left:20px;right:20px;bottom:0;width:auto;height:130px;margin:0}#documents>.hero>.wrap{padding:0 20px;min-height:0}#documents>.hero>.wrap:after{display:none}#admin>.hero>.wrap:after{height:115px}}
 @media(max-width:760px){#index>.hero>.wrap{padding:148px 20px 0}#index>.hero>.wrap:after{position:absolute;left:20px;right:20px;top:0;height:125px;margin:0;width:auto;background-size:auto 180px,cover}#index>.hero h1{font-size:29px}#architecture>.hero>.wrap,#case_hydraulic>.hero>.wrap,#method>.hero>.wrap,#admin>.hero>.wrap,#login>.hero>.wrap{padding:0 20px}#architecture>.hero>.wrap:after,#case_hydraulic>.hero>.wrap:after,#login>.hero>.wrap:after{position:relative;left:auto;right:auto;width:100%;height:160px;border-radius:12px}#method>.hero>.wrap:after{height:130px;background-size:auto 230px,cover}#cms>.hero>.wrap{padding:0 20px 150px}#cms>.hero>.wrap:after{position:absolute;left:20px;right:20px;bottom:0;width:auto;height:130px;margin:0}#documents>.hero>.wrap{padding:0 20px;min-height:0}#documents>.hero>.wrap:after{display:none}#admin>.hero>.wrap:after{height:115px}}
 </style><style id="guanlan-touch-targets">@media(max-width:760px){.tabs a,.return-dock a,.chips a,.btn,.x{min-height:44px;display:inline-flex;align-items:center;justify-content:center}.return-dock{flex-wrap:wrap}.foot{padding-bottom:115px}}</style></head><body>
 </style><style id="guanlan-touch-targets">@media(max-width:760px){.tabs a,.return-dock a,.chips a,.btn,.x{min-height:44px;display:inline-flex;align-items:center;justify-content:center}.return-dock{flex-wrap:wrap}.foot{padding-bottom:115px}}</style></head><body>
-<div class="top"><div class="logo">≋</div><div class="brand">观澜<small>WIND ASSET INTELLIGENCE</small></div><nav class="tabs"><a href="#index" data-nav="index">总览</a><a href="#architecture" data-nav="architecture">系统架构</a><a href="#method" data-nav="method">方法</a><a href="#findings" data-nav="findings">经验发现</a><a href="#case_hydraulic" data-nav="case_hydraulic">案例·液压</a><a href="#cms" data-nav="cms">振动·CMS</a><a href="#sim" data-nav="sim">仿真与回放</a><a href="#documents" data-nav="documents">交付文档</a><a href="#admin" data-nav="admin">系统状态</a><a href="http://127.0.0.1:18033/v2">登录</a></nav></div>
+<div class="top"><div class="logo">≋</div><div class="brand">观澜<small>WIND ASSET INTELLIGENCE</small></div><nav class="tabs"><a href="#index" data-nav="index">总览</a><a href="#architecture" data-nav="architecture">系统架构</a><a href="#method" data-nav="method">方法</a><a href="#findings" data-nav="findings">经验发现</a><a href="#case_hydraulic" data-nav="case_hydraulic">案例·液压</a><a href="#cms" data-nav="cms">振动·CMS</a><a href="#sim" data-nav="sim">仿真与回放</a><a href="#documents" data-nav="documents">交付文档</a><a href="#admin" data-nav="admin">系统状态</a><a href="#recalc" data-nav="recalc">数据重算</a><a href="http://127.0.0.1:18033/v2">登录</a></nav></div>
 <div class="pg" id="index" data-title="总览"><section class="hero"><div class="wrap"><div class="kicker">Wind asset intelligence · 机群诊断与分析服务</div><h1>从运行证据到检修决策</h1><p class="lead">观澜把运行数据、状态监测与检修记录落到正确的部件上,然后给出一个带证据与边界的决定。</p>
 <div class="pg" id="index" data-title="总览"><section class="hero"><div class="wrap"><div class="kicker">Wind asset intelligence · 机群诊断与分析服务</div><h1>从运行证据到检修决策</h1><p class="lead">观澜把运行数据、状态监测与检修记录落到正确的部件上,然后给出一个带证据与边界的决定。</p>
 <div class="chips"><a class="chip ev" href="#sim" data-nav="sim">进入仿真与回放 →</a><a class="chip ev" href="#cms" data-nav="cms">振动 · CMS 在线系统 →</a><a class="chip" href="#method" data-nav="method">读方法</a><a class="chip" href="#case_hydraulic" data-nav="case_hydraulic">看一个发现案例:液压系统</a><a class="chip" href="#documents" data-nav="documents">看交付文档</a></div></div></section>
 <div class="chips"><a class="chip ev" href="#sim" data-nav="sim">进入仿真与回放 →</a><a class="chip ev" href="#cms" data-nav="cms">振动 · CMS 在线系统 →</a><a class="chip" href="#method" data-nav="method">读方法</a><a class="chip" href="#case_hydraulic" data-nav="case_hydraulic">看一个发现案例:液压系统</a><a class="chip" href="#documents" data-nav="documents">看交付文档</a></div></div></section>
 <section><div class="wrap"><div class="kicker">证据到决定的系统</div><h2>多源证据分析与检修后验证</h2><p class="sub">十分钟与秒级 SCADA · 状态监测 · 报警 · 工单 · 油样与检查 · 技术文档</p>
 <section><div class="wrap"><div class="kicker">证据到决定的系统</div><h2>多源证据分析与检修后验证</h2><p class="sub">十分钟与秒级 SCADA · 状态监测 · 报警 · 工单 · 油样与检查 · 技术文档</p>
@@ -878,7 +878,7 @@ requestAnimationFrame(frame);
 <section><div class="wrap"><div class="kicker">数据层</div><h2>如东 · 五层原生 WPS 已装配</h2><div class="grid g4"><div class="card"><div class="k">10-min / 1-min</div><h3>1min 2025-01-01 → 2026-06-30 (99.87%)</h3><p>cnt 层至 2026-07-07;交叠 546 天;136 号文断点 2025 旧机制 / 2026 新机制。</p></div><div class="card"><div class="k">温度 / 电网 / 标志 / 日汇总</div><h3>native_scturtemp · scturgrid · scturflag · dailysummary</h3><p>parquet 7.1 GB,指纹文件同目录。</p></div><div class="card"><div class="k">告警 / 工单 / 油样</div><h3>alarm.tsv 2016→ · 工单 4,298 · 油样 506</h3><p>本体对象:告警码 555、工单 4,298、油样 506。</p></div><div class="card"><div class="k">CMS · TCM</div><h3>2026-01-27 单窗 + 06-29→08-11 五窗</h3><p>320 万记录 / 160 万谱 / 7,688 条波形。</p></div></div></div></section>
 <section><div class="wrap"><div class="kicker">数据层</div><h2>如东 · 五层原生 WPS 已装配</h2><div class="grid g4"><div class="card"><div class="k">10-min / 1-min</div><h3>1min 2025-01-01 → 2026-06-30 (99.87%)</h3><p>cnt 层至 2026-07-07;交叠 546 天;136 号文断点 2025 旧机制 / 2026 新机制。</p></div><div class="card"><div class="k">温度 / 电网 / 标志 / 日汇总</div><h3>native_scturtemp · scturgrid · scturflag · dailysummary</h3><p>parquet 7.1 GB,指纹文件同目录。</p></div><div class="card"><div class="k">告警 / 工单 / 油样</div><h3>alarm.tsv 2016→ · 工单 4,298 · 油样 506</h3><p>本体对象:告警码 555、工单 4,298、油样 506。</p></div><div class="card"><div class="k">CMS · TCM</div><h3>2026-01-27 单窗 + 06-29→08-11 五窗</h3><p>320 万记录 / 160 万谱 / 7,688 条波形。</p></div></div></div></section>
 <section><div class="wrap"><div class="kicker">模型档</div><h2>分析规则与模型配置</h2><div class="grid g4"><div class="card"><div class="k">本地 · 轻量</div><h3>qwen3:8b</h3><p>本体问答 8–10 s;接地闸通过才上屏。</p></div><div class="card"><div class="k">本地 · 标准</div><h3>qwen3.8:27b / qwen3:32b</h3><p>交叉审与长文;本地审核票只作初筛(已知硬伤命中 0/4)。</p></div><div class="card"><div class="k">本地 · 推理</div><h3>deepseek-r1:14b / 32b</h3><p>已下载,未标定。</p></div><div class="card"><div class="k">云端</div><h3>DeepSeek · 千问</h3><p>训练与提案侧;密钥从环境读,不落盘。</p></div></div></div></section>
 <section><div class="wrap"><div class="kicker">模型档</div><h2>分析规则与模型配置</h2><div class="grid g4"><div class="card"><div class="k">本地 · 轻量</div><h3>qwen3:8b</h3><p>本体问答 8–10 s;接地闸通过才上屏。</p></div><div class="card"><div class="k">本地 · 标准</div><h3>qwen3.8:27b / qwen3:32b</h3><p>交叉审与长文;本地审核票只作初筛(已知硬伤命中 0/4)。</p></div><div class="card"><div class="k">本地 · 推理</div><h3>deepseek-r1:14b / 32b</h3><p>已下载,未标定。</p></div><div class="card"><div class="k">云端</div><h3>DeepSeek · 千问</h3><p>训练与提案侧;密钥从环境读,不落盘。</p></div></div></div></section>
 <section><div class="wrap"><div class="kicker">审级与队列</div><h2>本体 19,321 个对象 · 写入闸 C1–C17</h2><div class="grid g4"><div class="card"><div class="k">审级</div><h3>云端跨厂商双票 = 最低审级</h3><p>上云 / 含经济数字 / 确诊 / 跨机系统性 触发。</p></div><div class="card"><div class="k">月度滚动</div><h3>新数据 → 重跑 → 本体增量 → 只报变化</h3><p>验收状态机:29# 修后首例。</p></div><div class="card"><div class="k">审计六项</div><h3>悬空引用 / 孤儿 / 链接密度 / 字段覆盖 / …</h3><p>第七项(引用未审定)待加。</p></div><div class="card"><div class="k">回归</div><h3>paradigm 14/14 · windscada 4/5 · ontology 25/28</h3><p>失败项 = 原始 CSV 硬路径缺 / 测试与对象库状态耦合,非功能故障。</p></div></div></div></section>
 <section><div class="wrap"><div class="kicker">审级与队列</div><h2>本体 19,321 个对象 · 写入闸 C1–C17</h2><div class="grid g4"><div class="card"><div class="k">审级</div><h3>云端跨厂商双票 = 最低审级</h3><p>上云 / 含经济数字 / 确诊 / 跨机系统性 触发。</p></div><div class="card"><div class="k">月度滚动</div><h3>新数据 → 重跑 → 本体增量 → 只报变化</h3><p>验收状态机:29# 修后首例。</p></div><div class="card"><div class="k">审计六项</div><h3>悬空引用 / 孤儿 / 链接密度 / 字段覆盖 / …</h3><p>第七项(引用未审定)待加。</p></div><div class="card"><div class="k">回归</div><h3>paradigm 14/14 · windscada 4/5 · ontology 25/28</h3><p>失败项 = 原始 CSV 硬路径缺 / 测试与对象库状态耦合,非功能故障。</p></div></div></div></section>
-</div><div class="pg" id="login" data-title="登录"><section class="hero" style="padding-bottom:64px"><div class="wrap"><div class="kicker">Wind asset intelligence</div><h1>欢迎来到观澜。</h1><p class="lead">面向复杂运营的 AI 辅助风电资产智能,为清晰决策而建。</p><div class="chips"><span class="chip ev">● 风电智能</span><span class="chip ev">● AI 辅助</span><span class="chip ev">● 访问受控</span></div></div></section>
+</div><div class="pg" id="recalc" data-title="数据重算"><section class="hero"><div class="wrap"><div class="kicker">运维 · 数据重算与产物</div><h1>数据重算</h1><p class="lead">从 <code>data/raw</code> 的现场数据重算全部产物:执行重算、查看产物在位状态、清除/恢复产物。这一页就是运维控制台的重算与产物两块(同一后端、同一把锁),页面顶部菜单直接到这里,不必再记 <code>/ops</code>。</p><div class="chips"><span class="chip ev">● 重算期间按钮全灰</span><span class="chip ev">● 后端拒绝 = 真生效(HTTP 409)</span><span class="chip ev">● 重算产物即时进页面</span></div></div></section><section><div class="wrap"><div class="opsframe" style="border:1px solid #dfe6ef;border-radius:10px;overflow:hidden;background:#fff"><iframe id="rcf" title="数据重算与产物" style="width:100%;height:1150px;border:0;display:block" loading="lazy"></iframe></div><p class="small muted" style="margin-top:10px">看不到内容?说明网关还没起来(本页由网关提供,组件服务停了也能用)。直接打开:<a href="http://127.0.0.1:28084/ops" target="_blank" rel="noopener">/ops 控制台</a>。</p></div></section></div><div class="pg" id="login" data-title="登录"><section class="hero" style="padding-bottom:64px"><div class="wrap"><div class="kicker">Wind asset intelligence</div><h1>欢迎来到观澜。</h1><p class="lead">面向复杂运营的 AI 辅助风电资产智能,为清晰决策而建。</p><div class="chips"><span class="chip ev">● 风电智能</span><span class="chip ev">● AI 辅助</span><span class="chip ev">● 访问受控</span></div></div></section>
 <section><div class="wrap" style="max-width:560px"><div class="form"><div class="kicker">受保护工作区</div><h2 style="margin-top:8px">欢迎回来</h2><p class="sub">登录以进入受保护的观澜工作区(本机演示:直接进入在线系统)。</p><label>用户名</label><input placeholder="username"><label>密码</label><input type="password" placeholder="••••••••"><a class="btn" href="http://127.0.0.1:18033/v2">进入详细分析 →</a></div></div></section>
 <section><div class="wrap" style="max-width:560px"><div class="form"><div class="kicker">受保护工作区</div><h2 style="margin-top:8px">欢迎回来</h2><p class="sub">登录以进入受保护的观澜工作区(本机演示:直接进入在线系统)。</p><label>用户名</label><input placeholder="username"><label>密码</label><input type="password" placeholder="••••••••"><a class="btn" href="http://127.0.0.1:18033/v2">进入详细分析 →</a></div></div></section>
 </div>
 </div>
 <div class="foot"><span>观澜中文系统 · 详细分析</span><span>判断在证据 · 模型只转述</span><span>本机交互模块按对应入口启用</span></div>
 <div class="foot"><span>观澜中文系统 · 详细分析</span><span>判断在证据 · 模型只转述</span><span>本机交互模块按对应入口启用</span></div>
@@ -886,7 +886,7 @@ requestAnimationFrame(frame);
 <!--@TEMPLATE:tpl-U6_变桨液压仿真台.html--><!--@TEMPLATE:tpl-standard_panel_zh.html--><!--@TEMPLATE:tpl-coverage_ch0_zh.html--><!--@TEMPLATE:tpl-如东传动链实际运行诊断_单文件版.html--><!--@TEMPLATE:tpl-U6_液压公共站健康报告_客户版.html--><!--@TEMPLATE:tpl-如东海上风电场_整机综合诊断与风险评估_主轴冲击深挖修订版V2.1.html-->
 <!--@TEMPLATE:tpl-U6_变桨液压仿真台.html--><!--@TEMPLATE:tpl-standard_panel_zh.html--><!--@TEMPLATE:tpl-coverage_ch0_zh.html--><!--@TEMPLATE:tpl-如东传动链实际运行诊断_单文件版.html--><!--@TEMPLATE:tpl-U6_液压公共站健康报告_客户版.html--><!--@TEMPLATE:tpl-如东海上风电场_整机综合诊断与风险评估_主轴冲击深挖修订版V2.1.html-->
 <script>
 <script>
 (function(){
 (function(){
-  function show(k){document.querySelectorAll('.pg').forEach(p=>p.classList.toggle('on',p.id===k));document.querySelectorAll('.tabs a').forEach(a=>a.classList.toggle('on',a.dataset.nav===k));window.scrollTo(0,0)}
+  function show(k){document.querySelectorAll('.pg').forEach(p=>p.classList.toggle('on',p.id===k));document.querySelectorAll('.tabs a').forEach(a=>a.classList.toggle('on',a.dataset.nav===k));/*数据重算: iframe 懒加载 —— 页面隐藏时不去轮询 /ops (它每 2 s 一次), 切到本页才挂 src*/if(k==='recalc'){var f=document.getElementById('rcf');if(f&&!f.getAttribute('src'))f.setAttribute('src','http://127.0.0.1:28084/ops/recalc')}window.scrollTo(0,0)}
   function route(){var k=(location.hash||'#index').slice(1);if(!document.getElementById(k))k='index';show(k)}
   function route(){var k=(location.hash||'#index').slice(1);if(!document.getElementById(k))k='index';show(k)}
   window.addEventListener('hashchange',route);route();
   window.addEventListener('hashchange',route);route();
   document.addEventListener('click',function(e){var a=e.target.closest('a[data-open]');if(!a)return;e.preventDefault();var t=document.getElementById('tpl-'+a.dataset.open);if(!t)return;document.getElementById('ovt').textContent=a.dataset.open;document.getElementById('ovf').srcdoc=t.content.textContent;/*guanlan-anchor-fix-v2*/(function(){var f=document.getElementById('ovf');if(!f||f._afix)return;f._afix=1;f.addEventListener('load',function(){try{var d=f.contentDocument;if(!d||d._afix)return;d._afix=1;d.addEventListener('click',function(ev){var a=ev.target&&ev.target.closest?ev.target.closest('a[href^="#"]'):null;if(!a)return;ev.preventDefault();var id=decodeURIComponent((a.getAttribute('href')||'#').slice(1));if(!id){d.defaultView.scrollTo(0,0);return}var el=d.getElementById(id)||(d.getElementsByName(id)||[])[0];if(el&&el.scrollIntoView)el.scrollIntoView({block:'start'})},true)}catch(e){}})})();document.getElementById('ov').classList.add('on')});
   document.addEventListener('click',function(e){var a=e.target.closest('a[data-open]');if(!a)return;e.preventDefault();var t=document.getElementById('tpl-'+a.dataset.open);if(!t)return;document.getElementById('ovt').textContent=a.dataset.open;document.getElementById('ovf').srcdoc=t.content.textContent;/*guanlan-anchor-fix-v2*/(function(){var f=document.getElementById('ovf');if(!f||f._afix)return;f._afix=1;f.addEventListener('load',function(){try{var d=f.contentDocument;if(!d||d._afix)return;d._afix=1;d.addEventListener('click',function(ev){var a=ev.target&&ev.target.closest?ev.target.closest('a[href^="#"]'):null;if(!a)return;ev.preventDefault();var id=decodeURIComponent((a.getAttribute('href')||'#').slice(1));if(!id){d.defaultView.scrollTo(0,0);return}var el=d.getElementById(id)||(d.getElementsByName(id)||[])[0];if(el&&el.scrollIntoView)el.scrollIntoView({block:'start'})},true)}catch(e){}})})();document.getElementById('ov').classList.add('on')});

+ 7 - 7
scripts/_ops_launch.py

@@ -24,6 +24,8 @@ import subprocess
 import sys
 import sys
 
 
 ROOT = pathlib.Path(__file__).resolve().parents[1]
 ROOT = pathlib.Path(__file__).resolve().parents[1]
+sys.path.insert(0, str(ROOT))
+from src import proc as _proc            # 无窗口子进程 (2026-09-16 用户令: 不弹命令窗口)
 
 
 
 
 def main() -> int:
 def main() -> int:
@@ -38,13 +40,11 @@ def main() -> int:
     log = pathlib.Path(a.log)
     log = pathlib.Path(a.log)
     log.parent.mkdir(parents=True, exist_ok=True)
     log.parent.mkdir(parents=True, exist_ok=True)
     env = dict(os.environ, PYTHONUTF8='1', PYTHONIOENCODING='utf-8', PYTHONUNBUFFERED='1')
     env = dict(os.environ, PYTHONUTF8='1', PYTHONIOENCODING='utf-8', PYTHONUNBUFFERED='1')
-    flags = 0
-    if os.name == 'nt':
-        flags = subprocess.CREATE_NEW_PROCESS_GROUP | getattr(subprocess, 'DETACHED_PROCESS', 0)
-    with open(log, 'ab') as f:
-        p = subprocess.Popen([sys.executable, 'scripts/_ops_run.py', '--tag', a.tag, '--'] + cmd,
-                             cwd=str(ROOT), env=env, stdout=f, stderr=subprocess.STDOUT,
-                             stdin=subprocess.DEVNULL, creationflags=flags, close_fds=True)
+    # ★ 用 src.proc.spawn (CREATE_NO_WINDOW + 新进程组), **不要**再加 DETACHED_PROCESS:
+    #   两者互斥 (前者=给一个没有窗口的控制台, 后者=不给控制台); 原先这里用 DETACHED,
+    #   而本启动器自己可能是从"无控制台"的网关侧被拉起的, 落地实测会看到黑窗 (2026-09-16)。
+    p = _proc.spawn([sys.executable, 'scripts/_ops_run.py', '--tag', a.tag, '--'] + cmd,
+                    log=log, env=env, cwd=ROOT)
     print(p.pid)          # ← 最后一行是执行器 pid, 父进程读它
     print(p.pid)          # ← 最后一行是执行器 pid, 父进程读它
     return 0
     return 0
 
 

+ 9 - 1
scripts/_ops_run.py

@@ -23,6 +23,8 @@ import threading
 import time
 import time
 
 
 ROOT = pathlib.Path(__file__).resolve().parents[1]
 ROOT = pathlib.Path(__file__).resolve().parents[1]
+sys.path.insert(0, str(ROOT))
+from src import proc as _proc            # 无窗口子进程 (2026-09-16 用户令: 不弹命令窗口)
 JOB = ROOT / 'run' / 'ops_job.json'
 JOB = ROOT / 'run' / 'ops_job.json'
 HEARTBEAT_S = 15
 HEARTBEAT_S = 15
 
 
@@ -64,7 +66,13 @@ def main() -> int:
     threading.Thread(target=_heartbeat, args=(stop_evt,), daemon=True).start()
     threading.Thread(target=_heartbeat, args=(stop_evt,), daemon=True).start()
     t0 = time.time()
     t0 = time.time()
     try:
     try:
-        rc = subprocess.run([sys.executable] + cmd, cwd=str(ROOT), env=dict(os.environ)).returncode
+        # 走 src.proc.run: 本进程是被无窗口起的 (没有可见控制台), 裸 spawn 会让 Windows
+        # 给子进程**新建一个可见控制台窗口** —— 从页面点"重算"就是这个窗口在跑 20 分钟 (2026-09-16 实测)。
+        # ★stdout/stderr **显式**传本进程的句柄: CREATE_NO_WINDOW 会新建控制台并把标准句柄重指过去,
+        #   不显式传的话, 子进程 (以及它的子进程) 的输出会掉进那个隐形控制台 —— 2026-09-16 实逮:
+        #   页面上"重算中"但 logs/ops_rebuild_*.log 里除了命令行一个字都没有。
+        rc = _proc.run([sys.executable] + cmd, cwd=str(ROOT), env=dict(os.environ),
+                       stdout=sys.stdout, stderr=sys.stderr).returncode
     except Exception as e:
     except Exception as e:
         print(f'[X] 命令起不来: {type(e).__name__}: {e}', flush=True)
         print(f'[X] 命令起不来: {type(e).__name__}: {e}', flush=True)
         rc = 99
         rc = 99

+ 4 - 1
scripts/_ops_start_and_open.py

@@ -20,6 +20,7 @@ ROOT = pathlib.Path(__file__).resolve().parents[1]
 sys.path.insert(0, str(ROOT))
 sys.path.insert(0, str(ROOT))
 
 
 import guanlan as G                                            # noqa: E402
 import guanlan as G                                            # noqa: E402
+from src import proc as _proc                                  # noqa: E402 无窗口子进程
 
 
 
 
 def main() -> int:
 def main() -> int:
@@ -27,7 +28,9 @@ def main() -> int:
     c = G.cfg()
     c = G.cfg()
     url = f"http://{c['host']}:{c['gateway']}/"
     url = f"http://{c['host']}:{c['gateway']}/"
     print(f'① 启动服务 (guanlan.py serve) …', flush=True)
     print(f'① 启动服务 (guanlan.py serve) …', flush=True)
-    rc = subprocess.run([sys.executable, 'guanlan.py', 'serve'], cwd=str(ROOT)).returncode
+    # src.proc.run: 本脚本常由**无控制台**的运维任务/网关侧拉起, 裸 spawn 会让 serve 拿到一个
+    # 可见控制台 (2026-09-16 实测: 现场会攒下标题为 .venv\Scripts\python.exe 的黑窗)。
+    rc = _proc.run([sys.executable, 'guanlan.py', 'serve'], cwd=str(ROOT)).returncode
     if rc == 2:                                   # serve 的退出码 2 = 网关 60 s 没就绪
     if rc == 2:                                   # serve 的退出码 2 = 网关 60 s 没就绪
         print(f'[X] 网关未就绪 (serve 退出码 {rc}) —— 看 logs/gateway.log')
         print(f'[X] 网关未就绪 (serve 退出码 {rc}) —— 看 logs/gateway.log')
         return rc
         return rc

+ 20 - 1
scripts/check_transferable.py

@@ -42,6 +42,11 @@ RE_ABS = re.compile(r"""(?<![\w:/])(?:[A-Za-z]:[\\/](?![\\/])|/Users/[A-Za-z0-9_
 TEXT_EXT = {'.py', '.bat', '.sh', '.ps1', '.cmd', '.json', '.md', '.txt', '.html', '.htm', '.yaml', '.yml', '.cfg', '.ini', '.toml'}
 TEXT_EXT = {'.py', '.bat', '.sh', '.ps1', '.cmd', '.json', '.md', '.txt', '.html', '.htm', '.yaml', '.yml', '.cfg', '.ini', '.toml'}
 SKIP_DIRS = {'.venv', '.git', '__pycache__', 'node_modules', '_products_off', 'logs', 'run'}
 SKIP_DIRS = {'.venv', '.git', '__pycache__', 'node_modules', '_products_off', 'logs', 'run'}
 SKIP_PATH_HINT = ('_products_off_prev',)
 SKIP_PATH_HINT = ('_products_off_prev',)
+# 扫描范围里要**排除现场原始件**: data/raw/** 是现场给的测量导出 (SCADA csv、Brande TCM JSON ——
+# 2026-09-12 落位后单这一个目录就是 2.5 万个 .json / 150 GB), 且**不随包分发** (pack_dist.py 的
+# EXCLUDE_GLOBS 含 'data')。文本闸门扫它: ①对"换机可用"零信息量 (那些文件里就算有绝对路径也不是我们的
+# 接线) ②每次验收要多读几十 GB 文本 (8 MB 以下的一律会读进内存), 把闸门从秒级拖成小时级。
+SKIP_RAW_PREFIX = ('data/raw/', 'data/_incoming/')
 # 允许的例外: 文档里明确在讲"路径约定"或演示用的占位
 # 允许的例外: 文档里明确在讲"路径约定"或演示用的占位
 ALLOW_LINE = ('portability-allow', '<安装目录>', '<现场包目录>', '<场站名称>', 'D:\\guanlan', '/opt/guanlan',
 ALLOW_LINE = ('portability-allow', '<安装目录>', '<现场包目录>', '<场站名称>', 'D:\\guanlan', '/opt/guanlan',
               'D:/guanlan', '示例', '例:', '例如')
               'D:/guanlan', '示例', '例:', '例如')
@@ -55,6 +60,10 @@ def walk_files():
         parts = set(p.relative_to(ROOT).parts)
         parts = set(p.relative_to(ROOT).parts)
         if parts & SKIP_DIRS or any(h in rel for h in SKIP_PATH_HINT):
         if parts & SKIP_DIRS or any(h in rel for h in SKIP_PATH_HINT):
             continue
             continue
+        if rel.startswith(SKIP_RAW_PREFIX):
+            continue
+        if rel.startswith(SKIP_RAW_PREFIX):
+            continue
         if p.stat().st_size > 8 * 1024 * 1024:      # 超大文本(如 20MB 门户)另算, 见下
         if p.stat().st_size > 8 * 1024 * 1024:      # 超大文本(如 20MB 门户)另算, 见下
             continue
             continue
         yield p
         yield p
@@ -248,7 +257,17 @@ def run_and_probe(dst: pathlib.Path, py: pathlib.Path, base_port=28084) -> int:
     print(f'   serve 退出码 {rs.returncode} (1 = 有模块未就绪, 属正常), 耗时 {time.time() - t0:.0f}s')
     print(f'   serve 退出码 {rs.returncode} (1 = 有模块未就绪, 属正常), 耗时 {time.time() - t0:.0f}s')
     ok = 0
     ok = 0
     try:
     try:
-        urls = [('/', 20225828), ('/detail/', 229643), ('/cms/', None), ('/ops', None), ('/healthz', None)]
+        # ★门户期望字节数**不写死** (2026-09-16): 原来常量 20225828 是某一版门户"服务出来"的字节数,
+        #   门户外壳一改 (例如按用户令在菜单里加「数据重算」) 它就过期, 于是核验会打印一条永远不成立的
+        #   "★与主包不同", 看起来像移植出了问题。改为**从正在跑的主实例现量**: 主实例没起就不比这一项。
+        exp_portal = None
+        try:
+            with urllib.request.urlopen(f'http://127.0.0.1:{base_port}/', timeout=10) as rr:
+                exp_portal = len(rr.read())
+                print(f'   参照: 主实例门户 {exp_portal:,} B (端口 {base_port})')
+        except Exception as e:
+            print(f'   参照: 主实例 ({base_port}) 未响应 → 门户字节数不做同字节比对 ({type(e).__name__})')
+        urls = [('/', exp_portal), ('/detail/', None), ('/cms/', None), ('/ops', None), ('/healthz', None)]
         for path, exp_size in urls:
         for path, exp_size in urls:
             try:
             try:
                 with urllib.request.urlopen(f'http://127.0.0.1:{gw}{path}', timeout=30) as rr:
                 with urllib.request.urlopen(f'http://127.0.0.1:{gw}{path}', timeout=30) as rr:

+ 8 - 3
scripts/guanlan_facts_contract.py

@@ -9,10 +9,15 @@
 import datetime as dt, hashlib, json, os, re, shutil, subprocess, sys, tempfile
 import datetime as dt, hashlib, json, os, re, shutil, subprocess, sys, tempfile
 from pathlib import Path
 from pathlib import Path
 
 
-ROOT = Path(__file__).resolve().parents[1]; OUT = ROOT / "outputs/rudong/guanlan"; DERIVED = OUT / "derived"
+ROOT = Path(__file__).resolve().parents[1]
+sys.path.insert(0, str(ROOT))
+from src import paths as P          # 路径唯一真源 (2026-09-16 统一: 原先写死 outputs/rudong/…)
+OUT = P.guanlan(); DERIVED = OUT / "derived"
 CONTRACT = OUT / "facts_contract_v0.json"
 CONTRACT = OUT / "facts_contract_v0.json"
-SRC = {"findings": ROOT / "outputs/rudong/sop/findings.json", "e5": ROOT / "outputs/rudong/paradigm_r1/experiments/E5_coverage_table/底稿.json",
-       "e3": ROOT / "outputs/rudong/paradigm_r1/experiments/E3_candidate_closure/底稿.json", "s0": ROOT / "outputs/rudong/paradigm_r1/experiments/E8_rudong_rebuild/底稿_s0.json"}
+SRC = {"findings": P.sop() / "findings.json",
+       "e5": P.paradigm() / "experiments/E5_coverage_table/底稿.json",
+       "e3": P.paradigm() / "experiments/E3_candidate_closure/底稿.json",
+       "s0": P.paradigm() / "experiments/E8_rudong_rebuild/底稿_s0.json"}
 sha = lambda b: hashlib.sha256(b).hexdigest(); sha16f = lambda p: sha(Path(p).read_bytes())[:16]; J = lambda p: json.loads(Path(p).read_text(encoding="utf-8"))
 sha = lambda b: hashlib.sha256(b).hexdigest(); sha16f = lambda p: sha(Path(p).read_bytes())[:16]; J = lambda p: json.loads(Path(p).read_text(encoding="utf-8"))
 SIX = ("定论", "准定论·预警", "候选", "参考", "INSUFFICIENT", "撤回")
 SIX = ("定论", "准定论·预警", "候选", "参考", "INSUFFICIENT", "撤回")
 BANNED = re.compile(r"中广核|华电|国电|大唐|华能|三峡|龙源|/Users/|[A-Za-z0-9._%+-]+@[A-Za-z0-9.-]+\.[a-z]{2,}|1[3-9]\d{9}")   # 业主/集团名 · 本机路径 · 邮箱 · 手机 (对外面孔闸)
 BANNED = re.compile(r"中广核|华电|国电|大唐|华能|三峡|龙源|/Users/|[A-Za-z0-9._%+-]+@[A-Za-z0-9.-]+\.[a-z]{2,}|1[3-9]\d{9}")   # 业主/集团名 · 本机路径 · 邮箱 · 手机 (对外面孔闸)

+ 14 - 4
scripts/guanlan_gateway.py

@@ -12,6 +12,7 @@
 import sys as _sys, pathlib as _plb
 import sys as _sys, pathlib as _plb
 _sys.path.insert(0, str(_plb.Path(__file__).resolve().parents[1]))
 _sys.path.insert(0, str(_plb.Path(__file__).resolve().parents[1]))
 from src import paths as P
 from src import paths as P
+from src import proc as _proc           # 无窗口子进程 (2026-09-16 用户令: 不弹命令窗口)
 import argparse, datetime as dt, hashlib, http.client, json, os, pathlib, re, socket, subprocess, sys, time, urllib.parse
 import argparse, datetime as dt, hashlib, http.client, json, os, pathlib, re, socket, subprocess, sys, time, urllib.parse
 from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer
 from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer
 from concurrent.futures import ThreadPoolExecutor
 from concurrent.futures import ThreadPoolExecutor
@@ -168,7 +169,9 @@ def healthz():
 def version():
 def version():
     man = json.loads(MANIFEST.read_text(encoding="utf-8")) if MANIFEST.exists() else {}
     man = json.loads(MANIFEST.read_text(encoding="utf-8")) if MANIFEST.exists() else {}
     b = man.get("detail_style_baseline", {}); c = man.get("content_contract", {})
     b = man.get("detail_style_baseline", {}); c = man.get("content_contract", {})
-    try: head = subprocess.run(["git", "-C", str(ROOT), "rev-parse", "--short", "HEAD"], capture_output=True, text=True, timeout=5).stdout.strip()
+    # src.proc.run_text: 网关无控制台, 裸 spawn git 会给它新建一个可见控制台窗口 ——
+    # 而 /api/version 会被页面反复请求 ⇒ 反复闪窗 (2026-09-16 实测)。
+    try: head = _proc.run_text(["git", "-C", str(ROOT), "rev-parse", "--short", "HEAD"], timeout=5).stdout.strip()
     except Exception: head = None
     except Exception: head = None
     con = json.loads(CONTRACT.read_text(encoding="utf-8")) if CONTRACT.exists() else {}
     con = json.loads(CONTRACT.read_text(encoding="utf-8")) if CONTRACT.exists() else {}
     try: psha = _portal_sha()                       # 缓存版: 同一份门户不重复读 20 MB
     try: psha = _portal_sha()                       # 缓存版: 同一份门户不重复读 20 MB
@@ -198,8 +201,15 @@ class H(BaseHTTPRequestHandler):
         b = json.dumps(obj, ensure_ascii=False, indent=1).encode("utf-8"); self.send_response(code); self.send_header("Content-Type", "application/json; charset=utf-8"); self.send_header("Content-Length", str(len(b))); self.send_header("Cache-Control", "no-store"); self.end_headers()
         b = json.dumps(obj, ensure_ascii=False, indent=1).encode("utf-8"); self.send_response(code); self.send_header("Content-Type", "application/json; charset=utf-8"); self.send_header("Content-Length", str(len(b))); self.send_header("Cache-Control", "no-store"); self.end_headers()
         if self.command != "HEAD": self.wfile.write(b)
         if self.command != "HEAD": self.wfile.write(b)
 
 
-    def _bytes(self, b, ctype, code=200):
-        self.send_response(code); self.send_header("Content-Type", ctype); self.send_header("Content-Length", str(len(b))); self.end_headers()
+    def _bytes(self, b, ctype, code=200, no_store=False):
+        self.send_response(code); self.send_header("Content-Type", ctype); self.send_header("Content-Length", str(len(b)))
+        if no_store:
+            # 运维控制台/数据重算页**必须**不缓存 (2026-09-16 实逮): 页面 HTML/脚本是随包更新的
+            # (例如"内嵌版按钮点不动"这次修复), 没有 no-store 时浏览器会用旧页面 —— 用户看到的是
+            # "改了还是没反应", 排查方向会被带偏。数据面同样: 状态必须每次现取。
+            self.send_header("Cache-Control", "no-store, must-revalidate")
+            self.send_header("Pragma", "no-cache")
+        self.end_headers()
         if self.command != "HEAD": self.wfile.write(b)
         if self.command != "HEAD": self.wfile.write(b)
 
 
     def _route(self, path):
     def _route(self, path):
@@ -226,7 +236,7 @@ class H(BaseHTTPRequestHandler):
             ops = _ops_module()
             ops = _ops_module()
             body = self._read_body() if self.command == "POST" else b""
             body = self._read_body() if self.command == "POST" else b""
             code, mime, data = ops.handle(self.command, path, body)
             code, mime, data = ops.handle(self.command, path, body)
-            return self._bytes(data, mime, code)
+            return self._bytes(data, mime, code, no_store=True)
         if path == "/local-ai/status":
         if path == "/local-ai/status":
             run, models = ollama_models(); return self._json(dict(running=run, models=models, endpoint="/local-ai/", note=None if run else "本机模型未启动"), 200 if run else 503)
             run, models = ollama_models(); return self._json(dict(running=run, models=models, endpoint="/local-ai/", note=None if run else "本机模型未启动"), 200 if run else 503)
         if path.startswith("/release/"):   # E9/E10 裁定: Release 层不合并主库, 由网关只读暴露: /release/ 列表; /release/r1/… (E9 契约层) /release/r2/… (E10 分层覆盖)
         if path.startswith("/release/"):   # E9/E10 裁定: Release 层不合并主库, 由网关只读暴露: /release/ 列表; /release/r1/… (E9 契约层) /release/r2/… (E10 分层覆盖)

+ 86 - 63
scripts/guanlan_ops.py

@@ -8,14 +8,14 @@ r"""观澜运维控制台的后端 (2026-09-12) —— 把"停服务 / 起服务
 本模块给网关加三个东西:
 本模块给网关加三个东西:
   · `GET  /ops`                  一个自适应控制台页面(单文件, 无外部依赖);
   · `GET  /ops`                  一个自适应控制台页面(单文件, 无外部依赖);
   · `GET  /ops/api/state`        真实状态: 各服务端口通不通 · 产物在不在 · 有没有任务在跑 · 上次结果;
   · `GET  /ops/api/state`        真实状态: 各服务端口通不通 · 产物在不在 · 有没有任务在跑 · 上次结果;
-  · `POST /ops/api/<动作>`       停服务 · 起服务 · 重算 · 清产物 · 恢复产物。
+  · `POST /ops/api/<动作>`       停服务 · 起服务 · 重算 · 清产物。
 
 
 ## 三条设计纪律
 ## 三条设计纪律
 
 
 1. **页面活着的服务不能被自己停掉。** 控制台由网关(28084)提供, 所以"停服务"默认**保留网关**,
 1. **页面活着的服务不能被自己停掉。** 控制台由网关(28084)提供, 所以"停服务"默认**保留网关**,
    否则按钮刚点完页面就没了、再也没法"启动服务"。要连网关一起停, 用 `include_gateway=True`,
    否则按钮刚点完页面就没了、再也没法"启动服务"。要连网关一起停, 用 `include_gateway=True`,
    这时用**分离进程**先停再起(页面会断开十几秒, 之后自动重连)。
    这时用**分离进程**先停再起(页面会断开十几秒, 之后自动重连)。
-2. **一次只允许一个动作。** 四个动作互相冲突(重算要重启服务、清产物要挪走产物), 所以有一把锁:
+2. **一次只允许一个动作。** 各动作互相冲突(重算要重启服务、清产物要删产物), 所以有一把锁:
    `run/ops_job.json` 里记着当前任务; 任务在跑时, 所有会冲突的按钮在后端**与前端都会被禁掉**
    `run/ops_job.json` 里记着当前任务; 任务在跑时, 所有会冲突的按钮在后端**与前端都会被禁掉**
    (后端拒绝 = 真生效, 前端禁用只是提示)。判断以后端为准。
    (后端拒绝 = 真生效, 前端禁用只是提示)。判断以后端为准。
 3. **日志与状态落盘。** 动作全部 `Popen` 到 `logs/ops_<动作>_<时间>.log`, 页面轮询 job 状态并把日志尾巴显示出来 ——
 3. **日志与状态落盘。** 动作全部 `Popen` 到 `logs/ops_<动作>_<时间>.log`, 页面轮询 job 状态并把日志尾巴显示出来 ——
@@ -26,8 +26,7 @@ r"""观澜运维控制台的后端 (2026-09-12) —— 把"停服务 / 起服务
     停服务    : 至少有一个组件服务在监听
     停服务    : 至少有一个组件服务在监听
     启动服务  : 至少有一个组件服务没在监听
     启动服务  : 至少有一个组件服务没在监听
     执行重算  : 没有任务在跑
     执行重算  : 没有任务在跑
-    清除产物  : 没有任务在跑 且 产物在位 且 暂存区没有同名目录(或允许 --archive-old)
-    恢复产物  : 没有任务在跑 且 产物被清掉过(暂存区清单存在)
+    清除产物  : 没有任务在跑 且 产物在位 (**直接删除, 不留备份** —— 2026-09-16 用户令; 无"恢复产物"按钮)
 """
 """
 from __future__ import annotations
 from __future__ import annotations
 
 
@@ -43,6 +42,7 @@ import time
 ROOT = pathlib.Path(__file__).resolve().parents[1]
 ROOT = pathlib.Path(__file__).resolve().parents[1]
 sys.path.insert(0, str(ROOT))
 sys.path.insert(0, str(ROOT))
 from src import paths as P                                        # noqa: E402
 from src import paths as P                                        # noqa: E402
+from src import proc as _proc                                     # noqa: E402 无窗口子进程
 
 
 RUN = ROOT / 'run'
 RUN = ROOT / 'run'
 LOGS = ROOT / 'logs'
 LOGS = ROOT / 'logs'
@@ -112,8 +112,9 @@ def job_running() -> tuple[bool, dict]:
             if os.name == 'nt':
             if os.name == 'nt':
                 # errors='replace': tasklist 输出是控制台代码页(GBK), 而本进程可能是 PYTHONUTF8=1
                 # errors='replace': tasklist 输出是控制台代码页(GBK), 而本进程可能是 PYTHONUTF8=1
                 # → 严格解码会失败并使 stdout 为 None (guanlan.py 的 alive() 踩过同一个坑)
                 # → 严格解码会失败并使 stdout 为 None (guanlan.py 的 alive() 踩过同一个坑)
-                out = subprocess.run(['tasklist', '/FI', f'PID eq {pid}'], capture_output=True,
-                                     text=True, errors='replace').stdout
+                # ★ 走 src.proc.run_text: 本模块跑在网关进程里 (无控制台), 裸 spawn tasklist 会让
+                #   Windows 新建可见控制台 —— 而 /ops 页面每 2 s 轮询一次 ⇒ **反复闪窗** (2026-09-16 实测)。
+                out = _proc.run_text(['tasklist', '/FI', f'PID eq {pid}']).stdout
                 alive = bool(out) and str(pid) in out
                 alive = bool(out) and str(pid) in out
             else:
             else:
                 os.kill(pid, 0); alive = True
                 os.kill(pid, 0); alive = True
@@ -148,8 +149,9 @@ def spawn(args: list, tag: str, detached=False) -> dict:
     RUN.mkdir(exist_ok=True); LOGS.mkdir(exist_ok=True)
     RUN.mkdir(exist_ok=True); LOGS.mkdir(exist_ok=True)
     log = LOGS / f'ops_{tag}_{time.strftime("%Y%m%d_%H%M%S")}.log'
     log = LOGS / f'ops_{tag}_{time.strftime("%Y%m%d_%H%M%S")}.log'
     env = dict(os.environ, PYTHONUTF8='1', PYTHONIOENCODING='utf-8', PYTHONUNBUFFERED='1')
     env = dict(os.environ, PYTHONUTF8='1', PYTHONIOENCODING='utf-8', PYTHONUNBUFFERED='1')
-    r = subprocess.run([PY, 'scripts/_ops_launch.py', '--tag', tag, '--log', str(log), '--'] + list(args),
-                       cwd=str(ROOT), env=env, capture_output=True, text=True, errors='replace', timeout=60)
+    # 走 src.proc.run_text: 网关进程无控制台, 裸 spawn 会让这次"拉启动器"也闪一下窗
+    r = _proc.run_text([PY, 'scripts/_ops_launch.py', '--tag', tag, '--log', str(log), '--'] + list(args),
+                       cwd=str(ROOT), env=env, timeout=60)
     pid = None
     pid = None
     for line in reversed((r.stdout or '').strip().splitlines()):
     for line in reversed((r.stdout or '').strip().splitlines()):
         if line.strip().isdigit():
         if line.strip().isdigit():
@@ -173,18 +175,15 @@ def products_state() -> dict:
             stores[d.name] = sum(1 for _ in d.rglob('*') if _.is_file())
             stores[d.name] = sum(1 for _ in d.rglob('*') if _.is_file())
     prov = fam / '_provenance.json'
     prov = fam / '_provenance.json'
     prov_d = json.loads(prov.read_text(encoding='utf-8')) if prov.is_file() else None
     prov_d = json.loads(prov.read_text(encoding='utf-8')) if prov.is_file() else None
-    stash_dirs = {}
-    if OFF.is_dir():
-        for d in sorted(x for x in OFF.iterdir() if x.is_dir() and x.name != '_baseline_kept'):
-            stash_dirs[d.name] = sum(1 for _ in (d / store.name).rglob('*') if _.is_file()) if (d / store.name).is_dir() else 0
-    kept = list((OFF / '_baseline_kept').rglob('*')) if (OFF / '_baseline_kept').is_dir() else []
+    # ★2026-09-16 用户令"不要备份清除的产物": 清除 = 真删除, 没有暂存区可列。
+    #   这里只报"有没有旧设计留下的 _products_off*"(有就提示手工删掉), 不再提供"恢复产物"。
+    legacy = sorted(x.name for x in ROOT.glob('_products_off*') if x.is_dir())
     return dict(
     return dict(
         fam=str(fam.relative_to(ROOT)).replace('\\', '/'), store=str(store.relative_to(ROOT)).replace('\\', '/'),
         fam=str(fam.relative_to(ROOT)).replace('\\', '/'), store=str(store.relative_to(ROOT)).replace('\\', '/'),
         files=n, in_place=n > 0, stores=stores,
         files=n, in_place=n > 0, stores=stores,
         cleared=n == 0,
         cleared=n == 0,
-        stash_has_same=any(v > 0 for v in stash_dirs.values()), stash=stash_dirs,
-        can_restore=OFF_MANIFEST.is_file() and n == 0,
-        baseline_kept=sum(1 for x in kept if x.is_file()),
+        no_backup=True,                                  # 现口径: 清除不留备份
+        legacy_backups=legacy,                           # 旧设计残留 (按用户令不自动删, 只提示)
         provenance=(dict(counts=prov_d.get('counts'), at=prov_d.get('at')) if prov_d else None),
         provenance=(dict(counts=prov_d.get('counts'), at=prov_d.get('at')) if prov_d else None),
     )
     )
 
 
@@ -214,7 +213,6 @@ def state() -> dict:
             start_services=any_comp_down and not running,
             start_services=any_comp_down and not running,
             rebuild=not running,
             rebuild=not running,
             products_off=pr['in_place'] and not running,
             products_off=pr['in_place'] and not running,
-            products_on=(not pr['in_place']) and not running,
         ),
         ),
         anchors=dict(alarms=_rows(f'{pr["fam"]}/windscada/alarms.parquet'),
         anchors=dict(alarms=_rows(f'{pr["fam"]}/windscada/alarms.parquet'),
                      workorders=_rows(f'{pr["fam"]}/windscada/workorders.parquet'),
                      workorders=_rows(f'{pr["fam"]}/windscada/workorders.parquet'),
@@ -309,38 +307,47 @@ def act_rebuild(body: dict) -> tuple[int, dict]:
 
 
 
 
 def act_products_off(body: dict) -> tuple[int, dict]:
 def act_products_off(body: dict) -> tuple[int, dict]:
+    """清除产物 —— 2026-09-16 用户令: **直接删除, 不留备份**。
+
+    原实现是 `products_state.py --off`(移动到 `_products_off/`, 可 `--on` 还原)。现口径: 真删,
+    故这里传 `--yes`(脚本对不可恢复操作要求显式确认; 前端已 confirm 过一次)。
+    恢复缺失的**随包件**改用交付包补齐:
+    `python scripts/products_restore_missing.py --stash <交付包.zip>`。
+    """
     if (e := _guard('products_off')):
     if (e := _guard('products_off')):
         return 409, dict(err=e)
         return 409, dict(err=e)
     pr = products_state()
     pr = products_state()
     if not pr['in_place']:
     if not pr['in_place']:
         return 409, dict(err='产物已经是清空状态, 无需再清')
         return 409, dict(err='产物已经是清空状态, 无需再清')
-    args = ['scripts/products_state.py', '--off']
-    if pr['stash_has_same'] or body.get('archive_old'):
-        args.append('--archive-old')      # 暂存区已有同名产物时必须加, 否则会嵌套(见该脚本注释)
-    j = spawn(args, 'products_off')
-    return 200, dict(ok=True, job=j, note='正在把产物挪到 _products_off/ (挪完请点"启动服务"让页面呈现空状态)')
-
-
-def act_products_on(body: dict) -> tuple[int, dict]:
-    if (e := _guard('products_on')):
-        return 409, dict(err=e)
-    pr = products_state()
-    if pr['in_place']:
-        return 409, dict(err='产物已在位, 无需恢复')
-    if not OFF_MANIFEST.is_file():
-        return 409, dict(err='没有 _products_off/manifest.json (没清过或清单已删), 无法恢复')
-    j = spawn(['scripts/products_state.py', '--on'], 'products_on')
-    return 200, dict(ok=True, job=j, note='正在恢复产物; 恢复后请点"启动服务"')
+    j = spawn(['scripts/products_state.py', '--off', '--yes'], 'products_off')
+    return 200, dict(ok=True, job=j, note='正在**删除**产物(不留备份, 不可恢复); 清完请点"启动服务"让页面呈现空状态')
 
 
 
 
 ACTIONS = dict(stop_services=act_stop, start_services=act_start, rebuild=act_rebuild,
 ACTIONS = dict(stop_services=act_stop, start_services=act_start, rebuild=act_rebuild,
-               products_off=act_products_off, products_on=act_products_on)
+               products_off=act_products_off)
 
 
 
 
 def handle(method: str, path: str, body: bytes) -> tuple[int, str, bytes]:
 def handle(method: str, path: str, body: bytes) -> tuple[int, str, bytes]:
-    """网关调用入口: → (http code, content-type, body bytes)。"""
+    """网关调用入口: → (http code, content-type, body bytes)。
+
+    路由:
+      `/ops`           运维控制台整页 (服务 + 重算 + 产物 + 动作进度)
+      `/ops/recalc`    **内嵌版**: 只保留 重算 + 产物 (+ 动作进度), 给门户菜单「数据重算」用
+                       (2026-09-16 用户令: 把 /ops 的重算与产物搬到门户菜单里)。
+                       走独立路由而不是 `?embed=…`: 网关转给本模块的是 `u.path` (查询串被丢掉),
+                       用查询串会变成"看起来支持、实际不生效"的静默坑。
+      `/ops/api/...`   运维动作 JSON API (GET state / POST 各动作)
+    """
     if path in ('/ops', '/ops/'):
     if path in ('/ops', '/ops/'):
         return 200, 'text/html; charset=utf-8', PAGE.encode('utf-8')
         return 200, 'text/html; charset=utf-8', PAGE.encode('utf-8')
+    if path in ('/ops/recalc', '/ops/recalc/'):
+        import re as _re
+        html = _re.sub(r'<!--HIDE_IN_EMBED-->.*?<!--/HIDE_IN_EMBED-->', '', PAGE, flags=_re.S)
+        html = (html.replace('<title>观澜 · 运维控制台</title>', '<title>观澜 · 数据重算</title>')
+                    .replace('</style>', '.wrap{max-width:100%;padding:8px 10px 24px}'
+                                        'body{background:transparent}'
+                                        '.card{margin:0 0 12px}</style>', 1))
+        return 200, 'text/html; charset=utf-8', html.encode('utf-8')
     if path == '/ops/api/state':
     if path == '/ops/api/state':
         return 200, 'application/json; charset=utf-8', json.dumps(state(), ensure_ascii=False).encode('utf-8')
         return 200, 'application/json; charset=utf-8', json.dumps(state(), ensure_ascii=False).encode('utf-8')
     if path.startswith('/ops/api/'):
     if path.startswith('/ops/api/'):
@@ -385,10 +392,13 @@ th{color:var(--mut);font-weight:600}
 .busy{background:#FFF7E6;border:1px solid #F0D9A8;color:#7A5600;padding:8px 12px;border-radius:8px;margin:0 0 12px}
 .busy{background:#FFF7E6;border:1px solid #F0D9A8;color:#7A5600;padding:8px 12px;border-radius:8px;margin:0 0 12px}
 label{font-size:13.5px;color:var(--mut);display:flex;gap:6px;align-items:center}
 label{font-size:13.5px;color:var(--mut);display:flex;gap:6px;align-items:center}
 </style></head><body><div class="wrap">
 </style></head><body><div class="wrap">
+<!--HIDE_IN_EMBED-->
 <h1>观澜 · 运维控制台</h1>
 <h1>观澜 · 运维控制台</h1>
 <p class="sub">停/启服务 · 执行重算 · 清除产物 —— 按钮按真实状态启用; 不可用的动作后端也会拒绝。本页由网关(端口 <span id="gw"></span>)提供。</p>
 <p class="sub">停/启服务 · 执行重算 · 清除产物 —— 按钮按真实状态启用; 不可用的动作后端也会拒绝。本页由网关(端口 <span id="gw"></span>)提供。</p>
+<!--/HIDE_IN_EMBED-->
 <div id="busy"></div>
 <div id="busy"></div>
 
 
+<!--HIDE_IN_EMBED-->
 <div class="card"><h2>服务</h2><div id="svc"></div>
 <div class="card"><h2>服务</h2><div id="svc"></div>
   <div class="row" style="margin-top:12px">
   <div class="row" style="margin-top:12px">
     <button id="b_start" class="ghost">启动服务(并打开门户)</button>
     <button id="b_start" class="ghost">启动服务(并打开门户)</button>
@@ -397,6 +407,7 @@ label{font-size:13.5px;color:var(--mut);display:flex;gap:6px;align-items:center}
   </div>
   </div>
   <p class="hint" id="svc_hint"></p>
   <p class="hint" id="svc_hint"></p>
 </div>
 </div>
+<!--/HIDE_IN_EMBED-->
 
 
 <div class="card"><h2>重算</h2>
 <div class="card"><h2>重算</h2>
   <div class="row">
   <div class="row">
@@ -412,14 +423,14 @@ label{font-size:13.5px;color:var(--mut);display:flex;gap:6px;align-items:center}
 
 
 <div class="card"><h2>产物</h2><div id="prod"></div>
 <div class="card"><h2>产物</h2><div id="prod"></div>
   <div class="row" style="margin-top:12px">
   <div class="row" style="margin-top:12px">
-    <button id="b_on" class="ghost">恢复产物</button>
-    <button id="b_off" class="danger">清除产物(挪到 _products_off,可恢复)</button>
+    <button id="b_off" class="danger">清除产物(直接删除,不留备份)</button>
   </div>
   </div>
-  <p class="hint">清除后门户首页仍可打开(它是静态交付页), 但工作台会显示"无产物"; 恢复后请点"启动服务"。</p>
+  <p class="hint">按用户令(2026-09-16)**清除不留备份、不可恢复**;清完请点"启动服务"让页面呈现空状态。
+     要补回"包内没有生成端"的随包件:<code>python scripts/products_restore_missing.py --stash &lt;交付包.zip&gt;</code>(从交付包按需补齐)。</p>
 </div>
 </div>
 
 
 <div class="card"><h2>最近一次动作</h2><div id="job">(无)</div><pre id="log"></pre></div>
 <div class="card"><h2>最近一次动作</h2><div id="job">(无)</div><pre id="log"></pre></div>
-<p class="hint">命令行等价物: <code>guanlan.py stop/serve</code> · <code>scripts/rebuild_all.py</code> · <code>scripts/products_state.py --off/--on</code>(手册 §0/§5b)。</p>
+<p class="hint">命令行等价物: <code>guanlan.py stop/serve</code> · <code>scripts/rebuild_all.py</code> · <code>scripts/products_state.py --off --yes/--status</code>(手册 §0/§5b)。清除产物**不留备份**(用户令 2026-09-16);要补回缺失的随包件用 <code>scripts/products_restore_missing.py --stash &lt;交付包.zip&gt;</code>。</p>
 </div><script>
 </div><script>
 const $ = s => document.querySelector(s);
 const $ = s => document.querySelector(s);
 let S = null, busyTimer = null;
 let S = null, busyTimer = null;
@@ -428,35 +439,37 @@ async function api(path, body){
   return [r.status, await r.json().catch(()=>({err:'非 JSON 响应'}))];
   return [r.status, await r.json().catch(()=>({err:'非 JSON 响应'}))];
 }
 }
 function pill(up){ return `<span class="pill ${up?'up':'down'}">${up?'运行中':'未运行'}</span>`; }
 function pill(up){ return `<span class="pill ${up?'up':'down'}">${up?'运行中':'未运行'}</span>`; }
+/* set: 元素可能不存在 —— 内嵌版 (/ops/recalc) 会去掉"服务"卡与页头, 直接 $() 取值再赋值会抛错,
+   而这里一抛整个 render() 就断, 页面看着像"卡住不动"(2026-09-16 加内嵌版时踩过)。 */
+function set(sel, fn){ const e = $(sel); if(e) fn(e); }
 function render(){
 function render(){
   const s = S; if(!s) return;
   const s = S; if(!s) return;
-  $('#gw').textContent = s.services.gateway ? s.services.gateway.port : '?';
+  set('#gw', e=>e.textContent = s.services.gateway ? s.services.gateway.port : '?');
   const rows = Object.entries(s.services).map(([n,v]) =>
   const rows = Object.entries(s.services).map(([n,v]) =>
      `<tr><td>${n}${v.is_gateway?' <span class="mut">(控制台本体)</span>':''}</td><td>${v.port}</td><td>${pill(v.up)}</td></tr>`).join('');
      `<tr><td>${n}${v.is_gateway?' <span class="mut">(控制台本体)</span>':''}</td><td>${v.port}</td><td>${pill(v.up)}</td></tr>`).join('');
-  $('#svc').innerHTML = `<table><tr><th>服务</th><th>端口</th><th>状态</th></tr>${rows}</table>`;
+  set('#svc', e=>e.innerHTML = `<table><tr><th>服务</th><th>端口</th><th>状态</th></tr>${rows}</table>`);
   const p = s.products;
   const p = s.products;
   const st = Object.entries(p.stores||{}).map(([k,v])=>`${k} ${v}`).join(' · ') || '(空)';
   const st = Object.entries(p.stores||{}).map(([k,v])=>`${k} ${v}`).join(' · ') || '(空)';
-  $('#prod').innerHTML = `<table>
+  set('#prod', e=>e.innerHTML = `<table>
     <tr><th>产物仓</th><td>${p.store}</td></tr>
     <tr><th>产物仓</th><td>${p.store}</td></tr>
     <tr><th>件数</th><td>${p.files} 件 ${p.in_place?'<span class="pill up">在位</span>':'<span class="pill down">已清空</span>'}</td></tr>
     <tr><th>件数</th><td>${p.files} 件 ${p.in_place?'<span class="pill up">在位</span>':'<span class="pill down">已清空</span>'}</td></tr>
     <tr><th>分布</th><td class="mut">${st}</td></tr>
     <tr><th>分布</th><td class="mut">${st}</td></tr>
     <tr><th>来源台账</th><td class="mut">${p.provenance?`raw 重算 ${p.provenance.counts['raw-derived']} 件 · 随包补齐 ${p.provenance.counts.shipped} 件 (${p.provenance.at})`:'(无 _provenance.json)'}</td></tr>
     <tr><th>来源台账</th><td class="mut">${p.provenance?`raw 重算 ${p.provenance.counts['raw-derived']} 件 · 随包补齐 ${p.provenance.counts.shipped} 件 (${p.provenance.at})`:'(无 _provenance.json)'}</td></tr>
-    <tr><th>暂存区</th><td class="mut">${Object.keys(p.stash||{}).length?Object.entries(p.stash).map(([k,v])=>`${k}:${v} 件`).join(' · '):'(空)'}${p.stash_has_same?' <b>← 有同名产物, 清除时会自动存档上一代</b>':''}</td></tr>
+    <tr><th>备份</th><td class="mut">${p.no_backup?'**不留备份**(用户令 2026-09-16: 清除 = 直接删除)':''}${(p.legacy_backups&&p.legacy_backups.length)?` <b>← 旧设计残留 ${p.legacy_backups.join(' · ')} —— 确认不需要后手工删掉</b>`:''}</td></tr>
     <tr><th>验收锚点</th><td class="mut">报警 ${s.anchors.alarms??'—'} 行 · 工单 ${s.anchors.workorders??'—'} · temp_monthly ${s.anchors.temp_monthly??'—'} · 本体 ${s.anchors.objects??'—'} 对象</td></tr>
     <tr><th>验收锚点</th><td class="mut">报警 ${s.anchors.alarms??'—'} 行 · 工单 ${s.anchors.workorders??'—'} · temp_monthly ${s.anchors.temp_monthly??'—'} · 本体 ${s.anchors.objects??'—'} 对象</td></tr>
-  </table>`;
+  </table>`);
   const b = s.buttons, run = s.job && s.job.status==='running';
   const b = s.buttons, run = s.job && s.job.status==='running';
-  $('#b_start').disabled   = !b.start_services;
-  $('#b_stop').disabled    = !b.stop_services;
-  $('#b_restart').disabled = !b.stop_services;
-  $('#b_rebuild').disabled = !b.rebuild;
-  $('#b_off').disabled     = !b.products_off;
-  $('#b_on').disabled      = !b.products_on;
-  $('#svc_hint').textContent = b.start_services ? '有服务未运行 → 可启动。' : '全部组件服务已在运行 → 启动按钮已禁用。';
-  $('#rebuild_hint').textContent = b.rebuild ? '' : '(有任务在跑, 重算按钮已禁用)';
-  $('#busy').innerHTML = run ? `<div class="busy">正在执行: <b>${s.job.kind}</b>(起于 ${s.job.started})${s.job.note?' — '+s.job.note:''} · 完成后本页自动刷新</div>` : '';
+  set('#b_start',   e=>e.disabled = !b.start_services);
+  set('#b_stop',    e=>e.disabled = !b.stop_services);
+  set('#b_restart', e=>e.disabled = !b.stop_services);
+  set('#b_rebuild', e=>e.disabled = !b.rebuild);
+  set('#b_off',     e=>e.disabled = !b.products_off);
+  set('#svc_hint', e=>e.textContent = b.start_services ? '有服务未运行 → 可启动。' : '全部组件服务已在运行 → 启动按钮已禁用。');
+  set('#rebuild_hint', e=>e.textContent = b.rebuild ? '' : '(有任务在跑, 重算按钮已禁用)');
+  set('#busy', e=>e.innerHTML = run ? `<div class="busy">正在执行: <b>${s.job.kind}</b>(起于 ${s.job.started})${s.job.note?' — '+s.job.note:''} · 完成后本页自动刷新</div>` : '');
   const j = s.job;
   const j = s.job;
-  $('#job').innerHTML = j ? `<div>${j.kind} · <b>${j.status==='running'?'执行中':(j.rc===0?'完成 (退出码 0)':'结束 (退出码 '+j.rc+')')}</b> · ${j.started||''}${j.seconds?' · 耗时 '+j.seconds+'s':''}</div><div class="mut">${j.cmd||''}</div>${j.note?'<div class="mut">'+j.note+'</div>':''}` : '(无)';
-  $('#log').textContent = (j && j.log_tail && j.log_tail.length) ? j.log_tail.join('\n') : '';
+  set('#job', e=>e.innerHTML = j ? `<div>${j.kind} · <b>${j.status==='running'?'执行中':(j.rc===0?'完成 (退出码 0)':'结束 (退出码 '+j.rc+')')}</b> · ${j.started||''}${j.seconds?' · 耗时 '+j.seconds+'s':''}</div><div class="mut">${j.cmd||''}</div>${j.note?'<div class="mut">'+j.note+'</div>':''}` : '(无)');
+  set('#log', e=>e.textContent = (j && j.log_tail && j.log_tail.length) ? j.log_tail.join('\n') : '');
 }
 }
 async function refresh(){ const [c,d] = await api('/ops/api/state'); if(c===200){ S=d; render(); } }
 async function refresh(){ const [c,d] = await api('/ops/api/state'); if(c===200){ S=d; render(); } }
 async function fire(name, body){
 async function fire(name, body){
@@ -465,12 +478,22 @@ async function fire(name, body){
   else { if(d.note) console.log(d.note); }
   else { if(d.note) console.log(d.note); }
   await refresh();
   await refresh();
 }
 }
-$('#b_start').onclick   = ()=>fire('start_services', {open_browser:true});
-$('#b_stop').onclick    = ()=>fire('stop_services', {include_gateway:false});
-$('#b_restart').onclick = ()=>{ if(confirm('完整重启会连网关一起停, 本页会断开约 15 秒后自动恢复。继续?')) fire('stop_services', {include_gateway:true}); };
-$('#b_rebuild').onclick = ()=>{ if(confirm('开始重算? 期间服务会被重启, 页面可能短暂打不开。')) fire('rebuild', {skip_scada: !$('#f_scada').checked, with_verify: $('#f_verify').checked}); };
-$('#b_off').onclick     = ()=>{ if(confirm('清除产物(挪到 _products_off,可恢复)?')) fire('products_off', {}); };
-$('#b_on').onclick      = ()=>{ if(confirm('恢复产物?')) fire('products_on', {}); };
-refresh(); setInterval(refresh, 2000);
+/* on: 与 set 同理 —— 内嵌版没有"服务"卡, 直接 $('#b_start').onclick=… 会抛 TypeError,
+   而**这一抛会让后面所有绑定与 refresh() 全不执行** (2026-09-16 用户实测: 门户 #recalc 页里
+   "执行重算 / 清除产物"点了没反应)。容错绑定 + 下面 try/catch 兜底 = 同类问题不再静默。 */
+function on(sel, fn){ const e = $(sel); if(e) e.onclick = fn; }
+try {
+  on('#b_start',   ()=>fire('start_services', {open_browser:true}));
+  on('#b_stop',    ()=>fire('stop_services', {include_gateway:false}));
+  on('#b_restart', ()=>{ if(confirm('完整重启会连网关一起停, 本页会断开约 15 秒后自动恢复。继续?')) fire('stop_services', {include_gateway:true}); });
+  on('#b_rebuild', ()=>{ if(confirm('开始重算? 期间服务会被重启, 页面可能短暂打不开。')) fire('rebuild', {skip_scada: !$('#f_scada').checked, with_verify: $('#f_verify').checked}); });
+  on('#b_off',     ()=>{ if(confirm('清除产物 = **直接删除, 不留备份, 不可恢复**。继续?')) fire('products_off', {}); });
+  refresh(); setInterval(refresh, 2000);
+} catch (err) {
+  /* 界面脚本自己坏了也要看得见 (否则表现就是"按钮点了没反应", 排查全靠猜) */
+  const b = document.getElementById('busy');
+  if (b) b.innerHTML = '<div class="busy">⚠ 本页脚本出错, 按钮可能不可用: ' + err + ' (请反馈该行)</div>';
+  else alert('本页脚本出错: ' + err);
+}
 </script></body></html>
 </script></body></html>
 """
 """

+ 150 - 0
scripts/guanlan_start_hidden.py

@@ -0,0 +1,150 @@
+#!/usr/bin/env python3
+# -*- coding: utf-8 -*-
+"""**无窗口**启动观澜 (2026-09-16 用户令: "启动观澜系统, 弹出的命令窗口, 改为不弹出方式")。
+
+为什么单开一个脚本: `guanlan.py serve` 是**前台阻塞**命令 (它在控制台里打启动横幅、等 healthz),
+双击 start.bat 就会一直挂着一个黑窗; 而各组件服务其实早已是 `DETACHED_PROCESS` 起的 (不弹窗)。
+所以只需要一个"把 serve 藏到后台、等就绪、打开浏览器、自己退出"的入口。
+
+做法:
+  1. 网关已在 → 只打开浏览器 (重复点不重复起, 幂等);
+  2. 否则用 `CREATE_NO_WINDOW | DETACHED_PROCESS` 起 `guanlan.py serve`, stdout/stderr 落
+     `logs/serve.log` (有窗口才有地方看日志, 藏起来就必须落文件);
+  3. 轮询 `/healthz` 直到就绪 (最多 `--wait` 秒), 就绪后按 `--open/--no-open` 决定是否开浏览器;
+  4. 全过程写 `logs/start_hidden.log` (含退出码与失败原因) —— **不静默**:
+     失败时不但写日志, 还会弹一个消息框 (无窗口模式下唯一能让人看见的方式)。
+
+由 `start_hidden.vbs` 以 `pythonw.exe` 调用 (窗口样式 0 = 完全不显示)。
+
+用法:
+    .venv\\Scripts\\pythonw.exe scripts\\guanlan_start_hidden.py          # 隐藏启动 + 开浏览器
+    .venv\\Scripts\\python.exe  scripts\\guanlan_start_hidden.py --no-open  # 隐藏启动, 不开浏览器
+    .venv\\Scripts\\python.exe  scripts\\guanlan_start_hidden.py --wait 180 # 加长等待
+"""
+from __future__ import annotations
+
+import argparse
+import datetime as dt
+import json
+import os
+import pathlib
+import socket
+import subprocess
+import sys
+import time
+import urllib.request
+
+ROOT = pathlib.Path(__file__).resolve().parents[1]
+sys.path.insert(0, str(ROOT))
+from src import paths as P          # noqa: E402
+from src.console import soft        # noqa: E402
+
+soft()
+LOG = P.LOGS / 'start_hidden.log'
+
+
+def log(msg: str) -> None:
+    P.LOGS.mkdir(parents=True, exist_ok=True)
+    line = f'{dt.datetime.now():%Y-%m-%d %H:%M:%S} {msg}'
+    with open(LOG, 'a', encoding='utf-8') as f:
+        f.write(line + '\n')
+
+
+def port_up(host: str, port: int, t: float = 0.4) -> bool:
+    try:
+        with socket.create_connection((host, port), timeout=t):
+            return True
+    except OSError:
+        return False
+
+
+def healthz(url: str, timeout=6.0) -> bool:
+    try:
+        with urllib.request.urlopen(url, timeout=timeout) as r:
+            return r.status == 200
+    except Exception:
+        return False
+
+
+def notify(title: str, text: str) -> None:
+    """无窗口模式下唯一能让人看见失败的通道 (Windows 消息框; 非 Windows 走 stderr)。"""
+    if os.name != 'nt':
+        print(f'{title}: {text}', file=sys.stderr)
+        return
+    try:
+        import ctypes
+        ctypes.windll.user32.MessageBoxW(0, text, title, 0x30)      # MB_ICONWARNING
+    except Exception:
+        pass
+
+
+def main() -> int:
+    ap = argparse.ArgumentParser()
+    ap.add_argument('--host', default='127.0.0.1')
+    ap.add_argument('--gateway', type=int, default=None, help='网关端口 (默认读 configs/serve.json)')
+    ap.add_argument('--wait', type=int, default=120, help='等 healthz 就绪的最长秒数')
+    ap.add_argument('--no-open', action='store_true', help='就绪后不开浏览器')
+    ap.add_argument('--quiet-fail', action='store_true', help='失败时不弹消息框 (只写日志)')
+    a = ap.parse_args()
+
+    gw = a.gateway
+    if gw is None:
+        try:
+            cfg = json.loads(P.SERVE_JSON.read_text(encoding='utf-8-sig'))
+            gw = int(cfg.get('gateway') or 28084)
+        except Exception:
+            gw = 28084
+    url = f'http://{a.host}:{gw}/'
+    log(f'--- 无窗口启动请求 (端口 {gw}, wait={a.wait}s) ---')
+
+    if port_up(a.host, gw):
+        log(f'网关 {gw} 已在运行 → 只打开页面')
+        if not a.no_open:
+            import webbrowser
+            webbrowser.open(url)
+        return 0
+
+    py = P.venv_python() or pathlib.Path(sys.executable)
+    cmd = [str(py), 'guanlan.py', 'serve']
+    flags = 0
+    if os.name == 'nt':
+        flags = (getattr(subprocess, 'CREATE_NO_WINDOW', 0x08000000)
+                 | getattr(subprocess, 'CREATE_NEW_PROCESS_GROUP', 0)
+                 | getattr(subprocess, 'DETACHED_PROCESS', 0))
+    P.LOGS.mkdir(parents=True, exist_ok=True)
+    lf = open(P.LOGS / 'serve.log', 'ab')
+    kw = dict(cwd=str(ROOT), stdout=lf, stderr=subprocess.STDOUT, stdin=subprocess.DEVNULL)
+    if os.name == 'nt':
+        kw['creationflags'] = flags
+    else:
+        kw['start_new_session'] = True
+    try:
+        pid = subprocess.Popen(cmd, **kw).pid
+    except Exception as e:
+        log(f'✘ 起 serve 失败: {type(e).__name__}: {e}')
+        if not a.quiet_fail:
+            notify('观澜启动失败', f'无法启动服务: {e}\n详见 logs\\start_hidden.log')
+        return 3
+    log(f'起 serve pid={pid} (无窗口; 日志 logs/serve.log)')
+
+    t0 = time.time()
+    while time.time() - t0 < a.wait:
+        if healthz(f'{url}healthz'):
+            log(f'✔ 就绪 ({time.time() - t0:.0f}s) → {url}')
+            if not a.no_open:
+                try:
+                    import webbrowser
+                    webbrowser.open(url)
+                except Exception as e:
+                    log(f'⚠ 打开浏览器失败: {e}')
+            return 0
+        time.sleep(1.5)
+    log(f'✘ {a.wait}s 内 healthz 未就绪 (serve pid={pid} 可能已退出)')
+    if not a.quiet_fail:
+        notify('观澜启动未就绪', f'{a.wait} 秒内网关未就绪, 请查看:\n'
+                                  f'logs\\serve.log\nlogs\\start_hidden.log')
+    return 2
+
+
+if __name__ == '__main__':
+    sys.exit(main())

+ 17 - 4
scripts/ingest_ops_2025.py

@@ -21,8 +21,10 @@
   重算成熟度时必须按实际非空率判, 不能因为"字段有了"就宣布该维可评估 —
   重算成熟度时必须按实际非空率判, 不能因为"字段有了"就宣布该维可评估 —
   那会把 19/818 的覆盖说成能力达标。
   那会把 19/818 的覆盖说成能力达标。
 """
 """
+import pathlib as _pl, sys as _sys
+_sys.path.insert(0, str(_pl.Path(__file__).resolve().parents[1]))   # 2026-09-16: 与其它摄入脚本一致
 from src import paths as _P
 from src import paths as _P
-import argparse, json, pathlib, sys, warnings
+import argparse, json, os, pathlib, sys, warnings
 warnings.filterwarnings('ignore')
 warnings.filterwarnings('ignore')
 import pandas as pd
 import pandas as pd
 
 
@@ -44,7 +46,12 @@ def _ledger_range(raw):
         return None, None
         return None, None
 
 
 
 
-OBJ = pathlib.Path('outputs/rudong/ontology/objects.json')
+OBJ = _P.objects_json()
+# ★2026-09-16 修: 原为 `pathlib.Path('outputs/rudong/ontology/objects.json')` —— **cwd 相对** + 写死场名。
+#   本文件在模块级用它, 于是: ① 从别处调用/换工作目录就会写到别处; ② 它还是全仓唯一没有
+#   `sys.path.insert(0, str(ROOT))` 的摄入脚本, 干净环境下连 `src.paths` 都导不进来。
+#   注意: 该脚本此前因为缺 `import os` (上面那行用 os.environ) 在模块级就 NameError, 从未真正跑过 ——
+#   所以"补 import os"必须和"改掉 cwd 相对路径 + 走 Store 落盘"同一次做完, 否则只是把一把静默的凶器上膛。
 OPS = pathlib.Path(os.environ.get('WINDSCADA_OPS_LIB') or (_P.RAW_ROOT / '工作库' / '20_structured' / 'ops'))
 OPS = pathlib.Path(os.environ.get('WINDSCADA_OPS_LIB') or (_P.RAW_ROOT / '工作库' / '20_structured' / 'ops'))
                           # 原为 mac 盘路径 (本机不存在): 改为 env 可配 + 安装目录相对默认位置
                           # 原为 mac 盘路径 (本机不存在): 改为 env 可配 + 安装目录相对默认位置
 
 
@@ -297,8 +304,14 @@ def main():
     if a.dry_run:
     if a.dry_run:
         print('  (dry-run, 未写盘)')
         print('  (dry-run, 未写盘)')
         return 0
         return 0
-    OBJ.write_text(json.dumps(db, ensure_ascii=False, indent=1), encoding='utf-8')
-    print('  已写入')
+    # 走本体 store 落盘 (2026-09-16): 原来直接 `OBJ.write_text(...)` —— 绕过了 store 的
+    # ① 全库 validate 闸 ② F12 原子写 (tmp + os.replace)。对象库里混进一个坏对象时, 直写会
+    # 让**整库**在下次读取时才炸, 而 store.save() 会当场报出是哪个对象不合法。
+    from src.ontology.store import Store
+    st = Store(OBJ)
+    st.objects = db
+    st.save()
+    print(f'  已写入 {_P.rel(OBJ)} (经 Store 闸)')
     return 0
     return 0
 
 
 
 

+ 349 - 0
scripts/inventory_products.py

@@ -0,0 +1,349 @@
+#!/usr/bin/env python3
+# -*- coding: utf-8 -*-
+r"""产物清点与「输入 ↔ 产物」呼应校验 (2026-09-16 用户令 1、2)。
+
+用户令:
+  1. 基于观澜源码, 认真检查、梳理**所有产物**(产物类型/功用/输出路径)确保无遗漏; 输出路径不统一的改源码统一;
+     完善到 `<安装目录>\docs\系统设计说明.md`。
+  2. 确保 `data/raw` 下的输入数据**可用于重算**, 且**产物与输入数据呼应**。
+
+本脚本是三件事的**单一实现** (别在文档里手抄第二份, 会飘):
+  · 清点: 每个产物仓的件数/大小/类型, 以及**逐件来源**(raw-derived / shipped, 读 `_provenance.json`
+    与 `_derived_manifest.json`);
+  · 呼应: 对每类输入算它的**数据跨度/条数**, 与它喂出来的产物**逐项对拍** (产物跨度必须落在输入跨度内,
+    且输入更新时产物不该明显落后);
+  · 写档: `--write-doc` 把清单块写进 `docs\系统设计说明.md` 的
+    `<!-- INVENTORY:BEGIN --> … <!-- INVENTORY:END -->` 之间 (文档其余部分手工维护)。
+
+用法:
+    python scripts/inventory_products.py                # 打清单 + 呼应结论 (人看)
+    python scripts/inventory_products.py --check        # 只做呼应校验 (有问题 → 退出码 5)
+    python scripts/inventory_products.py --write-doc    # 把清单块写进 docs/系统设计说明.md
+    python scripts/inventory_products.py --json         # 机器可读
+"""
+from __future__ import annotations
+
+import argparse
+import datetime as dt
+import json
+import pathlib
+import re
+import sys
+from collections import OrderedDict
+
+ROOT = pathlib.Path(__file__).resolve().parents[1]
+sys.path.insert(0, str(ROOT))
+from src.console import soft  # noqa: E402
+from src import paths as P    # noqa: E402
+
+soft()
+DOC = P.DOCS / '系统设计说明.md'
+MARK_B, MARK_E = '<!-- INVENTORY:BEGIN', '<!-- INVENTORY:END -->'
+
+# ── 产物仓: 路径助手 → (仓名, 功用, 谁生成, 谁消费) ───────────────────────────────
+# ★路径一律用 src/paths.py 的助手 (唯一真源); 这里写的是**相对安装根**的形式, 便于跨平台显示。
+STORES = OrderedDict([
+    ('windscada',  dict(helper='P.store()',            purpose='L0 标准仓: SCADA/台账/派生分析的全部 parquet (页面主取数处)',
+                        producer='rebuild_from_raw.py (三门台账) + --scada (10 个构建器) + windscada_monthly_build.py',
+                        consumer='scripts/windscada_serve.py 各视图 · src/windscada/taxonomy · subsys/fusion')),
+    ('ontology',   dict(helper='P.ont()',              purpose='本体对象库: 码表/手册/工单展开/失效树 + 检索索引 + 实机参数',
+                        producer='python -m src.ontology.kb_ingest → populate → chain_ingest → trend_ingest → retrieval.build',
+                        consumer='脚本 windscada_serve.py 本体页/问答 · scripts/guanlan_facts_contract.py')),
+    ('windcms',    dict(helper='P.cms()',              purpose='CMS 振动诊断产物: 状态评估报告/逐台页/工作台页/知识库 + 厂家报告转录',
+                        producer='scripts/windcms.py report/kb · scripts/vib_reports_build.py',
+                        consumer='自服务 :18020 (每请求现读) · taxonomy.system_matrix (转录设备状态)')),
+    ('m5_cms_tcm', dict(helper='P.m5()',               purpose='振动线出件与窗级分析: handoff 接口 + 窗索引/谱库 + TCM 兼容件',
+                        producer='scripts/vib_raw_build.py (窗索引/谱) · 振动线出件 (handoff, 随包快照)',
+                        consumer='src/windscada/subsys/fusion.py · src/windcms/data.py · scripts/windcms.py report')),
+    ('tcm_compatible_replay', dict(helper='P.tcm_replay()', purpose='TCM 兼容链回放资产 (模型表/掩码阈值/裁决记录)',
+                        producer='随包快照 (无生成端)', consumer='src/windcms/report*.py · config.mask_thresholds')),
+    ('sop',        dict(helper='P.sop()',              purpose='SOP 中间件/评审/台账与事实契约底稿',
+                        producer='随包快照 (无生成端)', consumer='scripts/guanlan_facts_contract.py · 门户结论段')),
+    ('guanlan',    dict(helper='P.guanlan()',          purpose='事实契约与对外派生 (可上云面孔)',
+                        producer='scripts/guanlan_facts_contract.py', consumer='门户 #findings · /api/facts')),
+    ('pitch',      dict(helper='P.pitch()',            purpose='变桨侧派生件 (零位/日粒度)',
+                        producer='随包快照 + rebuild_from_raw --scada', consumer='脚本 windscada_serve.py 变桨面')),
+    ('paradigm_r1', dict(helper='(无助手: P.out_root()/paradigm_r1)', purpose='范式实验件 (E3/E5/E8 底稿, 事实契约输入)',
+                        producer='随包快照 (无生成端)', consumer='scripts/guanlan_facts_contract.py')),
+])
+
+# ── 输入类 → (放什么, 喂哪些产物, 跨度判据) ────────────────────────────────────────
+# span: 'csv_col' = 逐台 CSV 的首/末数据行; 'fname' = 从文件名年份/日期推; 'dir' = 目录名 (年/月)
+INPUTS = OrderedDict([
+    ('scada_10min', dict(what='SCADA 10min 导出 (逐台 WTG01.csv…WTG38.csv)', span='csv_col', time_col='Time',
+                         feeds=['windscada/powercurve*.parquet', 'windscada/loss_monthly.parquet', 'windscada/temp_monthly.parquet',
+                                'windscada/temp_bins.parquet', 'windscada/stop_events.parquet', 'windscada/yaw_daily.parquet',
+                                'windscada/hydraulic_accum.parquet', 'windscada/thermal_chain.parquet',
+                                'windscada/curve_*.parquet', 'windscada/control_*.parquet', 'windscada/system_aux.parquet'])),
+    ('scada_1min', dict(what='1min 导出 (逐台)', span='csv_col', time_col='occur_time',
+                        feeds=['windscada/watch_channels_monthly.parquet (watch 月度)'])),
+    ('故障报警',   dict(what='报警事件导出 (SpreadsheetML *.xls)', span='fname_year',
+                        feeds=['windscada/alarms.parquet'])),
+    ('风机故障记录', dict(what='检修工单台账 (*.xls/xlsx)', span='fname_year',
+                        feeds=['windscada/workorders.parquet', 'windscada/mblub_monthly.parquet'])),
+    ('油样报告',   dict(what='油液化验报告 (*.pdf)', span='fname_date',
+                        feeds=['windscada/oil_samples_index.parquet'])),
+    ('windcms',    dict(what='CMS 原始测量导出 (Brande TCM *_decode.json) + 厂家评估报告', span='dir',
+                        feeds=['m5_cms_tcm/windows/<窗>/{index.parquet,spectra/*}', 'windcms/厂家报告提取_*.json'])),
+    ('m5_cms_tcm', dict(what='振动线出件与 TCM 侧报告', span='none',
+                        feeds=['m5_cms_tcm/handoff_vibration_v2.json (现场正本优先)', 'windcms/报告_TCM传动链振动分析_*.md'])),
+])
+# 机理层 (不在场站目录下)
+TECH_INPUT = ('西门子4.0技术资料', '厂商技术资料 (故障处理手册/维护 WI/图纸/对译表)',
+              ['ontology/objects.json', 'ontology/turbine_params.parquet', 'ontology/retrieval_index.json'])
+
+
+def _rows_of(path: pathlib.Path, probe: int = 1):
+    """只读首/末若干行, 不整表载入 (scada csv 单文件几百 MB)。"""
+    try:
+        with open(path, 'rb') as f:
+            head = f.readline()
+            f.seek(max(0, path.stat().st_size - 65536))
+            tail = f.read().splitlines()
+        return head, (tail[-1] if tail else b'')
+    except Exception:
+        return b'', b''
+
+
+def input_span(station: pathlib.Path, sub: str, spec: dict):
+    """→ (起, 止, 件数, 说明, 粒度)。粒度 ∈ {'日','月','年'}: 只有'日/月'才参与"产物落后"判定 ——
+    按文件名年份推出来的跨度天然是**年粒度**, 拿它当"输入到 2026-12"会造出假缺口 (本脚本第一版就这么
+    误报过 2 条: 报警/工单"落后 5 个月"), 故显式带粒度、按粒度决定能不能比。"""
+    d = station / sub
+    if not d.is_dir():
+        return None, None, 0, '目录不存在', '年'
+    files = [p for p in d.rglob('*') if p.is_file()]
+    kind = spec.get('span')
+    if kind == 'csv_col':
+        col = spec.get('time_col')
+        starts, ends = [], []
+        for p in files[:60]:
+            try:
+                with open(p, 'rb') as f:
+                    first = f.readline()          # 表头
+                    first = f.readline()          # 第一条数据行
+                    f.seek(max(0, p.stat().st_size - 65536))
+                    last = f.read().splitlines()[-1]
+                hdr = first.decode('utf-8', 'replace').strip().split(',')
+                i = hdr.index(col) if col in hdr else 0
+                starts.append(hdr[i])
+                ends.append(last.decode('utf-8', 'replace').split(',')[0])
+            except Exception:
+                continue
+        if starts and ends:
+            return min(starts), max(ends), len(files), f'{len(files)} 件 (抽样 {len(starts)} 件首/末数据行)', '日'
+        return None, None, len(files), 'CSV 首/末行解析失败', '年'
+    if kind == 'fname_year':
+        ys = sorted({int(m.group()) for p in files for m in [re.search(r'(20\d\d)', p.name)] if m})
+        return (f'{ys[0]}-01' if ys else None), (f'{ys[-1]}-12' if ys else None), len(files), \
+               f'{len(files)} 件 (文件名年份, 年粒度)', '年'
+    if kind == 'fname_date':
+        ds = []
+        for p in files:
+            m = re.search(r'(\d{2})(\d{2})(20\d\d)', p.name)
+            if m:
+                ds.append(f'{m.group(3)}-{m.group(2)}-{m.group(1)}')
+        return (min(ds) if ds else None), (max(ds) if ds else None), len(files), \
+               f'{len(files)} 件 (文件名日期)', '日'
+    if kind == 'dir':
+        months = sorted({f'{m.group(1)}-{m.group(2)}' for p in d.rglob('*')
+                         for m in [re.search(r'[/\\](20\d\d)[/\\](\d{2})[/\\]', str(p))] if m})
+        return (months[0] if months else None), (months[-1] if months else None), len(files), \
+               f'{len(files)} 件 (目录年月 {", ".join(months) or "未识别"})', '月'
+    return None, None, len(files), '按件数清点 (无跨度判据)', '年'
+
+
+def product_stats(farm: str):
+    """→ {仓: dict(n, bytes, files: [(名, 类型, 大小)], kinds)}。"""
+    out = {}
+    for name in STORES:
+        d = P.out_root(farm) / name
+        if not d.is_dir():
+            out[name] = dict(n=0, bytes=0, kinds={}, newest=None)
+            continue
+        files = [p for p in d.rglob('*') if p.is_file()]
+        kinds = {}
+        for p in files:
+            k = p.suffix.lower() or '(无扩展名)'
+            kinds[k] = kinds.get(k, 0) + 1
+        newest = max((p.stat().st_mtime for p in files), default=None)
+        out[name] = dict(n=len(files), bytes=sum(p.stat().st_size for p in files), kinds=kinds,
+                         newest=(dt.datetime.fromtimestamp(newest).strftime('%Y-%m-%d %H:%M') if newest else None))
+    return out
+
+
+def provenance(farm: str):
+    root = P.out_root(farm)
+    prov = {}
+    for fn in ('_provenance.json', '_derived_manifest.json'):
+        f = root / fn
+        if not f.exists():
+            continue
+        try:
+            d = json.loads(f.read_text(encoding='utf-8'))
+        except Exception:
+            continue
+        for rel, v in (d.get('files') or {}).items():
+            src = v.get('source') if isinstance(v, dict) else None
+            prov[rel] = dict(source=src or 'raw-derived',
+                             builder=(v.get('builder') if isinstance(v, dict) else str(v)) or '')
+    return prov
+
+
+def span_check(farm: str, station: pathlib.Path, verbose=True):
+    """逐类输入: 算输入跨度 + 对应产物的跨度/条数, 报联动结论。→ (rows, problems)"""
+    import pandas as pd
+    rows, problems = [], []
+    ST = P.store(farm)
+
+    def pspan(p: pathlib.Path, col=None, path_cols=()):
+        if p is None or not p.exists():
+            return None, None, 0
+        try:
+            if p.suffix == '.parquet':
+                d = pd.read_parquet(p)
+                n = len(d)
+                c = col if col in d.columns else next((c for c in path_cols if c in d.columns), None)
+                if c is None:
+                    return None, None, n
+                s = pd.to_datetime(d[c], errors='coerce')
+                return ((str(s.min())[:10] if s.notna().any() else None),
+                        (str(s.max())[:10] if s.notna().any() else None), n)
+        except Exception:
+            return None, None, 0
+        return None, None, 0
+
+    # 输入类 → 它喂出来的产物 (路径, 时间列)
+    PAIRS = {
+        'scada_10min': [('windscada/temp_monthly.parquet', ST / 'temp_monthly.parquet', 'month'),
+                        ('windscada/loss_monthly.parquet', ST / 'loss_monthly.parquet', 'month'),
+                        ('windscada/powercurve_bins.parquet', ST / 'powercurve_bins.parquet', None)],
+        '故障报警': [('windscada/alarms.parquet', ST / 'alarms.parquet', 't_on')],
+        '风机故障记录': [('windscada/workorders.parquet', ST / 'workorders.parquet', 't_report')],
+        '油样报告': [('windscada/oil_samples_index.parquet', ST / 'oil_samples_index.parquet', 'date')],
+    }
+    for sub, spec in INPUTS.items():
+        if sub == 'm5_cms_tcm':
+            continue                                     # 出件不是"测量输入", 无跨度判据
+        i_s, i_e, i_n, note, gran = input_span(station, sub, spec)
+        comparable = gran in ('日', '月')                 # 年粒度不参与"落后"判定 (见 input_span 注释)
+        if sub == 'windcms':
+            # ★glob 必须写 `w[0-9][0-9][0-9][0-9]`: `w[????]` 是"字符类里含 ? 之一", 匹配不到任何窗名
+            #  (本脚本第一版就这么写, windcms 那行因此整行不出现 —— 静默漏项, 不是报错)
+            wins = [w for w in sorted((P.m5(farm) / 'windows').glob('w[0-9][0-9][0-9][0-9]'))
+                    if (w / 'index.parquet').exists()]
+            if wins:
+                w = wins[-1]
+                p_s, p_e, p_n = pspan(w / 'index.parquet', 'trigger_time')
+                rows.append(dict(input=sub, note=note, in_span=f'{i_s} ~ {i_e}', files=i_n,
+                                 product=f'm5_cms_tcm/windows/{w.name}/index.parquet',
+                                 prod_span=f'{p_s} ~ {p_e}', rows_n=p_n))
+                if p_s and i_s and p_s[:7] < i_s[:7]:
+                    problems.append(f'{sub}: 窗 {w.name} 索引起点 {p_s} 早于输入最早月 {i_s} (跨窗混入?)')
+            continue
+        for label, path, col in PAIRS.get(sub, []):
+            p_s, p_e, p_n = pspan(path, col)
+            rows.append(dict(input=sub, note=note, in_span=f'{i_s} ~ {i_e}', files=i_n, gran=gran,
+                             product=label, prod_span=f'{p_s} ~ {p_e}', rows_n=p_n))
+            if comparable and i_e and p_e and _months_between(p_e[:7], i_e[:7]) >= 2:
+                problems.append(f'{sub}: 产物 {label} 只到 {p_e}, 输入到 {i_e} '
+                                f'(落后 {_months_between(p_e[:7], i_e[:7])} 个月) → 未重算 或 该产物无生成端')
+            elif not comparable and p_e:
+                pass                                     # 年粒度: 跨度只作参考, 不判落后 (避免假缺口)
+    if verbose:
+        for r in rows:
+            print(f"  {r['input']:12s} 输入 {r['in_span']:26s} {r['files']:6d} 件 | "
+                  f"{r['product']:52s} {r['rows_n']:8d} 行 {r['prod_span']}")
+    return rows, problems
+
+
+def _months_between(a: str, b: str) -> int:
+    try:
+        ya, ma = int(a[:4]), int(a[5:7])
+        yb, mb = int(b[:4]), int(b[5:7])
+        return (yb - ya) * 12 + (mb - ma)
+    except Exception:
+        return 0
+
+
+def doc_block(farm: str) -> str:
+    """生成写进 docs/系统设计说明.md 的清单块 (Markdown)。"""
+    ps = product_stats(farm)
+    prov = provenance(farm)
+    from collections import Counter
+    by_store = Counter()
+    for rel, v in prov.items():
+        by_store[(rel.split('/')[0], v['source'])] += 1
+    L = [f'<!-- INVENTORY:BEGIN (由 scripts/inventory_products.py --write-doc 生成, 勿手改) -->',
+         f'*自动生成于 {dt.datetime.now():%Y-%m-%d %H:%M};数据源: `outputs/{farm}/_provenance.json` + `_derived_manifest.json` + 实际文件*', '',
+         '| 产物仓 | 输出路径 (相对安装根) | 功用 | 生成端 | 消费端 | 件数 | 大小 | 来源(raw 重算/随包) |',
+         '|---|---|---|---|---|---|---|---|']
+    for name, meta in STORES.items():
+        st = ps[name]
+        rd = by_store.get((name, 'raw-derived'), 0)
+        sh = by_store.get((name, 'shipped'), 0)
+        path = f'outputs/{farm}/{name}/'
+        L.append(f"| `{name}` | `{path}` | {meta['purpose']} | {meta['producer']} | {meta['consumer']} | "
+                 f"{st['n']} | {st['bytes'] / 1048576:.1f} MB | {rd} / {sh} |")
+    L += ['', f'合计 {sum(s["n"] for s in ps.values())} 件 / {sum(s["bytes"] for s in ps.values()) / 1048576:.1f} MB;'
+              f'其中 raw 重算 {sum(v for (_, k), v in by_store.items() if k == "raw-derived")} 件、'
+              f'随包补齐 {sum(v for (_, k), v in by_store.items() if k == "shipped")} 件。', MARK_E]
+    return '\n'.join(L)
+
+
+def write_doc(farm: str) -> int:
+    block = doc_block(farm)
+    txt = DOC.read_text(encoding='utf-8') if DOC.exists() else ''
+    if MARK_B in txt and MARK_E in txt:
+        head = txt.split(MARK_B)[0]
+        tail = txt.split(MARK_E)[1]
+        new = head + block + tail
+    else:
+        new = (txt + '\n\n## 附: 产物清单 (自动生成)\n\n' + block + '\n') if txt else block + '\n'
+    DOC.parent.mkdir(parents=True, exist_ok=True)
+    DOC.write_text(new, encoding='utf-8')
+    print(f'已写产物清单 → {P.rel(DOC)}')
+    return 0
+
+
+def main() -> int:
+    ap = argparse.ArgumentParser()
+    ap.add_argument('--farm', default=os_farm())
+    ap.add_argument('--check', action='store_true', help='只做呼应校验 (有问题退出码 5)')
+    ap.add_argument('--write-doc', action='store_true', help='把清单块写进 docs/系统设计说明.md')
+    ap.add_argument('--json', action='store_true')
+    a = ap.parse_args()
+    from src.windscada.config import farm as _farm, raw_station_dir
+    station = pathlib.Path(raw_station_dir(a.farm))
+    ps = product_stats(a.farm)
+
+    if a.write_doc:
+        return write_doc(a.farm)
+
+    print(f'=== 产物仓 ({len(STORES)} 个) @ outputs/{a.farm}/ ===')
+    for name, meta in STORES.items():
+        st = ps[name]
+        kinds = ' '.join(f'{k}×{v}' for k, v in sorted(st['kinds'].items(), key=lambda kv: -kv[1])[:5])
+        print(f"  {name:24s} {st['n']:6d} 件 {st['bytes'] / 1048576:9.1f} MB  {meta['helper']:28s} {kinds}")
+        print(f"     功用: {meta['purpose']}")
+
+    print(f'\n=== 输入 ↔ 产物 呼应 (输入根 {P.rel(station)}) ===')
+    rows, problems = span_check(a.farm, station)
+    if not rows:
+        print('  (无可对拍项)')
+    print(f'\n=== 结论: {"呼应正常" if not problems else str(len(problems)) + " 项要处理"} ===')
+    for p in problems:
+        print(f'  [!] {p}')
+    if a.json:
+        print(json.dumps(dict(products={k: {kk: vv for kk, vv in v.items() if kk != "kinds"} for k, v in ps.items()},
+                              spans=rows, problems=problems), ensure_ascii=False, indent=1))
+    return 5 if (a.check and problems) else 0
+
+
+def os_farm() -> str:
+    import os
+    return os.environ.get('WINDSCADA_FARM') or 'rudong'
+
+
+if __name__ == '__main__':
+    sys.exit(main())

+ 61 - 21
scripts/pack_dist.py

@@ -12,25 +12,33 @@ r"""打一个"拷到别的电脑能装、能跑"的分发包 (2026-09-12)。
 ## 包含 / 排除 (规则即文档)
 ## 包含 / 排除 (规则即文档)
 
 
 **包含**: 程序(`src/` `scripts/` `guanlan.py`) · 配置(`configs/`) · 离线安装件(`wheels/win_amd64/` `vendor/python/`)
 **包含**: 程序(`src/` `scripts/` `guanlan.py`) · 配置(`configs/`) · 离线安装件(`wheels/win_amd64/` `vendor/python/`)
-· 页面与交付(`release/` `resources/` `reference/`) · **产物**(`outputs/`, 页面取数靠它) · 文档(`docs/` `README_先读我.txt`
-`测试须知.txt` `_修复记录_20260911/`) · 安装/起停脚本(`install.bat` `install.sh` `install.ps1` `check.bat` `start.bat`
-`stop.bat`) · `requirements.txt`。
+· 页面与交付(`release/` `resources/` `reference/`) · **产物**(`outputs/`, 页面取数靠它; 加 `--no-products`
+则不含) · 文档(`docs/` `README_先读我.txt` `测试须知.txt`) · 安装/起停脚本(`install.bat` `install.sh`
+`install.ps1` `check.bat` `start.bat` `stop.bat`) · `requirements.txt`。
+
+**不含**修复记录/临时目录(`_修复记录_*` 之类): 本包已纳入 git 管理, 变更历史由版本库承载 (2026-09-12 用户令)。
 
 
 **排除**(每条都写了理由):
 **排除**(每条都写了理由):
     .git/ .venv/ .github/            版本库与虚拟环境 —— venv 换机必失效, 目标机安装时重建
     .git/ .venv/ .github/            版本库与虚拟环境 —— venv 换机必失效, 目标机安装时重建
-    _products_off*/                  本机"清除产物"的暂存与存档(每份含整份产物); 目标机用不到
+    _products_off*/                  历史遗留的"清除产物"暂存档(2026-09-16 起清除=真删, 不再产生此类目录)
     logs/ run/                       本机日志与 run/pids.json(旧 PID, 到新机器上是无效引用)
     logs/ run/                       本机日志与 run/pids.json(旧 PID, 到新机器上是无效引用)
-    data/raw/                        31 GB 现场原始件(约定"原始件不随包分发"); 要一起交付用 --with-data
+    data/raw/                        现场原始件(约定"原始件不随包分发"); 要一起交付用 --with-data
     __pycache__/ *.pyc               解释器缓存
     __pycache__/ *.pyc               解释器缓存
     *.zip(顶层)                      旧的交付压缩包; 避免包中包
     *.zip(顶层)                      旧的交付压缩包; 避免包中包
 
 
 ## 用法
 ## 用法
 
 
     python scripts/pack_dist.py --dry-run                  # 只报会打什么/多大, 不写文件
     python scripts/pack_dist.py --dry-run                  # 只报会打什么/多大, 不写文件
-    python scripts/pack_dist.py                            # 打包 → <父目录>/guanlan-rudong-v2_0.2.0_dist_<日期>.zip
+    python scripts/pack_dist.py                            # 打包 → <父目录>/guanlan-v<VERSION>_dist_<日期>.zip
+    python scripts/pack_dist.py --no-products              # **不含产物**(用户令 2026-09-16): 开箱页面显示"无产物"
     python scripts/pack_dist.py --with-data                # 连 data/raw 一起打 (30 GB, 慎用)
     python scripts/pack_dist.py --with-data                # 连 data/raw 一起打 (30 GB, 慎用)
     python scripts/pack_dist.py --out D:\out\x.zip         # 指定输出
     python scripts/pack_dist.py --out D:\out\x.zip         # 指定输出
     python scripts/pack_dist.py --verify <zip>             # **开箱验证**: 解压 → 离线安装 → 起服务核验页面 → 清理
     python scripts/pack_dist.py --verify <zip>             # **开箱验证**: 解压 → 离线安装 → 起服务核验页面 → 清理
+
+## 排除结果写进包
+
+包里 `dist-manifest.json` 记下 version / no_data / no_products / no_logs 与逐条排除理由;
+`guanlan.py check` 读它 —— 打了"不含产物"的包, 产物那几行检查按 `[--]` 说明而不判 FAIL。
 """
 """
 from __future__ import annotations
 from __future__ import annotations
 
 
@@ -48,8 +56,13 @@ import zipfile
 ROOT = pathlib.Path(__file__).resolve().parents[1]
 ROOT = pathlib.Path(__file__).resolve().parents[1]
 sys.path.insert(0, str(ROOT))
 sys.path.insert(0, str(ROOT))
 
 
-INCLUDE_DIRS = ['src', 'scripts', 'configs', 'release', 'resources', 'reference', 'outputs', 'docs',
-                'wheels', 'vendor', '_修复记录_20260911']
+INCLUDE_DIRS = ['src', 'scripts', 'configs', 'release', 'resources', 'reference', 'docs',
+                'wheels', 'vendor']
+PRODUCTS_DIR = 'outputs'          # 产物仓: 默认打进包 (页面开箱有数); --no-products 时排除
+# ★ 不打包"修复记录/临时/脚手架"类目录 (2026-09-12 用户令): 本包已纳入 git 管理, 变更历史由版本库承载,
+#   不必在交付物里再带一份 `_修复记录_<日期>\`。原先这里含 `_修复记录_20260911` (该目录的**内容**已并入
+#   `docs\振动数据接入_v0.1.md` 与 `docs\数据目录结构与落位约定_v0.2.md`, 目录本身按用户令删除)。
+VERSION = '0.4.0'                 # 包版本 (2026-09-16): 写进 dist-manifest.json 与默认文件名
 INCLUDE_FILES = ['guanlan.py', 'install.bat', 'install.ps1', 'install.sh', 'check.bat', 'start.bat', 'stop.bat',
 INCLUDE_FILES = ['guanlan.py', 'install.bat', 'install.ps1', 'install.sh', 'check.bat', 'start.bat', 'stop.bat',
                  'requirements.txt', 'README_先读我.txt', '测试须知.txt']
                  'requirements.txt', 'README_先读我.txt', '测试须知.txt']
 EXCLUDE_DIRS = {'.git', '.venv', '.github', 'logs', 'run', '__pycache__'}
 EXCLUDE_DIRS = {'.git', '.venv', '.github', 'logs', 'run', '__pycache__'}
@@ -69,13 +82,25 @@ def _xfer():
     return m
     return m
 
 
 
 
-def plan(with_data=False):
-    """→ (要打包的顶层条目 list, 排除说明 list)"""
+def plan(with_data=False, no_products=False):
+    """→ (要打包的顶层条目 list, 排除说明 list)
+
+    2026-09-16 用户令: 打包**不含 输入数据 / 产物 / 日志**。
+      · 输入数据 `data/`   —— 一直默认排除 (要用 --with-data 才带);
+      · 产物     `outputs/` —— 新增 --no-products 排除 (默认仍带: 页面开箱即有数);
+      · 日志     `logs/`    —— 一直在 EXCLUDE_DIRS 里 (连同 run/ 旧 PID)。
+    排除产物时, 包内 `dist-manifest.json` 会记 `no_products: true`,
+    目标机 `guanlan.py check` 据此把"产物缺失"显示为 `--`(待重算) 而不是 FAIL。
+    """
     inc = []
     inc = []
     for d in INCLUDE_DIRS:
     for d in INCLUDE_DIRS:
         p = ROOT / d
         p = ROOT / d
         if p.is_dir():
         if p.is_dir():
             inc.append(p)
             inc.append(p)
+    if not no_products:
+        p = ROOT / PRODUCTS_DIR
+        if p.is_dir():
+            inc.append(p)
     if with_data:
     if with_data:
         inc.append(ROOT / 'data')
         inc.append(ROOT / 'data')
     for f in INCLUDE_FILES:
     for f in INCLUDE_FILES:
@@ -83,8 +108,15 @@ def plan(with_data=False):
         if p.is_file():
         if p.is_file():
             inc.append(p)
             inc.append(p)
     skipped = [('.git/', '版本库'), ('.venv/', '虚拟环境(换机必失效, 目标机安装时重建)')]
     skipped = [('.git/', '版本库'), ('.venv/', '虚拟环境(换机必失效, 目标机安装时重建)')]
-    skipped += [(g + '/', '本机痕迹/超大体量, 见脚本头部说明') for g in EXCLUDE_GLOBS]
-    skipped += [('logs/, run/', '本机日志与旧 PID')]
+    _glob_why = {
+        '_products_off*': '历史遗留的"清除产物"暂存档 (2026-09-16 起清除=真删, 不再产生; 老机器上若有可手工删)',
+        'data': '现场原始输入件 (约定"原始件不随包分发"; 要一起交付用 --with-data) —— 用户令: 输入数据不进包',
+    }
+    skipped += [(g + '/', _glob_why.get(g, '本机痕迹/超大体量, 见脚本头部说明')) for g in EXCLUDE_GLOBS]
+    skipped += [('logs/, run/', '本机日志与旧 PID —— 用户令: 日志不进包')]
+    if no_products and (ROOT / PRODUCTS_DIR).is_dir():
+        skipped += [('outputs/', '用户令: 产物不进包 (目标机放数据后 rebuild_all.py 重算; '
+                                 '或在目标机也用 --no-products 的同一口径)')]
     return inc, skipped
     return inc, skipped
 
 
 
 
@@ -98,8 +130,8 @@ def size_of(p: pathlib.Path):
     return n, s
     return n, s
 
 
 
 
-def build(out: pathlib.Path, with_data=False, progress=True) -> int:
-    inc, skipped = plan(with_data)
+def build(out: pathlib.Path, with_data=False, progress=True, no_products=False) -> int:
+    inc, skipped = plan(with_data, no_products)
     tot_n = tot_s = 0
     tot_n = tot_s = 0
     print('== 要打进包的内容 ==')
     print('== 要打进包的内容 ==')
     for p in inc:
     for p in inc:
@@ -112,10 +144,15 @@ def build(out: pathlib.Path, with_data=False, progress=True) -> int:
         print(f'   {name:26s} {why}')
         print(f'   {name:26s} {why}')
 
 
     out.parent.mkdir(parents=True, exist_ok=True)
     out.parent.mkdir(parents=True, exist_ok=True)
-    man = dict(built=time.strftime('%Y-%m-%d %H:%M:%S'), root=str(ROOT), includes=[p.relative_to(ROOT).as_posix() for p in inc],
+    man = dict(version=VERSION, built=time.strftime('%Y-%m-%d %H:%M:%S'), root=str(ROOT),
+               includes=[p.relative_to(ROOT).as_posix() for p in inc],
                excluded={n: w for n, w in skipped}, files=tot_n, bytes=tot_s,
                excluded={n: w for n, w in skipped}, files=tot_n, bytes=tot_s,
-               note='本包由 scripts/pack_dist.py 生成; 目标机解压后跑 install.bat(Windows) 或 sh install.sh(Linux/macOS), '
-                    '再 check → start。.venv 不在包内(换机必失效, 安装时重建)。')
+               no_data=not with_data, no_products=bool(no_products), no_logs=True,
+               note='本包由 scripts/pack_dist.py 生成; **不含 输入数据(data/) · 产物(outputs/) · 日志(logs/, run/)** '
+                    '(2026-09-16 用户令)。目标机解压后跑 install.bat(Windows) 或 sh install.sh(Linux/macOS), '
+                    '再 check → start。要让页面有数: 把现场包放好后跑 '
+                    '`scripts/place_raw_data.py --src <现场包目录> --scope full` + `scripts/rebuild_all.py`。'
+                    '.venv 不在包内(换机必失效, 安装时重建)。')
     t0 = time.time()
     t0 = time.time()
     with zipfile.ZipFile(out, 'w', zipfile.ZIP_DEFLATED, compresslevel=6) as z:
     with zipfile.ZipFile(out, 'w', zipfile.ZIP_DEFLATED, compresslevel=6) as z:
         z.writestr('dist-manifest.json', json.dumps(man, ensure_ascii=False, indent=1))
         z.writestr('dist-manifest.json', json.dumps(man, ensure_ascii=False, indent=1))
@@ -175,17 +212,20 @@ def main() -> int:
     ap = argparse.ArgumentParser()
     ap = argparse.ArgumentParser()
     ap.add_argument('--out', default=None)
     ap.add_argument('--out', default=None)
     ap.add_argument('--with-data', action='store_true', help='连 data/raw 一起打(约 31 GB, 慎用)')
     ap.add_argument('--with-data', action='store_true', help='连 data/raw 一起打(约 31 GB, 慎用)')
+    ap.add_argument('--no-products', action='store_true',
+                    help='**不含产物** (outputs/): 用户令 2026-09-16 "打包不含 输入数据/产物/日志"')
     ap.add_argument('--dry-run', action='store_true')
     ap.add_argument('--dry-run', action='store_true')
     ap.add_argument('--verify', default=None, help='对已打好的 zip 做开箱验证(解压→安装→起服务核验)')
     ap.add_argument('--verify', default=None, help='对已打好的 zip 做开箱验证(解压→安装→起服务核验)')
     ap.add_argument('--keep', action='store_true', help='--verify 后保留解压目录')
     ap.add_argument('--keep', action='store_true', help='--verify 后保留解压目录')
     a = ap.parse_args()
     a = ap.parse_args()
     if a.verify:
     if a.verify:
         return verify(pathlib.Path(a.verify), a.keep)
         return verify(pathlib.Path(a.verify), a.keep)
-    out = pathlib.Path(a.out) if a.out else ROOT.parent / f'guanlan-rudong-v2_0.2.0_dist_{time.strftime("%Y%m%d")}.zip'
+    out = (pathlib.Path(a.out) if a.out
+           else ROOT.parent / f'guanlan-v{VERSION}_dist_{time.strftime("%Y%m%d")}.zip')
     if a.dry_run:
     if a.dry_run:
-        inc, skipped = plan(a.with_data)
+        inc, skipped = plan(a.with_data, a.no_products)
         tot_n = tot_s = 0
         tot_n = tot_s = 0
-        print('== 会打进包 ==')
+        print(f'== 会打进包 (version {VERSION}) ==')
         for p in inc:
         for p in inc:
             n, s = size_of(p); tot_n += n; tot_s += s
             n, s = size_of(p); tot_n += n; tot_s += s
             print(f'   {p.relative_to(ROOT).as_posix():26s} {n:6d} 件  {s/1e6:9.1f} MB')
             print(f'   {p.relative_to(ROOT).as_posix():26s} {n:6d} 件  {s/1e6:9.1f} MB')
@@ -195,7 +235,7 @@ def main() -> int:
             print(f'   {name:26s} {why}')
             print(f'   {name:26s} {why}')
         print('\n(dry-run, 未写文件)')
         print('\n(dry-run, 未写文件)')
         return 0
         return 0
-    return build(out, a.with_data)
+    return build(out, a.with_data, no_products=a.no_products)
 
 
 
 
 if __name__ == '__main__':
 if __name__ == '__main__':

+ 133 - 33
scripts/place_raw_data.py

@@ -45,20 +45,55 @@ A2 定了四项数据源的落位: data/raw/<场站名称>/{scada_10min, 故障
             `src_farm_names=['如海','如东']` 同源(如海=如东项目)。**包内没有消费者**: 它补的是
             `src_farm_names=['如海','如东']` 同源(如海=如东项目)。**包内没有消费者**: 它补的是
             数据层完整性, 不改变任何页面数值 —— 落它是因为"扫描辨识"会把它报成缺口。
             数据层完整性, 不改变任何页面数值 —— 落它是因为"扫描辨识"会把它报成缺口。
 
 
+## 振动侧 (--scope vib, 2026-09-12 用户令)
+
+用户令: 「数据层里的「CMS 振动评估报告」应遵循 `<安装目录>\\data\\raw\\如东\\windcms`、
+「振动线 handoff」应遵循 `<安装目录>\\data\\raw\\如东\\m5_cms_tcm` 存放; 修改系统支持振动数据参与
+运行、重算; 自现场包提取相应振动数据存放至上述目录」。
+
+  <场站>/windcms/CMS_RuDong_CGN_202603-04/measurement/<年>/<月>/<WTGxx>/*_decode.json
+      ← CMS_RuDong_CGN_202603-04.zip 的 measurement/** 全部成员 (25,679 件, 解压 153.7 GB)
+      依据: 这就是 `src/windcms/pipeline.py` 的 `detect_input()` 认的 **`tcm_decoded_json`**
+            (Brande TCM Enterprise 导出, SiteName=CGN Rudong, LocationName=WTGxx)。原样保留
+            包内 `measurement/` 这一层, 于是"哪个包来的哪一层"仍可追溯; 摄入按 rglob 找
+            `*_decode.json`, 套不套这层都能吃。**这是振动侧唯一能重算的源**: 六层链全部
+            输入都是它 (索引/谱 → 扫线 → 能量占比 → 模型 → 融合 → 报告)。
+            ★ 2026-09-11 版的 SKIPPED 里写着"振动分析报告是成品牌报告, windcms 要的测点索引
+              包里没有" —— 那句在当时成立; 2026-09-12 用户把 CMS 原始导出补进现场包后不再成立,
+              故本范围落地, 并在下面改注。
+
+  <场站>/windcms/厂家报告/上海电气_月度/     ← 如东风场数据.zip:
+      `数据收集/更新/(8)风机振动数据、油液分析记录(2025.1-至今)/振动分析报告/**`
+      (12 份月度用印版 PDF, 2025-06…2026-05) + 现场目录散装件
+      `中广核如东海上风电场2026年07月振动分析报告_上海电气.docx`
+      依据: 数据层「CMS 振动评估报告」的**厂商评估报告**侧 (VDI3834 / NB/T 31129-2018 判级)。
+            12 份 PDF 是纯扫描件 (无文本层, 实测 /Font=0、pdftotext 类抽取为空) → 只作归档证据,
+            数值不参与判级; docx 有文本层, 由 scripts/vib_reports_build.py 摄入。
+
+  <场站>/m5_cms_tcm/厂家报告/                ← 现场目录散装件
+      `中广核如东海上风电场传动链振动分析报告_大生科技_20260311(1).docx`
+      依据: 数据层「振动线 handoff」的**TCM 侧深度分析报告** (依托机组自带 TCM M-system 数据)。
+            若干现场给了 handoff 正本 (`handoff_vibration_v2.json` / `component_history.json`),
+            也放本目录 → 摄入优先采用现场正本, 不再用包内 shipped 快照。
+
 用法:
 用法:
     python scripts/place_raw_data.py --dry-run                # 只报要落什么, 不写盘 (默认 a2)
     python scripts/place_raw_data.py --dry-run                # 只报要落什么, 不写盘 (默认 a2)
     python scripts/place_raw_data.py                          # 真落位 (已存在则覆盖)
     python scripts/place_raw_data.py                          # 真落位 (已存在则覆盖)
     python scripts/place_raw_data.py --scope mech --dry-run    # 看机理层/1min 会落什么
     python scripts/place_raw_data.py --scope mech --dry-run    # 看机理层/1min 会落什么
-    python scripts/place_raw_data.py --scope full              # 两组一起
+    python scripts/place_raw_data.py --scope vib               # 落振动侧 (CMS 原始导出 153.7 GB + 厂家报告)
+    python scripts/place_raw_data.py --scope vib --limit 300   # 冒烟: 只落 300 件 (试跑/校验用)
+    python scripts/place_raw_data.py --scope full              # 三组一起
     python scripts/place_raw_data.py --src D:\\别的现场数据目录
     python scripts/place_raw_data.py --src D:\\别的现场数据目录
 
 
-## 故意不做的事 (两组范围共有的判断)
+## 故意不做的事 (各范围共有的判断)
 
 
   · 不落 `(3)现场检修记录/2025年检修记录` 与 `2026年检修记录`: 与 工作/风机故障记录/2025年故障记录、
   · 不落 `(3)现场检修记录/2025年检修记录` 与 `2026年检修记录`: 与 工作/风机故障记录/2025年故障记录、
     2026年故障记录 **逐件同名同大小**(19/19 件), 是同一批月度汇总表的副本。落两遍会让"按年目录"
     2026年故障记录 **逐件同名同大小**(19/19 件), 是同一批月度汇总表的副本。落两遍会让"按年目录"
     摄入看到两份重复台账 (2026-08-31 长停台账虚高 35 倍那类事故的同款成因: 快照重复必须归并, 不能叠加)。
     摄入看到两份重复台账 (2026-08-31 长停台账虚高 35 倍那类事故的同款成因: 快照重复必须归并, 不能叠加)。
-  · 不落 `(8)…/振动分析报告/`(12 份月度用印版 PDF): 是振动线 (CMS) 的**成品牌报告**, 不是可再加工的
-    测量数据; 而 `windcms`/`m5_cms_tcm` 要的是 CMS 测点索引(handoff), 包里没有。
+  · 振动报告只从 `(8)…/振动分析报告/` 落一遍 (不在 a2/mech 范围): 包内另有两处**同件副本** ——
+    根目录 `中广核如东海上风电场2026年5月振动分析报告用印版(2).pdf`(与 2026年05月 件同尺寸)
+    与嵌套包 `如东海上振动报告11份.zip`(其 11 份与 `振动分析报告/` 的 11 份逐件同尺寸)。
+    落三遍会得到 3 份同名报告, 摄入时会互相覆盖或重复计数 —— 故只在 vib 范围落目录里那一份。
   · 不落 `scada数据(如东)/**`(19 个月 × 12 个通道组 zip, 2.4 GB): 它是 `scada_10min/*.csv` 的**上游**
   · 不落 `scada数据(如东)/**`(19 个月 × 12 个通道组 zip, 2.4 GB): 它是 `scada_10min/*.csv` 的**上游**
     原始通道导出, 包内没有任何脚本读它(构建器读的是已经平铺好的 10min CSV) —— 落了也不参与重算。
     原始通道导出, 包内没有任何脚本读它(构建器读的是已经平铺好的 10min CSV) —— 落了也不参与重算。
   · 不落 `fastlog数据/`(4 件 WTG0x.xls): 全库搜 `fastlog` 只有 2 处**注释**提到它("运行态见证"),
   · 不落 `fastlog数据/`(4 件 WTG0x.xls): 全库搜 `fastlog` 只有 2 处**注释**提到它("运行态见证"),
@@ -105,6 +140,27 @@ RULES_MECH = [
     ('1分钟数据.zip', '', 'scada_1min', True, 'station'),
     ('1分钟数据.zip', '', 'scada_1min', True, 'station'),
 ]
 ]
 
 
+# 振动侧 (--scope vib): 数据层两条振动行的源件 —— CMS 原始测量导出 + 厂商评估报告
+ZIP_CMS = 'CMS_RuDong_CGN_202603-04.zip'
+DIR_SA_REPORTS = '如东风场数据/数据收集/更新/(8)风机振动数据、油液分析记录(2025.1-至今)/振动分析报告/'
+RULES_VIB = [
+    # ① CMS 原始测量导出 (Brande TCM Enterprise, SiteName=CGN Rudong): 25,679 件 *_decode.json。
+    #    strip=False 保留包内 measurement/ 层 → <包名>/measurement/2026/03/WTGxx/*.json, 来源可追溯。
+    (ZIP_CMS, 'measurement/', f'windcms/{ZIP_CMS[:-4]}', False, 'station'),
+    # ② 上海电气月度振动分析报告 (12 份扫描件 PDF): 数据层「CMS 振动评估报告」的厂商报告侧
+    ('如东风场数据.zip', DIR_SA_REPORTS, 'windcms/厂家报告/上海电气_月度', True, 'station'),
+]
+
+# 现场目录下的散装件 (不是压缩包成员): (源文件名, 目标相对 data/raw/<场站>/ 的路径)
+RULES_VIB_LOOSE = [
+    # 上海电气 2026年07月报告: 有文本层, 由 scripts/vib_reports_build.py 摄入成评估报告
+    ('中广核如东海上风电场2026年07月振动分析报告_上海电气.docx',
+     'windcms/厂家报告/上海电气_月度'),
+    # 大生科技传动链振动分析报告 (TCM M-system 数据): 数据层「振动线 handoff」的 TCM 侧
+    ('中广核如东海上风电场传动链振动分析报告_大生科技_20260311(1).docx',
+     'm5_cms_tcm/厂家报告'),
+]
+
 # 说明性的"故意不落", 只在报告里列出来
 # 说明性的"故意不落", 只在报告里列出来
 SKIPPED = [
 SKIPPED = [
     ('如东风场数据.zip', '如东风场数据/数据收集/更新/(3)现场检修记录(2025.1-至今)/2025年检修记录/',
     ('如东风场数据.zip', '如东风场数据/数据收集/更新/(3)现场检修记录(2025.1-至今)/2025年检修记录/',
@@ -112,7 +168,13 @@ SKIPPED = [
     ('如东风场数据.zip', '如东风场数据/数据收集/更新/(3)现场检修记录(2025.1-至今)/2026年检修记录/',
     ('如东风场数据.zip', '如东风场数据/数据收集/更新/(3)现场检修记录(2025.1-至今)/2026年检修记录/',
      '与 工作/风机故障记录/2026年故障记录 逐件同名校验相同 (副本)'),
      '与 工作/风机故障记录/2026年故障记录 逐件同名校验相同 (副本)'),
     ('如东风场数据.zip', '如东风场数据/数据收集/更新/(8)风机振动数据、油液分析记录(2025.1-至今)/振动分析报告/',
     ('如东风场数据.zip', '如东风场数据/数据收集/更新/(8)风机振动数据、油液分析记录(2025.1-至今)/振动分析报告/',
-     '振动线成品牌报告, 不是可再加工的测量数据; windcms 要的测点索引包里没有'),
+     '※ 2026-09-12 起**已改为落位** (vib 范围, → windcms/厂家报告/上海电气_月度)。原判断"成品牌报告+'
+     '包内没有测点索引"在当时成立; 用户补入 CMS_RuDong_CGN_202603-04.zip 后, 索引源件已具备, '
+     '报告作为数据层证据一并落位。此处保留记录以免后人以为漏了'),
+    ('如东风场数据.zip', '如东风场数据/中广核如东海上风电场2026年5月振动分析报告用印版(2).pdf',
+     '与 振动分析报告/…2026年05月…用印版.pdf 同尺寸 (副本, 只落目录里那一份)'),
+    ('如东风场数据.zip', '如东风场数据/如东海上振动报告11份.zip',
+     '其 11 份与 振动分析报告/ 的 11 份逐件同尺寸 (副本包, 只落目录里那一份)'),
     ('如东风场数据.zip', '如东风场数据/scada数据(如东)/',
     ('如东风场数据.zip', '如东风场数据/scada数据(如东)/',
      'scada_10min/*.csv 的上游原始通道导出 (2.4 GB), 包内无脚本读它 → 不参与重算'),
      'scada_10min/*.csv 的上游原始通道导出 (2.4 GB), 包内无脚本读它 → 不参与重算'),
     ('如东风场数据.zip', '如东风场数据/fastlog数据/',
     ('如东风场数据.zip', '如东风场数据/fastlog数据/',
@@ -140,8 +202,11 @@ def human(n: float) -> str:
     return f'{n:.1f} GB'
     return f'{n:.1f} GB'
 
 
 
 
-def plan(src: pathlib.Path, station: pathlib.Path, rules):
-    """→ [(zip, 目标文件, 包内条目名, 解压后大小, 目标根, 目标名)]; 只读压缩包目录, 不解压。"""
+def plan(src: pathlib.Path, station: pathlib.Path, rules, loose_rules=(), limit: int = 0):
+    """→ [(zip, 目标文件, 包内条目名, 解压后大小, 目标根, 目标名)]; 只读压缩包目录, 不解压。
+
+    zip=None 表示散装件 (现场目录下的单个文件, 如两份 docx 振动报告不在任何压缩包里)。
+    limit>0 时只取前 limit 件 (冒烟/试跑用; 大包 2.5 万件跑一遍要几十分钟, 得能小样验证)。"""
     out = []
     out = []
     for zname, prefix, target, strip, root in rules:
     for zname, prefix, target, strip, root in rules:
         zp = src / zname
         zp = src / zname
@@ -165,6 +230,13 @@ def plan(src: pathlib.Path, station: pathlib.Path, rules):
                         continue
                         continue
                     rest = prefix.rstrip('/').rsplit('/', 1)[-1]
                     rest = prefix.rstrip('/').rsplit('/', 1)[-1]
                 out.append((zname, base / target / rest, name, info.file_size, root, target))
                 out.append((zname, base / target / rest, name, info.file_size, root, target))
+    for fname, target in loose_rules:
+        fp = src / fname
+        if not fp.exists():
+            raise SystemExit(f'缺散装源件: {fp}')
+        out.append((None, station / target / fname, fname, fp.stat().st_size, 'station', target))
+    if limit:
+        out = out[:limit]
     return out
     return out
 
 
 
 
@@ -172,8 +244,10 @@ def main() -> int:
     ap = argparse.ArgumentParser()
     ap = argparse.ArgumentParser()
     ap.add_argument('--src', default=str(DEFAULT_SRC), help='现场数据目录 (默认 %(default)s)')
     ap.add_argument('--src', default=str(DEFAULT_SRC), help='现场数据目录 (默认 %(default)s)')
     ap.add_argument('--dry-run', action='store_true', help='只报计划, 不写盘')
     ap.add_argument('--dry-run', action='store_true', help='只报计划, 不写盘')
-    ap.add_argument('--scope', choices=('a2', 'mech', 'full'), default='a2',
-                    help='a2=只落四项数据层(A3 语义, 默认) · mech=机理层源件+1min · full=两组一起')
+    ap.add_argument('--scope', choices=('a2', 'mech', 'vib', 'full'), default='a2',
+                    help='a2=四项数据层(A3 语义, 默认) · mech=机理层源件+1min · '
+                         'vib=振动侧(CMS 原始导出+厂家报告) · full=三组一起')
+    ap.add_argument('--limit', type=int, default=0, help='只落前 N 件 (冒烟/试跑; 默认 0=全落)')
     a = ap.parse_args()
     a = ap.parse_args()
 
 
     src = pathlib.Path(a.src)
     src = pathlib.Path(a.src)
@@ -186,8 +260,13 @@ def main() -> int:
     print(f'场站目录: {station}   (来自 src/windscada/config.py raw_station={cfg.get("raw_station")!r})')
     print(f'场站目录: {station}   (来自 src/windscada/config.py raw_station={cfg.get("raw_station")!r})')
     print(f'现场数据: {src}   范围: --scope {a.scope}\n')
     print(f'现场数据: {src}   范围: --scope {a.scope}\n')
 
 
-    rules = {'a2': RULES, 'mech': RULES_MECH, 'full': RULES + RULES_MECH}[a.scope]
-    items = plan(src, station, rules)
+    rules, loose = {
+        'a2': (RULES, ()),
+        'mech': (RULES_MECH, ()),
+        'vib': (RULES_VIB, RULES_VIB_LOOSE),
+        'full': (RULES + RULES_MECH + RULES_VIB, RULES_VIB_LOOSE),
+    }[a.scope]
+    items = plan(src, station, rules, loose, limit=a.limit)
     by_target = {}
     by_target = {}
     for zname, dst, entry, size, root, target in items:
     for zname, dst, entry, size, root, target in items:
         d = by_target.setdefault((root, target), [0, 0])
         d = by_target.setdefault((root, target), [0, 0])
@@ -196,12 +275,14 @@ def main() -> int:
     print(f'== 计划落位 (scope={a.scope}) ==')
     print(f'== 计划落位 (scope={a.scope}) ==')
     for (root, target), (n, sz) in sorted(by_target.items(), key=lambda kv: kv[0][1]):
     for (root, target), (n, sz) in sorted(by_target.items(), key=lambda kv: kv[0][1]):
         where = (station / target) if root == 'station' else (station.parent / target)
         where = (station / target) if root == 'station' else (station.parent / target)
-        print(f'  {target:18s} {n:5d} 件  {human(sz):>10s}   → {where}')
+        print(f'  {target:44s} {n:6d} 件  {human(sz):>10s}   → {where}')
     print(f'  合计 {len(items)} 件, {human(sum(i[3] for i in items))}')
     print(f'  合计 {len(items)} 件, {human(sum(i[3] for i in items))}')
 
 
-    print('\n== 故意不落 ==')
-    for zname, prefix, why in SKIPPED:
-        print(f'  {prefix}\n      理由: {why}')
+    if a.scope in ('vib', 'full'):
+        print('\n== 故意不落 (振动侧同件副本) ==')
+        for zname, prefix, why in SKIPPED:
+            if '振动' in prefix or '报告11份' in prefix:
+                print(f'  {prefix}\n      理由: {why}')
 
 
     if a.dry_run:
     if a.dry_run:
         print('\n(dry-run, 未写盘)')
         print('\n(dry-run, 未写盘)')
@@ -210,24 +291,43 @@ def main() -> int:
     print('\n== 落位 ==')
     print('\n== 落位 ==')
     done = 0
     done = 0
     skipped = 0
     skipped = 0
-    for zname, dst, entry, size, root, target in items:
-        # 已有同尺寸文件 = 已经是最新 → 跳过。重跑一次不该把 15 GB 原样再抄一遍
-        # (mech 范围 15.5 GB, a2 范围 14.7 GB; 2026-09-11 修单文件规则时就是靠这个避免整盘重写)。
-        if dst.exists() and dst.stat().st_size == size:
-            skipped += 1
-            continue
-        dst.parent.mkdir(parents=True, exist_ok=True)
-        with zipfile.ZipFile(src / zname) as zf:
-            info = next(i for i in zf.infolist() if gbk_name(i).replace('\\', '/') == entry)
-            with zf.open(info) as fsrc, open(dst, 'wb') as fdst:
-                shutil.copyfileobj(fsrc, fdst, 1024 * 1024 * 4)
-        got = dst.stat().st_size
-        if got != size:
-            raise SystemExit(f'写出大小不符: {dst} 期望 {size} 实得 {got}')
-        done += 1
-        if done % 10 == 0 or size > 50 * 1024 * 1024:
-            print(f'  [{done}/{len(items)}] {human(size):>10s}  {dst.relative_to(station.parent)}', flush=True)
-    print(f'\n完成 (scope={a.scope}): 新写/更新 {done} 件, 已是最新跳过 {skipped} 件, 共 {len(items)} 件')
+    written = 0
+    # 按包分组 + **每包只开一次**。旧写法在循环体里 with ZipFile(...) 且用 next(...) 线性找条目 ——
+    # 对 2.5 万件的大包是 25,679 次开包 × 25,759 次名字比较 ≈ 6.6 亿次比较, 实测不可接受。
+    # 改成: 每包一次性建 {条目名: ZipInfo} 映射, 顺序流式写出。
+    by_zip = {}
+    for it in items:
+        by_zip.setdefault(it[0], []).append(it)
+    for zname, group in by_zip.items():
+        zf = zipfile.ZipFile(src / zname) if zname else None
+        try:
+            imap = {gbk_name(i).replace('\\', '/'): i for i in zf.infolist()} if zf else {}
+            for _, dst, entry, size, root, target in group:
+                # 已有同尺寸文件 = 已经是最新 → 跳过。重跑一次不该把 150 GB 原样再抄一遍
+                # (vib 范围解压 153.7 GB; a2 14.7 GB / mech 15.5 GB; 2026-09-11 修单文件规则时
+                #  就是靠这个避免整盘重写)。
+                if dst.exists() and dst.stat().st_size == size:
+                    skipped += 1
+                    continue
+                dst.parent.mkdir(parents=True, exist_ok=True)
+                if zf is None:
+                    with open(src / entry, 'rb') as fsrc, open(dst, 'wb') as fdst:
+                        shutil.copyfileobj(fsrc, fdst, 1024 * 1024 * 4)
+                else:
+                    with zf.open(imap[entry]) as fsrc, open(dst, 'wb') as fdst:
+                        shutil.copyfileobj(fsrc, fdst, 1024 * 1024 * 4)
+                got = dst.stat().st_size
+                if got != size:
+                    raise SystemExit(f'写出大小不符: {dst} 期望 {size} 实得 {got}')
+                done += 1
+                written += size
+                # 大包进度: 每 500 件报一次 (2.5 万件逐件打印会把日志刷爆); 单件 >200 MB 也报
+                if done % 500 == 0 or size > 200 * 1024 * 1024:
+                    print(f'  [{done}/{len(items)}] 已写 {human(written)}  {dst.name}', flush=True)
+        finally:
+            if zf is not None:
+                zf.close()
+    print(f'\n完成 (scope={a.scope}): 新写/更新 {done} 件 ({human(written)}), 已是最新跳过 {skipped} 件, 共 {len(items)} 件')
     return 0
     return 0
 
 
 
 

+ 37 - 0
scripts/portal_build.py

@@ -46,6 +46,7 @@ import pathlib
 import re
 import re
 import sys
 import sys
 import tempfile
 import tempfile
+import time
 
 
 ROOT = pathlib.Path(__file__).resolve().parents[1]
 ROOT = pathlib.Path(__file__).resolve().parents[1]
 sys.path.insert(0, str(ROOT))
 sys.path.insert(0, str(ROOT))
@@ -227,6 +228,39 @@ def do_verify() -> int:
     return 3
     return 3
 
 
 
 
+def do_rebaseline() -> int:
+    """把当前 shell.html / portal.html 记为新的基线 (manifest 里的 sha256)。
+
+    为什么需要它: manifest 的 sha256 是"漂移检测"的基准。**有意**改了门户外壳 (例如 2026-09-16
+    按用户令在菜单里加「数据重算」) 之后, `--check` 会如实报漂移 —— 这是对的, 但不能让人只有
+    "手改 JSON" 或 "跑 --extract 重新拆包" 两条路: 后者会用 portal.html 反过来覆盖 shell.html
+    (而装配过程会注入契约结论段/内嵌模板), 一不小心就把手工维护的外壳冲掉。
+    本命令只改 manifest 里的两个基准值, 并记下是谁、什么时候、为什么重基线。
+    """
+    if not MANIFEST.is_file():
+        print(f'[X] 缺 {P.rel(MANIFEST)}; 先跑 --extract')
+        return 1
+    if not (SHELL.is_file() and PORTAL.is_file()):
+        print('[X] shell.html / portal.html 不全, 不能重基线')
+        return 1
+    man = json.loads(rd(MANIFEST))
+    old_shell = man.get('shell', {}).get('sha256', '?')[:16]
+    old_portal = (man.get('expected_portal_sha256') or '?')[:16]
+    shell_b = SHELL.read_bytes()
+    portal_b = PORTAL.read_bytes()
+    man['shell'] = dict(file='shell.html', bytes=len(shell_b), sha256=sha(shell_b))
+    man['expected_portal_sha256'] = sha(portal_b)
+    man['notes'] = list(man.get('notes') or []) + [
+        f'重基线 {time.strftime("%Y-%m-%d %H:%M")}: shell {old_shell}→{sha(shell_b)[:16]}, '
+        f'portal {old_portal}→{sha(portal_b)[:16]} (有意改动门户外壳后重设漂移基准)']
+    wr(MANIFEST, json.dumps(man, ensure_ascii=False, indent=1))
+    print(f'已重基线 {P.rel(MANIFEST)}')
+    print(f'  shell.html  {old_shell} → {sha(shell_b)[:16]}  ({len(shell_b):,} B)')
+    print(f'  portal.html {old_portal} → {sha(portal_b)[:16]}  ({len(portal_b):,} B)')
+    print('  下一步: --check 应显示"全部与 manifest 一致 ✔"')
+    return 0
+
+
 def do_check() -> int:
 def do_check() -> int:
     if not MANIFEST.is_file():
     if not MANIFEST.is_file():
         print(f'[X] 缺 {P.rel(MANIFEST)}; 先跑 --extract')
         print(f'[X] 缺 {P.rel(MANIFEST)}; 先跑 --extract')
@@ -268,7 +302,10 @@ if __name__ == '__main__':
     g.add_argument('--extract', action='store_true', help='从现有 portal.html 拆出 portal_src/ (一次性)')
     g.add_argument('--extract', action='store_true', help='从现有 portal.html 拆出 portal_src/ (一次性)')
     g.add_argument('--verify', action='store_true', help='装配到临时文件并与现有门户逐字节比对')
     g.add_argument('--verify', action='store_true', help='装配到临时文件并与现有门户逐字节比对')
     g.add_argument('--check', action='store_true', help='源件与 manifest 的 sha256 漂移检查')
     g.add_argument('--check', action='store_true', help='源件与 manifest 的 sha256 漂移检查')
+    g.add_argument('--rebaseline', action='store_true',
+                   help='**有意**改了 shell.html/portal.html 后, 把当前状态记为新的漂移基准 (只改 manifest 的两个 sha, 不碰文件)')
     ap.add_argument('--no-claims', action='store_true', help='装配时不注入契约结论段')
     ap.add_argument('--no-claims', action='store_true', help='装配时不注入契约结论段')
     a = ap.parse_args()
     a = ap.parse_args()
     sys.exit(do_extract() if a.extract else do_verify() if a.verify else
     sys.exit(do_extract() if a.extract else do_verify() if a.verify else
+             do_rebaseline() if a.rebaseline else
              do_check() if a.check else do_build(claims=not a.no_claims))
              do_check() if a.check else do_build(claims=not a.no_claims))

+ 78 - 13
scripts/products_restore_missing.py

@@ -81,7 +81,8 @@ def stash_dir(explicit=None) -> pathlib.Path:
     含它的目录才是"标准答案"来源。原先只看第一个候选, 于是 --on 还原后暂存区空了 → 台账被写成 0/0。
     含它的目录才是"标准答案"来源。原先只看第一个候选, 于是 --on 还原后暂存区空了 → 台账被写成 0/0。
     """
     """
     if explicit:
     if explicit:
-        return pathlib.Path(explicit)
+        # 必须 resolve: 传相对路径时下面 `stash.relative_to(ROOT)` 会 ValueError (2026-09-12 实逮)
+        return pathlib.Path(explicit).resolve()
     marker = 'windscada/temp_monthly.parquet'
     marker = 'windscada/temp_monthly.parquet'
     cands = []
     cands = []
     for pat in ('_products_off/*', '_products_off_prev_*/*'):
     for pat in ('_products_off/*', '_products_off_prev_*/*'):
@@ -93,7 +94,21 @@ def stash_dir(explicit=None) -> pathlib.Path:
     scored.sort(key=lambda x: (-x[0], -x[1]))
     scored.sort(key=lambda x: (-x[0], -x[1]))
     if scored:
     if scored:
         return scored[0][2]
         return scored[0][2]
-    raise SystemExit('找不到随包产物原件 (_products_off/** 或 _products_off_prev_**); 用 --stash 指定')
+    raise StashMissing('找不到随包产物原件 (_products_off/** 或 _products_off_prev_**); 用 --stash 指定')
+
+
+class StashMissing(SystemExit):
+    """暂存区不存在 —— **不是崩溃, 是"这一步没得做"**。
+
+    2026-09-16 用户令"系统运行中/未运行都能重算"时实逮: `rebuild_all.py` 的第⑤步 (补齐"包内没有生成端"
+    的产物) 在**没有暂存区**的检出上必抛 SystemExit → 重建链当场中断, 后面的 ⑦本体链 与 ⑧审计**全不执行**
+    (脚本的中断语义是"后面的步骤依赖它")。而暂存区只有"清除产物"动作会产生 —— 首次全量重算、或
+    清完又还原过的机器上就都没有, 于是"从零重算"在交付包上跑不完。
+    处置: 本类带退出码 6, 由 rebuild_all.py 显式容忍并打印"跳过随包补齐" —— 缺的是"随包件"这一路,
+    不是重算本身; 如实报出来比假装通过好, 也比把整条链打断好。
+    """
+
+    exit_code = 6
 
 
 
 
 def main() -> int:
 def main() -> int:
@@ -101,16 +116,25 @@ def main() -> int:
     ap.add_argument('--dry-run', action='store_true')
     ap.add_argument('--dry-run', action='store_true')
     ap.add_argument('--stash', default=None)
     ap.add_argument('--stash', default=None)
     a = ap.parse_args()
     a = ap.parse_args()
-    stash = stash_dir(a.stash)
-    dest_root = P.STORE if hasattr(P, 'STORE') else ROOT / 'outputs' / 'rudong'
-    print(f'随包件: {stash.relative_to(ROOT)}')
+    try:
+        stash = stash_dir(a.stash)
+    except StashMissing as e:
+        # 没有暂存区 → 这一步没得做。**不许静默、也不许把整条重算链打断**(见 StashMissing 的说明)。
+        print(f'[跳过] {e}')
+        print('       处置: ① 从**交付包 zip** 按需补齐: --stash <交付包.zip> (推荐; 2026-09-16 用户令'
+              '"清除产物不留备份"之后, 交付包就是随包件的来源); '
+              '② --stash <随包产物目录> 显式指定一个目录; ③ 只想让页面有数 → 产物就在位, 无需本步。')
+        return StashMissing.exit_code
+    # ★2026-09-16: 原写法 `P.STORE if hasattr(P,'STORE') else ROOT/'outputs'/'rudong'` —— P.STORE 根本
+    #   不存在, 于是**恒**落到硬编码的 rudong, 多场部署下会把 rudong 的随包件补进别的场 (静默串场)。
+    dest_root = P.out_root()
+    print(f'随包件: {P.rel(stash)}')
     print(f'产物仓: {dest_root.relative_to(ROOT)}\n')
     print(f'产物仓: {dest_root.relative_to(ROOT)}\n')
 
 
     prov = {}
     prov = {}
-    for src in sorted(stash.rglob('*')):
-        if not src.is_file():
-            continue
-        rel = src.relative_to(stash).as_posix()
+
+    def take(rel: str, src_path=None, src_bytes=None):
+        """补齐一件 (或在位则只记账)。src_path=目录来源, src_bytes=zip 来源。"""
         dst = dest_root / rel
         dst = dest_root / rel
         if dst.exists():
         if dst.exists():
             # 已存在的件分两类: 我们自己重算的(raw-derived) 与 上轮已补齐的(shipped)。
             # 已存在的件分两类: 我们自己重算的(raw-derived) 与 上轮已补齐的(shipped)。
@@ -120,17 +144,58 @@ def main() -> int:
                 prov[rel] = dict(source='raw-derived', builder=RAW_DERIVED[rel])
                 prov[rel] = dict(source='raw-derived', builder=RAW_DERIVED[rel])
             else:
             else:
                 prov[rel] = dict(source='shipped', why='包内无生成端 / 规则未复现 → 随包件补齐 (本轮已在位)')
                 prov[rel] = dict(source='shipped', why='包内无生成端 / 规则未复现 → 随包件补齐 (本轮已在位)')
-            continue
+            return
         prov[rel] = dict(source='shipped', why='包内无生成端 / 规则未复现 → 用随包件补齐')
         prov[rel] = dict(source='shipped', why='包内无生成端 / 规则未复现 → 用随包件补齐')
-        if not a.dry_run:
-            dst.parent.mkdir(parents=True, exist_ok=True)
-            shutil.copy2(src, dst)
+        if a.dry_run:
+            return
+        dst.parent.mkdir(parents=True, exist_ok=True)
+        if src_bytes is not None:
+            dst.write_bytes(src_bytes)
+        else:
+            shutil.copy2(src_path, dst)
+
+    if stash.is_file():
+        # 交付包 zip 形态 (2026-09-16): 按需从包里取 outputs/<场>/** 补齐 ——
+        # 这是"清除不留备份"之后**唯一**的随包件来源, 所以把它做成一等输入。
+        import zipfile
+        with zipfile.ZipFile(stash) as zf:
+            pref = f'outputs/{P.farm()}/'
+            members = [n for n in zf.namelist() if n.startswith(pref) and not n.endswith('/')]
+            print(f'  从交付包读取 {len(members)} 件 (前缀 {pref})')
+            for name in sorted(members):
+                if a.dry_run:
+                    take(name[len(pref):])
+                else:
+                    take(name[len(pref):], src_bytes=zf.read(name))
+    else:
+        for src in sorted(stash.rglob('*')):
+            if not src.is_file():
+                continue
+            take(src.relative_to(stash).as_posix(), src_path=src)
 
 
     # 重算件里有些**不在随包件里**(如我们新造的 turbine_params.parquet), 也要记进台账
     # 重算件里有些**不在随包件里**(如我们新造的 turbine_params.parquet), 也要记进台账
     for rel, builder in RAW_DERIVED.items():
     for rel, builder in RAW_DERIVED.items():
         if rel not in prov and (dest_root / rel).exists():
         if rel not in prov and (dest_root / rel).exists():
             prov[rel] = dict(source='raw-derived', builder=builder)
             prov[rel] = dict(source='raw-derived', builder=builder)
 
 
+    # 构建脚本的**自登记** (src/derived_manifest.py): 振动侧的产物名随窗/分片变, 写不进上面的精确表,
+    # 按名字通配又会误伤同名旧件 → 由"谁算的谁登记", 这里只认登记。缺失不影响其余记账。
+    try:
+        from src.derived_manifest import load as _load_manifest, prune as _prune_manifest
+        _gone = _prune_manifest(dest_root)      # 先清掉"登记了但盘上已删"的条目 (如被删的 _reimport 窗)
+        if _gone:
+            print(f'  自登记清理: {_gone} 条指向已删文件的条目')
+        _man = (_load_manifest(dest_root).get('files') or {})
+    except Exception:
+        _man = {}
+    for rel, info in _man.items():
+        if (dest_root / rel).exists():
+            prov[rel] = dict(source='raw-derived',
+                             builder=f"{info.get('builder', '?')} (自登记 by {info.get('by', '?')})")
+        else:
+            prov[rel] = dict(source='raw-derived',
+                             builder=f"{info.get('builder', '?')} (自登记, 但该件当前不在盘上)")
+
     by_store = {}
     by_store = {}
     for rel, m in prov.items():
     for rel, m in prov.items():
         store = rel.split('/')[0]
         store = rel.split('/')[0]

+ 55 - 130
scripts/products_state.py

@@ -1,40 +1,42 @@
 #!/usr/bin/env python3
 #!/usr/bin/env python3
 # -*- coding: utf-8 -*-
 # -*- coding: utf-8 -*-
-r"""产物开关 —— 把 outputs\<场>\ 下的产物整体挪走 / 挪回 (人工检查页面空状态用)。
+r"""产物开关 —— 把 outputs\<场>\ 下的产物**直接清掉** / 看当前状态 (人工检查页面空状态用)。
 
 
-## 为什么做成开关而不是删除
+## 2026-09-16 用户令: 清除产物**不留备份**
 
 
-产物是"上一次算出来的结果", 人工检查"没数据时页面长什么样"必须把它们清掉; 但清掉之后一定要能**一条命令还原**,
-否则就得重跑整条摄入链(而且像油样那 102 行华创合并报告根本重算不出来)。所以: **移动**(同盘瞬间完成) + 清单留痕。
+原设计是"移动到 `_products_off/<场>/` + 清单留痕, 一条命令还原"。用户令改为**不备份**:
+清掉就是清掉。于是本脚本:
+  · `--off --yes`  真删除 (rmtree), 打印删了什么; **不再产生 `_products_off*` 目录**;
+  · `--status`     看产物在位/已清 (清掉后就是"已清", 没有备份清单可列);
+  · `--on`         已**移除**: 没有备份就没有"挪回"。需要恢复缺失的**随包件**时, 用
+                   `python scripts/products_restore_missing.py --stash <交付包.zip>`
+                   (从交付包 zip 里按需补齐 —— 那是"从交付件恢复", 不是"留一份备份")。
 
 
 ## 保留不动的
 ## 保留不动的
 
 
 - `data\raw\` 现场数据 (检查的是"产物没了会怎样", 不是"数据没了会怎样")
 - `data\raw\` 现场数据 (检查的是"产物没了会怎样", 不是"数据没了会怎样")
-- `outputs\<场>\windscada\_pre_rebuild_20260911\` 随包基线备份 (它是备份, 不是产物; `rebuild_from_raw.py --verify` 要用)
 - `release\` 门户与三维/交付件 (那是外壳, 不是数据产物)
 - `release\` 门户与三维/交付件 (那是外壳, 不是数据产物)
 - `logs\` `run\` 运行期文件
 - `logs\` `run\` 运行期文件
+- ★ `outputs\<场>\windscada\_pre_rebuild_20260911\` **不再特殊保留**: 它曾是"随包基线备份"而在清除时豁免,
+  现在按"不留备份"一并删掉 —— 需要等价验收时, 用 `products_restore_missing.py --stash <交付包.zip>`
+  把这个目录从交付包里补回来即可。
 
 
 用法:
 用法:
-    python scripts/products_state.py --off      # 挪走全部产物, 打印清单
-    python scripts/products_state.py --status   # 看当前是开还是关
-    python scripts/products_state.py --on       # 全部挪回
+    python scripts/products_state.py --status            # 看当前产物状态
+    python scripts/products_state.py --off --yes         # 清掉全部产物 (真删除, 不可恢复; --yes 是防手滑)
 """
 """
 from __future__ import annotations
 from __future__ import annotations
 
 
 import argparse
 import argparse
-import json
 import pathlib
 import pathlib
 import shutil
 import shutil
 import sys
 import sys
-import time
 
 
 ROOT = pathlib.Path(__file__).resolve().parents[1]
 ROOT = pathlib.Path(__file__).resolve().parents[1]
 sys.path.insert(0, str(ROOT))
 sys.path.insert(0, str(ROOT))
 from src import paths as P                                        # noqa: E402
 from src import paths as P                                        # noqa: E402
 
 
-OFF = ROOT / '_products_off'                                       # 挪走后的存放处 (仓库外无需, 放安装目录下便于还原)
-MANIFEST = OFF / 'manifest.json'
-KEEP = ('_pre_rebuild_20260911',)                                  # 随包基线备份, 不算产物
+LEGACY_OFF = ROOT / '_products_off'                                # 旧版暂存区 (只用于提示, 不再写入)
 
 
 
 
 def farms() -> list[pathlib.Path]:
 def farms() -> list[pathlib.Path]:
@@ -42,15 +44,9 @@ def farms() -> list[pathlib.Path]:
     return sorted(p for p in out.iterdir() if p.is_dir()) if out.is_dir() else []
     return sorted(p for p in out.iterdir() if p.is_dir()) if out.is_dir() else []
 
 
 
 
-def plan() -> list[tuple[pathlib.Path, pathlib.Path]]:
-    """→ [(原位置, 挪走后的位置)]"""
-    items = []
-    for fam in farms():
-        for d in sorted(fam.iterdir()):
-            if d.name in KEEP:
-                continue
-            items.append((d, OFF / fam.name / d.name))
-    return items
+def plan() -> list[pathlib.Path]:
+    """→ 待清除的产物目录列表 (产物仓本身)。"""
+    return [d for fam in farms() for d in sorted(fam.iterdir())]
 
 
 
 
 def count(p: pathlib.Path) -> tuple[int, float]:
 def count(p: pathlib.Path) -> tuple[int, float]:
@@ -60,131 +56,60 @@ def count(p: pathlib.Path) -> tuple[int, float]:
     return len(fs), sum(f.stat().st_size for f in fs) / 1024 / 1024
     return len(fs), sum(f.stat().st_size for f in fs) / 1024 / 1024
 
 
 
 
-def archive_old() -> bool:
-    """把**上一代**暂存区改名存档, 给这一次挪走腾出干净的落点。
-
-    为什么必须做: 挪走的落点是 `_products_off/<场>/<目录>`; 若该目录已存在(shutil.move 的语义是
-    "目的地是已存在目录 → 把源挪进去"), 会嵌套成 `_products_off/rudong/windscada/windscada/…`,
-    随后 `--on` 就会把新旧两代混在一起还原 —— 2026-09-12 实逮这个隐患。
-    存档时**保留** `_baseline_kept/`(随包基线, `rebuild_from_raw.py --verify` 要用), 只动场站目录与清单。
-    """
-    arch = ROOT / f'_products_off_prev_{time.strftime("%Y%m%d_%H%M%S")}'
-    moved = []
-    for fam in farms():
-        src = OFF / fam.name
-        if src.exists():
-            arch.mkdir(parents=True, exist_ok=True)
-            shutil.move(str(src), str(arch / fam.name))
-            moved.append(fam.name)
-    if MANIFEST.exists():
-        arch.mkdir(parents=True, exist_ok=True)
-        shutil.move(str(MANIFEST), str(arch / 'manifest.json'))
-    if not moved and not arch.exists():
-        return False
-    print(f'  上一代暂存区已存档 → {P.rel(arch)}/  (场站: {", ".join(moved) or "无"}; 基线 _baseline_kept 留在原地)')
-    return True
-
-
-def do_off(archive=False) -> int:
-    items = plan()
+def do_off(assume_yes=False) -> int:
+    items = [p for p in plan() if p.exists()]
     if not items:
     if not items:
-        print('产物已经全部挪走了 (用 --status 看清单)')
+        print('产物已经是空的 (用 --status 看)')
         return 0
         return 0
-    OFF.mkdir(parents=True, exist_ok=True)
-    conflicts = [dst for _, dst in items if dst.exists()]
-    if conflicts:
-        if not archive:
-            print(f'[X] 暂存区里已经有同名产物 ({len(conflicts)} 个, 例: {P.rel(conflicts[0])}) —— 直接挪会嵌套成'
-                  f' <目录>/<目录>/, 之后 --on 会把两代混在一起。\n'
-                  f'    二选一: ① 加 --archive-old 把上一代存档后重跑 (推荐);\n'
-                  f'           ② 手工把整个 {P.rel(OFF)} 改名后再跑 (例: Move-Item _products_off _products_off_旧)。')
-            return 1
-        archive_old()
-    man = []
-    # 随包基线备份藏在 windscada 里面: 先摘出来, 它不是产物 (rebuild_from_raw.py --verify 要用)
-    baseline = P.store() / KEEP[0]
-    baseline_kept = False
-    kept_at = OFF / '_baseline_kept' / KEEP[0]
-    if baseline.exists():
-        if not kept_at.exists():
-            kept_at.parent.mkdir(parents=True, exist_ok=True)
-            n, mb = count(baseline)
-            shutil.move(str(baseline), str(kept_at))
-            print(f'  保留 {P.rel(baseline):34s} {n:5d} 件 {mb:8.1f} MB  (随包基线, 不算产物)')
-    kept_at = OFF / '_baseline_kept' / KEEP[0]
-    baseline_kept = kept_at.exists()          # ★必须在 move **之后**判 (2026-09-11 曾写成之前 → 恒 False, --on 就不还原基线)
-    for src, dst in items:
-        if not src.exists():
-            continue
+    total_n = total_mb = 0
+    rows = []
+    for src in items:
         n, mb = count(src)
         n, mb = count(src)
-        dst.parent.mkdir(parents=True, exist_ok=True)
-        shutil.move(str(src), str(dst))
-        man.append(dict(rel=P.rel(src), to=P.rel(dst), files=n, mb=round(mb, 1)))
-        print(f'  挪走 {P.rel(src):34s} {n:5d} 件 {mb:8.1f} MB')
-    MANIFEST.write_text(json.dumps(dict(at=time.strftime('%Y-%m-%d %H:%M:%S'),
-                                        baseline_kept=baseline_kept, items=man),
-                                   ensure_ascii=False, indent=1), encoding='utf-8')
-    print(f'\n共 {len(man)} 项已挪到 {P.rel(OFF)} (清单: {P.rel(MANIFEST)})')
-    print(f'还原: python scripts/products_state.py --on')
-    return 0
-
-
-def do_on() -> int:
-    if not MANIFEST.exists():
-        print('没有清单 (没挪过, 或清单已删)')
-        return 1
-    man = json.loads(MANIFEST.read_text(encoding='utf-8'))
-    n = 0
-    for it in man['items']:
-        src, dst = P.resolve(it['to']), P.resolve(it['rel'])
-        if not src.exists():
-            print(f'  [!] 备份里没有 {it["to"]}, 跳过')
-            continue
-        dst.parent.mkdir(parents=True, exist_ok=True)
-        if dst.exists():
-            shutil.rmtree(dst) if dst.is_dir() else dst.unlink()
-        shutil.move(str(src), str(dst))
-        n += 1
-        print(f'  还原 {it["rel"]:34s} {it["files"]:5d} 件')
-    if man.get('baseline_kept') or (OFF / '_baseline_kept' / KEEP[0]).exists():
-        b = OFF / '_baseline_kept' / KEEP[0]        # 不看清单标记, 直接看备份在不在 (更稳)
-        if b.exists():
-            target = P.store() / KEEP[0]
-            target.parent.mkdir(parents=True, exist_ok=True)
-            if target.exists():
-                shutil.rmtree(target)
-            shutil.move(str(b), str(target))
-            print(f'  还原 {P.rel(target):34s} (随包基线)')
-    if OFF.exists() and not any(OFF.rglob('*')):
-        shutil.rmtree(OFF, ignore_errors=True)
-    print(f'\n共还原 {n} 项')
+        rows.append((P.rel(src), n, mb))
+        total_n += n
+        total_mb += mb
+    print('即将**删除**以下产物 (按用户令: 不留备份, 删除后不可恢复):')
+    for rel, n, mb in rows:
+        print(f'  {rel:34s} {n:5d} 件 {mb:8.1f} MB')
+    print(f'  合计 {len(rows)} 项 / {total_n} 件 / {total_mb:.1f} MB')
+    print('  恢复办法: python scripts/products_restore_missing.py --stash <交付包.zip> (从交付包补缺失件)')
+    if not assume_yes:
+        print('\n[X] 未执行: 这是不可恢复操作, 请显式加 --yes (运维控制台里点按钮时已由前端确认过一次)')
+        return 2
+    for src in items:
+        shutil.rmtree(src, ignore_errors=True) if src.is_dir() else src.unlink(missing_ok=True)
+    if LEGACY_OFF.exists():
+        print(f'  [!] 检测到旧版暂存区 {P.rel(LEGACY_OFF)} —— 按"不留备份"的口径, 本次**未删除**它; '
+              f'确认不需要后请手工删掉 (它是旧设计留下的产物备份)')
+    print(f'\n已删除 {len(rows)} 项 / {total_n} 件 / {total_mb:.1f} MB (无备份)')
     return 0
     return 0
 
 
 
 
 def do_status() -> int:
 def do_status() -> int:
-    items = plan()
+    items = [p for p in plan() if p.exists()]
     if not items:
     if not items:
-        print('当前: 产物**已全部挪走** (页面应呈空状态)')
-        if MANIFEST.exists():
-            man = json.loads(MANIFEST.read_text(encoding='utf-8'))
-            print(f'  挪走时间 {man["at"]}, 共 {len(man["items"])} 项')
-            for it in man['items']:
-                print(f'    {it["rel"]:34s} {it["files"]:5d} 件 {it["mb"]:8.1f} MB')
+        print('当前: 产物**已清空** (页面应呈空状态)')
+        print('  恢复: python scripts/products_restore_missing.py --stash <交付包.zip>')
         return 0
         return 0
     print(f'当前: 产物**在位** ({len(items)} 项)')
     print(f'当前: 产物**在位** ({len(items)} 项)')
-    for src, _ in items:
+    for src in items:
         n, mb = count(src)
         n, mb = count(src)
         print(f'    {P.rel(src):34s} {n:5d} 件 {mb:8.1f} MB')
         print(f'    {P.rel(src):34s} {n:5d} 件 {mb:8.1f} MB')
+    print('  说明: 按 2026-09-16 用户令, 清除产物不留备份 (没有 _products_off/ 可还原)')
     return 0
     return 0
 
 
 
 
 if __name__ == '__main__':
 if __name__ == '__main__':
     ap = argparse.ArgumentParser()
     ap = argparse.ArgumentParser()
     g = ap.add_mutually_exclusive_group(required=True)
     g = ap.add_mutually_exclusive_group(required=True)
-    g.add_argument('--off', action='store_true', help='挪走全部产物')
-    g.add_argument('--on', action='store_true', help='全部挪回')
+    g.add_argument('--off', action='store_true', help='清掉全部产物 (**直接删除, 不留备份**)')
     g.add_argument('--status', action='store_true', help='看当前状态')
     g.add_argument('--status', action='store_true', help='看当前状态')
-    ap.add_argument('--archive-old', action='store_true',
-                    help='--off 前把上一代暂存区存档到 _products_off_prev_<时间戳>/ (暂存区已有同名产物时必须加)')
+    g.add_argument('--on', action='store_true', help='[已移除] 旧版的"挪回" —— 见文件头说明')
+    ap.add_argument('--yes', action='store_true', help='确认执行不可恢复的清除 (--off 必带)')
     a = ap.parse_args()
     a = ap.parse_args()
-    sys.exit(do_off(archive=a.archive_old) if a.off else do_on() if a.on else do_status())
+    if a.on:
+        print('[X] --on 已移除: 按用户令"清除产物不留备份", 没有备份可挪回。\n'
+              '    要从交付包补回缺失的随包件: python scripts/products_restore_missing.py --stash <交付包.zip>\n'
+              '    要重新算出数据面产物: python scripts/rebuild_all.py (SCADA 侧加 --scada)')
+        sys.exit(3)
+    sys.exit(do_off(assume_yes=a.yes) if a.off else do_status())

+ 33 - 2
scripts/rebuild_all.py

@@ -53,8 +53,35 @@ def build_plan(a) -> list:
     plan.append(step_cmd('② 三门台账', [PY, 'scripts/rebuild_from_raw.py']))
     plan.append(step_cmd('② 三门台账', [PY, 'scripts/rebuild_from_raw.py']))
     if not a.skip_scada:
     if not a.skip_scada:
         plan.append(step_cmd('③ SCADA 侧 10 个构建器 (约 15 分钟)', [PY, 'scripts/rebuild_from_raw.py', '--scada']))
         plan.append(step_cmd('③ SCADA 侧 10 个构建器 (约 15 分钟)', [PY, 'scripts/rebuild_from_raw.py', '--scada']))
-    plan.append(step_cmd('④ 月度派生件', [PY, 'scripts/windscada_monthly_build.py']))
-    plan.append(step_cmd('⑤ 补齐"包内没有生成端"的产物', [PY, 'scripts/products_restore_missing.py']))
+    # ④ 月度派生件: `windscada_monthly_build.py` 会拿随包基线**逐值比对**, 没有基线时它返回 5
+    #    ("无法比对 —— 这不是通过", 这是它的诚实口径, 不要改它的退出码)。
+    #    ★但"重算链"与"等价验收"是两件事: 基线只是**验收标准答案**, 缺它不该让整条链停在第④步
+    #    (2026-09-16 实逮: 基线目录被清理后, 从页面点"执行重算" 387 s 后 rc=1 断在这里, ⑦⑧ 全不执行)。
+    #    故这里容忍 5 并写清"跳过了什么"; 要真做验收, 先把随包件恢复成暂存区/基线(见 products_state.py)。
+    plan.append(step_cmd('④ 月度派生件', [PY, 'scripts/windscada_monthly_build.py'],
+                         tolerate=(5,),
+                         note='rc=5 = 找不到随包基线可比对 ⇒ **跳过等价验收**(不是通过)。'
+                              '要恢复验收能力: 把随包产物放回 _products_off*/<场>/ 或 '
+                              'outputs/<场>/windscada/_pre_rebuild_20260911/'))
+    if not a.skip_vib:
+        # 振动侧 (2026-09-12 用户令"振动数据参与运行、重算"): 原始导出的 Brande TCM JSON
+        # (data/raw/<场站>/windcms/**/*_decode.json) → 窗索引 + 谱库。消费者 (src/windcms/data.py)
+        # 自动发现新窗, 所以这一步之后 CMS 的谱图/标量/页面就带上了新数据。
+        # ★**不跑** `windcms.py report` —— 本包缺六层链的 model_run/fusion 产物, 重生成会用残缺输入
+        #   覆盖随包快照 (report.md 16,857 B → 467 B, overview.html −15%; 实测见 vib_raw_build.py 文件头)。
+        #   六层链补齐后应把该步改为带 --with-report。
+        # 现场没有振动原始件时本步**空跑退出 0** (不是每台机器都有振动导出, 不该报失败)。
+        plan.append(step_cmd('④b 振动侧摄入 (原始导出 → 窗索引/谱)',
+                             [PY, 'scripts/vib_raw_build.py'] + (['--with-report'] if a.vib_report else []),
+                             note='没有振动原始件时空跑属正常; 若解析出错会以非零退出 (不静默)'))
+    # ⑤ 补齐"包内没有生成端"的产物: 源件是**暂存区**(清除产物时挪出来的那份随包件)。
+    #    暂存区不存在时该步返回 6 (不是崩溃) —— 首次全量重算/清完又还原过的机器上都没有暂存区,
+    #    若不容忍就会把后面的 ⑦本体链 与 ⑧审计 一起打断 ("从零重算"整条链跑不完; 2026-09-16 实逮)。
+    plan.append(step_cmd('⑤ 补齐"包内没有生成端"的产物', [PY, 'scripts/products_restore_missing.py'],
+                         tolerate=(6,),
+                         note='rc=6 = 没有随包件来源, 该步跳过。★2026-09-16 用户令"清除产物不留备份"之后, '
+                              '随包件不再保留 _products_off/ 暂存区; 要补齐请显式给交付包: '
+                              'scripts/products_restore_missing.py --stash <交付包.zip> (或在 --src 的现场包场景下先补齐)'))
     if not a.no_restart:
     if not a.no_restart:
         # 只重启**组件服务**, 网关留着 —— 两个理由:
         # 只重启**组件服务**, 网关留着 —— 两个理由:
         #  ① 控制台页面是网关提供的, 连它一起停的话, 从页面点的重算会把自己的界面(甚至自己)弄没;
         #  ① 控制台页面是网关提供的, 连它一起停的话, 从页面点的重算会把自己的界面(甚至自己)弄没;
@@ -84,6 +111,10 @@ def main() -> int:
     ap = argparse.ArgumentParser()
     ap = argparse.ArgumentParser()
     ap.add_argument('--src', default=None, help='现场数据包目录 (给了就先跑 place_raw_data --scope full)')
     ap.add_argument('--src', default=None, help='现场数据包目录 (给了就先跑 place_raw_data --scope full)')
     ap.add_argument('--skip-scada', action='store_true', help='跳过 SCADA 侧 10 个构建器 (没换 10min 数据时用)')
     ap.add_argument('--skip-scada', action='store_true', help='跳过 SCADA 侧 10 个构建器 (没换 10min 数据时用)')
+    ap.add_argument('--skip-vib', action='store_true', help='跳过振动侧摄入 (没换 CMS 原始导出时用)')
+    ap.add_argument('--vib-report', action='store_true',
+                    help='振动侧额外重生成 CMS 报告/知识库 —— ★默认关: 本包缺六层链的 model_run/fusion '
+                         '产物, 重生成会掉内容 (report.md −97%), 详见 scripts/vib_raw_build.py 文件头')
     ap.add_argument('--no-restart', action='store_true', help='不自动重启服务')
     ap.add_argument('--no-restart', action='store_true', help='不自动重启服务')
     ap.add_argument('--with-verify', action='store_true', help='末尾加等价验收')
     ap.add_argument('--with-verify', action='store_true', help='末尾加等价验收')
     ap.add_argument('--dry-run', action='store_true', help='只打印计划')
     ap.add_argument('--dry-run', action='store_true', help='只打印计划')

+ 383 - 0
scripts/rudong_tcm_index.py

@@ -0,0 +1,383 @@
+#!/usr/bin/env python3
+# -*- coding: utf-8 -*-
+"""TCM 解码 JSON → 窗索引 index.parquet (2026-09-12 补; `src/windcms/pipeline.py` 的 tcm_decoded_json 第一步).
+
+## 它是谁、为什么在这里
+
+`src/windcms/pipeline.py::ingest()` 对"含 `*_decode.json` 的目录"跑两步:
+    ① rudong_tcm_index.py   --root <dir> --out <m5>/windows/<w>/index.parquet
+    ② rudong_tcm_spectra.py --root <dir> --out <m5>/windows/<w>/spectra
+这两个脚本在 v0.2.0 里**缺失**(振动线分支 claude/vibration-data-diagnosis-32b69e 上的东西没随包),
+于是"振动数据参与重算"这条路是断的: 只有产物 `tcm_index.parquet`(w0127 汇总) 在, 没有生成端。
+2026-09-12 用户补入 CMS 原始导出 (`CMS_RuDong_CGN_202603-04.zip`, 25,679 件 Brande TCM 导出) 后,
+按消费端契约把这两步补齐。
+
+## 输入
+
+Brande TCM Enterprise 导出 (西门子机组自带 M-system), 每个文件是一次 API 响应:
+    {"expiry", "buildId", "method", "controller", "serial", "requestInfo",
+     "body": {"body": {"<ISO时间戳>": [{"Record": {...}}, ...]}}}
+`Record` 下才有真数据: `Site` / `Location` / `ConfigurationSettings` / `Sensor` / `Measurement`
+(后者再带 `Conditions` 与 `DataSets`)。文件布局实测两种, 本脚本都认:
+    measurement/<YYYY>/<MM>/<WTGxx>/<WTGxx>_<uuid>_decode.json     (本包; --root 指到 measurement 或更上层都行)
+    <WTGxx>/<YYYY>/<MM>/<WTGxx>_<uuid>_decode.json                 (w0127 那批的布局)
+
+## 输出契约 (消费者 = src/windcms/data.py)
+
+`_read_index()` 只挑 `turbine, sensor_name, meas_name, trigger_time, rpm, condition_key, alarm_type,
+ds_size, scalar_value, overload`; `load_alarms` 读 `turbine, alarm_type`; 其余列是谱/工况/报警阈值元数据,
+供风电场页面与后续六层链用。**列名与列义必须与包内 `outputs/<场>/m5_cms_tcm/tcm_index.parquet`
+(330,308 行 × 54 列, 2026-01-27~02-03 窗) 逐列对齐** —— 那份是同一摄取逻辑的产物, 是唯一的格式基准;
+本脚本的 54 列与它同名同义 (见 COLSPEC), 因此新旧窗可以拼接进入同一分析集。
+
+## 并行与可重入
+
+150 GB / 2.5 万件的单线程 json 解析要 ~1 小时, 机器 14 核 → 按机组切片并行 (`--jobs`, 默认 8)。
+每个子进程写自己的分片 parquet (`<out>.parts/p<i>.parquet`), 父进程合并后删除分片。
+单文件/单记录异常不中断整窗: 记进 `parse_error` 列 (响亮留痕), 文件级失败计数并在末尾汇总报告。
+
+用法:
+    python scripts/rudong_tcm_index.py --root data/raw/如东/windcms/CMS_RuDong_CGN_202603-04/measurement \
+        --out outputs/rudong/m5_cms_tcm/windows/w0317/index.parquet
+    python scripts/rudong_tcm_index.py --root <dir> --out <index.parquet> --jobs 12 --turbines WTG01,WTG02
+    python scripts/rudong_tcm_index.py --root <dir> --out <index.parquet> --limit 200 --dry-run
+"""
+from __future__ import annotations
+
+import argparse
+import json
+import math
+import os
+import pathlib
+import shutil
+import sys
+import time
+import zipfile
+
+ROOT = pathlib.Path(__file__).resolve().parents[1]
+sys.path.insert(0, str(ROOT))
+from src.console import soft  # noqa: E402
+
+soft()
+
+# 厂内 JSON 层 → 索引列 (缺键给 None; 表里写的就是 JSON 里的键名, 便于逐列对拍)
+REC_MAP = {
+    'serial': ('Location', 'SerialNumber'),
+    'turbine': ('Location', 'LocationName'),
+    'config_name': ('ConfigurationSettings', 'ConfigurationName'),
+    'config_time': ('ConfigurationSettings', 'ConfigurationTime'),
+    'turbine_type': ('ConfigurationSettings', 'TurbineType'),
+    'recording_time_s': ('ConfigurationSettings', 'RecordingTime_s'),
+    'monitoring_cycle_min': ('ConfigurationSettings', 'MonitoringCycle-Minutes'),
+    'sensor_name': ('Sensor', 'SensorName'),
+    'sensor_type': ('Sensor', 'SensorType'),
+    'sensor_addr': ('Sensor', 'SensorAddress'),
+    'sensor_sn': ('Sensor', 'Sensor_Sn'),
+    'sensitivity_mvpg': ('Sensor', 'Sensitivity_mVpG'),
+    'sensor_unit': ('Sensor', 'SensorUnit'),
+    'hw_serial': ('Sensor', 'HwSerial'),
+    'meas_type': ('Measurement', 'MeasurementType'),
+    'meas_name': ('Measurement', 'MeasurementName'),
+    'meas_source': ('Measurement', 'Source'),
+    'category': ('Measurement', 'Category'),
+    'meas_key': ('Measurement', 'MeasurementKey'),
+    'trigger_time': ('Measurement', 'Trigger_Time'),
+    'trigger_ms': ('Measurement', 'Trigger_Time_ms'),
+    'duration_s': ('Measurement', 'Measurement_Duration-s'),
+    'rpm': ('Measurement', 'RPM'),
+    'nominal_freq_hz': ('Measurement', 'NominalFrequency-Hz'),
+    'min_freq_hz': ('Measurement', 'MinimumFrequency-Hz'),
+    'max_freq_hz': ('Measurement', 'MaximumFrequency-Hz'),
+    'lower_freq_hz': ('Measurement', 'LowerFrequency-Hz'),
+    'upper_freq_hz': ('Measurement', 'UpperFrequency-Hz'),
+    'overload': ('Measurement', 'Overload'),
+    'n_averages': ('Measurement', 'NumberOfAverages'),
+    'bandwidth_hz': ('Measurement', 'Bandwidth-Hz'),
+    'lines': ('Measurement', 'Lines'),
+    'integration': ('Measurement', 'Integration'),
+    'condition_name': ('Measurement', 'Conditions', 'ConditionsName'),
+    'condition_key': ('Measurement', 'ConditionKey'),
+    'alarm_type': ('Measurement', 'AlarmType'),
+    'red_alarm': ('Measurement', 'RedAlarm'),
+    'yellow_alarm': ('Measurement', 'YellowAlarm'),
+    'blue_alarm': ('Measurement', 'BlueAlarm'),
+    'red_hys': ('Measurement', 'RedHysAlarm'),
+    'yellow_hys': ('Measurement', 'YellowHysAlarm'),
+    'trend_hys': ('Measurement', 'TrendHysAlarm'),
+    'fault_freq': ('Measurement', 'FaultFreq'),
+}
+# DataSets 下的列。★层级别踩错 (2026-09-12 对拍逮到): `Size`/`Dimension`/`Values` 在
+# `Measurement.DataSets.DataSet` 里, 而 X/Y 轴四件在**上一层** `Measurement.DataSets` 里 ——
+# 一开始全按 DataSet 取, 结果 x_offset/x_delta/x_unit/y_unit 四列整列为空,
+# 而 "x_delta == 带宽/lines" 的内部一致性检查比例是 0.000 (本该 1.000), 就是这里露的马脚。
+DS_MAP = {
+    'ds_dim': 'Dimension',
+    'ds_size': 'Size',
+}
+AXIS_MAP = {
+    'x_offset': 'X-axisOffset',
+    'x_delta': 'X-axisDelta',
+    'x_unit': 'X-axisUnit',
+    'y_unit': 'Y-axisUnit',
+}
+TXT_COLS = ('file', 'serial', 'turbine', 'ts_key', 'config_name', 'config_time', 'turbine_type',
+            'sensor_name', 'sensor_type', 'sensor_addr', 'sensor_sn', 'sensor_unit', 'hw_serial',
+            'meas_type', 'meas_name', 'meas_source', 'category', 'meas_key', 'trigger_time',
+            'integration', 'condition_name', 'condition_key', 'alarm_type', 'red_hys', 'yellow_hys',
+            'trend_hys', 'x_unit', 'y_unit', 'parse_error')
+NUM_COLS = ('rec_i', 'recording_time_s', 'monitoring_cycle_min', 'sensitivity_mvpg', 'trigger_ms',
+            'duration_s', 'rpm', 'nominal_freq_hz', 'min_freq_hz', 'max_freq_hz', 'lower_freq_hz',
+            'upper_freq_hz', 'overload', 'n_averages', 'bandwidth_hz', 'lines', 'fault_freq',
+            'red_alarm', 'yellow_alarm', 'blue_alarm', 'ds_dim', 'ds_size', 'x_offset', 'x_delta',
+            'scalar_value')
+# 54 列的**列序**逐字抄自包内 `outputs/<场>/m5_cms_tcm/tcm_index.parquet` (振动线 v0.2.0 产物,
+# 330,308 行) —— 消费者按列名取数, 但"列序也要一致"是为了让新旧窗在人工对拍/并排打印时能直接比。
+# (文本列与数值列是交错的, 所以不能用 TXT_COLS + NUM_COLS 拼。)
+COLSPEC = ['file', 'serial', 'turbine', 'ts_key', 'rec_i', 'config_name', 'config_time', 'turbine_type',
+           'recording_time_s', 'monitoring_cycle_min', 'sensor_name', 'sensor_type', 'sensor_addr',
+           'sensor_sn', 'sensitivity_mvpg', 'sensor_unit', 'hw_serial', 'meas_type', 'meas_name',
+           'meas_source', 'category', 'meas_key', 'trigger_time', 'trigger_ms', 'duration_s', 'rpm',
+           'nominal_freq_hz', 'min_freq_hz', 'max_freq_hz', 'lower_freq_hz', 'upper_freq_hz', 'overload',
+           'n_averages', 'bandwidth_hz', 'lines', 'integration', 'condition_name', 'condition_key',
+           'alarm_type', 'red_alarm', 'yellow_alarm', 'blue_alarm', 'red_hys', 'yellow_hys', 'trend_hys',
+           'fault_freq', 'ds_dim', 'ds_size', 'x_offset', 'x_delta', 'x_unit', 'y_unit', 'scalar_value',
+           'parse_error']
+
+
+def _dig(d, path):
+    """按键路径取 (任一环缺失返回 None, 不抛)。"""
+    for k in path:
+        if not isinstance(d, dict):
+            return None
+        d = d.get(k)
+    return d
+
+
+def _f(v):
+    """→ float; 空/非数 → NaN (索引列必须能进 pandas float 列, 字符串混进去会让整列变 object)。"""
+    if v is None or v == '':
+        return math.nan
+    try:
+        return float(v)
+    except (TypeError, ValueError):
+        return math.nan
+
+
+def _s(v):
+    return None if v is None else str(v)
+
+
+def _datasets(m):
+    """Measurement.DataSets → (DataSets 层, DataSet 记录)。DataSet 可能是 dict 也可能是 list。"""
+    dss = m.get('DataSets') or {}
+    if not isinstance(dss, dict):
+        return {}, {}
+    ds = dss.get('DataSet')
+    if isinstance(ds, list):
+        ds = ds[0] if ds else {}
+    return dss, (ds if isinstance(ds, dict) else {})
+
+
+def rows_of_doc(doc, relpath, ts_key, recs, rows):
+    """一个 (文件, 时间戳) 下的一批记录 → 追加进 rows。"""
+    for rec_i, item in enumerate(recs):
+        rec = item.get('Record') if isinstance(item, dict) else None
+        if not isinstance(rec, dict):
+            rows.append(dict(file=relpath, ts_key=ts_key, rec_i=rec_i,
+                             parse_error='记录非 Record 结构: ' + str(type(item).__name__)))
+            continue
+        try:
+            m = rec.get('Measurement') or {}
+            dss, ds = _datasets(m)
+            vals = ds.get('Values')
+            size = _f(ds.get('Size'))
+            row = dict(file=relpath, ts_key=ts_key, rec_i=rec_i)
+            for col, path in REC_MAP.items():
+                row[col] = _dig(rec, path)
+            for col, key in DS_MAP.items():
+                row[col] = ds.get(key)
+            for col, key in AXIS_MAP.items():          # X/Y 轴在 DataSets 层, 不在 DataSet 里
+                row[col] = dss.get(key)
+            # 标量 (Size==1): Values 字符串本身就是标量值; 谱/波形 (Size>1) 的 Values 是长串, 值不进索引
+            row['scalar_value'] = _f(vals) if (size == 1 and isinstance(vals, (str, int, float))) else None
+            row['parse_error'] = None
+            rows.append(row)
+        except Exception as exc:                      # 单记录异常不许断整窗
+            rows.append(dict(file=relpath, ts_key=ts_key, rec_i=rec_i,
+                             parse_error=f'{type(exc).__name__}: {exc}'))
+
+
+def rows_of_file(raw: bytes, relpath: str, rows: list):
+    doc = json.loads(raw.decode('utf-8', 'replace'))
+    body = ((doc.get('body') or {}).get('body')) or {}
+    if not isinstance(body, dict):
+        rows.append(dict(file=relpath, parse_error='body.body 非 dict'))
+        return
+    for ts_key, recs in body.items():
+        if isinstance(recs, list):
+            rows_of_doc(doc, relpath, ts_key, recs, rows)
+        else:
+            rows.append(dict(file=relpath, ts_key=ts_key, parse_error='记录集非 list'))
+
+
+def relpath_of(name: str) -> str:
+    """包内成员名 → `file` 列。去掉 `measurement/` 这一层 (包本身的组织层, 不是数据层)。"""
+    n = name.replace('\\', '/')
+    if n.startswith('measurement/'):
+        n = n[len('measurement/'):]
+    return n
+
+
+def discover(root: pathlib.Path, turbines, limit):
+    """→ [(成员名/相对路径, 完整路径|None, zip 路径|None)]。目录与 zip 都支持。"""
+    out = []
+    if root.is_file() and root.suffix.lower() == '.zip':
+        with zipfile.ZipFile(root) as zf:
+            for i in zf.infolist():
+                if i.is_dir() or not i.filename.endswith('_decode.json'):
+                    continue
+                out.append((i.filename, None, str(root)))
+    else:
+        for p in sorted(root.rglob('*_decode.json')):
+            out.append((str(p.relative_to(root)).replace('\\', '/'), str(p), None))
+    if turbines:
+        want = {t.upper() for t in turbines}
+        out = [x for x in out if any(t in x[0].upper() for t in want)]
+    if limit:
+        out = out[:limit]
+    return out
+
+
+def flatten(items, rows):
+    """逐文件解析 (目录: 直接读; zip: 按需打开一次)。"""
+    cur_zip = None
+    cur_path = None
+    try:
+        for name, full, zpath in items:
+            rel = relpath_of(name)
+            try:
+                if zpath:
+                    if cur_zip is None or cur_path != zpath:
+                        if cur_zip:
+                            cur_zip.close()
+                        cur_zip = zipfile.ZipFile(zpath)
+                        cur_path = zpath
+                    raw = cur_zip.read(name)
+                else:
+                    raw = pathlib.Path(full).read_bytes()
+                rows_of_file(raw, rel, rows)
+            except Exception as exc:
+                rows.append(dict(file=rel, parse_error=f'文件级失败 {type(exc).__name__}: {exc}'))
+    finally:
+        if cur_zip:
+            cur_zip.close()
+
+
+def worker(payload):
+    """子进程: 解析分片 → 写分片 parquet → 回 (分片路径, 行数, 文件数, 失败数)。"""
+    part_path, items = payload
+    import pandas as pd
+    rows = []
+    flatten(items, rows)
+    df = pd.DataFrame(rows)
+    for c in COLSPEC:
+        if c not in df.columns:
+            df[c] = None
+    df = df[list(COLSPEC)]
+    for c in NUM_COLS:
+        # 与 shipped 一致的 float64 (整列同型; 有 NaN 的列本就会升为 float, 没有 NaN 的列不升则出现 int64)
+        df[c] = pd.to_numeric(df[c], errors='coerce').astype('float64')
+    for c in TXT_COLS:
+        df[c] = df[c].astype('str')          # pandas 3 的 'str' dtype (shipped 同款)
+    bad = int(df['parse_error'].notna().sum()) if 'parse_error' in df else 0
+    df.to_parquet(part_path, index=False)
+    return part_path, len(df), len(items), bad
+
+
+def main() -> int:
+    ap = argparse.ArgumentParser()
+    ap.add_argument('--root', required=True, help='含 *_decode.json 的目录, 或导出 zip')
+    ap.add_argument('--out', required=True, help='index.parquet 输出路径')
+    ap.add_argument('--jobs', type=int, default=min(8, (os.cpu_count() or 4)),
+                    help='并行进程数 (按机组切片; 默认 %(default)s)')
+    ap.add_argument('--turbines', default=None, help='只处理这些机组 (逗号分隔, 如 WTG01,WTG02)')
+    ap.add_argument('--limit', type=int, default=0, help='只处理前 N 个文件 (冒烟; 默认 0=全部)')
+    ap.add_argument('--dry-run', action='store_true', help='只报将处理多少文件, 不解析')
+    a = ap.parse_args()
+
+    root = pathlib.Path(a.root)
+    if not root.exists():
+        raise SystemExit(f'输入不存在: {root}')
+    out = pathlib.Path(a.out)
+    turbines = [x.strip() for x in a.turbines.split(',')] if a.turbines else None
+    t0 = time.time()
+    items = discover(root, turbines, a.limit)
+    if not items:
+        raise SystemExit(f'没找到 *_decode.json: {root}')
+    size = sum((pathlib.Path(f).stat().st_size if f else 0) for _, f, _ in items)
+    print(f'输入: {root}')
+    print(f'发现 {len(items)} 个 *_decode.json  目录内合计 {size / 1073741824:.2f} GB(仅目录模式可量)')
+    if a.dry_run:
+        print('(dry-run, 未解析)')
+        return 0
+
+    import pandas as pd
+    import concurrent.futures as cf
+
+    # 切片: 优先按机组 (一个子进程只碰自己那几台 → 负载均匀且写入互不干扰)
+    groups = {}
+    for it in items:
+        key = next((seg for seg in it[0].replace('\\', '/').split('/') if seg.upper().startswith('WTG')), '_other')
+        groups.setdefault(key, []).append(it)
+    chunks = []
+    n = max(1, a.jobs)
+    per = math.ceil(len(items) / n)
+    cur = []
+    for k in sorted(groups):
+        cur.extend(groups[k])
+        if len(cur) >= per:
+            chunks.append(cur)
+            cur = []
+    if cur:
+        chunks.append(cur)
+
+    parts_dir = out.parent / (out.stem + '.parts')
+    if parts_dir.exists():
+        shutil.rmtree(parts_dir)
+    parts_dir.mkdir(parents=True, exist_ok=True)
+    payloads = [(str(parts_dir / f'p{i:02d}.parquet'), ch) for i, ch in enumerate(chunks)]
+    print(f'并行 {min(n, len(payloads))} 进程 × {len(payloads)} 片  →  {out}')
+
+    done_files = 0
+    total_rows = 0
+    total_bad = 0
+    part_files = []
+    with cf.ProcessPoolExecutor(max_workers=min(n, len(payloads))) as ex:
+        for part, nrow, nfile, nbad in ex.map(worker, payloads):
+            part_files.append(part)
+            done_files += nfile
+            total_rows += nrow
+            total_bad += nbad
+            print(f'  [{done_files}/{len(items)} 文件] 累计 {total_rows} 行, 异常 {total_bad}  '
+                  f'({time.time() - t0:.0f}s, {part})', flush=True)
+
+    df = pd.concat([pd.read_parquet(p) for p in sorted(part_files)], ignore_index=True)
+    out.parent.mkdir(parents=True, exist_ok=True)
+    df.to_parquet(out, index=False)
+    shutil.rmtree(parts_dir, ignore_errors=True)
+
+    print(f'\n完成: {len(df)} 行 × {df.shape[1]} 列 → {out}')
+    print(f'  机组 {df.turbine.nunique()} 台; 传感器 {df.sensor_name.nunique()} 种; 测量名 {df.meas_name.nunique()} 种')
+    if 'trigger_time' in df:
+        tt = pd.to_datetime(df.trigger_time, errors='coerce')
+        print(f'  时间窗 {tt.min()} → {tt.max()}')
+    print(f'  标量行(Size=1) {int((df.ds_size == 1).sum())}; 谱/波形行(Size>1) {int((df.ds_size > 1).sum())}; '
+          f'解析异常 {int(df.parse_error.notna().sum())}')
+    if total_bad:
+        print(f'  ⚠ 有 {total_bad} 条记录带 parse_error (见该列), 未静默丢弃')
+    print(f'  用时 {time.time() - t0:.0f}s')
+    return 0
+
+
+if __name__ == '__main__':
+    sys.exit(main())

+ 334 - 0
scripts/rudong_tcm_spectra.py

@@ -0,0 +1,334 @@
+#!/usr/bin/env python3
+# -*- coding: utf-8 -*-
+"""TCM 解码 JSON → 谱库 (npz 分片 + spectra_meta.parquet) —— pipeline.py 的 tcm_decoded_json 第二步.
+
+## 契约 (消费者 = src/windcms/data.py::spectra_meta / spectrum)
+
+    spectra_meta.parquet 每行一条谱: turbine, sensor, meas_name, trigger_time, shard, shard_row,
+                                    x_offset, x_delta (+ rpm/y_unit/alarm_type/n_points 等附列)
+    <store>/<shard>.npz  key='values' = 二维数组 (n 条 × n 点); 第 shard_row 行是该条谱
+    取数: x = x_offset + arange(len(v)) * x_delta,  v = npz['values'][shard_row]
+
+`data.spectra_meta()` 找 meta 的两条路径 (都写, 免得换 `--out` 就失联):
+    <window>/spectra/spectra_meta.parquet   ← 单窗 (本脚本 --out 指 windows/<w>/spectra)
+    <store>/spectra_meta.parquet            ← 合并库 (--out 指 m5/spectra 时)
+
+## 输入
+
+同上一步 (Brande TCM 导出 JSON)。**谱在 `Record.Measurement.DataSets.DataSet.Values` 里, 是空格
+分隔的数值字符串** (不是数组! 实测 6401 点 ~ 60 KB 字符串/条), Size = Lines + 1, X-axisDelta =
+带宽/Lines, X-axisUnit='Hz'。标量记录 Size==1 (其 Values 就是标量值, 归索引脚本处理, 这里跳过)。
+
+## 只转需要的测量
+
+原始导出里绝大部分字节是 `Time_*` 波形 (65536/200000 点), 而页面/判据用的是 `FFT_*` 谱。
+默认 `--meas FFT_` 只转谱; 要全转用 `--meas ALL` (磁盘会显著变大: 实测单个窗的谱 npz 约 11 GB 量级)。
+
+## 落盘与并行
+
+按 (测量名, 点数) 分组装分片, 每片 `--shard-rows` 条 (默认 256) 写一个 npz (float32, 见 --dtype)。
+并行按机组切片, 子进程写自己的 `p<NN>/` 命名空间 (shard 是相对 store 的路径, 带子目录合法),
+末尾父进程合并各分片的 meta。分片原子落盘 (先 .tmp 再改名), 中断不会留半截 npz。
+
+用法:
+    python scripts/rudong_tcm_spectra.py --root <measurement 目录> --out outputs/rudong/m5_cms_tcm/windows/w0317/spectra
+    python scripts/rudong_tcm_spectra.py --root <dir> --out <store> --meas FFT_ --jobs 8 --turbines WTG09
+"""
+from __future__ import annotations
+
+import argparse
+import json
+import math
+import os
+import pathlib
+import shutil
+import sys
+import time
+import zipfile
+
+ROOT = pathlib.Path(__file__).resolve().parents[1]
+sys.path.insert(0, str(ROOT))
+from src.console import soft  # noqa: E402
+
+soft()
+
+
+def _f(v):
+    if v is None or v == '':
+        return math.nan
+    try:
+        return float(v)
+    except (TypeError, ValueError):
+        return math.nan
+
+
+def _datasets(m):
+    """Measurement.DataSets → (DataSets 层, DataSet 记录)。X/Y 轴在 DataSets 层, Size/Values 在 DataSet 里。"""
+    dss = m.get('DataSets') or {}
+    if not isinstance(dss, dict):
+        return {}, {}
+    ds = dss.get('DataSet')
+    if isinstance(ds, list):
+        ds = ds[0] if ds else {}
+    return dss, (ds if isinstance(ds, dict) else {})
+
+
+def values_to_array(vals):
+    """Values (空格分隔字符串 / list) → float 数组; 不合法返回 None。"""
+    import numpy as np
+    if isinstance(vals, list):
+        try:
+            return np.asarray(vals, dtype=float)
+        except (TypeError, ValueError):
+            return None
+    if not isinstance(vals, str):
+        return None
+    try:
+        arr = np.fromstring(vals, sep=' ')          # C 级解析, 6401 点 ~ 0.2 ms
+    except Exception:
+        return None
+    if arr.size == 0:
+        try:
+            arr = np.asarray(vals.split(), dtype=float)
+        except (TypeError, ValueError):
+            return None
+    return arr if arr.size else None
+
+
+class ShardWriter:
+    """按 (meas_name, 点数) 攒批写 npz。"""
+
+    def __init__(self, store: pathlib.Path, ns: str, shard_rows: int, dtype):
+        self.store = store
+        self.ns = ns
+        self.shard_rows = shard_rows
+        self.dtype = dtype
+        self.buf = {}          # key → list[np.ndarray]
+        self.idx = {}          # key → 该 key 已写出片数
+        self.meta = []
+        self.shards = 0
+
+    def add(self, key, turbine, sensor, meas, trigger, x_off, x_delta, extra, arr):
+        rows = self.buf.setdefault(key, [])
+        rows.append(arr)
+        row_in_shard = len(rows) - 1
+        shard = f'{self.ns}/{self._stem(key, self.idx.get(key, 0))}'
+        self.meta.append(dict(turbine=turbine, sensor=sensor, meas_name=meas, trigger_time=trigger,
+                              shard=shard, shard_row=row_in_shard, x_offset=x_off, x_delta=x_delta,
+                              n_points=int(arr.size), **extra))
+        if len(rows) >= self.shard_rows:
+            self.flush_key(key)
+
+    def _stem(self, key, i):
+        meas, npts = key
+        safe = ''.join(c if (c.isalnum() or c in '-_.') else '_' for c in meas)
+        return f'{safe}_n{npts}_{i:04d}.npz'
+
+    def flush_key(self, key):
+        import numpy as np
+        rows = self.buf.get(key)
+        if not rows:
+            return
+        i = self.idx.get(key, 0)
+        rel = f'{self.ns}/{self._stem(key, i)}'
+        dst = self.store / rel
+        dst.parent.mkdir(parents=True, exist_ok=True)
+        # 不同测量/点数的分片点数一致 (同 key 同长度), 可直接 stack
+        arr = np.stack(rows).astype(self.dtype, copy=False)
+        tmp = dst.with_suffix('.npz.tmp')
+        with open(tmp, 'wb') as fh:
+            np.savez_compressed(fh, values=arr)
+        os.replace(tmp, dst)
+        self.shards += 1
+        self.idx[key] = i + 1
+        self.buf[key] = []
+
+    def close(self):
+        for key in list(self.buf):
+            self.flush_key(key)
+        return self.meta, self.shards
+
+
+def rows_of_file(raw: bytes, meas_prefix: str, writer: ShardWriter, stats: dict):
+    doc = json.loads(raw.decode('utf-8', 'replace'))
+    body = ((doc.get('body') or {}).get('body')) or {}
+    if not isinstance(body, dict):
+        stats['bad'] += 1
+        return
+    for _ts, recs in body.items():
+        if not isinstance(recs, list):
+            continue
+        for item in recs:
+            rec = item.get('Record') if isinstance(item, dict) else None
+            if not isinstance(rec, dict):
+                stats['bad'] += 1
+                continue
+            try:
+                m = rec.get('Measurement') or {}
+                meas = str(m.get('MeasurementName') or '?')
+                if meas_prefix != 'ALL' and not meas.startswith(meas_prefix):
+                    stats['skipped'] += 1
+                    continue
+                dss, ds = _datasets(m)
+                size = _f(ds.get('Size'))
+                if not (size and size > 1):
+                    stats['skipped'] += 1          # 标量行归索引脚本
+                    continue
+                arr = values_to_array(ds.get('Values'))
+                if arr is None or arr.size <= 1:
+                    stats['bad'] += 1
+                    continue
+                loc = (rec.get('Location') or {}).get('LocationName') or '?'
+                sens = (rec.get('Sensor') or {}).get('SensorName') or '?'
+                extra = dict(rpm=_f(m.get('RPM')), y_unit=dss.get('Y-axisUnit'),
+                             x_unit=dss.get('X-axisUnit'), alarm_type=m.get('AlarmType'),
+                             condition_key=m.get('ConditionKey'), ds_size=size,
+                             meas_key=m.get('MeasurementKey'))
+                writer.add((meas, int(arr.size)), loc, sens, meas, m.get('Trigger_Time'),
+                           _f(dss.get('X-axisOffset')), _f(dss.get('X-axisDelta')), extra, arr)
+                stats['rows'] += 1
+            except Exception:
+                stats['bad'] += 1
+
+
+def worker(payload):
+    part_path, items, store, ns, shard_rows, meas, dtype = payload
+    import pandas as pd
+    writer = ShardWriter(pathlib.Path(store), ns, shard_rows, dtype)
+    stats = dict(rows=0, bad=0, skipped=0)
+    cur_zip = None
+    cur_path = None
+    try:
+        for name, full, zpath in items:
+            try:
+                if zpath:
+                    if cur_zip is None or cur_path != zpath:
+                        if cur_zip:
+                            cur_zip.close()
+                        cur_zip = zipfile.ZipFile(zpath)
+                        cur_path = zpath
+                    raw = cur_zip.read(name)
+                else:
+                    raw = pathlib.Path(full).read_bytes()
+                rows_of_file(raw, meas, writer, stats)
+            except Exception:
+                stats['bad'] += 1
+    finally:
+        if cur_zip:
+            cur_zip.close()
+    meta, shards = writer.close()
+    pd.DataFrame(meta).to_parquet(part_path, index=False)
+    row = dict(part=part_path, rows=stats['rows'], bad=stats['bad'], skipped=stats['skipped'],
+               shards=shards, files=len(items))
+    return row
+
+
+def discover(root: pathlib.Path, turbines, limit):
+    out = []
+    if root.is_file() and root.suffix.lower() == '.zip':
+        with zipfile.ZipFile(root) as zf:
+            for i in zf.infolist():
+                if not i.is_dir() and i.filename.endswith('_decode.json'):
+                    out.append((i.filename, None, str(root)))
+    else:
+        for p in sorted(root.rglob('*_decode.json')):
+            out.append((str(p.relative_to(root)).replace('\\', '/'), str(p), None))
+    if turbines:
+        want = {t.upper() for t in turbines}
+        out = [x for x in out if any(t in x[0].upper() for t in want)]
+    return out[:limit] if limit else out
+
+
+def main() -> int:
+    ap = argparse.ArgumentParser()
+    ap.add_argument('--root', required=True, help='含 *_decode.json 的目录, 或导出 zip')
+    ap.add_argument('--out', required=True, help='谱库目录 (通常 <window>/spectra)')
+    ap.add_argument('--meas', default='FFT_', help="只转这些测量 (前缀匹配); 'ALL' = 全转 (默认 %(default)s)")
+    ap.add_argument('--jobs', type=int, default=min(8, (os.cpu_count() or 4)), help='并行进程数')
+    ap.add_argument('--shard-rows', type=int, default=256, help='每个 npz 装多少条谱 (默认 %(default)s)')
+    ap.add_argument('--dtype', default='float32', choices=('float32', 'float64'),
+                    help='谱点存储精度 (默认 float32; 判据标量不从这里算, 谱图显示足够)')
+    ap.add_argument('--turbines', default=None, help='只处理这些机组 (逗号分隔)')
+    ap.add_argument('--limit', type=int, default=0, help='只处理前 N 个文件 (冒烟)')
+    ap.add_argument('--dry-run', action='store_true')
+    a = ap.parse_args()
+
+    root = pathlib.Path(a.root)
+    if not root.exists():
+        raise SystemExit(f'输入不存在: {root}')
+    out = pathlib.Path(a.out)
+    turbines = [x.strip() for x in a.turbines.split(',')] if a.turbines else None
+    import numpy as np
+    dtype = np.float32 if a.dtype == 'float32' else np.float64
+
+    t0 = time.time()
+    items = discover(root, turbines, a.limit)
+    if not items:
+        raise SystemExit(f'没找到 *_decode.json: {root}')
+    print(f'输入: {root}\n发现 {len(items)} 个 *_decode.json; 测量过滤 {a.meas}; 谱库 → {out}')
+    if a.dry_run:
+        print('(dry-run, 未解析)')
+        return 0
+
+    import pandas as pd
+    import concurrent.futures as cf
+
+    groups = {}
+    for it in items:
+        key = next((seg for seg in it[0].replace('\\', '/').split('/') if seg.upper().startswith('WTG')), '_other')
+        groups.setdefault(key, []).append(it)
+    n = max(1, a.jobs)
+    per = math.ceil(len(items) / n)
+    chunks, cur = [], []
+    for k in sorted(groups):
+        cur.extend(groups[k])
+        if len(cur) >= per:
+            chunks.append(cur)
+            cur = []
+    if cur:
+        chunks.append(cur)
+
+    parts_dir = out.parent / (out.name + '.meta.parts')
+    if parts_dir.exists():
+        shutil.rmtree(parts_dir)
+    parts_dir.mkdir(parents=True, exist_ok=True)
+    payloads = [(str(parts_dir / f'p{i:02d}.parquet'), ch, str(out), f'p{i:02d}', a.shard_rows, a.meas, dtype)
+                for i, ch in enumerate(chunks)]
+    print(f'并行 {min(n, len(payloads))} 进程 × {len(payloads)} 片')
+
+    rows_total = bad = skipped = shards = files_done = 0
+    parts = []
+    with cf.ProcessPoolExecutor(max_workers=min(n, len(payloads))) as ex:
+        for r in ex.map(worker, payloads):
+            parts.append(r['part'])
+            rows_total += r['rows']
+            bad += r['bad']
+            skipped += r['skipped']
+            shards += r['shards']
+            files_done += r['files']
+            print(f"  [{files_done}/{len(items)} 文件] 谱 {rows_total} 条, 分片 {shards}, "
+                  f"跳过(非目标测量/标量) {skipped}, 异常 {bad}  ({time.time() - t0:.0f}s)", flush=True)
+
+    metas = [pd.read_parquet(p) for p in sorted(parts) if pathlib.Path(p).stat().st_size > 0]
+    meta = pd.concat(metas, ignore_index=True) if metas else pd.DataFrame(
+        columns=['turbine', 'sensor', 'meas_name', 'trigger_time', 'shard', 'shard_row', 'x_offset', 'x_delta'])
+    for p in parts:
+        pathlib.Path(p).unlink(missing_ok=True)
+    parts_dir.rmdir()
+
+    out.mkdir(parents=True, exist_ok=True)
+    meta.to_parquet(out / 'spectra_meta.parquet', index=False)
+    # 单窗约定: data.spectra_meta() 对 <window>/spectra 优先读 <window>/spectra_meta.parquet
+    if out.name == 'spectra':
+        meta.to_parquet(out.parent / 'spectra_meta.parquet', index=False)
+    print(f'\n完成: {len(meta)} 条谱 → {out} (+ {out.parent / "spectra_meta.parquet" if out.name == "spectra" else "—"})')
+    print(f'  分片文件 {shards} 个; 机组 {meta.turbine.nunique() if len(meta) else 0} 台; '
+          f'测量 {meta.meas_name.nunique() if len(meta) else 0} 种')
+    if len(meta):
+        print('  逐测量条数:\n' + meta.meas_name.value_counts().head(20).to_string())
+    print(f'  用时 {time.time() - t0:.0f}s')
+    return 0
+
+
+if __name__ == '__main__':
+    sys.exit(main())

+ 271 - 0
scripts/vib_raw_build.py

@@ -0,0 +1,271 @@
+#!/usr/bin/env python3
+# -*- coding: utf-8 -*-
+"""振动侧一键摄入: data/raw/<场站>/{windcms,m5_cms_tcm} 的原始件 → 窗索引/谱库 → (可选)CMS 报告与知识库.
+
+## 为什么需要它 (2026-09-12 用户令)
+
+用户令: 「数据层里的「CMS 振动评估报告」应遵循 `<安装目录>\\data\\raw\\如东\\windcms`、
+「振动线 handoff」应遵循 `<安装目录>\\data\\raw\\如东\\m5_cms_tcm` 存放; 修改系统支持振动数据参与
+运行、重算」。此前振动侧只有**产物**(`outputs/<场>/m5_cms_tcm` 90 件 + `outputs/<场>/windcms` 53 件),
+源件不在 data/raw, 生成端脚本 (振动线分支的 rudong_tcm_*.py 七件) 也没随包 ⇒ 从零重算时振动侧是断的。
+
+两个事实把这件事限定得很清楚:
+  · `src/windcms/pipeline.py` 认的输入就是"含 `*_decode.json` 的目录"(Brande TCM Enterprise 导出),
+    这一步**必须能跑**, 否则数据进了 data/raw 也只是躺着;
+  · 同分支的六层链 (oem_scan / energy_share / model_run / fusion) 脚本仍未随包 —— 本脚本**不假装**
+    能把那几步跑出来: 它只做"索引 + 谱 + (可选)报告/知识库", 其余在报告里如实写"缺失"。
+
+## 这一步跑完, 振动数据就真的参与了运行
+
+    data/raw/<场站>/windcms/.../*_decode.json
+        → outputs/<场>/m5_cms_tcm/windows/<窗>/index.parquet     (54 列, 与包内 tcm_index.parquet 同构)
+        → outputs/<场>/m5_cms_tcm/windows/<窗>/spectra/*.npz + spectra_meta.parquet
+    消费者 (无需改一行代码, 窗是自动发现的):
+        src/windcms/data.py::windows()      → 新窗进分析集
+        src/windcms/data.py::load_scalars() → 标量进 CMS 报告/页面
+        src/windcms/data.py::spectrum()     → 谱图能取到 (此前 spectra 库整个缺失, 谱图是死的)
+        scripts/windcms.py report / kb      → 报告与知识库重生成
+
+## ★ 为什么"重生成 CMS 报告"默认**不跑** (2026-09-12 实测)
+
+跑一次 `scripts/windcms.py report` 在本包会**让产物变差**, 不是变好:
+
+| 件 | 随包快照 | 本包重生成后 | 变化 |
+|---|---|---|---|
+| `windcms/report.md` | 16,857 B (含逐台融合级表 + L4 过闸谱线) | **467 B** (融合级表空、L4 写"无") | **−97.2%** |
+| `windcms/overview.html` | 646,218 B | 548,466 B | −15.1% |
+| `windcms/index_eng.html` | 443,841 B | 408,112 B | −8.0% |
+
+根因不在数据, 在**缺件**: `src/windcms/data.py::load_model()` 要读
+`m5/model_run_l6.parquet` 与 `m5/fusion_38.csv`, 而这两件属六层链的 `model_run` / `fusion` 两步 ——
+**那四步脚本没随包**(见 `missing_chain`), 产物也不在。于是重生成 = 用残缺输入覆盖完整快照。
+
+所以: 摄入(索引/谱)默认跑, **报告/知识库默认不跑**; 要跑得显式 `--with-report`, 并且跑之前**先备份**
+`outputs/<场>/windcms/`。六层链补齐后这个默认值应当翻过来 (那时重生成才会≥快照)。
+
+## 用法
+
+    python scripts/vib_raw_build.py --farm rudong                     # 摄入(索引+谱), 不动报告
+    python scripts/vib_raw_build.py --window w0316                    # 指定窗名
+    python scripts/vib_raw_build.py --jobs 12 --meas FFT_             # 并行/测量过滤
+    python scripts/vib_raw_build.py --with-report                     # 额外重生成 CMS 报告/知识库(见上, 慎用)
+    python scripts/vib_raw_build.py --dry-run                         # 只报会做什么
+"""
+from __future__ import annotations
+
+import argparse
+import json
+import os
+import pathlib
+import shutil
+import subprocess
+import sys
+import time
+
+ROOT = pathlib.Path(__file__).resolve().parents[1]
+sys.path.insert(0, str(ROOT))
+from src.console import soft  # noqa: E402
+
+soft()
+PY = sys.executable
+
+
+def human(n):
+    for u in ('B', 'KB', 'MB', 'GB'):
+        if n < 1024 or u == 'GB':
+            return f'{n:.1f} {u}'
+        n /= 1024
+
+
+def find_roots(station: pathlib.Path):
+    """→ 振动侧原始件根列表 (含 *_decode.json 的目录或导出 zip)。"""
+    roots = []
+    for name in ('windcms', 'm5_cms_tcm'):
+        d = station / name
+        if not d.is_dir():
+            continue
+        for z in sorted(d.rglob('*.zip')):
+            if z.name.upper().startswith('CMS') or 'decode' in z.name.lower():
+                roots.append(z)
+        for sub in sorted({p.parent for p in d.rglob('*_decode.json')}):
+            # 取**最上层**那个含 decode 文件的目录 (measurement/), 避免每台机组一个 root
+            top = sub
+            while top.parent != d and any(top.parent.rglob('*_decode.json')):
+                top = top.parent
+            if top not in roots:
+                roots.append(top)
+    return roots
+
+
+def run(step, cmd, log):
+    t0 = time.time()
+    print(f'[{step}] {" ".join(str(c) for c in cmd)}', flush=True)
+    p = subprocess.run([str(c) for c in cmd], cwd=str(ROOT), capture_output=True, text=True,
+                       encoding='utf-8', errors='replace')
+    tail = (p.stdout or '') + (p.stderr or '')
+    rec = dict(step=step, cmd=' '.join(str(c) for c in cmd), rc=p.returncode,
+               seconds=round(time.time() - t0, 1), tail=tail[-1200:])
+    log.append(rec)
+    print(f'    rc={p.returncode}  {rec["seconds"]}s', flush=True)
+    if p.returncode != 0:
+        print(tail[-2000:], flush=True)
+        raise SystemExit(f'步骤 {step} 失败 (rc={p.returncode})')
+    return rec
+
+
+def _rels_of_window(win: pathlib.Path, m5: pathlib.Path) -> dict:
+    """某窗的产物 → {相对产物仓的路径: 构建器说明}; 供正常摄入与 --register-only 共用 (单一实现)。"""
+    rels = {f'm5_cms_tcm/windows/{win.name}/index.parquet':
+            'scripts/rudong_tcm_index.py (54 列, 与包内 tcm_index.parquet 同列名列序)',
+            f'm5_cms_tcm/windows/{win.name}/spectra_meta.parquet':
+            'scripts/rudong_tcm_spectra.py',
+            'm5_cms_tcm/vib_raw_manifest.json': 'scripts/vib_raw_build.py'}
+    sp = win / 'spectra'
+    if sp.is_dir():
+        for p in sp.rglob('*.npz'):
+            rels[f'm5_cms_tcm/windows/{win.name}/spectra/{p.relative_to(sp).as_posix()}'] = \
+                'scripts/rudong_tcm_spectra.py (npz 分片: DataSets.DataSet.Values 空格串 → float 数组)'
+        # 谱库目录里那份 meta 是同一内容的两条读取路径之一, 也登记
+        rels[f'm5_cms_tcm/windows/{win.name}/spectra/spectra_meta.parquet'] = \
+            'scripts/rudong_tcm_spectra.py (与窗根同名件同内容: 兼顾两种读取约定)'
+    return rels
+
+
+def main() -> int:
+    ap = argparse.ArgumentParser()
+    ap.add_argument('--farm', default='rudong')
+    ap.add_argument('--window', default=None, help='窗名 (默认按数据起始日推 wMMDD)')
+    ap.add_argument('--jobs', type=int, default=min(8, (os.cpu_count() or 4)))
+    ap.add_argument('--meas', default='FFT_', help="谱摄入的测量过滤 (默认 FFT_; ALL=全转)")
+    ap.add_argument('--skip-spectra', action='store_true', help='只做索引 (谱很占盘)')
+    ap.add_argument('--with-report', action='store_true',
+                    help='额外跑 windcms.py report/kb —— ★默认关: 本包缺六层链的 model_run/fusion 产物, '
+                         '重生成会让 report.md/overview.html 掉内容 (见文件头实测表); 跑前先备份 windcms\\')
+    ap.add_argument('--register-only', action='store_true',
+                    help='不摄入, 只把**已存在**的窗体件补进来源自登记 (幂等修复: 例如先前的摄入跑在自登记'
+                         '功能之前, 台账就漏记了那些件)')
+    ap.add_argument('--limit', type=int, default=0, help='冒烟: 只摄入前 N 个文件')
+    ap.add_argument('--dry-run', action='store_true')
+    a = ap.parse_args()
+
+    from src.windscada.config import farm, raw_station_dir
+    from src import paths as P
+    cfg = farm(a.farm)
+    station = pathlib.Path(raw_station_dir(a.farm))
+    m5 = P.m5(a.farm)
+    roots = find_roots(station)
+    print(f'场站原始件目录: {station}')
+    print(f'振动侧产物目录: {m5}')
+
+    if a.register_only:
+        # 只补登记: 找已存在的窗 (--window 指定, 否则取名字最大的那个), 逐件登记。
+        # 放在"源件存在性检查"之前 —— 补登记不需要重新读原始件。
+        cands = sorted(p for p in (m5 / 'windows').glob('w[0-9][0-9][0-9][0-9]')
+                       if (p / 'index.parquet').exists())
+        if a.window:
+            cands = [p for p in cands if p.name == a.window]
+        if not cands:
+            print('没有已存在的窗可登记 (先正常跑一次摄入)')
+            return 0
+        win = cands[-1]
+        rels = _rels_of_window(win, m5)
+        from src.derived_manifest import record as _rec
+        p = _rec(P.out_root(a.farm), rels, by='scripts/vib_raw_build.py')
+        print(f'已登记 {len(rels)} 件 (窗 {win.name}) → {p}')
+        return 0
+
+    if not roots:
+        print('没找到振动原始件 (*_decode.json 目录或 CMS 导出 zip)。'
+              '先用 scripts/place_raw_data.py --scope vib 从现场包落位。')
+        return 0
+    for r in roots:
+        n = len(list(r.rglob('*_decode.json'))) if r.is_dir() else '(zip)'
+        sz = sum(p.stat().st_size for p in r.rglob('*_decode.json')) if r.is_dir() else r.stat().st_size
+        print(f'  源件根: {r}   {n} 个 decode 文件, {human(sz)}')
+    if a.dry_run:
+        print('(dry-run, 未摄入)')
+        return 0
+
+    windows_dir = m5 / 'windows'
+    staging = windows_dir / '_staging_ingest'          # 'test'/'_reimport' 之外的名字, 会被 data.windows() 看见 → 立即改名
+    if staging.exists():
+        shutil.rmtree(staging)
+    staging.mkdir(parents=True, exist_ok=True)
+
+    log = []
+    t0 = time.time()
+    root = roots[0] if len(roots) == 1 else station / 'windcms'   # 多个根时交给各自目录 (index 支持 rglob)
+    idx = staging / 'index.parquet'
+    run('index', [PY, str(ROOT / 'scripts/rudong_tcm_index.py'), '--root', str(root), '--out', str(idx),
+                  '--jobs', str(a.jobs)] + (['--limit', str(a.limit)] if a.limit else []), log)
+
+    n_spectra = 0
+    if not a.skip_spectra:
+        run('spectra', [PY, str(ROOT / 'scripts/rudong_tcm_spectra.py'), '--root', str(root),
+                        '--out', str(staging / 'spectra'), '--meas', a.meas, '--jobs', str(a.jobs)]
+            + (['--limit', str(a.limit)] if a.limit else []), log)
+        sm = staging / 'spectra' / 'spectra_meta.parquet'
+        if sm.exists():
+            import pandas as pd
+            n_spectra = len(pd.read_parquet(sm))
+
+    # 窗名: 用数据自身的起始日 (wMMDD), 与既有 w0127/w0707/w0811 同口径
+    import pandas as pd
+    d = pd.read_parquet(idx)
+    tt = pd.to_datetime(d.trigger_time, errors='coerce')
+    tmin, tmax = tt.min(), tt.max()
+    win = a.window or f'w{tmin:%m%d}'
+    final = windows_dir / win
+    if final.exists():
+        final = windows_dir / f'{win}_reimport_{time.strftime("%m%d%H%M")}'
+        print(f'⚠ 窗 {win} 已存在 → 本次摄入落到 {final.name} (data.EXCLUDE_DEFAULT 会排除 _reimport 窗, 不污染生产集)')
+    staging.rename(final)
+    print(f'\n窗: {final.name}   数据窗 {tmin} → {tmax}   行 {len(d)}   谱 {n_spectra}')
+
+    if a.with_report:
+        run('report', [PY, str(ROOT / 'scripts/windcms.py'), 'report', '--farm', a.farm], log)
+        run('kb', [PY, str(ROOT / 'scripts/windcms.py'), 'kb', '--farm', a.farm], log)
+    else:
+        print('(跳过 CMS 报告/知识库重生成 —— 本包缺 model_run/fusion 产物, 重生成会掉内容; '
+              '要跑用 --with-report, 且先备份 windcms\\)')
+
+    man = dict(farm=a.farm, window=final.name, window_dir=str(final.relative_to(ROOT)),
+               # 源件路径写**相对安装根**的形态: 这文件会随包分发, 绝对路径换机后就是死链
+               # (check_transferable.py 会把 outputs 下的绝对路径算作"机器相关路径")
+               sources=[str(r.relative_to(ROOT)) if str(r).startswith(str(ROOT)) else str(r) for r in roots],
+               sources_note='路径相对<安装目录>', rows=int(len(d)), spectra=int(n_spectra),
+               time_min=str(tmin), time_max=str(tmax),
+               turbines=int(d.turbine.nunique()), sensors=int(d.sensor_name.nunique()),
+               meas_names=int(d.meas_name.nunique()),
+               steps=[{k: v for k, v in r.items() if k != 'tail'} for r in log],
+               finished=time.strftime('%Y-%m-%d %H:%M:%S'), seconds=round(time.time() - t0, 1),
+               missing_chain=['rudong_tcm_oem_scan', 'rudong_line_energy_share', 'rudong_model_run',
+                              'rudong_fusion_run'],
+               missing_note='六层链的四步脚本未随包 (振动线分支 claude/vibration-data-diagnosis-32b69e); '
+                            '本脚本只做索引/谱/报告/知识库, 不冒充跑过那四步')
+    (m5 / 'vib_raw_manifest.json').write_text(json.dumps(man, ensure_ascii=False, indent=1), encoding='utf-8')
+
+    # 来源自登记: 让 _provenance.json 把这些件记成 raw-derived (见 src/derived_manifest.py 的说明)
+    try:
+        from src.derived_manifest import record as _rec
+        rels = _rels_of_window(final, m5)
+        if a.with_report:
+            for p in (P.cms(a.farm)).glob('报告_CMS振动状态评估报告_*.md'):
+                # 只登记"今天生成的"这一件 (随包/厂家转录的那些不冒充自算)
+                if time.strftime('%Y-%m-%d') in p.name:
+                    rels[f'windcms/{p.name}'] = 'scripts/windcms.py report (数据窗: ' + final.name + ')'
+        _rec(P.out_root(a.farm), rels, by='scripts/vib_raw_build.py')
+        print(f'  自登记 {len(rels)} 件产物 → {P.out_root(a.farm) / "_derived_manifest.json"}')
+    except Exception as _e:
+        print(f'  ⚠ 来源自登记失败 ({type(_e).__name__}: {_e}) —— 台账会漏记这些件, 但不影响产物本身')
+
+    print(f'\n完成: 用时 {time.time() - t0:.0f}s; 清单 → {m5 / "vib_raw_manifest.json"}')
+    print(f'  新窗 {final.name} 已进分析集: windcms 的谱图/标量/报告页会自动收录它 '
+          f'(消费者读取时现扫 m5\\windows\\w????\\index.parquet, 无需改代码)')
+    print('  windscada 融合面读 handoff (m5_cms_tcm/handoff_vibration_v2.json) —— 那是振动线出件, 本次不动它')
+    print('  看效果: python scripts/windcms.py serve --port 8020  → http://127.0.0.1:8020/')
+    return 0
+
+
+if __name__ == '__main__':
+    sys.exit(main())

+ 357 - 0
scripts/vib_reports_build.py

@@ -0,0 +1,357 @@
+#!/usr/bin/env python3
+# -*- coding: utf-8 -*-
+"""厂家振动报告 (docx) → 结构化提取 + 消费端格式的评估报告 (2026-09-12 用户令).
+
+## 范围与诚实边界 (先读这段, 再看代码)
+
+数据层「CMS 振动评估报告」的源件有两类, 处理方式**不同**:
+
+  · CMS 原始测量导出 (Brande TCM `*_decode.json`)  → 可重算, 走 scripts/vib_raw_build.py
+    (索引/谱/六层链), 这一路是本目录下 `report_CMS振动状态评估报告_*.md` 的**正路**。
+  · 厂家月度评估报告 (上海电气 12 份用印版 PDF + 2026年07月 docx; 大生科技传动链 docx)
+    → **PDF 全部是扫描件** (2026-09-12 用 pypdf 抽检: 12 页 12 图, extract_text 长度 0),
+      没有 OCR 就取不出数值 ⇒ 只登记归档, 不进判级;
+      **docx 有文本层**, 本脚本把里面的逐台判级**逐字转录**成结构化件与一份 report md。
+
+**转录纪律** (照抄本包既有做法, 防"看起来像算出来的"): 每个产物的 meta 里写清源文件+sha256+表头+行数;
+取值一律来自报告原文, 不补不猜; 报告没给的列写 '—' 并在说明里点明"厂家未给"; 聚合量(如综合级)
+必须标注它是**取严规则**得出的, 不是厂家判的。
+
+## 输出的两份东西 (都在 outputs/<场>/ 下, 与既有消费者兼容)
+
+  windcms/厂家报告提取_<报告期>.json    逐台原文 + provenance (给机器读/复核)
+  windcms/报告_CMS振动状态评估报告_<报告期>.md
+       与 windcms 自产报告**同构**: `## 附录 A` 下 `| WTGxx | 主轴承前 | 主轴承后 | 齿轮箱 | 发电机 | 综合 | …`
+       (消费者: src/windscada/subsys/fusion.py::windcms_grades 取第 6 列=综合,
+        src/windscada/taxonomy.py 取第 2..5 列)。日期用**报告期**, 于是它不会顶掉更新的自产报告,
+       但随包自产报告缺失时它就是最新版 (从零重算的机器上正好用得上)。
+  m5_cms_tcm/厂家报告提取_<报告期>.json + 报告_TCM传动链振动分析_<报告期>.md
+       大生科技 (TCM M-system 数据) 逐台状态等级/分析/建议, 原文透传; **不**冒充 handoff 判级。
+
+用法:
+    python scripts/vib_reports_build.py --dump     # 只打印解析结果 (人工核对用, 不写盘)
+    python scripts/vib_reports_build.py            # 摄入并写产物
+"""
+from __future__ import annotations
+
+import argparse
+import hashlib
+import json
+import pathlib
+import re
+import sys
+import time
+
+ROOT = pathlib.Path(__file__).resolve().parents[1]
+sys.path.insert(0, str(ROOT))
+from src.console import soft  # noqa: E402
+
+soft()
+
+STATE_ORDER = ['危险', '报警', '良好', '优秀', '不可判']      # 取严: 越靠前越严重 (报告用四级 + 不可判)
+NO_DATA_MARKS = ('--', '—', '无数据', '通讯中断', '无通讯')
+
+
+def sha256(p: pathlib.Path) -> str:
+    h = hashlib.sha256()
+    with open(p, 'rb') as f:
+        for chunk in iter(lambda: f.read(1 << 20), b''):
+            h.update(chunk)
+    return h.hexdigest()[:16]
+
+
+def grid_of(table):
+    """docx 表 → 二维网格, **正确还原合并单元格**。
+
+    为什么不能直接用 `row.cells`: python-docx 对合并区会把同一个 `tc` 在相邻列/相邻行重复给出 ——
+    照抄等于把 "优秀 | 优秀 | 优秀" 压成 "优秀", 列就对不齐了 (第一版手工 dump 就吃了这个亏:
+    WTG-01 行显示成 `WTG-01 | 优秀 | 无`, 而实际是 主轴承=优秀 齿轮箱=优秀 发电机=优秀 结论=无)。
+    规则: 同一行里 tc 与左邻相同 = 横向合并 (沿用左值); tc 与上方同行同列相同 = 纵向合并 (沿用上值)。
+    """
+    ncol = len(table.columns)
+    out, tc_above = [], [None] * ncol
+    for row in table.rows:
+        vals, tcs, prev = [], [], None
+        for ci, cell in enumerate(row.cells):
+            tc = cell._tc
+            if tc is prev:                      # 横向合并的续格
+                vals.append(vals[-1] if vals else '')
+                tcs.append(tc)
+                continue
+            if tc is tc_above[ci] and ci < len(tc_above):   # 纵向合并的续格
+                vals.append(out[-1][ci] if out else '')
+            else:
+                vals.append(cell.text.strip().replace('\n', ' '))
+            tcs.append(tc)
+            prev = tc
+        tc_above = tcs
+        out.append(vals)
+    return out
+
+
+def find_table(doc, must_have, header_hint=None):
+    """按表头关键字找表 → (index, grid)。must_have: 表头行必须都含这些词。"""
+    import docx
+    for i, tb in enumerate(doc.tables):
+        g = grid_of(tb)
+        if g and all(any(k in c for c in g[0]) for k in must_have):
+            return i, g
+    return None, None
+
+
+def norm_turbine(s):
+    m = re.search(r'WTG[\s_-]*0*(\d{1,2})', str(s).upper())
+    return f'WTG{int(m.group(1)):02d}' if m else None
+
+
+def state_of(s):
+    """文本 → 五级之一; 认不出给 '不可判' (宁可说不可判, 不冒充优秀)。"""
+    t = str(s or '').strip()
+    if any(k in t for k in NO_DATA_MARKS) or t == '':
+        return '不可判'
+    for v in ('危险', '报警', '预警', '良好', '优秀'):
+        if v in t:
+            return {'预警': '报警'}.get(v, v)          # 上海电气四级里没有"预警", 大生科技有; 此处按四级归一
+    if '异常' in t:
+        return '不可判'
+    return '不可判'
+
+
+def worst(*states):
+    """取严 (综合列的口径; 厂家未给综合列时必须写明这是我们的聚合规则)。"""
+    live = [s for s in states if s]
+    return min(live, key=lambda s: STATE_ORDER.index(s) if s in STATE_ORDER else 99) if live else '不可判'
+
+
+def parse_sa(path: pathlib.Path):
+    """上海电气月度振动分析报告: 表「机组号|主轴承|齿轮箱|发电机|诊断结论和维护建议」+ 总览句校验和。"""
+    import docx
+    d = docx.Document(str(path))
+    ti, g = find_table(d, ['机组号', '主轴承', '齿轮箱'])
+    if g is None:
+        return None
+    head = g[0]
+    idx = {name: next((i for i, c in enumerate(head) if name in c), None)
+           for name in ('机组号', '主轴承', '齿轮箱', '发电机', '诊断')}
+    rows = []
+    for r in g[1:]:
+        t = norm_turbine(r[idx['机组号']] if idx['机组号'] is not None else '')
+        if not t:
+            continue
+        rows.append(dict(turbine=t,
+                         主轴承=state_of(r[idx['主轴承']]) if idx['主轴承'] is not None else '不可判',
+                         齿轮箱=state_of(r[idx['齿轮箱']]) if idx['齿轮箱'] is not None else '不可判',
+                         发电机=state_of(r[idx['发电机']]) if idx['发电机'] is not None else '不可判',
+                         厂家结论=(r[idx['诊断']] if idx['诊断'] is not None and idx['诊断'] < len(r) else '')))
+    # 总览句 = 报告的**自带校验和** (2026-07 报告原文: 优秀34 良好3 预警0 无数据1 测点异常1)。
+    # ★它在**表格单元格**里, 不是段落 (第一版只扫 paragraphs → 拿到空串, 于是"校验和"形同虚设)。
+    overview, sums = '', {}
+    texts = [p.text for p in d.paragraphs]
+    for tb in d.tables:
+        for row in tb.rows:
+            for c in row.cells:
+                texts.append(c.text)
+    for t in texts:
+        if '检测结果' in t and '机组' in t and '台' in t:
+            overview = t.strip()
+            break
+    for k in ('优秀', '良好', '预警', '无数据', '测点异常'):
+        # 原话是"运行状态优秀机组34台" —— 等级词与数字之间夹着"机组"二字, 只写 `{k}\s*(\d+)` 会漏掉优秀
+        m = re.search(rf'{k}[^0-9]{{0,6}}(\d+)\s*台', overview)
+        if m:
+            sums[k] = int(m.group(1))
+    period = re.search(r'(20\d\d)\s*年\s*(\d{1,2})\s*月', path.name)
+    # 报告期: 文件名 2026年07月 → 用**月末** (报告覆盖整个月; 用月初会让它比同类报告显得更新)
+    import calendar
+    if period:
+        y, mo = int(period.group(1)), int(period.group(2))
+        period = f'{y:04d}-{mo:02d}-{calendar.monthrange(y, mo)[1]:02d}'
+    else:
+        period = time.strftime('%Y-%m-%d')
+    return dict(kind='上海电气月度', file=path.name, sha256=sha256(path), table_index=ti,
+                header=[c for c in head], period=period, overview=overview, overview_counts=sums,
+                rows=rows)
+
+
+def parse_ds(path: pathlib.Path):
+    """大生科技传动链振动分析报告: 表「机组号|状态等级|分析与结论|建议」+ 数据窗 (报告正文)。"""
+    import docx
+    d = docx.Document(str(path))
+    ti, g = find_table(d, ['机组号', '状态等级', '分析'])
+    if g is None:
+        return None
+    head = g[0]
+    idx = {name: next((i for i, c in enumerate(head) if name in c), None)
+           for name in ('机组号', '状态等级', '分析', '建议')}
+    rows = []
+    for r in g[1:]:
+        t = norm_turbine(r[idx['机组号']] if idx['机组号'] is not None else '')
+        if not t:
+            continue
+        rows.append(dict(turbine=t,
+                         状态等级=str(r[idx['状态等级']]).strip() if idx['状态等级'] is not None else '',
+                         分析与结论=r[idx['分析']] if idx['分析'] is not None else '',
+                         建议=r[idx['建议']] if idx['建议'] is not None else ''))
+    txt = '\n'.join(p.text for p in d.paragraphs)
+    # 数据窗: 正文原话 "本次分析采集了 A 至 B 期间的振动监测数据" (两个时间戳都要带时刻)
+    win = re.search(r'(\d{4}年\d{1,2}月\d{1,2}日\s*\d{1,2}:\d{2}:\d{2})\s*至\s*'
+                    r'(\d{4}年\d{1,2}月\d{1,2}日\s*\d{1,2}:\d{2}:\d{2})', txt)
+    rep = re.search(r'报告日期[::]\s*([\d.]{8,10})', txt)
+    period = (rep.group(1).replace('.', '-') if rep else time.strftime('%Y-%m-%d'))
+    if re.fullmatch(r'\d{4}-\d{1,2}-\d{1,2}', period):
+        y, m, dd = period.split('-')
+        period = f'{y}-{int(m):02d}-{int(dd):02d}'
+    return dict(kind='大生科技传动链', file=path.name, sha256=sha256(path), table_index=ti,
+                header=[c for c in head], period=period,
+                data_window=(f'{win.group(1)} ~ {win.group(2)}' if win else '未标'),
+                rows=rows)
+
+
+def md_sa(ext, src_note):
+    """上海电气提取 → 与 windcms 自产报告同构的评估报告 (附录 A 表)。"""
+    L = [f'# CMS 振动状态评估报告 (厂家件转录) {ext["period"]}', '',
+         f'> 来源: `{ext["file"]}` (sha256 {ext["sha256"]}, 表 {ext["table_index"]}) —— {src_note}', '',
+         f'> 报告原文总览: {ext["overview"] or "(未找到总览句)"}', '',
+         '> ★ 本表是从**厂家月度振动分析报告**逐字转录的(上海电气, 依据 VDI3834 与 NB/T 31129-2018):',
+         '> ① 厂家按「主轴承 / 齿轮箱 / 发电机」给级, **主轴承不分前后** —— 本表两列填同一值, 不是我们分的;',
+         '> ② 厂家**未给综合列**, 本表「综合」= 三项**取严**(危险>报警>良好>优秀>不可判) 的聚合规则结果;',
+         '> ③ 厂家未给「融合级/CMS 红黄/行动等级建议」, 一律写 `—` (不猜);',
+         '> ④ 判级细节(谱/门槛)以厂家报告与 CMS 原始导出分析为准, 本表只承接"哪台什么级"。', '',
+         '## 附录 A 38 台状态等级表(厂家报告转录,不代表确诊数量)', '',
+         '| 机组号 | 主轴承前 | 主轴承后 | 齿轮箱 | 发电机 | 综合 | 融合级(模型) | 厂家结论与建议 |',
+         '|---|---|---|---|---|---|---|---|']
+    for r in sorted(ext['rows'], key=lambda x: x['turbine']):
+        zh = worst(r['主轴承'], r['齿轮箱'], r['发电机'])
+        note = (r['厂家结论'] or '—').replace('|', '/')
+        L.append(f"| {r['turbine']} | {r['主轴承']} | {r['主轴承']} | {r['齿轮箱']} | {r['发电机']} | "
+                 f"{zh} | — | {note} |")
+    L += ['', '## 说明事项', '',
+          '1. 本件是**厂家报告的转录**, 不是观澜自算结果; 观澜自算报告 (CMS 原始导出 → 六层链) 见同目录 `报告_CMS振动状态评估报告_*.md`。',
+          '2. 厂家报告里的「无数据/通讯中断/测点异常」在本表记为 `不可判` —— 没测到 ≠ 正常。',
+          f'3. 报告期为 {ext["period"]} (文件名月份, 取月末)。', '']
+    return '\n'.join(L)
+
+
+def md_ds(ext, src_note):
+    L = [f'# TCM 传动链振动分析报告 (厂家件转录) {ext["period"]}', '',
+         f'> 来源: `{ext["file"]}` (sha256 {ext["sha256"]}, 表 {ext["table_index"]}) —— {src_note}', '',
+         f'> 数据窗: {ext["data_window"]}    报告日期: {ext["period"]}', '',
+         '> ★ 原文透传: 状态等级/分析与结论/建议按厂家报告逐字记录。**本件不构成 handoff 判级** ——',
+         '> handoff 的定谳链判级 (候选以上/证据窗) 是振动线接口件, 本报告只作 TCM 侧证据与交叉参考。', '',
+         '| 机组 | 状态等级 | 分析与结论 | 建议 |', '|---|---|---|---|']
+    for r in sorted(ext['rows'], key=lambda x: x['turbine']):
+        cells = [r['turbine'], r['状态等级'], r['分析与结论'] or '—', r['建议'] or '—']
+        L.append('| ' + ' | '.join(str(c).replace('|', '/') for c in cells) + ' |')
+    L += ['', '## 说明事项', '',
+          '1. 判据/图谱在厂家报告原文 (本包只落可解析的 docx; 若同批有 PDF 扫描件, 扫描件只归档不解析)。',
+          '2. 状态等级为厂家四级口径 (优秀/良好/预警/报警) + "数据异常"(系统故障无通讯), 与观澜五级不是同一套, 消费时别直接比。', '']
+    return '\n'.join(L)
+
+
+def main() -> int:
+    ap = argparse.ArgumentParser()
+    ap.add_argument('--farm', default='rudong')
+    ap.add_argument('--dump', action='store_true', help='只打印, 不写盘')
+    a = ap.parse_args()
+    from src.windscada.config import raw_station_dir
+    from src import paths as P
+
+    station = pathlib.Path(raw_station_dir(a.farm))
+    sa_dir = station / 'windcms' / '厂家报告' / '上海电气_月度'
+    ds_dir = station / 'm5_cms_tcm' / '厂家报告'
+    out_cms = P.cms(a.farm)
+    out_m5 = P.m5(a.farm)
+
+    sa_files = sorted(sa_dir.glob('*.docx')) if sa_dir.is_dir() else []
+    ds_files = sorted(ds_dir.glob('*.docx')) if ds_dir.is_dir() else []
+    print(f'上海电气 docx: {len(sa_files)} 件   {sa_dir}')
+    for p in sa_files:
+        print(f'   {p.name}  {p.stat().st_size / 1048576:.2f} MB')
+    print(f'大生科技 docx: {len(ds_files)} 件   {ds_dir}')
+    for p in ds_files:
+        print(f'   {p.name}  {p.stat().st_size / 1048576:.2f} MB')
+    pdfs = sorted(sa_dir.glob('*.pdf')) if sa_dir.is_dir() else []
+    if pdfs:
+        print(f'扫描件 PDF: {len(pdfs)} 件 (登记归档, 不解析)')
+    if not sa_files and not ds_files:
+        print('没有可解析的 docx 厂家报告 (只有扫描件 PDF 属正常) → 空跑退出')
+        return 0
+
+    written = []
+    for p in sa_files:
+        ext = parse_sa(p)
+        if not ext:
+            print(f'⚠ {p.name}: 没找到"机组检测结果列表"表 (跳过, 不猜)')
+            continue
+        zh = [worst(r['主轴承'], r['齿轮箱'], r['发电机']) for r in ext['rows']]
+        print(f'\n== {p.name} ({ext["kind"]}, 表{ext["table_index"]}) ==')
+        print(f'   报告期 {ext["period"]}  台数 {len(ext["rows"])}  原文总览: {ext["overview"]}')
+        from collections import Counter
+        got = Counter(zh)
+        print(f'   转录结果取严分布: {dict(got)}')
+        print(f'   厂家总览句: {ext["overview_counts"]}')
+        if a.dump:
+            for r in ext['rows'][:6]:
+                print('   ', r)
+            continue
+        out_cms.mkdir(parents=True, exist_ok=True)
+        jp = out_cms / f'厂家报告提取_{ext["period"]}.json'
+        jp.write_text(json.dumps(ext, ensure_ascii=False, indent=1), encoding='utf-8')
+        mp = out_cms / f'报告_CMS振动状态评估报告_{ext["period"]}.md'
+        mp.write_text(md_sa(ext, '现场包 厂家月度评估报告'), encoding='utf-8')
+        written += [jp, mp]
+        print(f'   → {jp.name}\n   → {mp.name}')
+    for p in ds_files:
+        ext = parse_ds(p)
+        if not ext:
+            print(f'⚠ {p.name}: 没找到"信号分析与结论描述"表 (跳过, 不猜)')
+            continue
+        print(f'\n== {p.name} ({ext["kind"]}, 表{ext["table_index"]}) ==')
+        print(f'   报告期 {ext["period"]}  数据窗 {ext["data_window"]}  台数 {len(ext["rows"])}')
+        from collections import Counter
+        print(f'   状态等级分布: {dict(Counter(r["状态等级"] for r in ext["rows"]))}')
+        if a.dump:
+            for r in ext['rows'][:4]:
+                print('   ', {k: (v[:60] + '…' if isinstance(v, str) and len(v) > 60 else v) for k, v in r.items()})
+            continue
+        out_m5.mkdir(parents=True, exist_ok=True)
+        jp = out_m5 / f'厂家报告提取_{ext["period"]}.json'
+        jp.write_text(json.dumps(ext, ensure_ascii=False, indent=1), encoding='utf-8')
+        mp = out_m5 / f'报告_TCM传动链振动分析_{ext["period"]}.md'
+        mp.write_text(md_ds(ext, '现场包 大生科技 TCM 传动链分析报告 (docx)'), encoding='utf-8')
+        written += [jp, mp]
+        print(f'   → {jp.name}\n   → {mp.name}')
+
+    if not a.dump and pdfs:
+        reg = dict(note='扫描件登记: 无文本层 → 只归档, 数值不进系统; 需数值须 OCR 或取厂家电子件',
+                   verified='2026-09-12 用 pypdf 抽检一件: 12 页 / 每页 1 图 / extract_text() 长度 0',
+                   files=[dict(name=q.name, bytes=q.stat().st_size, sha256=sha256(q)) for q in pdfs])
+        rp = sa_dir / '_扫描件清单.json'
+        rp.write_text(json.dumps(reg, ensure_ascii=False, indent=1), encoding='utf-8')
+        written.append(rp)
+        print(f'\n扫描件登记 → {rp}')
+
+    # 来源自登记 (见 src/derived_manifest.py): 转录件也是"从 data/raw 来的", 台账要记 raw-derived
+    if not a.dump and written:
+        try:
+            from src.derived_manifest import record as _rec
+            from src import paths as _P
+            rels = {}
+            for p in written:
+                if p.suffix.lower() not in ('.json', '.md') or '_扫描件清单' in p.name:
+                    continue
+                if p.parent == out_cms:
+                    rels[f'windcms/{p.name}'] = 'scripts/vib_reports_build.py (厂家 docx 报告逐台转录)'
+                elif p.parent == out_m5:
+                    rels[f'm5_cms_tcm/{p.name}'] = 'scripts/vib_reports_build.py (厂家 docx 报告逐台转录)'
+            if rels:
+                _rec(_P.out_root(a.farm), rels, by='scripts/vib_reports_build.py')
+                print(f'自登记 {len(rels)} 件 → {_P.out_root(a.farm) / "_derived_manifest.json"}')
+        except Exception as _e:
+            print(f'⚠ 来源自登记失败 ({type(_e).__name__}: {_e})')
+    print(f'\n完成: 写出 {len(written)} 件')
+    return 0
+
+
+if __name__ == '__main__':
+    sys.exit(main())

+ 113 - 27
scripts/windscada_serve.py

@@ -9,6 +9,7 @@ import sys as _sys, pathlib as _plb
 _sys.path.insert(0, str(_plb.Path(__file__).resolve().parents[1]))
 _sys.path.insert(0, str(_plb.Path(__file__).resolve().parents[1]))
 from src import paths as P
 from src import paths as P
 import sys, json, pathlib, re, threading, urllib.parse, collections
 import sys, json, pathlib, re, threading, urllib.parse, collections
+import hashlib, time
 sys.path.insert(0, str(pathlib.Path(__file__).resolve().parents[1]))     # 安装根入 sys.path (先于 src.* 导入)
 sys.path.insert(0, str(pathlib.Path(__file__).resolve().parents[1]))     # 安装根入 sys.path (先于 src.* 导入)
 from src import paths as _P                                              # 路径唯一真源 (与 cwd 无关)
 from src import paths as _P                                              # 路径唯一真源 (与 cwd 无关)
 import numpy as np, pandas as pd
 import numpy as np, pandas as pd
@@ -511,6 +512,56 @@ LOCK = threading.Lock()
 _CACHE = {}
 _CACHE = {}
 
 
 
 
+# ── 产物指纹 → **运行中自动重载** (2026-09-16 用户令: "重算后的产物可即时用于页面呈现") ──────
+# 背景 (实测): _load() 把产物**一次性**读进内存, 之后只在启动时读一次 ⇒ 重算完成后页面仍显示旧数,
+# 必须重启组件服务才更新。判定依据: 改产物后 /detail/api/fleet 等仍回旧值, 重启组件后才变。
+# 做法: 取"页面取数依赖的产物文件"的 (路径, mtime_ns, size) 摘要当指纹; _load() 前比一次
+# (带 TTL, 免得每请求都 stat 一遍), 指纹变了就重载 ⇒ 重算产物**不再需要重启服务**。
+# 也提供 /api/reload 显式重载 (运维控制台重算结束后可点一下)。
+_STAMP_TTL = 2.0          # 秒: 指纹有效期 (stat ~150 个文件 ≈ 1~3 ms, 不值得每请求都做)
+_STAMP = {'val': None, 'at': 0.0}
+
+
+def _product_files():
+    """页面取数依赖的产物文件清单 (顺序稳定: 供指纹; 只列**产物**, 不含 reference/ 随包契约)。
+
+    覆盖 _load() 直读的件 + 它经 taxonomy/temp_nbm/hydraulic/fusion 间接读的件:
+      windscada/*.parquet|csv|json  · ontology/*.json  · pitch/*.parquet
+      m5_cms_tcm/{handoff_vibration_v2,component_history,baseline_38}.json  · windcms/报告_CMS*.md
+    """
+    try:
+        from src.windcms.config import cms_out as _cms_out     # 与写侧同一解析口 (WINDCMS_OUT)
+        _cms_dir = _cms_out()
+    except Exception:
+        _cms_dir = _P.cms()
+    pats = ((ST, ('*.parquet', '*.csv', '*.json')),
+            (_P.ont(), ('*.json',)),
+            (_P.pitch(), ('*.parquet',)),
+            (_P.m5(), ('handoff_vibration_v2.json', 'component_history.json', 'baseline_38.json')),
+            (_cms_dir, ('报告_CMS振动状态评估报告_*.md',)))
+    files = []
+    for d, ps in pats:
+        for pat in ps:
+            files.extend(sorted(d.glob(pat)))
+    return files
+
+
+def products_stamp(force=False):
+    """产物指纹 (sha1 前 16 位)。force=True 时忽略 TTL 立即重算。"""
+    now = time.time()
+    if not force and _STAMP['val'] is not None and (now - _STAMP['at']) < _STAMP_TTL:
+        return _STAMP['val']
+    h = hashlib.sha1()
+    for p in _product_files():
+        try:
+            st = p.stat()
+            h.update(f'{_P.rel(p)}|{st.st_mtime_ns}|{st.st_size}\n'.encode('utf-8'))
+        except OSError:
+            h.update(f'{_P.rel(p)}|MISSING\n'.encode('utf-8'))
+    _STAMP.update(val=h.hexdigest()[:16], at=now)
+    return _STAMP['val']
+
+
 class ProductsMissing(RuntimeError):
 class ProductsMissing(RuntimeError):
     """产物缺失 (outputs/<场>/… 被清空或尚未生成)。
     """产物缺失 (outputs/<场>/… 被清空或尚未生成)。
 
 
@@ -524,34 +575,55 @@ class ProductsMissing(RuntimeError):
         self.what, self.path = name, path
         self.what, self.path = name, path
 
 
 
 
+def reload_products(reason=''):
+    """显式清缓存 (下一次请求重载)。返回清前的指纹, 供日志/接口回显。"""
+    with LOCK:
+        old = _CACHE.get('__stamp')
+        _CACHE.clear()
+    print(f'[reload] 清产物缓存 ({reason or "手动"}) 旧指纹={old}', flush=True)
+    return old
+
+
 def _load():
 def _load():
     with LOCK:
     with LOCK:
-        if _CACHE.get('__loaded'): return
+        stamp = products_stamp()
+        if _CACHE.get('__loaded'):
+            if _CACHE.get('__stamp') == stamp:
+                return
+            # 产物变了 (典型: 刚跑完重算) → 重载。**先建后换**: 下面任何一步抛异常都不会破坏旧缓存,
+            # 请求照旧能用旧数 (降级但不空白), 同时日志留痕。
+            print(f'[reload] 产物指纹变化 {_CACHE.get("__stamp")} → {stamp}, 重载', flush=True)
         tmp = {}
         tmp = {}
-        for key, name in (('tm', 'temp_monthly.parquet'), ('al', 'alarms.parquet'),
-                          ('lm', 'loss_monthly.parquet'), ('bins', 'powercurve_bins.parquet'),
-                          ('pcd', 'powercurve_dev.parquet')):
-            f = ST / name
-            if not f.exists():
-                raise ProductsMissing(name, f)
-            tmp[key] = pd.read_parquet(f)
-        tmp['al']['month'] = tmp['al']['t_on'].dt.to_period('M').astype(str)
-        tmp['pcd'] = tmp['pcd'].set_index('turbine')
-        for key, name in (('duty', 'duty_monthly.parquet'),):
-            f = ST / name
-            tmp[key] = pd.read_parquet(f) if f.exists() else None
         try:
         try:
-            tmp['sysmx'] = taxonomy.system_matrix()
-            tmp['treg'] = temp_nbm.registry()
-            tmp['treg'] = tmp['treg'][0] if isinstance(tmp['treg'], tuple) else tmp['treg']
-            tmp['hyd'], _ = hydraulic.registry()
-            tmp['hyd'] = tmp['hyd'].set_index('turbine')
-        except FileNotFoundError as e:            # 这些派生件同样在产物仓里; 缺了就按"无产物"处理
-            raise ProductsMissing(getattr(e, 'filename', '派生产物'), getattr(e, 'filename', ST))
-        zp = _P.pitch() / 'pitch_zero_monthly.parquet'
-        tmp['zero'] = pd.read_parquet(zp) if zp.exists() else None
-        tmp['__loaded'] = True
-        _CACHE.update(tmp)                        # 原子提交: 失败时不留下半截缓存
+            for key, name in (('tm', 'temp_monthly.parquet'), ('al', 'alarms.parquet'),
+                              ('lm', 'loss_monthly.parquet'), ('bins', 'powercurve_bins.parquet'),
+                              ('pcd', 'powercurve_dev.parquet')):
+                f = ST / name
+                if not f.exists():
+                    raise ProductsMissing(name, f)
+                tmp[key] = pd.read_parquet(f)
+            tmp['al']['month'] = tmp['al']['t_on'].dt.to_period('M').astype(str)
+            tmp['pcd'] = tmp['pcd'].set_index('turbine')
+            for key, name in (('duty', 'duty_monthly.parquet'),):
+                f = ST / name
+                tmp[key] = pd.read_parquet(f) if f.exists() else None
+            try:
+                tmp['sysmx'] = taxonomy.system_matrix()
+                tmp['treg'] = temp_nbm.registry()
+                tmp['treg'] = tmp['treg'][0] if isinstance(tmp['treg'], tuple) else tmp['treg']
+                tmp['hyd'], _ = hydraulic.registry()
+                tmp['hyd'] = tmp['hyd'].set_index('turbine')
+            except FileNotFoundError as e:            # 这些派生件同样在产物仓里; 缺了就按"无产物"处理
+                raise ProductsMissing(getattr(e, 'filename', '派生产物'), getattr(e, 'filename', ST))
+            zp = _P.pitch() / 'pitch_zero_monthly.parquet'
+            tmp['zero'] = pd.read_parquet(zp) if zp.exists() else None
+            tmp['__stamp'] = stamp
+            tmp['__loaded'] = True
+            _CACHE.update(tmp)                        # 原子提交: 失败时不留下半截缓存
+        except Exception:
+            if _CACHE.get('__loaded'):
+                print('[reload] 重载失败 → 继续用上一份缓存 (页面不会空白, 但数是旧的; 看上面的异常)', flush=True)
+            raise
 
 
 WINDOWS = ['近30日', '近90日', '2026年', '2026H1', '2025H2', '全程']
 WINDOWS = ['近30日', '近90日', '2026年', '2026H1', '2025H2', '全程']
 def months_of(win):
 def months_of(win):
@@ -1397,7 +1469,10 @@ def _ask_worker(q, model, lang='zh'):
             cmd.append(q + ' Answer in English only. End every conclusion sentence with [object/id]. Finally run claim_check.')
             cmd.append(q + ' Answer in English only. End every conclusion sentence with [object/id]. Finally run claim_check.')
         else:
         else:
             cmd.append(q + ' 每个结论句末尾用[对象id]标注来源。最后调用claim_check。')
             cmd.append(q + ' 每个结论句末尾用[对象id]标注来源。最后调用claim_check。')
-        r = subprocess.run(cmd, capture_output=True, text=True, timeout=600)
+        # 2026-09-16: 走 src.proc.run —— 本函数跑在 detail 服务进程里 (无可见控制台), 裸 spawn
+        # 会让 Windows 给这条"云端档问答"新建可见控制台窗口 (本地档不走这里)。
+        from src import proc as _proc
+        r = _proc.run(cmd, capture_output=True, text=True, timeout=600)
         ans = r.stdout.strip() or ('(空输出) stderr: ' + r.stderr[-500:])
         ans = r.stdout.strip() or ('(空输出) stderr: ' + r.stderr[-500:])
         out.write_text(ans, encoding='utf-8')
         out.write_text(ans, encoding='utf-8')
         sys.path.insert(0, str(pathlib.Path('scripts').resolve()))
         sys.path.insert(0, str(pathlib.Path('scripts').resolve()))
@@ -5031,6 +5106,14 @@ class H(BaseHTTPRequestHandler):
                         self._send('404', code=404)
                         self._send('404', code=404)
                 else:
                 else:
                     self._send('403', code=403)
                     self._send('403', code=403)
+            elif u.path == '/api/reload':
+                # 产物重载 (2026-09-16 用户令: 重算产物即时可用于页面呈现)。
+                # 平时**不需要**调它 —— _load() 每次请求都会比产物指纹, 变了自动重载;
+                # 这个端点给"重算结束想立刻生效"的显式动作, 也便于运维界面/脚本确认当前指纹。
+                _old = reload_products(urllib.parse.parse_qs(u.query).get('why', ['manual'])[0])
+                self._send(_jdump(dict(ok=True, old_stamp=_old, stamp=products_stamp(force=True),
+                                       files=len(_product_files())), ensure_ascii=False),
+                           'application/json; charset=utf-8')
             else:
             else:
                 self._send('404', code=404)
                 self._send('404', code=404)
         except ProductsMissing as e:
         except ProductsMissing as e:
@@ -5038,8 +5121,10 @@ class H(BaseHTTPRequestHandler):
             self._send(_jdump(dict(kind='multiline', title='无产物', unit='', months=[], series=[],
             self._send(_jdump(dict(kind='multiline', title='无产物', unit='', months=[], series=[],
                                    err='no_products',
                                    err='no_products',
                                    note=f'{e.what} 不存在 → {_P.rel(e.path)}; 先跑 '
                                    note=f'{e.what} 不存在 → {_P.rel(e.path)}; 先跑 '
-                                        f'scripts/rebuild_from_raw.py (SCADA 侧加 --scada) 生成产物, '
-                                        f'或用 scripts/products_state.py --on 还原'), ensure_ascii=False),
+                                        f'scripts/rebuild_from_raw.py (SCADA 侧加 --scada) 生成产物; '
+                                        f'"包内没有生成端"的件从交付包补齐: '
+                                        f'python scripts/products_restore_missing.py --stash <交付包.zip> '
+                                        f'(2026-09-16 起清除产物不留备份, 故不再有 --on 还原)'), ensure_ascii=False),
                        'application/json; charset=utf-8')
                        'application/json; charset=utf-8')
         except Exception as e:
         except Exception as e:
             import traceback
             import traceback
@@ -5075,6 +5160,7 @@ if __name__ == '__main__':
         _me.CFG = _cfg.farm()
         _me.CFG = _cfg.farm()
         _me.ST = pathlib.Path(_me.CFG['store'])
         _me.ST = pathlib.Path(_me.CFG['store'])
         _CACHE.clear()
         _CACHE.clear()
+        _STAMP['val'] = None            # 换场后产物指纹必须重算, 否则新场的缓存判为"未变化"
         print(f'场: {_me.CFG["name"]} ({a.farm}, {_me.CFG["n_turbines"]} 台, 仓 {_me.ST})')
         print(f'场: {_me.CFG["name"]} ({a.farm}, {_me.CFG["n_turbines"]} 台, 仓 {_me.ST})')
     else:
     else:
         print(f'场: {CFG["name"]} (默认; 可用 {list(__import__("src.windscada.config", fromlist=["available"]).available())})')
         print(f'场: {CFG["name"]} (默认; 可用 {list(__import__("src.windscada.config", fromlist=["available"]).available())})')

+ 82 - 0
src/derived_manifest.py

@@ -0,0 +1,82 @@
+# -*- coding: utf-8 -*-
+"""产物来源**自登记**: 构建脚本落盘后把"这一件是我从 data/raw 算出来的"记进 `outputs/<场>/_derived_manifest.json`。
+
+## 为什么要有它
+
+`outputs/<场>/_provenance.json` 是**逐件来源台账**(raw-derived = 由 data/raw 重算 / shipped = 包内无生成端,
+用随包件补齐)。它由 `scripts/products_restore_missing.py` 生成, 而那个脚本是按**随包快照**逐件走一遍的 ——
+于是**新造的、快照里根本没有的产物**不会自动进台账 (既不算 raw-derived 也不算 shipped)。
+
+早期的做法是在 `products_restore_missing.py` 里维护一张 `RAW_DERIVED` 精确路径表。对"件数少、名字固定"
+的产物够用; 但振动侧的产物是 `<窗>/index.parquet` + `<窗>/spectra/*.npz`(分片名带序号) —— 窗名与分片数
+都随数据变, 写不进精确表, 而**按名字通配**又会误伤同名旧件 (例如 `报告_CMS振动状态评估报告_*.md`
+既有随包/自产的、也有厂家报告转录的, 名字形态一样)。
+
+所以改成**自登记**: 谁算的谁登记, 台账只认这份登记。名字对不上不是问题, 因为登记的是**相对路径本身**。
+
+用法 (构建脚本内):
+    from src.derived_manifest import record
+    record(P.out_root('rudong'), {rel: 'scripts/rudong_tcm_index.py (54 列, 与包内 tcm_index.parquet 同构)'},
+           by='scripts/vib_raw_build.py')
+"""
+from __future__ import annotations
+
+import json
+import pathlib
+import time
+
+FILENAME = '_derived_manifest.json'
+
+
+def path_of(store_root) -> pathlib.Path:
+    return pathlib.Path(store_root) / FILENAME
+
+
+def load(store_root) -> dict:
+    p = path_of(store_root)
+    if not p.exists():
+        return {}
+    try:
+        return json.loads(p.read_text(encoding='utf-8'))
+    except Exception:
+        return {}
+
+
+def prune(store_root) -> int:
+    """删掉**登记了但盘上已不存在**的条目, 返回删除数。
+
+    为什么需要 (2026-09-16 实逮): 振动摄入对同一批数据重跑时会落 `<窗>_reimport_<时分>` 窗
+    (设计如此, 该窗被 `data.EXCLUDE_DEFAULT` 排除在生产集外), 而登记是**追加式**的 ——
+    只补不删。重算几次后 `_derived_manifest.json` 里就攒了成百上千条指向已删目录的条目,
+    `_provenance.json` 的 raw-derived 计数随之虚增 (实测 1,740 → 3,444, 而盘上并没有多出这些件)。
+    台账是本包的"来源正本", 虚高等于说假话 ⇒ 每次生成台账前先 prune。
+    """
+    store_root = pathlib.Path(store_root)
+    cur = load(store_root)
+    files = cur.get('files') or {}
+    keep = {rel: v for rel, v in files.items() if (store_root / rel).exists()}
+    gone = len(files) - len(keep)
+    if gone:
+        cur['files'] = keep
+        cur['pruned'] = f'{time.strftime("%Y-%m-%d %H:%M")} 清理 {gone} 条不在盘的登记'
+        path_of(store_root).write_text(json.dumps(cur, ensure_ascii=False, indent=1), encoding='utf-8')
+    return gone
+
+
+def record(store_root, files: dict, by: str) -> pathlib.Path:
+    """把 {相对产物仓的路径: 构建器说明} 合并进登记 (幂等: 同路径后写覆盖先写)。
+
+    幂等很关键 —— 重跑摄入不该让登记无限膨胀; 同时**不删**别的构建器登记的条目
+    (振动摄入与厂家报告摄入是两个脚本, 各登各的)。"""
+    store_root = pathlib.Path(store_root)
+    cur = load(store_root)
+    entries = cur.get('files') or {}
+    for rel, builder in files.items():
+        entries[pathlib.Path(rel).as_posix()] = dict(builder=builder, by=by,
+                                                     at=time.strftime('%Y-%m-%d %H:%M:%S'))
+    cur = dict(note='产物来源自登记: 由构建脚本落盘后写入; _provenance.json 生成时把这些件记为 raw-derived',
+               at=time.strftime('%Y-%m-%d %H:%M:%S'), files=entries)
+    p = path_of(store_root)
+    p.parent.mkdir(parents=True, exist_ok=True)
+    p.write_text(json.dumps(cur, ensure_ascii=False, indent=1), encoding='utf-8')
+    return p

+ 21 - 16
src/ontology/maintenance.py

@@ -15,18 +15,15 @@ from .. import paths as P
 import re
 import re
 import json, os, pathlib, sys, datetime as _dt
 import json, os, pathlib, sys, datetime as _dt
 
 
-ROOT = pathlib.Path(__file__).resolve().parents[2]   # v2: 不再依赖 cwd (xzy 测试报告)
+ROOT = P.ROOT                     # 2026-09-16 统一: 原为 `parents[2]` 自推一份 (与本模块的 P.ROOT 同值,
+                                  # 但属"影子真源"; 同一文件里还并存自写的 disp()/_RAW_STR —— 一并归到 src/paths.py)
 ST = P.store()
 ST = P.store()
 ONT = P.ont()
 ONT = P.ont()
 CMS = P.cms()
 CMS = P.cms()
-# 原始件目录: 由 src.windscada.config 统一给 (env WINDSCADA_RUDONG_SRC / configs/serve.json raw_dir 可覆盖;
-# 默认 <安装目录>/data/raw)。2026-09-08 xzy 测试逮: 原为开发机绝对路径 /Users/yuanying/rudong/…,
+# 原始件目录: 唯一真源就是 src/paths.py 的 RAW_ROOT (env WINDSCADA_RUDONG_SRC / configs/serve.json
+# 的 raw_dir 最终都经它生效)。2026-09-08 xzy 测试逮: 原为开发机绝对路径 /Users/yuanying/rudong/…,
 # 在 Windows 上既非法也不存在, 页面照着抄必然补不了数据。
 # 在 Windows 上既非法也不存在, 页面照着抄必然补不了数据。
-try:
-    from src.windscada.config import RUDONG_SRC as _RAW_STR
-except Exception:
-    _RAW_STR = os.environ.get('WINDSCADA_RUDONG_SRC') or str(ROOT / 'data' / 'raw')
-RAW = pathlib.Path(_RAW_STR)
+RAW = P.RAW_ROOT
 # ★2026-09-11 用户令 A2 + 场站扫描辨识: 数据层源在 data/raw/<场站名称>/ 下。场站目录名不再写死在
 # ★2026-09-11 用户令 A2 + 场站扫描辨识: 数据层源在 data/raw/<场站名称>/ 下。场站目录名不再写死在
 #   代码里 —— config.farm() 扫 data/raw 的下一级目录辨识 (辨识依据见 station_note), 这里取结果用。
 #   代码里 —— config.farm() 扫 data/raw 的下一级目录辨识 (辨识依据见 station_note), 这里取结果用。
 try:
 try:
@@ -44,11 +41,11 @@ TECH = RAW / '西门子4.0技术资料'
 
 
 
 
 def disp(p) -> str:
 def disp(p) -> str:
-    """给人看的路径: 安装目录内的写成 <安装目录>/…, 其余给绝对路径; 分隔符随本机 (Windows 上是反斜杠).
+    """给人看的路径 (安装目录内 → `<安装目录>/…`) —— 2026-09-16 起直接转发 src/paths.py 的同名函数。
+
+    原来本模块自己写了一份 (语义相同但少一步 resolve), 与 P.disp 并存 = 同一份"显示规则"两个实现;
     2026-09-09 用户看页面反馈: 直接摆一串本机绝对路径, 客户读不出该往哪放。"""
     2026-09-09 用户看页面反馈: 直接摆一串本机绝对路径, 客户读不出该往哪放。"""
-    p = pathlib.Path(p)
-    try: return str(pathlib.Path('<安装目录>') / p.relative_to(ROOT))
-    except ValueError: return str(p)
+    return P.disp(p)
 
 
 
 
 def _mtime(p):
 def _mtime(p):
@@ -114,11 +111,19 @@ SOURCES = [
          说明='SGS + 华标两家, 报告按台号/部件分目录; 文件名自带日期/台号/部件/sample_id; '
          说明='SGS + 华标两家, 报告按台号/部件分目录; 文件名自带日期/台号/部件/sample_id; '
               '新一轮到货后重跑, 时效胶囊自动翻绿'),
               '新一轮到货后重跑, 时效胶囊自动翻绿'),
     dict(层='数据', 名='CMS 振动评估报告', 产物=CMS / '报告_CMS振动状态评估报告_*.md', 时间列=None,
     dict(层='数据', 名='CMS 振动评估报告', 产物=CMS / '报告_CMS振动状态评估报告_*.md', 时间列=None,
-         位置=disp(CMS), 摄入='windcms analyze (振动线会话)',
-         频率='月/事件驱动', 责任='振动线', 说明='设备状态五级来源; 观澜自动取最新版不写死日期'),
+         位置=P.disp_dir(STATION / 'windcms'),
+         摄入='python scripts/vib_raw_build.py  (索引→谱; 重生成 CMS 报告用 --with-report, 见 docs 振动数据接入 §3b)',
+         频率='月/事件驱动', 责任='振动线',
+         说明='本目录两类源件: ①CMS 原始测量导出 (Brande TCM *_decode.json → 索引/谱/六层链, 可重算); '
+              '②厂商月度评估报告 (上海电气 12 份用印版 PDF 实测为扫描件无文本层 ⇒ 只作归档证据, 不进判级)。'
+              '设备状态五级来源; 观澜自动取最新版不写死日期'),
     dict(层='数据', 名='振动线 handoff', 产物=P.m5() / 'handoff_vibration_v2.json', 时间列=None,
     dict(层='数据', 名='振动线 handoff', 产物=P.m5() / 'handoff_vibration_v2.json', 时间列=None,
-         位置=disp(P.m5()), 摄入='振动线出件 → fusion 自动读',
-         频率='事件驱动', 责任='振动线', 说明='定谳链判级 (候选以上); 与 CMS 报告是两条轴, 取严合并'),
+         位置=P.disp_dir(STATION / 'm5_cms_tcm'),
+         摄入='振动线出件 → fusion 自动读 (现场正本落本目录则优先采用)',
+         频率='事件驱动', 责任='振动线',
+         说明='定谳链判级 (候选以上); 与 CMS 报告是两条轴, 取严合并。本目录放 TCM 侧深度分析报告 '
+              '(大生科技) 与现场给的接口正本 (handoff_vibration_v2.json / component_history.json); '
+              '包内那份 json 目前是 shipped 快照 (振动线分支的产物未随包), 现场给正本后即改由现场件驱动'),
     # ── 机理层 ──
     # ── 机理层 ──
     dict(层='机理', 名='报警码表与处置手册', 产物=ONT / 'objects.json', 时间列=None,
     dict(层='机理', 名='报警码表与处置手册', 产物=ONT / 'objects.json', 时间列=None,
          位置=disp(TECH / '故障处理/故障处理手册.xlsx'), 摄入='python -m src.ontology.kb_ingest',
          位置=disp(TECH / '故障处理/故障处理手册.xlsx'), 摄入='python -m src.ontology.kb_ingest',

+ 18 - 0
src/paths.py

@@ -107,6 +107,24 @@ def pitch(name: str | None = None) -> pathlib.Path:
     return out_root(name) / 'pitch'
     return out_root(name) / 'pitch'
 
 
 
 
+def paradigm(name: str | None = None) -> pathlib.Path:
+    """范式实验件 (E3/E5/E8 底稿) —— 事实契约的输入之一 (2026-09-16 补: 原先直接用
+    `ROOT/'outputs'/'rudong'/'paradigm_r1'` 拼, 既写死场名又绕过了本模块)。"""
+    return out_root(name) / 'paradigm_r1'
+
+
+def report_dir(name: str | None = None) -> pathlib.Path:
+    """报告交付件目录 (`交接_振动→状态评估报告_*.md` / `现场单_*.md`) —— 由振动线出件,
+    并被 `src/windcms/config.py` 的 knowledge_docs 引用 (2026-09-16 补: 该目录在 v0.2.0 里
+    没有明确归属, 一直以 `out_root()/'report'` 的裸拼形式出现)。"""
+    return out_root(name) / 'report'
+
+
+def cloud(name: str | None = None) -> pathlib.Path:
+    """可上云面孔 (脱敏后的契约/派生件/页面) —— `scripts/guanlan_cloud_*.py` 的落点。"""
+    return guanlan(name) / 'cloud'
+
+
 def contract(name: str | None = None) -> pathlib.Path:
 def contract(name: str | None = None) -> pathlib.Path:
     """场契约 (机型判据参数), 属 reference 侧, 不在 outputs。"""
     """场契约 (机型判据参数), 属 reference 侧, 不在 outputs。"""
     return REFERENCE / farm(name) / 'windscada_contract.yaml'
     return REFERENCE / farm(name) / 'windscada_contract.yaml'

+ 156 - 0
src/proc.py

@@ -0,0 +1,156 @@
+# -*- coding: utf-8 -*-
+"""子进程创建的统一口径 —— **不弹命令窗口** (2026-09-16 用户令: 启动/操作观澜时不弹命令窗口)。
+
+## 为什么要收敛到一处
+
+Windows 上"**无控制台的父进程** + 裸 spawn 一个控制台程序" = 系统给子进程**新建一个可见控制台窗口**。
+本项目里这类父进程很多:
+  · 网关 `guanlan_gateway.py` 自己是被 `DETACHED_PROCESS` 起来的 (无控制台) → 它调的 `tasklist` / `git` 会闪窗;
+  · 运维动作进程 (`scripts/_ops_launch.py` → `_ops_run.py`) 同样无控制台 → 从页面点"重算"时,
+    动作全过程都在一个可见窗口里跑 (最长 20 分钟), 页面每 2 s 轮询 `tasklist` 还会**反复闪窗**;
+  · 组件服务原先有的地方用 `DETACHED_PROCESS` (子进程干脆没有控制台), 有的地方什么都不加 (于是弹窗),
+    两种写法混用 —— 实测现场会攒下多个标题为 `.venv\\Scripts\\python.exe` 的黑窗。
+
+统一到本模块后: 只认 `NO_WINDOW` (`CREATE_NO_WINDOW`) 一种写法。
+★ `CREATE_NO_WINDOW` 与 `DETACHED_PROCESS` **互斥**, 不要叠加 —— 前者是"给一个没有窗口的控制台",
+  后者是"不给控制台"; 叠在一起行为依赖 Windows 版本。需要"子进程活过父进程"时用
+  `NEW_GROUP` (`CREATE_NEW_PROCESS_GROUP`) + 不共享控制台即可, 不需要 DETACHED。
+
+## 日志去哪了 (hide 窗口不等于看不见)
+
+窗口藏起来后, 子进程的 stdout/stderr 一律重定向到 `logs/<name>.log` (`spawn(log=…)`),
+`/ops` 页面也会显示任务日志尾巴 —— 排障路径不变, 只是不再靠一个黑窗。
+
+## 用法
+
+    from src.proc import spawn, run, NO_WINDOW
+    pid = spawn([py, 'scripts/x.py'], log=P.LOGS / 'x.log', env=e, cwd=ROOT)   # 后台, 不弹窗
+    r = run(['tasklist', '/FI', f'PID eq {pid}'], capture_output=True, text=True)  # 等待, 不弹窗
+"""
+from __future__ import annotations
+
+import os
+import pathlib
+import subprocess
+import sys
+
+WIN = os.name == 'nt'
+# CREATE_NO_WINDOW = 0x08000000: 给子进程一个**没有窗口**的控制台 (stdio 仍可重定向)
+NO_WINDOW = 0x08000000 if WIN else 0
+# CREATE_NEW_PROCESS_GROUP = 0x00000200: 子进程不受父进程 Ctrl-C 影响, 也不共享父的控制台事件
+NEW_GROUP = 0x00000200 if WIN else 0
+
+ROOT = pathlib.Path(__file__).resolve().parents[1]
+
+
+def flags(*, new_group: bool = True) -> int:
+    """本模块的统一 creationflags (非 Windows 返回 0, 调用方无需分支)。"""
+    f = NO_WINDOW
+    if new_group:
+        f |= NEW_GROUP
+    return f
+
+
+def _kw(kw: dict, *, new_group: bool = True) -> dict:
+    if WIN:
+        kw.setdefault('creationflags', flags(new_group=new_group))
+    else:
+        kw.setdefault('start_new_session', True)
+    return kw
+
+
+def _inherit_stdio(kw: dict) -> dict:
+    """把父进程**当前**的标准句柄显式交给子进程 (STARTF_USESTDHANDLES)。
+
+    为什么必须显式 (2026-09-16 实逮, 是本模块第一版引入的坑): Windows 上 `CREATE_NO_WINDOW` 会给
+    子进程新建一个"没有窗口的控制台", 而**新建控制台会把该进程的标准句柄重指到新控制台的缓冲区** ——
+    于是"父进程 stdout 已被重定向到日志文件"这件事, 在**孙子辈**就丢了。
+    现场表现: 从页面点"执行重算" → `_ops_launch`(显式 stdout=日志) → `_ops_run`(输出进了日志 ✔)
+    → `rebuild_all.py`(没显式句柄 → 输出掉进那个隐形控制台, 页面只剩"运行中"、日志里一个字都没有)。
+    显式传 `sys.stdout/stderr` 即可让整条链写进同一个日志; `sys.stdout is None` (pythonw) 时跳过。
+    """
+    if kw.get('capture_output'):
+        # ★ `capture_output=True` 与显式 `stdout=` **互斥**, 同给会抛 ValueError。
+        #   2026-09-16 实逮: 忘了这一步 → guanlan_ops.job_running() 里的 `tasklist` 每次都抛,
+        #   异常被 except 吞成 alive=False ⇒ **任何在跑的重算都被立刻改写成"被强杀"**,
+        #   页面显示完成、重算按钮重新可点(可能并发起两个重算)。
+        #   这里什么都不加就对了; 若调用方自己又传了 stdout, 让 Python 照旧抛错 (不替它吞)。
+        return kw
+    if 'stdout' not in kw:
+        s = sys.stdout
+        if s is not None and hasattr(s, 'fileno'):
+            try:
+                s.fileno()
+                kw['stdout'] = s
+            except Exception:
+                pass
+    if 'stderr' not in kw:
+        s = sys.stderr
+        if s is not None and hasattr(s, 'fileno'):
+            try:
+                s.fileno()
+                kw['stderr'] = s            # 不合并到 stdout: 让调用方自己决定 (spawn(log=) 才合并)
+            except Exception:
+                pass
+    return kw
+
+
+def spawn(cmd, log=None, env=None, cwd=None, *, new_group: bool = True, stdin_devnull: bool = True, **popen_kw):
+    """后台起一个**无窗口**子进程。`log` 给路径时把 stdout/stderr 追加进该文件。
+
+    `**popen_kw` 透传给 `Popen` —— 调用方要自己给 `stdout=`/`stderr=` 句柄时 (例如"日志由子进程
+    自己写、父进程只留 fd") 用得上; 给了就**不覆盖**它。没给则继承父进程当前的标准句柄 (见 _inherit_stdio)。
+    → Popen 对象 (拿 `.pid`)。
+    """
+    kw = dict(cwd=str(cwd or ROOT), env=env)
+    if stdin_devnull and 'stdin' not in popen_kw:
+        kw['stdin'] = subprocess.DEVNULL
+    fh = None
+    if log is not None and 'stdout' not in popen_kw:
+        log = pathlib.Path(log)
+        log.parent.mkdir(parents=True, exist_ok=True)
+        fh = open(log, 'ab')
+        kw['stdout'] = fh
+        kw['stderr'] = subprocess.STDOUT
+    kw.update(popen_kw)
+    _inherit_stdio(kw)
+    _kw(kw, new_group=new_group)
+    p = subprocess.Popen([str(c) for c in cmd], **kw)
+    if fh is not None:
+        # 父进程不一定等子进程结束; 句柄由子进程持有, 这里不要 close 掉它 —— 交给 GC/进程退出即可。
+        p._dsh_log = fh          # 仅作引用保存, 避免过早回收
+    return p
+
+
+def run(cmd, *, new_group: bool = True, **kw):
+    """前台等待的 `subprocess.run`, 但**不弹窗** (tasklist / git / 短命令都该走这里)。
+
+    ★ 文本解码请显式带 `errors='replace'`: `tasklist` 的输出是**控制台代码页**(中文 Windows = GBK),
+      而本进程可能是 PYTHONUTF8=1 起的 (默认文本编码 UTF-8) → 解码失败会让 `r.stdout` 变成 None,
+      调用方再 `str(pid) in r.stdout` 就 TypeError (2026-09-12 实逮过)。
+    ★ 未显式给 stdout/stderr 时继承父进程当前句柄 (见 _inherit_stdio) —— 否则 `CREATE_NO_WINDOW`
+      新建的控制台会把子进程输出"吸走", 日志里什么都看不到。
+    """
+    _inherit_stdio(kw)
+    _kw(kw, new_group=new_group)
+    return subprocess.run([str(c) for c in cmd], **kw)
+
+
+def run_text(cmd, **kw):
+    """`run` + text=True + errors='replace' (控制台输出来源的默认姿势)。"""
+    kw.setdefault('capture_output', True)
+    kw.setdefault('text', True)
+    kw.setdefault('errors', 'replace')
+    return run(cmd, **kw)
+
+
+def python_exe() -> str:
+    """跑脚本用的解释器 (venv 优先)。放这里是为了让调用方一处取值, 别再各写一遍。"""
+    try:
+        from src import paths as P
+        v = P.venv_python()
+        if v:
+            return str(v)
+    except Exception:
+        pass
+    return sys.executable

+ 29 - 2
src/sop/wrapup.py

@@ -154,9 +154,36 @@ if __name__ == '__main__':
 
 
     # 柱2 day_costs: cleaned parquet 场中位功率 × price.py 参数化价 (config 驱动 + per月季节 + provenance)
     # 柱2 day_costs: cleaned parquet 场中位功率 × price.py 参数化价 (config 驱动 + per月季节 + provenance)
     day_costs, provenance = None, None
     day_costs, provenance = None, None
-    pq = P.out_root(a.farm) / 'cleaned' / 'turbine.parquet'
+    # ★ 写读要配对 (2026-09-16 修): 写侧 `src/sop/contract_gate.py:131` 落的是 `cleaned/<section>.parquet`
+    #   (section 来自契约配置, 可能是 turbine / turbine_10min_cnt / …), 而这里原先**硬编码** 'turbine.parquet'
+    #   ⇒ section 不是 turbine 时 pq.exists() 恒假 → day_costs 静默变 None、柱2 悄悄降级成 INSUFFICIENT,
+    #   一句报错都没有。改为: 优先读 clean_gate.json 里记的 section, 否则取该目录下唯一的 parquet;
+    #   仍然取不到就**响亮**说明 (而不是让下游以为"这场的经济性算不出来")。
+    _cdir = P.out_root(a.farm) / 'cleaned'
+    pq, _why = None, ''
+    _gate = _cdir / 'clean_gate.json'
+    try:
+        if _gate.exists():
+            import json as _j
+            _sec = (_j.loads(_gate.read_text(encoding='utf-8')) or {}).get('section')
+            if _sec and (_cdir / f'{_sec}.parquet').exists():
+                pq = _cdir / f'{_sec}.parquet'
+    except Exception:
+        pass
+    if pq is None:
+        _cands = sorted(_cdir.glob('*.parquet')) if _cdir.is_dir() else []
+        if len(_cands) == 1:
+            pq = _cands[0]
+        elif _cands:
+            pq, _why = _cands[0], f'目录里有 {len(_cands)} 个 parquet, 取第一个 {_cands[0].name}'
+        else:
+            _why = f'{P.rel(_cdir)} 下没有 cleaned parquet (清洗闸未落盘? persist=False?)'
+    if pq is None:
+        print(f'[wrapup] 柱2 无 cleaned parquet → day_costs 保持 None: {_why}', flush=True)
+    elif _why:
+        print(f'[wrapup] 柱2 cleaned parquet 选取说明: {_why}', flush=True)
     vcfg = ROOT / 'configs' / a.farm / 'value_assumptions.yaml'
     vcfg = ROOT / 'configs' / a.farm / 'value_assumptions.yaml'
-    if pq.exists():
+    if pq is not None and pq.exists():
         import pandas as pd
         import pandas as pd
         import yaml
         import yaml
         from src.sop.price import build_price_by_hour
         from src.sop.price import build_price_by_hour

+ 11 - 0
src/windcms/config.py

@@ -95,6 +95,17 @@ FARMS = {
     }
     }
 }
 }
 
 
+def cms_out():
+    """CMS 产物目录的**唯一解析口** —— 读侧与写侧必须走同一个 (2026-09-16 统一)。
+
+    原状: 只有写侧认 `WINDCMS_OUT` (上面 FARMS 里的 `'out'`), 而读侧三处硬绑 `P.cms()` ——
+    `src/windscada/taxonomy.py` 的报告转录、`src/windscada/subsys/fusion.py::windcms_grades`、
+    `scripts/windscada_serve.py` 的产物指纹。设了该环境变量就会"**写到旁路、页面读生产**",
+    表现为"重算完了页面还是旧数", 且**一句报错都没有** (冻结构建自检就活在这个组合上)。
+    """
+    return Path(os.environ.get('WINDCMS_OUT') or P.cms())
+
+
 def farm(name='rudong'):
 def farm(name='rudong'):
     if name not in FARMS:
     if name not in FARMS:
         raise SystemExit(f'未知场 {name}; 可用: {list(FARMS)}')
         raise SystemExit(f'未知场 {name}; 可用: {list(FARMS)}')

+ 34 - 4
src/windcms/data.py

@@ -68,19 +68,46 @@ def load_alarm_counts(cfg):
 _META = {}
 _META = {}
 
 
 
 
+def _mtime_stamp(paths) -> str:
+    """一组文件的 (路径, mtime_ns, size) 摘要 —— 用来判断"产物是否换过"。
+
+    2026-09-16 用户令"重算后的产物可即时用于页面呈现": 本模块的 _META 是**进程内**缓存,
+    长期跑着的服务 (windcms serve / 观澜组件) 在摄入新窗之后会一直用旧 meta (实测同类问题
+    在 scripts/windscada_serve.py 的 _CACHE 上出现过)。把文件指纹并进缓存键即可自动失效,
+    不必重启服务。"""
+    h = hashlib.sha1()
+    for p in paths:
+        try:
+            st = p.stat()
+            h.update(f'{p}|{st.st_mtime_ns}|{st.st_size}\n'.encode('utf-8'))
+        except OSError:
+            h.update(f'{p}|MISSING\n'.encode('utf-8'))
+    return h.hexdigest()[:12]
+
+
 def spectra_meta(cfg):
 def spectra_meta(cfg):
-    """合并谱元数据: 合并库 (cfg.spectra_dir, 覆盖 w0127) + 各窗 spectra 库; 加列 store (npz 所在目录)."""
-    key = str(cfg['out'])
+    """合并谱元数据: 合并库 (cfg.spectra_dir, 覆盖 w0127) + 各窗 spectra 库; 加列 store (npz 所在目录).
+
+    缓存键 = 产物目录 + **各窗 index/spectra_meta 的指纹** → 重算/新摄入后自动失效 (见 _mtime_stamp)。"""
+    base = cfg['spectra_dir']
+    wins = windows(cfg)
+    stamp_files = [base / 'spectra_meta.parquet']
+    for _w, info in wins.items():
+        stamp_files.append(info['index'])
+        if info['spectra_dir']:
+            sd = info['spectra_dir']
+            stamp_files.append(sd.parent / 'spectra_meta.parquet')
+            stamp_files.append(sd / 'spectra_meta.parquet')
+    key = f'{cfg["out"]}|{_mtime_stamp(stamp_files)}'
     if key in _META:
     if key in _META:
         return _META[key]
         return _META[key]
     parts = []
     parts = []
-    base = cfg['spectra_dir']
     if (base / 'spectra_meta.parquet').exists():
     if (base / 'spectra_meta.parquet').exists():
         m = pd.read_parquet(base / 'spectra_meta.parquet')
         m = pd.read_parquet(base / 'spectra_meta.parquet')
         m['store'] = str(base)
         m['store'] = str(base)
         m['window'] = 'w0127'
         m['window'] = 'w0127'
         parts.append(m)
         parts.append(m)
-    for w, info in windows(cfg).items():
+    for w, info in wins.items():
         sd = info['spectra_dir']
         sd = info['spectra_dir']
         if not sd:
         if not sd:
             continue
             continue
@@ -92,6 +119,9 @@ def spectra_meta(cfg):
             parts.append(m)
             parts.append(m)
     M = pd.concat(parts, ignore_index=True) if parts else pd.DataFrame(columns=['turbine', 'sensor', 'meas_name', 'trigger_time', 'store'])
     M = pd.concat(parts, ignore_index=True) if parts else pd.DataFrame(columns=['turbine', 'sensor', 'meas_name', 'trigger_time', 'store'])
     M['trigger_time'] = M['trigger_time'].astype(str)
     M['trigger_time'] = M['trigger_time'].astype(str)
+    # 只留最近 4 个指纹键, 免得长期运行的服务把每代 meta 都攒在内存里
+    if len(_META) > 4:
+        _META.clear()
     _META[key] = M
     _META[key] = M
     return M
     return M
 
 

+ 30 - 1
src/windcms/pipeline.py

@@ -59,10 +59,39 @@ def window_year_gate(m5_root, window, max_age_days=400):
     return tmax
     return tmax
 
 
 
 
+def _zip_has_decoded_json(path) -> bool:
+    """压缩包里是否已经是**解码后**的 `*_decode.json` (2026-09-16)。
+
+    为什么要有这个判断: `detect_input()` 把任何 .zip/.rar/.7z 都归为 `tcm_archive`, 而那条路要调
+    `scripts/rudong_tcm_ingest_raw.py` (解原始 base64+XML 导出) —— 该脚本**没随包** ⇒ 现场直接
+    把 CMS 导出的 zip 指过来必然崩在 `FileNotFoundError`。但其实"zip 里就是 `*_decode.json`"时
+    根本不需要它: 本包的 `rudong_tcm_index.py` / `rudong_tcm_spectra.py` **支持把 zip 当 --root**
+    (直接按成员读), 走它们即可。这里只做"是不是解码后包"的判定, 不假装支持原始导出。
+    """
+    import zipfile
+    try:
+        with zipfile.ZipFile(path) as zf:
+            return any(not i.is_dir() and i.filename.endswith('_decode.json') for i in zf.infolist())
+    except Exception:
+        return False
+
+
 def ingest(cfg, path, window, log, env):
 def ingest(cfg, path, window, log, env):
     kind = detect_input(path)
     kind = detect_input(path)
+    if kind == 'tcm_archive' and _zip_has_decoded_json(path):
+        # 解码后的导出包 (zip 形态): 直接走索引/谱, 不必经缺失的 rudong_tcm_ingest_raw.py
+        print(f'[ingest] {Path(path).name}: 内含 *_decode.json → 按"已解码导出"处理 (zip 当 root)', flush=True)
+        kind = 'tcm_decoded_json'
     if kind == 'tcm_archive':
     if kind == 'tcm_archive':
-        _run('ingest:tcm_archive', [PY, str(ROOT / 'scripts/rudong_tcm_ingest_raw.py'), '--rar', str(path), '--window', window], env, log)
+        raw = ROOT / 'scripts/rudong_tcm_ingest_raw.py'
+        if not raw.is_file():
+            raise RuntimeError(
+                f'输入 {path} 是**原始** TCM 导出包 (rar/7z, 需先解 base64+XML), 而解码脚本 '
+                f'{raw.name} 未随包 (振动线分支产物)。两条可走的路: '
+                f'① 现场若能直接给"含 *_decode.json 的导出"(zip 或目录), 把它放 data/raw/<场>/windcms/ 下, '
+                f'用 scripts/vib_raw_build.py 摄入 (本包支持 zip 当 root); '
+                f'② 向振动线索取 scripts/rudong_tcm_ingest_raw.py 后再跑本步。')
+        _run('ingest:tcm_archive', [PY, str(raw), '--rar', str(path), '--window', window], env, log)
     elif kind == 'tcm_decoded_json':
     elif kind == 'tcm_decoded_json':
         out = cfg['m5'] / 'windows' / window
         out = cfg['m5'] / 'windows' / window
         out.mkdir(parents=True, exist_ok=True)
         out.mkdir(parents=True, exist_ok=True)

+ 5 - 1
src/windcms/plugins.py

@@ -369,8 +369,12 @@ def analyze(ctx, input='', window='', steps='', confirm=False):
     ctx['cfg']['out'].mkdir(parents=True, exist_ok=True)
     ctx['cfg']['out'].mkdir(parents=True, exist_ok=True)
     # stdout= 只用到这个文件的 fd (内容由子进程自己写), 所以这里不 with-close:
     # stdout= 只用到这个文件的 fd (内容由子进程自己写), 所以这里不 with-close:
     # 句柄留着, 免得父进程提前关掉; 显式 encoding 是为了不留"缺 encoding"的门禁告警。
     # 句柄留着, 免得父进程提前关掉; 显式 encoding 是为了不留"缺 encoding"的门禁告警。
+    # ★2026-09-16: 本函数跑在 **CMS 服务进程**里 (由 guanlan.py 以无窗口方式拉起, 自己没有可见控制台),
+    #   裸 Popen 会让 Windows 给全链分析**新建一个可见控制台窗口** (从页面点"分析"就会看到黑窗)。
+    #   统一走 src.proc.spawn (CREATE_NO_WINDOW)。
+    from src import proc as _proc
     _logf = open(logp, 'w', encoding='utf-8')
     _logf = open(logp, 'w', encoding='utf-8')
-    p = subprocess.Popen(cmd, cwd=str(ROOT), stdout=_logf, stderr=subprocess.STDOUT)
+    p = _proc.spawn(cmd, cwd=ROOT, stdout=_logf, stderr=subprocess.STDOUT)
     lock.write_text(str(p.pid), encoding='utf-8')
     lock.write_text(str(p.pid), encoding='utf-8')
     return dict(kind='text', data=dict(pid=p.pid, log=str(logp)), text=f'全链已在后台启动 (pid {p.pid}); 进度看 job_status; 日志 {logp}. 完成后产物/报告/知识库自动刷新 (界面需刷新页面).')
     return dict(kind='text', data=dict(pid=p.pid, log=str(logp)), text=f'全链已在后台启动 (pid {p.pid}); 进度看 job_status; 日志 {logp}. 完成后产物/报告/知识库自动刷新 (界面需刷新页面).')
 
 

+ 8 - 3
src/windscada/config.py

@@ -46,10 +46,15 @@ REQUIRED = ('name', 'n_turbines', 'turbines', 'src_10min', 'src_alarm', 'store',
 RUDONG_SRC = str(P.RAW_ROOT)
 RUDONG_SRC = str(P.RAW_ROOT)
 RAW_ROOT = P.RAW_ROOT
 RAW_ROOT = P.RAW_ROOT
 
 
-# 场站目录下的**约定子目录名** — 这五个名字是摄入侧的接口, 改名等于换接口, 要同步 README 与维护页
-STATION_SUBDIRS = ('scada_10min', 'scada_1min', '故障报警', '风机故障记录', '油样报告')
+# 场站目录下的**约定子目录名** — 这些名字是摄入侧的接口, 改名等于换接口, 要同步 README 与维护页
+# 振动侧两目录 (2026-09-12 用户令: 「CMS 振动评估报告」遵循 data/raw/如东/windcms、
+# 「振动线 handoff」遵循 data/raw/如东/m5_cms_tcm) —— 它们此前只在 outputs/ 下以**产物**形式存在,
+# 源件不在 data/raw, 于是"从零重算"时振动侧无源可算。加进来后扫描/落位/数据层页都能认它们。
+STATION_SUBDIRS = ('scada_10min', 'scada_1min', '故障报警', '风机故障记录', '油样报告',
+                   'windcms', 'm5_cms_tcm')
 _SRC_KEYS = (('src_10min', 'scada_10min'), ('src_1min', 'scada_1min'), ('src_alarm', '故障报警'),
 _SRC_KEYS = (('src_10min', 'scada_10min'), ('src_1min', 'scada_1min'), ('src_alarm', '故障报警'),
-             ('src_workorder', '风机故障记录'), ('src_oil', '油样报告'))
+             ('src_workorder', '风机故障记录'), ('src_oil', '油样报告'),
+             ('src_windcms', 'windcms'), ('src_m5', 'm5_cms_tcm'))
 
 
 _BUILTIN = {
 _BUILTIN = {
     'rudong': {
     'rudong': {

+ 2 - 1
src/windscada/subsys/fusion.py

@@ -322,7 +322,8 @@ def windcms_grades():
         return _WCMS
         return _WCMS
     _WCMS = {}
     _WCMS = {}
     try:
     try:
-        rp = sorted((P.cms()).glob('报告_CMS振动状态评估报告_*.md'))[-1]
+        from src.windcms.config import cms_out     # 与写侧同一解析口 (WINDCMS_OUT; 见 windcms.config.cms_out)
+        rp = sorted(cms_out().glob('报告_CMS振动状态评估报告_*.md'))[-1]
         md = rp.read_text(encoding='utf-8')
         md = rp.read_text(encoding='utf-8')
         for l in md.split('## 附录 A')[1].splitlines():
         for l in md.split('## 附录 A')[1].splitlines():
             if not l.startswith('| WTG'):
             if not l.startswith('| WTG'):

+ 8 - 4
src/windscada/subsys/pitch.py

@@ -29,9 +29,12 @@ ALARM_FAM = {  # M4b 报警轴 (windscada alarms.parquet; 分册监测量对齐)
 }
 }
 
 
 
 
-def _alarm_counts(win_start, win_end):
+def _alarm_counts(win_start, win_end, store=None):
     import re as _re
     import re as _re
-    ap = P.store() / 'alarms.parquet'
+    # ★store 显式传入 (2026-09-16 修): 原写 `P.store()` —— 它跟的是环境变量 WINDSCADA_FARM,
+    #   而调用方 registry(cfg) 拿的是**显式选定的场**; 多场部署下这里会读到 rudong 的 alarms.parquet
+    #   而其余输入都来自当前场 ⇒ 跨场串数据, 且因为"文件存在、列名兼容"而完全不报错。
+    ap = (pathlib.Path(store) if store else P.store()) / 'alarms.parquet'
     if not ap.exists():
     if not ap.exists():
         return None
         return None
     al = pd.read_parquet(ap)
     al = pd.read_parquet(ap)
@@ -43,7 +46,8 @@ def _alarm_counts(win_start, win_end):
     return out
     return out
 
 
 
 
-def registry():
+def registry(cfg=None):
+    """变桨面登记表。`cfg` 给定时, 其 `store` 决定读哪个场的 alarms (见 _alarm_counts 的注释)。"""
     daily, zero = load()
     daily, zero = load()
     daily['date'] = pd.to_datetime(daily['date']).dt.date
     daily['date'] = pd.to_datetime(daily['date']).dt.date
     cur = _cur_win(daily).copy()
     cur = _cur_win(daily).copy()
@@ -73,7 +77,7 @@ def registry():
             zsum[t] = dict(months=int(len(g)), dev_med=float(g['zero_dev'].median()), last=float(g['zero_dev'].iloc[-1]),
             zsum[t] = dict(months=int(len(g)), dev_med=float(g['zero_dev'].median()), last=float(g['zero_dev'].iloc[-1]),
                            n_alarm=int((g['zero_dev'].abs() >= ZERO_ALARM).sum()), last_month=str(g['month'].iloc[-1]))
                            n_alarm=int((g['zero_dev'].abs() >= ZERO_ALARM).sum()), last_month=str(g['month'].iloc[-1]))
     win_end = daily['date'].max(); win_start = pd.Timestamp(win_end) - pd.Timedelta(days=60)
     win_end = daily['date'].max(); win_start = pd.Timestamp(win_end) - pd.Timedelta(days=60)
-    ac = _alarm_counts(win_start, win_end)
+    ac = _alarm_counts(win_start, win_end, store=(cfg or {}).get('store') if cfg else None)
     rows, detail = [], {}
     rows, detail = [], {}
     for t, r in fleet.iterrows():
     for t, r in fleet.iterrows():
         d = dict(维度={}, 依据={})
         d = dict(维度={}, 依据={})

+ 2 - 1
src/windscada/taxonomy.py

@@ -211,7 +211,8 @@ def system_matrix(cfg=None):
             _WCOL = {'主轴承': ('主轴承前', '主轴承后'), '齿轮箱': ('齿轮箱',), '发电机': ('发电机',)}
             _WCOL = {'主轴承': ('主轴承前', '主轴承后'), '齿轮箱': ('齿轮箱',), '发电机': ('发电机',)}
             _ORD = ['危险', '报警', '良好', '优秀', '不可判']
             _ORD = ['危险', '报警', '良好', '优秀', '不可判']
             import pathlib as _pl, re as _re3
             import pathlib as _pl, re as _re3
-            _wmd = sorted(P.cms().glob('报告_CMS振动状态评估报告_*.md'))
+            from src.windcms.config import cms_out as _cms_out     # 读写同一口 (WINDCMS_OUT 不能被读侧忽略)
+            _wmd = sorted(_cms_out().glob('报告_CMS振动状态评估报告_*.md'))
             _wdate = _wmd[-1].stem.rsplit('_', 1)[-1] if _wmd else '?'
             _wdate = _wmd[-1].stem.rsplit('_', 1)[-1] if _wmd else '?'
             _wpart = {}
             _wpart = {}
             if _wmd:
             if _wmd:

+ 45 - 1
start.bat

@@ -1,7 +1,38 @@
 @echo off
 @echo off
+rem ============================================================================
+rem  观澜 start.bat
+rem  2026-09-16 用户令: 启动观澜系统时"弹出的命令窗口"改为不弹出方式。
+rem
+rem  默认 (双击本文件)      : 走**无窗口**启动 —— 交给 start_hidden.vbs (窗口样式 0),
+rem                           后台起服务、等就绪、自动打开浏览器, 不留任何命令窗口。
+rem  start.bat console      : 前台模式 (旧行为) —— 在当前窗口里跑 serve 并实时打印日志,
+rem                           只在排障时用; 关掉这个窗口等于停掉 serve。
+rem  start.bat help         : 说明。
+rem
+rem  为什么分开: 各组件服务本来就是 DETACHED_PROCESS 起的 (不弹窗), 唯一会弹的就是
+rem  "在控制台里前台跑 guanlan.py serve" 这件事本身。排障又确实需要看见实时输出,
+rem  所以保留一个显式的 console 档, 而不是把日志彻底藏掉。
+rem ============================================================================
 chcp 65001 >nul
 chcp 65001 >nul
 set PYTHONUTF8=1
 set PYTHONUTF8=1
 cd /d "%~dp0"
 cd /d "%~dp0"
+
+if /i "%~1"=="help" goto help
+if /i "%~1"=="console" goto console
+if /i "%~1"=="-c" goto console
+
+if not exist ".venv\Scripts\pythonw.exe" (
+  echo [X] Not installed yet: .venv\Scripts\pythonw.exe not found
+  echo     Run install.bat first.
+  pause
+  exit /b 2
+)
+
+rem 无窗口启动: wscript 以隐藏窗口方式调用 pythonw.exe 跑启动器, 本窗口立即退出。
+wscript.exe //nologo "%~dp0start_hidden.vbs"
+exit /b 0
+
+:console
 if not exist ".venv\Scripts\python.exe" (
 if not exist ".venv\Scripts\python.exe" (
   echo [X] Not installed yet: .venv\Scripts\python.exe not found
   echo [X] Not installed yet: .venv\Scripts\python.exe not found
   echo     Run install.bat first. If install printed errors, send that screen back
   echo     Run install.bat first. If install printed errors, send that screen back
@@ -9,6 +40,9 @@ if not exist ".venv\Scripts\python.exe" (
   pause
   pause
   exit /b 2
   exit /b 2
 )
 )
+echo [i] Foreground mode - this window stays open and shows live logs.
+echo     Close it to stop the server; for no-window start just run start.bat with no argument.
+echo.
 ".venv\Scripts\python.exe" guanlan.py serve
 ".venv\Scripts\python.exe" guanlan.py serve
 if errorlevel 1 (
 if errorlevel 1 (
   rem serve exits 1 in two different cases: the gateway never came up, or the gateway is
   rem serve exits 1 in two different cases: the gateway never came up, or the gateway is
@@ -27,7 +61,7 @@ if errorlevel 1 (
   echo [!] Gateway is up, but at least one module is degraded ^(the [DOWN] lines above^).
   echo [!] Gateway is up, but at least one module is degraded ^(the [DOWN] lines above^).
   echo     Usual cause: local Ollama is not running or has no models pulled, which only
   echo     Usual cause: local Ollama is not running or has no models pulled, which only
   echo     disables the Q^&A / local-review pages. Every other page works.
   echo     disables the Q^&A / local-review pages. Every other page works.
-  echo     Details: logs\ for the degraded module.
+  echo     Details: logs/ for the degraded module.
 )
 )
 start "" http://127.0.0.1:28084/
 start "" http://127.0.0.1:28084/
 echo.
 echo.
@@ -35,3 +69,13 @@ echo [i] Ops console - stop/start services, rebuild, clear products:
 echo     http://127.0.0.1:28084/ops
 echo     http://127.0.0.1:28084/ops
 echo     It stays available while component services are stopped (only the gateway must be up);
 echo     It stays available while component services are stopped (only the gateway must be up);
 echo     buttons follow real state and the backend refuses impossible actions (HTTP 409).
 echo     buttons follow real state and the backend refuses impossible actions (HTTP 409).
+exit /b 0
+
+:help
+echo Usage:
+echo   start.bat            no-window start (default, recommended): hidden services + browser
+echo   start.bat console    foreground start with live logs (troubleshooting)
+echo   stop.bat             stop everything
+echo   check.bat            self-check
+echo Logs: logs\serve.log (no-window mode), logs\start_hidden.log (launcher trace)
+exit /b 0

TEMPAT SAMPAH
start_hidden.vbs