Просмотр исходного кода

交付包 v0.4.0: 打包不含 输入数据/产物/日志 (用户令 2026-09-16)

zhouyang.xie 3 недель назад
Родитель
Сommit
d7e253bfc0
56 измененных файлов с 3733 добавлено и 1455 удалено
  1. 26 17
      README_先读我.txt
  2. 0 323
      _修复记录_20260911/README.md
  3. 0 249
      _修复记录_20260911/fix_guanlan_bom.py
  4. 0 140
      _修复记录_20260911/fix_guanlan_pythonpath.py
  5. 0 84
      _修复记录_20260911/fix_guanlan_startbat.py
  6. 0 67
      _修复记录_20260911/restore_release_rudong.py
  7. 0 50
      _修复记录_20260911/scan_portal_links.py
  8. 0 44
      _修复记录_20260911/verify_bug2.py
  9. 0 61
      _修复记录_20260911/verify_guanlan_simsys.py
  10. 238 0
      docs/振动数据接入_v0.1.md
  11. 69 9
      docs/数据目录结构与落位约定_v0.2.md
  12. 3 2
      docs/移植与独立运行_v0.1.md
  13. 294 0
      docs/系统设计说明.md
  14. 4 3
      docs/说明书_观澜如东样板v2_v0.2.md
  15. 48 41
      docs/重算操作手册_v0.1.md
  16. 24 9
      docs/重算缺口与补件清单_v0.1.md
  17. 47 3
      guanlan.py
  18. 1 1
      release/portal_src/README.md
  19. 5 4
      release/portal_src/manifest.json
  20. 4 4
      release/portal_src/shell.html
  21. 7 7
      scripts/_ops_launch.py
  22. 9 1
      scripts/_ops_run.py
  23. 4 1
      scripts/_ops_start_and_open.py
  24. 20 1
      scripts/check_transferable.py
  25. 8 3
      scripts/guanlan_facts_contract.py
  26. 14 4
      scripts/guanlan_gateway.py
  27. 86 63
      scripts/guanlan_ops.py
  28. 150 0
      scripts/guanlan_start_hidden.py
  29. 17 4
      scripts/ingest_ops_2025.py
  30. 349 0
      scripts/inventory_products.py
  31. 61 21
      scripts/pack_dist.py
  32. 133 33
      scripts/place_raw_data.py
  33. 37 0
      scripts/portal_build.py
  34. 78 13
      scripts/products_restore_missing.py
  35. 55 130
      scripts/products_state.py
  36. 33 2
      scripts/rebuild_all.py
  37. 383 0
      scripts/rudong_tcm_index.py
  38. 334 0
      scripts/rudong_tcm_spectra.py
  39. 271 0
      scripts/vib_raw_build.py
  40. 357 0
      scripts/vib_reports_build.py
  41. 113 27
      scripts/windscada_serve.py
  42. 82 0
      src/derived_manifest.py
  43. 21 16
      src/ontology/maintenance.py
  44. 18 0
      src/paths.py
  45. 156 0
      src/proc.py
  46. 29 2
      src/sop/wrapup.py
  47. 11 0
      src/windcms/config.py
  48. 34 4
      src/windcms/data.py
  49. 30 1
      src/windcms/pipeline.py
  50. 5 1
      src/windcms/plugins.py
  51. 8 3
      src/windscada/config.py
  52. 2 1
      src/windscada/subsys/fusion.py
  53. 8 4
      src/windscada/subsys/pitch.py
  54. 2 1
      src/windscada/taxonomy.py
  55. 45 1
      start.bat
  56. BIN
      start_hidden.vbs

+ 26 - 17
README_先读我.txt

@@ -1,8 +1,12 @@
-观澜·如东样板 v2 0.2.0 — 单包全量离线应用 (Windows / Linux / macOS)
+观澜·如东样板 v2 0.4.0 — 单包全量离线应用 (Windows / Linux / macOS)
 
 <PY> = Windows 的 .venv\Scripts\python.exe, 或 Linux·macOS 的 .venv/bin/python;
        命令里路径写 / 即可两平台通用; Windows 也可以只双击 .bat, 不必敲命令。
 
+★ 本包**不含 输入数据(data/) · 产物(outputs/) · 日志(logs/, run/)**(用户令 2026-09-16):
+  解压安装后页面会显示"无产物"; 把现场数据包放好后跑一次重算(第四节), 页面立刻有数。
+  随包内容 = 程序/配置/门户与三维资产/交付文档/离线依赖轮子/便携运行时 (没有数据与产物)。
+
 一、安装 (约 10 分钟; 安装目录不要有空格和中文)
   Windows : 解压到 D:\guanlan\app → 双击 install.bat        (自动选 Python → 建 .venv → 包内轮子离线装依赖)
   Linux   : 解压到 /opt/guanlan  → sh install.sh            (麒麟 V10 / 统信 UOS 20 / Ubuntu / CentOS / macOS;
@@ -17,26 +21,31 @@
   Linux   : <PY> guanlan.py check → <PY> guanlan.py serve           停止: <PY> guanlan.py stop
   三个入口: 门户 http://127.0.0.1:28084/ · 工作台 /detail/ · 运维控制台 /ops
 
-三、运维控制台 /ops — 一个按钮一件事: 停/启服务 · 执行重算 · 清除产物
-  按钮按真实状态启用(不能做的灰, 后端也拒绝, 直接打 API 返回 409); "启动服务"起来后自动打开门户;
-  "停止组件服务"保留控制台本身; 同时刻只允许一个动作, 页面上有进度/退出码/日志尾巴。
-  清除产物 = 挪到 _products_off(可恢复), 不动 data/raw、release、logs。(控制台与重算链只用标准库, 两平台同源)
+三、重算与产物 (门户菜单「数据重算」, 或直接开 http://127.0.0.1:28084/ops)
+  一个按钮一件事: 停/启服务 · 执行重算 · 清除产物。按钮按真实状态启用(不能做的灰, 后端也拒绝,
+  直接打 API 返回 409); "启动服务"起来后自动打开门户; "停止组件服务"保留控制台本身; 同时刻只允许
+  一个动作, 页面上有进度/退出码/实时日志尾巴。重算进行中按钮按设计禁用 —— 那不是坏了。
+  ★ 清除产物 = **直接删除, 不留备份**(用户令 2026-09-16; 旧的 --on 还原已移除), 不动 data/raw、release;
+    要补回"包内没有生成端"的随包件: <PY> scripts/products_restore_missing.py --stash <交付包.zip>
 
-四、换数据后重算 (一条命令; 控制台点"执行重算"等价)
+四、放数据后重算 (一条命令; 控制台点"执行重算"等价)
   <PY> scripts/place_raw_data.py --src <现场包目录> --scope full      放数据(同尺寸自动跳过, 可反复跑)
   <PY> scripts/rebuild_all.py [--skip-scada] [--src <现场包目录>]     默认含 SCADA 侧约 15 分钟;
        只换台账类数据加 --skip-scada(约 2 分钟); --dry-run 只看计划; 完事自动重启组件服务。
   算成功判据: /detail/ 左栏「系统维护」两屏, 或 <PY> -m src.ontology.maintenance
-  本包锚点: 报警 39211 · 工单 5876 · 油样 404 · temp_monthly 19494 · 本体 9702 对象(审计 0 问题)
+  重算后应达锚点: 报警 39211 · 工单 5876 · 油样 404 · temp_monthly 19494 · 本体 9702 对象(审计 0 问题)
 
 五、必读 (如实)
-  · 页面的数来自 outputs/rudong 下的产物, 产物由 data/raw 下的原始件算出来(下一级目录名=场站名, 扫描辨识)。
-  · 随包里有一批产物**没有生成端**(pitch_daily、pc_monthly_bins、CMS/TCM 链等), 重算后由随包件补齐;
-    逐件来源见 outputs/rudong/_provenance.json (raw-derived 20 件 / shipped 568 件)。
-  · 某页没数据: 先看控制台「最近一次动作」的退出码与 logs/, 再看 docs/重算缺口与补件清单_v0.1.md。
-
-包内含: 程序与配置 / 如东分析产物(含 _provenance.json 来源台账) / 门户与仿真 / 三维资产 release/viewer /
-        治理清单交付件 release/如东(客户交付物勿外传) / 离线依赖轮子 / 便携运行时
-
-详细说明: docs/说明书_观澜如东样板v2_v0.2.md · docs/重算操作手册_v0.1.md
-          docs/数据目录结构与落位约定_v0.2.md · docs/重算缺口与补件清单_v0.1.md · _修复记录_20260911/README.md
+  · 页面的数来自 outputs/<场站> 下的产物, 产物由 data/raw 下的原始件算出来(下一级目录名=场站名, 扫描辨识)。
+    本包两样都没带: 没有 data/raw 时重算会"没有源件可算", 没有 outputs/ 时页面显示"无产物"。
+  · 有一批产物**没有生成端**(pitch_daily、pc_monthly_bins、CMS/TCM 链等): 从交付包按需补齐
+    (products_restore_missing.py --stash <交付包.zip>), 补齐后逐件来源见 outputs/<场站>/_provenance.json。
+  · 某页没数据: 先看控制台的退出码与 logs/, 再看 docs/重算缺口与补件清单_v0.1.md。
+
+包内含: 程序与配置 / 门户与仿真 / 三维资产 release/viewer / 治理清单交付件 release/如东(客户交付物勿外传) /
+        离线依赖轮子 wheels/ / 便携运行时 vendor/ / 交付文档 docs/
+包内不含: data/(输入数据) · outputs/(产物) · logs/ run/(日志) · .venv(安装时重建) · .git(版本库)
+
+详细说明: docs/系统设计说明.md · docs/重算操作手册_v0.1.md
+          docs/数据目录结构与落位约定_v0.2.md · docs/重算缺口与补件清单_v0.1.md
+          docs/振动数据接入_v0.1.md(振动侧 CMS/TCM 落位与摄入) · docs/说明书_观澜如东样板v2_v0.2.md

+ 0 - 323
_修复记录_20260911/README.md

@@ -1,323 +0,0 @@
-# 观澜·如东样板 v2 (v0.2.0) 启动报错修复记录 — 2026-09-11
-
-本目录**不属于原始交付包**, 是现场修完之后留下的记录: 每个脚本都可重复运行 (幂等),
-用来复现 "改了什么" 或在别的机器/另一份拷贝上重放同一批修复。
-
-## 现场现象
-
-```
-powershell -ExecutionPolicy Bypass -File install.ps1     -> 装完最后一步崩:
-    json.decoder.JSONDecodeError: Unexpected UTF-8 BOM (decode using utf-8-sig)
-start.bat                                                -> 同样崩; 随后
-    [X] Server did not start (reason above). Not opening the browser
-```
-
-## 根因 (5 条, 前 3 条是同一个坑的不同后果)
-
-| # | 位置 | 问题 |
-|---|------|------|
-| 1 | `install.ps1` 第 4 步 `Set-Content configs\serve.json -Encoding UTF8` | **Windows PowerShell 5.1 的 `-Encoding UTF8` 写出带 BOM 的 UTF-8**; `guanlan.py` 用裸 `utf-8` 读 -> `JSONDecodeError`。启动器在解析配置时就死了, 所以 `check`/`serve` 全跑不起来。 |
-| 2 | 同一行的 `Get-Content configs\serve.json -Raw` | PS 5.1 对**无 BOM** 的 UTF-8 按 ANSI 代码页 (中文机 = GBK) 解码, 于是把文件里的中文注释读成乱码, 再 `Set-Content` 写回 -> `_note` / `_note_raw` 两个字段永久损坏 (部分字节被替换成 `?`, 不可逆)。 |
-| 3 | 机器全局 `PYTHONPATH=D:\Program Files\Python\Lib\site-packages;` | ① pip 认为系统那套包 "已满足", **没把传递依赖装进 `.venv`** -> venv 缺 `urllib3`/`polars`/`pyyaml`/`jinja2`/`python-dotenv` 等, 换个没设 PYTHONPATH 的 shell 直接 `ModuleNotFoundError`; ② `PYTHONPATH` 排在 sys.path 前面, 把 pin 住的版本顶掉 (实测 `requests` 2.33.0 顶掉 2.34.2)。 |
-| 4 | `wheels\win_amd64\` | 缺 `colorama` 轮子 (`tqdm` 在 Windows 上的依赖)。之前被根因 3 掩盖: pip 看到系统已有 colorama 就不下载, 于是离线轮子集不全; 一旦 PYTHONPATH 摘干净, `pip install --no-index` 直接 `ERROR: No matching distribution found for colorama`。 |
-| 5 | `release/sim_sys_server.py` 第 36 行 `from scrub_rules import scrub, residual` | **`scrub_rules.py` 根本没随包发出** (全盘搜不到, zip 里也没有) -> `/sim/sys/ 仿真·四系统合页` 502, `logs/sim_sys.log` 里是 `ModuleNotFoundError`; 网关因此 `degraded`, `guanlan.py serve` 退出码 1。 |
-
-## 改了什么
-
-| 文件 | 改动 |
-|------|------|
-| `guanlan.py` | 新增 `jload()` = `utf-8-sig` 读 JSON (兼容有/无 BOM), 5 处配置读取全部改走它; 配置坏了只打印提示并退回内置默认端口, 不再抛栈。新增 `_hermetic()`: 启动时把 `PYTHONPATH` 从 `sys.path` 与环境里摘掉 (子进程继承干净环境)。 |
-| `install.ps1` | 第 4 步不再用 PS 的 cmdlet 碰配置, 改由 `.venv` 里的 Python 读写 (读 `utf-8-sig` / 写无 BOM 的 UTF-8), 并加注释说明为什么不能改回去; 顶部加 `Remove-Item Env:PYTHONPATH`, 让 pip 老老实实把包装进 `.venv`。文件仍保持 **UTF-8 带 BOM + CRLF** (PS 5.1 解码中文的前提)。 |
-| `install.sh` | 同样加 `unset PYTHONPATH` (保持 LF/无 BOM)。 |
-| `configs/serve.json` | 去掉 BOM; 还原 `_note` / `_note_raw` 两段中文注释 (依据残留可逆部分 + 说明书 §7 的措辞)。 |
-| `wheels\win_amd64\colorama-0.4.6-py2.py3-none-any.whl` | 补上缺失的轮子, 离线安装才完整。 |
-| `release/scrub_rules.py` | **恢复缺失文件**: 不新写任何脱敏规则, 只把包内唯一那份规则表 `src/windscada/deid_public.py` 转出 `scrub` / `residual`(= `audit`) / `scrub_or_die`。若拿到交付方原版, 直接覆盖。 |
-| `start.bat` | `guanlan.py serve` 的退出码 1 有两种含义 (网关没起来 / 网关起来了但有模块降级)。现在先探一次 `/healthz` 区分: 真没起来才报 `[X]` 并停下; 起来了但有降级则打 `[!]` 说明, 浏览器照常打开。 |
-
-## 备份与回滚
-
-原文件都留了副本, 直接改名覆盖即可回到改前状态:
-
-```
-guanlan.py.bak-bomfix      install.ps1.bak-bomfix     configs\serve.json.bak-bomfix
-guanlan.py.bak-pyfix       install.ps1.bak-pyfix      install.sh.bak-pyfix
-start.bat.bak-healthz
-```
-
-## 重放 / 验证
-
-```
-.venv\Scripts\python.exe fix_guanlan_bom.py          # 1 2 (BOM + 乱码)
-.venv\Scripts\python.exe fix_guanlan_pythonpath.py   # 3
-.venv\Scripts\python.exe fix_guanlan_startbat.py     # start.bat 判定
-.venv\Scripts\python.exe verify_guanlan_simsys.py    # 5 个仿真页 200 + 脱敏回扫干净
-```
-
-三个 `fix_*` 都是幂等的 (已改过会打印 "已修过")。`wheels\colorama*.whl` 与
-`release\scrub_rules.py` 属于新增文件, 没有对应的 `fix_*` 脚本。
-
-## bug-2 · `#documents` 页「如东液压系统治理清单 · 网页版」报 `{"err":"not found"}`
-
-**现象**: 门户 `http://127.0.0.1:28084/#documents` 里那个 iframe 显示
-`{"err": "not found", "path": "/如东/如东治理清单_交付_20260901/02_治理清单/如东_液压系统治理清单_v1.2_2026-09-01.html"}`。
-
-**根因**: v0.2.0 发行包里**没有 `release/如东/`** —— 0.2.0 的 zip 里 `release/` 下只有 `viewer/`,
-而门户那个 iframe 指的就是 `release/如东/.../_液压系统治理清单_v1.2_2026-09-01.html`;
-网关 `guanlan_gateway.py::_static()` 在 `release/` 里找不到文件就回那段 404 JSON
-(见 `_static()` 里 `return self._json(dict(err="not found", path=path), 404)`)。
-顺带核过: 门户里指向 `release/如东/` 的引用**只有这一条**, 所以不是路径写错, 是交付件没随 v0.2.0 发出来。
-
-**修复**: 从本机的 v0.1.0 全量包 `<临时目录>/guanlan-rudong-v2_0.1.0_all.zip`(这份交付件在里面是齐的)
-把整个 `release/如东/` 恢复到 v0.2.0 安装目录 —— 75 个文件 / 51.0 MB
-(`如东治理清单_交付_20260901/` 33 件、`如东取数单_2026-08-21/` 18 件、`观澜离线版.app/` 10 件、根下散件 14 件)。
-`release/如东/` 按 `.gitignore` 的既有约定不入库(客户交付物只在受保护工作区), 所以 git 里看不到这次恢复。
-
-**验收**: iframe URL 走网关取 → **[200] text/html, 103760 字节**, 标题 `如东液压系统治理清单`,
-正文 0 个外部相对引用(自包含); 同页另外两个引用 (`127.0.0.1:18020` CMS、`127.0.0.1:18033/v2`) 均 200;
-全门户扫一遍含「如东 / release」的静态引用 → 坏链 0。
-
-脚本: `restore_release_rudong.py` (恢复) · `verify_bug2.py` (验收) · `scan_portal_links.py` (同类坏链全量扫描)
-
-
-
-## 路径跨平台化(①+②, 2026-09-11)
-
-**用户令**: 后续部署到其它电脑; 兼容 Windows / Linux; **系统内目录路径必须使用相对路径**。
-
-**三条口径**(写进 `docs\数据目录结构与落位约定_v0.2.md` §7, 并由门禁自动检查):
-① 只写相对路径(相对**安装根**), 不写机器绝对路径;② 相对基准是安装根、**不是 cwd**, 一律经 `src\paths.py` 解析;
-③ 写进产物/清单/页面的用 POSIX 相对形式(`P.rel()`), 显示给人看时才用本机分隔符(`P.disp_dir()`)。
-
-**新增**
-- `src\paths.py` —— 唯一路径真源(`ROOT/RAW_ROOT` + `store()/ont()/cms()/m5()/sop()/guanlan()/pitch()/tcm_replay()`
-  + `RELEASE/PORTAL/VIEWER/SIM_DIR/LOGS/RUN` + `rel()/disp()/disp_dir()/resolve()/venv_python()`)。
-  `ROOT` = env `WINDSCADA_ROOT` → 否则按本文件位置回溯 ⇒ **整个安装目录可整拷到别的电脑/别的盘**。
-- `scripts\check_portability.py` —— 静态门禁: 机器绝对路径 / cwd 相对字面量 / `os.sep` 拼库存字符串 = ERROR;
-  例外必须具名登记理由。**当前通过**。
-- `scripts\page_fingerprint.py` —— 10 个端点的回归指纹(`status` + 归一化 sha256, 自动剔除 `checked/ms/repo_head`
-  这类设计上会变的字段、并把安装根前缀归一成 `<ROOT>`)。换机验收也用它: 基准机 `--save`, 新机 `--diff`。
-
-**清理的机器绝对路径**(换机必失配, 且是静默故障): windcms `MAIN_ROOT` 默认 mac 检出、windcms/ontology 写死的
-mac venv 解释器、`kb_ingest` 的 mac 技术资料路径、`ingest_ops_2025` 的 `/Volumes/BIG/…`、`place_raw_data` 的
-`<临时目录>/…`、界面占位符里的 `/Users/…`。**清理的 cwd 相对路径**: 20+ 处 `Path('outputs/rudong/…')`(ontology 14 个
-模块、windscada 4 个、windcms、serve/gateway、sop 8 个)全部归到 `paths`; `sop` 里写死的 `rudong` 改为跟场走。
-
-**配置/安装器**: `configs\serve.json` 的 `python` 改为相对(`.venv/Scripts/python.exe`); 启动器 `py()` 支持
-"相对→按安装根解析 / 绝对 / 留空自动探测 venv"; `install.ps1|sh` 写配置时写相对路径。
-
-**验证**
-- 门禁: 运行期 0 处机器路径、0 处 cwd 相对, `os.sep` 只留在显示函数里 ✔
-- 回归指纹: 改造前后 **10/10 端点逐字节一致** ⇒ 只换了路径解析, 页面与取数面零变化 ✔
-- 换 cwd 实测: 从 `C:\` 且不设 `PYTHONPATH` 跑 `scan_stations.py` / `farm()` → 场站辨识与 `store/src_10min` 均正确 ✔
-- 过程中我自己引入并修掉两处: 服务端 `from src import paths` 插在 `sys.path` 引导之前(导致 detail 起不来),
-  以及 `subsys` 子包导入深度写成两点(应为三点)。这两处都是**靠"重启+指纹比对"逮到**的 —— 印证了那把尺子的价值。
-
-
-
-## 场站扫描辨识 + 两个验收场景 (2026-09-11)
-
-**用户令**: 离线数据统一放 `data\raw\`, **下一级目录约定为「场站名称」**, 观澜系统**扫描辨识**。
-
-- `src/windscada/config.py`: 新增 `scan_stations()` / `station_scan()` / `station_report()`; `_expand()`
-  按扫描结果**派生** `src_*`(外场配置若显式给了 `src_*` 则不覆盖)。辨识规则: ①`raw_station` 全等
-  → ②与 `src_farm_names` 互为子串 → ③只有一个场站目录(单站部署)→ ④都不中 = 不猜, 报"未识别"。
-- `scripts/scan_stations.py`: 打印扫到的场站、辨识依据(`how`)、五个约定子目录的件数与存量。
-- 维护页新增一行「场站数据目录 (扫描辨识)」: 扫到 → `覆盖=如东/存在=True`; 扫不到 → `存在=False/覆盖=—`。
-- `scripts/windscada_serve.py`: `adhoc_query` 在原始件缺失时不再抛给 HTTP 处理器(前端只看到 500),
-  改为 **HTTP 200 + 结构化无数据**(`err=no_source` + "本机应在 …\scada_10min\WTG01.csv" 的提示)。
-
-**场景① 清空 `data\raw` + 挪开产物** → 扫描"0 个场站目录 / how=none"; 维护页五项全部 `条数=0 / 存在=False`;
-实时接口从 500 变 200+`no_source`。 ✅
-
-**场景② 按约定放回** → 扫描"1 个场站目录 如东(592 件), how=raw_station", 子目录 38/16/134/404;
-摄入后 报警 **0→39211**、工单 **0→5876**、油样 **0→404**、loss_monthly 3729(powercurve 31s + availability 88s);
-维护页五项覆盖区间全部回来, 实时接口 `months=7`。 ✅
-
-**产物等价**(从 raw 重算 vs 随包基线): `loss_monthly` 3729/3729、`powercurve_dev` 38/38、
-`powercurve_bins` 912/912、`alarms` 39211/39211 **逐值完全一致**; `workorders` 5876 行仅 1 条记录
-**纳秒尾差**(同一时刻 `.0000015` vs `.0000010`); `oil_samples_index` **506→404** —— 清空重建会丢
-**102 行华标 2026-07 批**(源件 `BG-2026-07-YP013 …pdf` 不在现场数据包里, 且一份覆盖多台), 已恢复完整件
-并写入文档 §4 作为缺口。
-
-文档: `docs\数据目录结构与落位约定_v0.2.md`(目录结构、落位约定、消费方式、可重建性边界、两个场景实测)。
-
-## 数据链补齐 · 让系统真正吃 data/raw/如东 (2026-09-11)
-
-**问题**: 说明书 §11 写着「不含从原始数据重生成产物的链 (P0)」, 而维护页四条「摄入命令」指向的文件
-**在包里一个都不存在**; 于是现场把新台账放进 `data/raw/如东/…` 也不会更新, 工单台账一直停在 2024-11-21。
-
-**新增 4 个脚本** (都在 `scripts/`; 维护页四条命令现在全部「命令可用=True」):
-
-| 脚本 | 源 → 产物 | 关键实现 |
-|---|---|---|
-| `windscada_alarms_ingest.py` | `故障报警/*.xls` → `alarms.parquet` | `.xls` 实为 SpreadsheetML(XML); 累计快照件自动跳过 (2025年全年 = Q1..Q4 精确并集 20155 行) |
-| `windscada_workorder_ingest.py` | `风机故障记录/**` → `workorders.parquet` | 按场名过滤 (集团表里混着民勤/来福/宝力格等十来个场); 归并跨表重复 1085 行; 日期按 `src/sop/ledger_dates.py` 同一纪律解析, 解析不了的整行保留 |
-| `windscada_watch_channels_build.py` | `油样报告/**/*.pdf` → `oil_samples_index.parquet` | 文件名解析 (日期/台号/部件/sample_id); 合并语义: 认不出的老行原样保留 |
-| `rebuild_from_raw.py` | 总入口 | `--verify` 等价验收; `--scada` 再跑包内 10 个 SCADA 侧构建器 |
-
-**等价验收** (`python scripts/rebuild_from_raw.py --verify`, 对 `_pre_rebuild_20260911/` 随包基线):
-
-```
-报警事件       39211 行 / 39211 行   逐值完全一致            ✔
-油液化验         506 行 /   506 行   逐值完全一致            ✔
-检修工单台账    1574 种内容复现 1566 种; 8 种差异逐条定位:
-   · 7 种 = 复位运行时间: 随包 NA、本链有值 — 旧链没认源列名「复位时间时间」这个错别字
-   · 1 种 = 随包把 Excel 序列号当纳秒解析成 1970-01-01, 本链解析为 2023-05-13
-   · 随包 78 行副本重复已归并; 另多出 69 行 = 旧链丢弃的日期不可解析行
-```
-
-**效果** (`/detail/v2#tab=system` 数据层, 已重启服务实测):
-
-| 行 | 改前 | 改后 |
-|---|---|---|
-| 检修工单台账 | 2020-01-03 ~ 2024-11-21, 1652 行 | **2020-01-03 ~ 2026-07-08, 5876 行** |
-| 四条「摄入命令」命令可用 | False | **True** |
-
-**依赖**: `requirements.txt` 加 `xlrd==2.0.2` (2023/2024 年总表是真 BIFF `.xls`, openpyxl 读不了),
-轮子已放进 `wheels/win_amd64/` (`wheels/` 按 .gitignore 不入库, 随包分发)。
-
-**SCADA 侧**: `python scripts/rebuild_from_raw.py --scada` 把 10 个构建器按依赖顺序跑一遍
-(powercurve → loss_monthly → curves/control/stop_events/温度/偏航/液压/热链/系统辅助)。
-抽验过等价性 (WTG01/WTG02 的 `temp_bins` 与随包件 1296/1296 行逐值相同、med 差 0.0),
-但**未整体重跑** —— 要跑建议先备份 `outputs/rudong/windscada/` 再逐项比对覆盖区间。
-
-## 门户拆包 + 闸门退出码 (2026-09-11 收尾)
-
-**起因**: 有人问 "`release/portal.html` 20 MB 里到底是什么、能不能不进 git"。查下来 99.3% 是内嵌的
-交付文档正文 (28 个 `<template>`: 治理清单分册、单机/整机报告、仿真台面板), 只有 138 KB 是门户自己的壳。
-原先整份文件入库, 每轮重算产物都往 git 塞 20 MB。
-
-**做法 (用户选 ②)**: 拆成"受管外壳 + 不入库内嵌件 + 可验证装配":
-
-| | 入库 | 说明 |
-|---|---|---|
-| `release\portal_src\shell.html` | ✅ | 138,544 B, 内嵌件处留 `<!--@TEMPLATE:id-->`, 资料索引处留 `<!--@GOVERNANCE_SOURCES-->` |
-| `release\portal_src\manifest.json` | ✅ | 各件 sha256 + 期望门户 sha256 + 行尾约定 |
-| `release\portal_src\README.md` | ✅ | 重建说明与两个坑 |
-| `release\portal_src\templates\` | ❌ 产物 | 28 件 20.04 MB (交付件正文) |
-| `release\portal_src\governance_sources.json` | ❌ 产物 | 1,512 B 脱敏资料索引 |
-| `release\portal.html` | ❌ 产物 | 装配结果, 服务/网关照读 |
-
-新增 `scripts\portal_build.py`: `--extract` 拆 / 默认装配 / `--verify` 逐字节比对 / `--check` 漂移检查。
-**实测 `--verify` 两行同为 `sha256 9b6aabeb6ca18d15`, 20,226,052 B —— 拆→装回到同一个文件**; `--check` 全件一致。
-
-**两个坑**:
-
-1. **行尾**: 门户通体 LF (0 处 CRLF)。Python 文本模式在 Windows 上写文件会把 `\n` 变 `\r\n` —— 20 MB
-   整体改写, `/api/version` 的 `portal_sha256` 与页面指纹全变。两个就地注入器
-   (`guanlan_portal_fix_anchors.py` / `guanlan_portal_inject_claims.py`) 补 `newline=""`,
-   `portal_build.py` 全程字节级读写并在装配后自检 CRLF。第一次拆出来的 `portal_src/` 就是这么废的。
-2. **闸门退出码**: 同一类坑让"校验通过"被报成失败 —— 中文控制台代码页 936, `print('… 一致 ✔')`
-   抛 `UnicodeEncodeError` 使进程以 1 退出。`portal_build --verify` 与 `page_fingerprint --diff`
-   各踩一次 (后者是回归闸门: 10 个端点全部回到基线, 却报失败)。新增 `src\console.py` 的 `soft()`
-   把编码错误降级为 `?`, 接入 5 个"以退出码讲话"的脚本; 复跑四个闸门均正确返回 0。
-
-**回归**: 拆包后 10 个端点全部与基线一致 (`portal.home 64b4b81158d09d9c`, `detail.v2 a0abff29c000f570` …)。
-其中一度看到 `/cms/` 掉线 —— 原因是最后那次 `guanlan.py serve` 发生在产物被挪走时, CMS 启动即因
-`src\windcms\data.py` "No objects to concatenate" 退出 (本该如此); 产物还原后重启即恢复
-(414,139 B, 与基线一致)。**教训: 产物开关与重启顺序有关, 挪/还产物后必须重启一遍。**
-
-## 从零重算 (2026-09-11 下午, 用户令: 不要随包产物, 基于 data/raw 重算)
-
-**做法**: 产物全部挪走 (`products_state.py --off`), 只留 `data/raw` 重算, 一个随包产物都不吃。
-
-```bat
-.venv\Scripts\python.exe scripts\products_state.py --off          :: 产物挪走
-.venv\Scripts\python.exe scripts\rebuild_from_raw.py              :: 三门台账
-.venv\Scripts\python.exe scripts\rebuild_from_raw.py --scada      :: SCADA 侧 10 个构建器 (逐台读 14.7 GB CSV)
-.venv\Scripts\python.exe scripts\rebuild_from_raw.py --verify     :: 等价验收 (需临时放回随包基线)
-```
-
-**结果**: 报警 0→39211 行 · 工单 0→5876 行 · 油样 0→404 行; SCADA 侧 10 个构建器全跑通,
-L0 仓 16 件 1.78 MB (随包 118 件 —— 差额就是"包内没有生成端"那些件, 见 docs §4);
-本体层 `objects.json` 2340 个对象 + 检索索引 + `turbine_params.parquet` 1706 条。
-等价验收: **alarms 39211/39211 逐值完全一致**; workorders 8 行差异全部可归类(旧链时间解析缺陷)
-+ 69 行本链新增覆盖; 唯一人工项是油样 506→404。
-
-**用户追问"缺的数据从 <现场包目录> 抽"** → `place_raw_data.py --scope mech` 落 355 件 15.5 GB:
-`data\raw\西门子4.0技术资料\`(317 件, 本体层的源件) 与 `<场站>\scada_1min\`(38 件 12.8 GB)。
-抽完本体层就从"包内没有源件⛔"变成**可重算**。**仍然缺的**: 油样那 102 行的源件
-(`BG-2026-07-YP013 … .pdf`) 不在包里(整个包只有 `2025年油样` 一批) → 那 102 行找不回来。
-
-### 这一轮逮到的三个真 bug (都不是参数问题, 是静默失效)
-
-| # | 症状 | 根因 | 影响面 |
-|---|---|---|---|
-| 1 | `python -m src.ontology.kb_ingest` 崩: `UnicodeEncodeError: 'gbk' codec can't encode '\u200b'` | `src/ontology/store.py` 的 `read_text()/write_text()` 没写 `encoding=` → Python 文本 I/O 用**系统 locale 编码**, 中文 Windows = cp936。读会乱码/解码失败, 写遇到 GBK 编不出的字符就崩 | **本体层在中文 Windows 上根本不可用**(开发机是 UTF-8 locale 所以从没暴露)。已修 `src/ontology/` 全层 10 处; 存量 12 处已进 `check_portability.py` 的新 WARN 项 6 |
-| 2 | `place_raw_data.py` 一跑就 `NameError: name 'os' is not defined` | 模块级 `DEFAULT_SRC` 用了 `os.environ` 却从没 `import os` | **该脚本自 `d90da05`(路径跨平台化) 起一直跑不通**; A3 那次落位不是它干的 → 说明"写成表的落位映射"从未被真正执行过 |
-| 3 | A2 规则 `…/大部件维修记录.20240619143912557.xlsx` 从没落过盘 | 包内前缀恰好等于条目本身时, 去前缀后 `rest` 为空, 旧代码把它当"目录条目"`continue` 掉 | 单文件形态的规则**全都静默丢件**: 受害两例 = 上面那个 25.8 MB 工单源 + 本轮新加的对译表。修后按前缀的文件名落位 |
-
-> 教训: 这三个都是"跑一次就暴露、但没人跑过"的类型。**从零重算本身就是最好的回归测试** ——
-> 它不依赖任何随包产物, 因而能照出所有"靠旧产物遮掩"的失效。
-
-**顺带发现**: 落位后工单源件 79→80 张, 但总行数不变(5876)——新件的行与既有来源内容重复,
-按"累计快照/副本必须归并"的纪律合并, 脚本把这类情况逐条打了出来(不是静默丢)。
-
-**要补什么、找谁补**: 那次重算照出来的缺口整理成了独立清单 ——
-`docs\重算缺口与补件清单_v0.1.md`(按"源件缺失(现场能补)"与"生成端缺失(研发补脚本)"两类,
-每项带证据路径、影响面、补齐判据与验证命令)。
-
-**另一件同类收尾**: 把剩下的 18 处"文本读写缺 `encoding=`"也清了(见下), 门禁 WARN 项 6 归零。
-
-## 文本 I/O 的 encoding (2026-09-11 收尾)
-
-上面 bug ① 只是冰山一角: 全库搜下来共 **28 处**文本 I/O 没写 `encoding=`(本体层 10 + 其余 18)。
-已全部补上 `encoding='utf-8'`, 其中写侧把 `json.dump(x, open(p,'w'))` 这类匿名句柄改成 `with` 块(正确关文件),
-`src\windcms\plugins.py` 里给子进程当 stdout 的那个句柄**不 with-close**(父进程持引用, 免得提前关掉)。
-
-`check_portability.py` 的 WARN 项 6 同时扩了扫描面: 除 `read_text/write_text` 外, 新增
-`open(..., 'w')` 与 `open(p)`(无 mode = 文本读)两种形态(二进制模式不算), 并与 ERROR 扫描共用 `ALLOW` 例外表
-(门禁自己的规则/文档里必须写出这些形态, 整份豁免 —— 否则会误报 10 处)。**实测命中 0。**
-
-冒烟: 11 个文件 `py_compile` 通过 · windcms 五个模块 + fusion/taxonomy/scenario_29 导入成功 ·
-重启服务 5/7 ok · 页面字节与改前逐项一致(`/` 20225828 · `/detail/` 229643 · `/sim/` 50998 ·
-`/sim/sys/` 123455 · `/viewer/` 4986) · 维护页数据层四个数字不变(报警 39211 · 工单 5876 · 油样 404 · 月表 3729)。
-
-## 探活页提速: /healthz 6.5 s → 冷探 1.5 s / 命中缓存 ~10 ms (2026-09-11)
-
-收尾测页面时发现 `/healthz` 要 **6.5 秒**才回来(而 `/` 那个 20 MB 门户只要 0.04 s)。拆开看是三笔叠加:
-
-| 项 | 原实现 | 代价 |
-|---|---|---|
-| 6 个上游探活 | **串行** `for` 循环, 单探针默认 3 s | 4.4 s(其中 CMS / Ollama 两个掉线模块各卡满超时) |
-| Ollama 模型查询 | 探活之后再串一次 | +2 s |
-| 门户 sha256 | **每次** `PORTAL.read_bytes()` 重读 20 MB 再哈希 | 0.1–数十秒(磁盘忙时) |
-
-改法(返回结构一字未改): ① 全部探针**并发**(`ThreadPoolExecutor`, 总耗时 = 最慢那个); ② 单探针 1.5 s;
-③ 门户 sha 按 `(mtime_ns, size)` 缓存(值不变就不重读); ④ 整个结果缓存 5 s —— 连续调用近乎零成本。
-
-实测: 冷探 **6.5 s → 1.53 s**, 命中缓存 **15 ms**; 重启后首探 8.8 ms(`guanlan.py serve` 自己探过一遍,
-缓存已热)。结构逐键比对无差异(顶层键/模块集合/每模块键/`local_ai` 键/`status`/门户 `file_sha256` 全同),
-`/api/version` 的 `portal_sha256` 与磁盘门户逐字节一致(`9b6aabeb…`)。
-
-**★一台机器上的实测结论**(写进代码注释了): 这台机器**连任何已关闭的 `127.0.0.1` 端口都不回 RST**,
-而是把 SYN 丢掉等超时 —— 对照用的随机闭端口 49996/49997 同样 1.5 s 超时, 不是我们端口或进程的问题,
-是系统防火墙/安全软件行为。所以"有模块掉线"时冷探下限就是 `PROBE_TIMEOUT`; 想更快只能靠缓存,
-**别为此把超时压到健康模块也可能被误判的程度**(实测最慢的健康探针 `sim_sys` 345 ms, 其余 <25 ms,
-1.5 s 留了 4 倍余量)。
-
-## 非法转义序列 + 自检加一行 (2026-09-11)
-
-全量编译 `src/` + `scripts/`(168 个文件)把 `SyntaxWarning` 当错误报出来, 逮到最后一处潜伏炸点:
-
-- `src\windcms\serve.py:12,68` —— JS 正则写在普通字符串里: `/\*\*([^*]+)\*\*/g` 与 `/\d+/g`。
-  Python 3.12 只警告, **未来版本会变成 SyntaxError**; 而且开发机上永远看不见, 换解释器才炸。
-- **不能简单改成 raw 字符串**: 同一串里混着 `\\n`(有意转义成 JS 的 `\n`) 与 `\d`(想写成 JS 的 `\d`),
-  改 raw 会把 `\\n` 变成"两个反斜杠 + n", 语义就变了。正确改法是**把非法转义补成合法**: `\*` → `\\*`。
-- **证明没改语义**: 改前/改后各把该文件 237 个字符串常量的求值内容算指纹 —— 都是 `70bd628f72f72101`,
-  JS 文本一字未变; 改后全库非法转义 **0 处**, `-W error::SyntaxWarning` 编译通过。
-
-`guanlan.py check` 顺带加了一行自检: **源码可编译 (168 个文件, 无非法转义/语法错)** —— 换机器/换解释器
-之前先跑它, 这类问题不必等运行到那一页才发现。
-
-> 自检现状(原始重算后): 3 项 FAIL —— `outputs/rudong/sop/findings.json` 与
-> `outputs/rudong/guanlan/facts_contract_v0.json` 不在(缺口清单 A3), 本机 Ollama 无模型(环境项);
-> L0/L1 产物、本体对象库、门户、仿真、viewer、原始件目录都是 OK。
-
-## 仍然存在 (不是代码问题)
-
-- `/local-ai/` 仍是 DOWN: 本机 Ollama 没运行, 且 `%USERPROFILE%\.ollama` 下**没有任何模型**
-  (`manifests` 都没有), 所以问答与本地审核页不可用。其余 6 个页面不受影响。
-  启用: 启动 Ollama, 按 `configs\models.json` 的 `pull_commands` 拉模型 (需联网/大流量),
-  或把有网机器的 `%USERPROFILE%\.ollama\models` 整个目录拷过来。
-- `release\如东` (治理清单交付件) 本来就是可选项, 不影响页面。

+ 0 - 249
_修复记录_20260911/fix_guanlan_bom.py

@@ -1,249 +0,0 @@
-#!/usr/bin/env python3
-# -*- coding: utf-8 -*-
-"""一次性修复: 观澜·如东样板 v2 启动器的 UTF-8 BOM 崩溃 + serve.json 中文乱码。
-
-报错: guanlan.py cfg() -> json.loads(p.read_text(encoding="utf-8"))
-      json.decoder.JSONDecodeError: Unexpected UTF-8 BOM (decode using utf-8-sig)
-
-根因 (install.ps1 第 4 步 "写配置", Windows PowerShell 5.1):
-  1) `Set-Content configs\\serve.json -Encoding UTF8` 写出的是 **带 BOM** 的 UTF-8;
-     guanlan.py 用裸 utf-8 读 -> 直接 JSONDecodeError, 启动器连 check 都跑不了。
-  2) 紧随其前的 `Get-Content` 对无 BOM 的 UTF-8 按 ANSI 代码页 (中文机 = GBK) 解码,
-     把 serve.json 里的中文注释读成乱码再写回 -> 注释字段损坏 (部分字节已被替换成 '?', 不可逆)。
-
-修复:
-  A. guanlan.py     : 所有 JSON 配置改走 jload() = utf-8-sig (兼容有/无 BOM); 配置损坏时打印提示并退回
-                      内置默认端口, 不再抛栈 (hand-edit 配置写坏 BOM 也不该让产品起不来)。
-  B. install.ps1    : 第 4 步改由 venv 里的 Python 读写配置 (读 utf-8-sig / 写无 BOM 的 UTF-8),
-                      不再用 PS 的 cmdlet 碰 serve.json; 保持文件 UTF-8 带 BOM + CRLF 的既有约定。
-  C. serve.json     : 去掉 BOM, 还原两个 _note 中文注释 (依据原文件残留可逆部分 + 说明书 §7)。
-
-幂等: 重复运行只报 "已修过"。原文件备份为 *.bak-bomfix。
-"""
-from __future__ import annotations
-
-import json
-import pathlib
-import shutil
-import sys
-
-ROOT = pathlib.Path(r"F:\temp\guanlan-rudong-v2_0.2.0")
-GUANLAN = ROOT / "guanlan.py"
-INSTALL_PS1 = ROOT / "install.ps1"
-SERVE_JSON = ROOT / "configs" / "serve.json"
-BAK = ".bak-bomfix"
-
-report: list[str] = []
-
-
-def fail(msg: str) -> None:
-    print("  [X] " + msg)
-    sys.exit(1)
-
-
-def backup(p: pathlib.Path) -> None:
-    b = p.with_name(p.name + BAK)
-    if not b.exists():
-        shutil.copy2(p, b)
-        print(f"  备份 {p.name} -> {b.name}")
-
-
-def read_bytes(p: pathlib.Path) -> bytes:
-    if not p.exists():
-        fail(f"文件不存在: {p}")
-    return p.read_bytes()
-
-
-def has_bom(b: bytes) -> bool:
-    return b.startswith(b"\xef\xbb\xbf")
-
-
-# ---------------------------------------------------------------- A. guanlan.py
-def fix_guanlan() -> None:
-    raw = read_bytes(GUANLAN)
-    if has_bom(raw):
-        fail("guanlan.py 竟然带 BOM, 先人工确认再修")
-    src = raw.decode("utf-8")  # 该文件是 UTF-8 无 BOM + LF
-
-    if "def jload(" in src:
-        print("  [=] guanlan.py 已修过 (存在 jload), 跳过")
-        return
-    backup(GUANLAN)
-
-    anchor_cfg = (
-        'def cfg():\n'
-        '    c = dict(DEFAULT); p = ROOT / "configs/serve.json"\n'
-        '    if p.exists(): c.update(json.loads(p.read_text(encoding="utf-8")))\n'
-        '    return c\n'
-    )
-    new_cfg = (
-        'def jload(p: Path):\n'
-        '    """读 JSON 配置。必须用 utf-8-sig: 记事本和 PowerShell 5.1 的 `Set-Content -Encoding UTF8`\n'
-        '    写出的都是 **带 BOM** 的 UTF-8, 裸 utf-8 读会 JSONDecodeError (Unexpected UTF-8 BOM)。"""\n'
-        '    return json.loads(Path(p).read_text(encoding="utf-8-sig"))\n'
-        '\n'
-        '\n'
-        'def cfg():\n'
-        '    c = dict(DEFAULT); p = ROOT / "configs/serve.json"\n'
-        '    if p.exists():\n'
-        '        try: c.update(jload(p))\n'
-        '        except Exception as ex:  # 配置写坏不该让整个产品起不来: 退回内置默认端口, 但要说清楚\n'
-        '            print(f"  [X] {p} 读不了 ({ex.__class__.__name__}: {str(ex)[:100]}); 本次用内置默认端口。"\n'
-        '                  f" 重跑 install 脚本可重写该文件")\n'
-        '    return c\n'
-    )
-
-    # (旧串, 新串, 期望出现次数)
-    edits = [
-        (anchor_cfg, new_cfg, 1),
-        ('mcfg = json.loads((ROOT / "configs/models.json").read_text(encoding="utf-8"))',
-         'mcfg = jload(ROOT / "configs/models.json")', 1),
-        ('pids = json.loads(PIDS.read_text(encoding="utf-8")) if PIDS.exists() else {}',
-         'pids = jload(PIDS) if PIDS.exists() else {}', 2),
-        ('pids = json.loads(PIDS.read_text(encoding="utf-8"))', 'pids = jload(PIDS)', 1),
-    ]
-    for old, new, want in edits:
-        got = src.count(old)
-        if got != want:
-            fail(f"guanlan.py 锚点命中 {got} 次 (期望 {want}), 未改动: {old[:60]!r}")
-        src = src.replace(old, new)
-
-    left = src.count('read_text(encoding="utf-8")')  # 修完不该再有裸 utf-8 的 JSON 读
-    if left:
-        fail(f"guanlan.py 仍有 {left} 处裸 utf-8 读 JSON")
-    compile(src, str(GUANLAN), "exec")  # 语法自检, 不过就不落盘
-    GUANLAN.write_bytes(src.encode("utf-8"))
-    report.append("guanlan.py: JSON 配置改走 utf-8-sig; 配置损坏时给提示不抛栈")
-    print("  [OK] guanlan.py 已修")
-
-
-# --------------------------------------------------------------- B. install.ps1
-NEW_STEP4 = '''Say "== 4/5 写配置"
-# 别改回 Get-Content / Set-Content -Encoding UTF8 (2026-09-10 实机踩过):
-#   PS 5.1 的 Get-Content 把无 BOM 的 UTF-8 当 ANSI(GBK) 解码 -> 中文注释变乱码;
-#   Set-Content -Encoding UTF8 又写出 BOM -> Python 侧用裸 utf-8 json.loads 直接 JSONDecodeError。
-#   交给 venv 里的 Python 读写: 读 utf-8-sig (有 BOM 也认), 写无 BOM 的 UTF-8。
-#   下面 Python 代码故意只用单引号, 避开 PS 5.1 传参时的引号转义坑。
-$cfgFix = @'
-import json, pathlib, sys
-p = pathlib.Path('configs/serve.json')
-c = json.loads(p.read_text(encoding='utf-8-sig'))
-c['python'] = sys.executable
-p.write_text(json.dumps(c, ensure_ascii=False, indent=2) + '\\n', encoding='utf-8')
-print('   configs/serve.json -> ' + sys.executable)
-'@
-& $vpy -c $cfgFix
-if ($LASTEXITCODE -ne 0) { Say "   X 写 configs\\serve.json 失败"; exit 4 }
-'''
-
-OLD_STEP4 = (
-    'Say "== 4/5 写配置"\n'
-    '$cfg = Get-Content configs\\serve.json -Raw | ConvertFrom-Json\n'
-    '$cfg.python = $vpy\n'
-    '$cfg | ConvertTo-Json -Depth 5 | Set-Content configs\\serve.json -Encoding UTF8\n'
-)
-
-
-def fix_install_ps1() -> None:
-    raw = read_bytes(INSTALL_PS1)
-    if not has_bom(raw):
-        fail("install.ps1 丢了 BOM —— 文件头注释要求 UTF-8 带 BOM + CRLF, 先人工确认")
-    text = raw.decode("utf-8-sig").replace("\r\n", "\n")
-    if "cfgFix" in text:
-        print("  [=] install.ps1 已修过 (存在 cfgFix), 跳过")
-        return
-    backup(INSTALL_PS1)
-
-    if text.count(OLD_STEP4) != 1:
-        fail(f"install.ps1 第 4 步锚点命中 {text.count(OLD_STEP4)} 次, 未改动")
-    text = text.replace(OLD_STEP4, NEW_STEP4.replace("\n", "\n"))
-
-    # 恢复该文件的编码约定: UTF-8 带 BOM + CRLF
-    out = text.replace("\r\n", "\n").replace("\n", "\r\n")
-    INSTALL_PS1.write_bytes(out.encode("utf-8-sig"))
-    report.append("install.ps1: 第 4 步改由 venv Python 写配置 (utf-8-sig 读 / 无 BOM 写)")
-    print("  [OK] install.ps1 已修 (保持 UTF-8 BOM + CRLF)")
-
-
-# ------------------------------------------------------ C. configs/serve.json
-NOTE = ("端口只在这里改 (网关路由表 scripts/guanlan_gateway.py 里的上游端口须同步; "
-        "v0.1 仍是硬编码, 见说明书 §7)")
-NOTE_RAW = (r"原始 SCADA/技术资料目录 (相对本目录或绝对路径, 如 D:\guanlan\data\raw); "
-            r"启动器导出为 WINDSCADA_RUDONG_SRC; 原始件不随包分发")
-
-
-def fix_serve_json() -> None:
-    raw = read_bytes(SERVE_JSON)
-    was_bom = has_bom(raw)
-    try:
-        cfg = json.loads(raw.decode("utf-8-sig"))
-    except Exception as ex:
-        fail(f"serve.json 内容已经是坏 JSON ({ex}); 从 configs\\serve.json{BAK} 或重跑 install 恢复")
-    if not isinstance(cfg, dict):
-        fail("serve.json 不是对象")
-
-    if not was_bom and cfg.get("_note") == NOTE and cfg.get("_note_raw") == NOTE_RAW:
-        print("  [=] serve.json 已修过 (无 BOM 且注释正常), 跳过")
-        return
-    backup(SERVE_JSON)
-
-    cfg["_note"] = NOTE
-    cfg["_note_raw"] = NOTE_RAW
-    SERVE_JSON.write_text(json.dumps(cfg, ensure_ascii=False, indent=2) + "\n", encoding="utf-8")
-    report.append("serve.json: 去掉 BOM, 还原 _note/_note_raw 中文注释")
-    print(f"  [OK] serve.json 已修 (原 BOM={was_bom}, python={cfg.get('python')})")
-
-
-def verify() -> None:
-    print("\n== 校验")
-    # 1) serve.json: 无 BOM, 且裸 utf-8 也能解析 (即 guanlan.py 老写法也不会再炸)
-    b = SERVE_JSON.read_bytes()
-    print(f"  serve.json BOM={has_bom(b)} bytes={len(b)}")
-    if has_bom(b):
-        fail("serve.json 仍带 BOM")
-    cfg = json.loads(b.decode("utf-8"))
-    for k in ("_note", "_note_raw"):
-        if "\ufffd" in cfg[k] or "?" in cfg[k]:
-            fail(f"{k} 仍有乱码: {cfg[k]!r}")
-    if not cfg.get("python", "").lower().endswith(r".venv\scripts\python.exe"):
-        fail(f"python 路径不对: {cfg.get('python')!r}")
-    print(f"  python -> {cfg['python']}")
-    print(f"  _note     : {cfg['_note']}")
-    print(f"  _note_raw : {cfg['_note_raw']}")
-
-    # 2) guanlan.py: 编译 + 现场跑一遍 cfg()
-    src = GUANLAN.read_bytes().decode("utf-8")
-    compile(src, str(GUANLAN), "exec")
-    ns: dict = {"__name__": "not_main", "__file__": str(GUANLAN)}
-    exec(compile(src, str(GUANLAN), "exec"), ns)
-    c = ns["cfg"]()
-    print(f"  guanlan.cfg() -> gateway={c['gateway']} python={c['python']} (无异常)")
-
-    # 3) install.ps1: 字节约定
-    pb = INSTALL_PS1.read_bytes()
-    txt = pb.decode("utf-8-sig")
-    crlf, lf = txt.count("\r\n"), txt.count("\n")
-    print(f"  install.ps1 BOM={has_bom(pb)} CRLF={crlf} LF-only={lf - crlf}")
-    if not has_bom(pb) or crlf != lf:
-        fail("install.ps1 不再是 UTF-8 BOM + 全 CRLF")
-
-
-def main() -> int:
-    print(f"目标: {ROOT}")
-    for p in (GUANLAN, INSTALL_PS1, SERVE_JSON):
-        if not p.exists():
-            fail(f"缺文件 {p}")
-    print("\n== A guanlan.py"); fix_guanlan()
-    print("\n== B install.ps1"); fix_install_ps1()
-    print("\n== C configs/serve.json"); fix_serve_json()
-    verify()
-    print("\n完成:")
-    for r in report:
-        print("  - " + r)
-    if not report:
-        print("  (无需改动, 之前已修)")
-    return 0
-
-
-if __name__ == "__main__":
-    sys.exit(main())

+ 0 - 140
_修复记录_20260911/fix_guanlan_pythonpath.py

@@ -1,140 +0,0 @@
-#!/usr/bin/env python3
-# -*- coding: utf-8 -*-
-"""第二处修复: 机器全局 PYTHONPATH 污染 .venv。
-
-实测 (本机 PYTHONPATH="D:\\Program Files\\Python\\Lib\\site-packages;") :
-  pip 在 venv 里看到系统 site-packages 的包, 判定 "Requirement already satisfied ... outside environment",
-  于是没有把传递依赖装进 .venv: venv 里缺 urllib3 / polars / pyyaml / jinja2 / python-dotenv 等,
-  并且 sys.path 里 PYTHONPATH 排在 venv site-packages **前面**, requests 被系统 2.33.0 顶掉了 pin 住的 2.34.2。
-  后果: 一旦换个没设 PYTHONPATH 的 shell (或别的机器用户), import requests/polars/yaml 直接 ModuleNotFoundError。
-
-修复 (都不改变功能, 只要求 "跑的就是 .venv 里那一套"):
-  D. install.ps1 / install.sh : 装依赖前摘掉 PYTHONPATH, 让 pip 把 requirements.txt 真正装进 .venv。
-  E. guanlan.py                : 启动时把 PYTHONPATH 从 sys.path/b环境里摘掉 (check/serve/stop 与子进程都干净)。
-
-幂等: 重复运行报 "已修过"。
-"""
-from __future__ import annotations
-
-import pathlib
-import shutil
-import sys
-
-ROOT = pathlib.Path(r"F:\temp\guanlan-rudong-v2_0.2.0")
-GUANLAN = ROOT / "guanlan.py"
-INSTALL_PS1 = ROOT / "install.ps1"
-INSTALL_SH = ROOT / "install.sh"
-BAK = ".bak-pyfix"
-
-changed: list[str] = []
-
-
-def fail(msg: str) -> None:
-    print("  [X] " + msg)
-    sys.exit(1)
-
-
-def backup(p: pathlib.Path) -> None:
-    b = p.with_name(p.name + BAK)
-    if not b.exists():
-        shutil.copy2(p, b)
-        print(f"  备份 {p.name} -> {b.name}")
-
-
-def sub1(text: str, old: str, new: str, what: str) -> str:
-    n = text.count(old)
-    if n != 1:
-        fail(f"{what}: 锚点命中 {n} 次 (期望 1)")
-    return text.replace(old, new)
-
-
-# ------------------------------------------------------------------ D1. install.ps1
-PS1_ANCHOR = '$ErrorActionPreference = "Stop"; $env:PYTHONUTF8 = "1"\n'
-PS1_INSERT = PS1_ANCHOR + (
-    '# 机器上若全局设了 PYTHONPATH (例如指向 D:\\Program Files\\Python\\Lib\\site-packages), pip 会把那里的包\n'
-    '# 当成 "已满足", 于是不装进 .venv —— venv 里就缺 urllib3 / polars / pyyaml / jinja2 等传递依赖,\n'
-    '# 换个没设 PYTHONPATH 的 shell 立刻 import 失败; 且 PYTHONPATH 排在 venv 之前会顶掉 pin 住的版本。\n'
-    '# 先把环境变量摘掉, 让下面 pip 老老实实按 requirements.txt 装进 .venv。\n'
-    'Remove-Item Env:PYTHONPATH -ErrorAction SilentlyContinue\n'
-)
-
-
-def fix_install_ps1() -> None:
-    raw = INSTALL_PS1.read_bytes()
-    text = raw.decode("utf-8-sig").replace("\r\n", "\n")
-    if "Remove-Item Env:PYTHONPATH" in text:
-        print("  [=] install.ps1 已修过, 跳过")
-        return
-    backup(INSTALL_PS1)
-    text = sub1(text, PS1_ANCHOR, PS1_INSERT, "install.ps1")
-    INSTALL_PS1.write_bytes(text.replace("\r\n", "\n").replace("\n", "\r\n").encode("utf-8-sig"))
-    changed.append("install.ps1: 装依赖前 Remove-Item Env:PYTHONPATH")
-    print("  [OK] install.ps1 已修 (保持 UTF-8 BOM + CRLF)")
-
-
-# ------------------------------------------------------------------ D2. install.sh
-SH_ANCHOR = 'set -e\n'
-SH_INSERT = SH_ANCHOR + (
-    '# 同理: 全局 PYTHONPATH 会让 pip 误判依赖已满足而不装进 .venv (见 install.ps1 注释)\n'
-    'unset PYTHONPATH\n'
-)
-
-
-def fix_install_sh() -> None:
-    raw = INSTALL_SH.read_bytes()
-    if raw.startswith(b"\xef\xbb\xbf"):
-        fail("install.sh 不该有 BOM")
-    text = raw.decode("utf-8")
-    if "unset PYTHONPATH" in text:
-        print("  [=] install.sh 已修过, 跳过")
-        return
-    backup(INSTALL_SH)
-    text = sub1(text, SH_ANCHOR, SH_INSERT, "install.sh")
-    INSTALL_SH.write_bytes(text.encode("utf-8"))  # 保持 LF / 无 BOM
-    changed.append("install.sh: 装依赖前 unset PYTHONPATH")
-    print("  [OK] install.sh 已修 (保持 LF)")
-
-
-# ------------------------------------------------------------------ E. guanlan.py
-G_ANCHOR = 'WIN = os.name == "nt"\n'
-G_INSERT = G_ANCHOR + '''
-
-def _hermetic():
-    """摘掉机器全局的 PYTHONPATH: 本项目依赖都装在 .venv 里, 而 PYTHONPATH 会被插到 sys.path 前面,
-    顶掉 venv 里 pin 住的版本 (实测 requests 2.33.0 顶掉 2.34.2), 还会让 pip 少装传递依赖。
-    子进程从 env() 继承的是摘干净之后的环境。"""
-    for d in [x for x in os.environ.pop("PYTHONPATH", "").split(os.pathsep) if x]:
-        while d in sys.path: sys.path.remove(d)
-
-
-_hermetic()
-'''
-
-
-def fix_guanlan() -> None:
-    raw = GUANLAN.read_bytes()
-    text = raw.decode("utf-8")
-    if "_hermetic()" in text:
-        print("  [=] guanlan.py 已修过, 跳过")
-        return
-    backup(GUANLAN)
-    text = sub1(text, G_ANCHOR, G_INSERT, "guanlan.py")
-    compile(text, str(GUANLAN), "exec")
-    GUANLAN.write_bytes(text.encode("utf-8"))
-    changed.append("guanlan.py: 启动时摘掉 PYTHONPATH (自身与子进程都跑 .venv 的那套)")
-    print("  [OK] guanlan.py 已修")
-
-
-def main() -> int:
-    print(f"目标: {ROOT}")
-    print("\n== D1 install.ps1"); fix_install_ps1()
-    print("\n== D2 install.sh"); fix_install_sh()
-    print("\n== E guanlan.py"); fix_guanlan()
-    print("\n完成:")
-    for c in changed or ["(无需改动, 之前已修)"]:
-        print("  - " + c)
-    return 0
-
-
-if __name__ == "__main__":
-    sys.exit(main())

+ 0 - 84
_修复记录_20260911/fix_guanlan_startbat.py

@@ -1,84 +0,0 @@
-#!/usr/bin/env python3
-# -*- coding: utf-8 -*-
-"""第三处修复: start.bat 把 "网关已起但有模块降级" 误报成 "没起来"。
-
-现状: guanlan.py serve 的退出码 1 有**两种**含义 —— ① 网关 healthz 压根没起来;
-② 网关起来了, 但有模块 [DOWN] (本机 Ollama 没跑 / 没拉模型时必然如此)。
-start.bat 只看 errorlevel, 于是 ② 也走 "[X] Server did not start ... Not opening the browser",
-用户看到的是 "服务没起", 实际网关在跑、6/7 页面都正常 —— 与原始报障里那条
-"[X] Server did not start (reason above)" 同源。
-
-修复: 退出码非 0 时先探一次 /healthz 区分两种情形; 真没起来才报 [X] 并停;
-起来了但有降级则打 [!] 说明 (哪一类页面受影响), 浏览器照常打开。
-保持 .bat 的 CRLF / 无 BOM 约定。
-"""
-from __future__ import annotations
-
-import pathlib
-import shutil
-import sys
-
-ROOT = pathlib.Path(r"F:\temp\guanlan-rudong-v2_0.2.0")
-START = ROOT / "start.bat"
-
-OLD = (
-    '".venv\\Scripts\\python.exe" guanlan.py serve\n'
-    'if errorlevel 1 (\n'
-    '  echo.\n'
-    '  echo [X] Server did not start ^(reason above^). Not opening the browser -\n'
-    '  echo     a "site cannot be reached" page would look like a network problem.\n'
-    '  echo     See logs\\gateway.log\n'
-    '  pause\n'
-    '  exit /b 1\n'
-    ')\n'
-    'start "" http://127.0.0.1:28084/\n'
-)
-
-NEW = '''".venv\\Scripts\\python.exe" guanlan.py serve
-if errorlevel 1 (
-  rem serve exits 1 in two different cases: the gateway never came up, or the gateway is
-  rem up with at least one degraded module, e.g. Ollama not running. Probe healthz to tell
-  rem them apart, otherwise a working server gets reported as "did not start".
-  ".venv\\Scripts\\python.exe" -c "import urllib.request,sys; sys.exit(0 if urllib.request.urlopen('http://127.0.0.1:28084/healthz',timeout=6).status==200 else 1)" 2>nul
-  if errorlevel 1 (
-    echo.
-    echo [X] Gateway did not come up ^(reason above^). Not opening the browser -
-    echo     a "site cannot be reached" page would look like a network problem.
-    echo     See logs\\gateway.log
-    pause
-    exit /b 1
-  )
-  echo.
-  echo [!] Gateway is up, but at least one module is degraded ^(the [DOWN] lines above^).
-  echo     Usual cause: local Ollama is not running or has no models pulled, which only
-  echo     disables the Q^&A / local-review pages. Every other page works.
-  echo     Details: logs\\ for the degraded module.
-)
-start "" http://127.0.0.1:28084/
-'''
-
-
-def main() -> int:
-    raw = START.read_bytes()
-    if raw.startswith(b"\xef\xbb\xbf"):
-        print("  [X] start.bat 不该有 BOM")
-        return 1
-    text = raw.decode("utf-8").replace("\r\n", "\n")   # .bat 是 CRLF, 先归一化再匹配
-    if "Gateway did not come up" in text:
-        print("  [=] start.bat 已修过, 跳过")
-        return 0
-    if text.count(OLD) != 1:
-        print(f"  [X] start.bat 锚点命中 {text.count(OLD)} 次, 未改动")
-        return 1
-    b = START.with_name(START.name + ".bak-healthz")
-    if not b.exists():
-        shutil.copy2(START, b)
-        print(f"  备份 start.bat -> {b.name}")
-    out = text.replace(OLD, NEW).replace("\r\n", "\n").replace("\n", "\r\n")
-    START.write_bytes(out.encode("utf-8"))     # UTF-8 无 BOM + CRLF
-    print("  [OK] start.bat 已修 (CRLF / 无 BOM 保持不变)")
-    return 0
-
-
-if __name__ == "__main__":
-    sys.exit(main())

+ 0 - 67
_修复记录_20260911/restore_release_rudong.py

@@ -1,67 +0,0 @@
-#!/usr/bin/env python3
-# -*- coding: utf-8 -*-
-"""bug-2 修复: 从 v0.1.0 发行包把缺失的 release/如东/ 交付件恢复到 v0.2.0 安装目录。
-
-根因: v0.2.0 包**根本没带** release/如东/ (0.2.0 zip 里 release/ 下只有 viewer/),
-而 release/portal.html 的 #documents 区块用 iframe 指向
-  http://127.0.0.1:28084/如东/如东治理清单_交付_20260901/02_治理清单/如东_液压系统治理清单_v1.2_2026-09-01.html
-该路径由网关 _static() 从 release/ 目录取文件, 文件不在 → 404 {"err":"not found", "path": …},
-iframe 里就直接显示那段 JSON。v0.1.0 包里这份交付件是齐的 (76 项), 直接恢复。
-
-(release/如东/ 按 .gitignore 约定不入库 —— 客户交付物只在受保护工作区。)
-"""
-from __future__ import annotations
-
-import pathlib
-import shutil
-import sys
-import zipfile
-
-ZIP = pathlib.Path(r"F:\temp\guanlan-rudong-v2_0.1.0_all.zip")
-REL = pathlib.Path(r"F:\temp\guanlan-rudong-v2_0.2.0\release")
-PREFIX = "release/如东/"
-
-
-def nm(info):
-    raw = info.filename
-    for enc in ("gbk", "utf-8"):
-        try:
-            return raw.encode("cp437").decode(enc)
-        except Exception:
-            continue
-    return raw
-
-
-def main() -> int:
-    if not ZIP.is_file():
-        print(f"  [X] 缺源包 {ZIP}")
-        return 1
-    print(f"源: {ZIP}")
-    print(f"目标: {REL}\\如东\\\n")
-    n_files = 0
-    n_bytes = 0
-    with zipfile.ZipFile(ZIP) as zf:
-        items = [(nm(i), i) for i in zf.infolist()]
-        rel = [(n.replace("\\", "/"), i) for n, i in items if n.replace("\\", "/").startswith(PREFIX)]
-        if not rel:
-            print("  [X] 源包里没有 release/如东/")
-            return 1
-        for name, info in rel:
-            target = REL / name[len("release/"):]
-            if info.is_dir():
-                target.mkdir(parents=True, exist_ok=True)
-                continue
-            target.parent.mkdir(parents=True, exist_ok=True)
-            with zf.open(info) as fsrc, open(target, "wb") as fdst:
-                shutil.copyfileobj(fsrc, fdst, 1024 * 1024)
-            if target.stat().st_size != info.file_size:
-                print(f"  [X] 大小不符: {target}")
-                return 1
-            n_files += 1
-            n_bytes += info.file_size
-    print(f"恢复 {n_files} 个文件, {n_bytes/1024/1024:.1f} MB")
-    return 0
-
-
-if __name__ == "__main__":
-    sys.exit(main())

+ 0 - 50
_修复记录_20260911/scan_portal_links.py

@@ -1,50 +0,0 @@
-#!/usr/bin/env python3
-# -*- coding: utf-8 -*-
-"""全量核查 portal.html 里指向本机静态件的引用 (含相对写法), 找出同类"交付件没随包"的坏链。"""
-import pathlib
-import re
-import urllib.error
-import urllib.parse
-import urllib.request
-
-REL = pathlib.Path(r"F:\temp\guanlan-rudong-v2_0.2.0\release")
-html = (REL / "portal.html").read_text(encoding="utf-8", errors="replace")
-
-# 所有 href/src 里含 如东 或 release 的
-cand = set()
-for m in re.finditer(r'(?:href|src)="([^"]+)"', html):
-    u = m.group(1)
-    if ("如东" in u or "/release" in u or "release/" in u) and not u.startswith(("mailto:", "javascript:")):
-        cand.add(u)
-print(f"候选引用 {len(cand)} 个\n")
-
-
-def check(u: str):
-    if u.startswith("http"):
-        full = u
-    else:
-        base = "http://127.0.0.1:28084/" if u.startswith("/") else "http://127.0.0.1:28084/"
-        full = base + u.lstrip("/")
-    p = urllib.parse.urlsplit(full)
-    full = urllib.parse.urlunsplit((p.scheme, p.netloc, urllib.parse.quote(p.path), p.query, ""))
-    try:
-        with urllib.request.urlopen(full, timeout=30) as r:
-            r.read(1)
-            return r.status
-    except urllib.error.HTTPError as e:
-        return e.code
-    except Exception as e:
-        return f"{e.__class__.__name__}"
-
-
-bad = []
-for u in sorted(cand):
-    s = check(u)
-    tag = "ok  " if s == 200 else "BAD "
-    if s != 200:
-        bad.append(u)
-    print(f"  [{s}] {tag} {u}")
-
-print(f"\n坏链 {len(bad)} 个")
-for u in bad:
-    print("   - " + u)

+ 0 - 44
_修复记录_20260911/verify_bug2.py

@@ -1,44 +0,0 @@
-#!/usr/bin/env python3
-# -*- coding: utf-8 -*-
-"""bug-2 验收: 走网关取那个 iframe URL, 并把它自身的相对引用也逐条取一遍。"""
-import re
-import urllib.parse
-import urllib.request
-
-URL = ("http://127.0.0.1:28084/如东/如东治理清单_交付_20260901/"
-       "02_治理清单/如东_液压系统治理清单_v1.2_2026-09-01.html")
-
-
-def get(u):
-    req = urllib.request.Request(u, headers={'User-Agent': 'verify'})
-    try:
-        with urllib.request.urlopen(req, timeout=30) as r:
-            return r.status, r.headers.get('Content-Type'), r.read()
-    except urllib.error.HTTPError as e:
-        return e.code, e.headers.get('Content-Type'), e.read()
-
-
-def q(u):
-    """URL 里中文需编码, 网关会 unquote。"""
-    p = urllib.parse.urlsplit(u)
-    return urllib.parse.urlunsplit((p.scheme, p.netloc,
-                                    urllib.parse.quote(p.path), p.query, p.fragment))
-
-
-st, ct, body = get(q(URL))
-print(f"iframe URL : [{st}] {ct} {len(body)} 字节")
-txt = body.decode('utf-8', 'replace')
-print(f"首行       : {txt.splitlines()[0][:120] if txt.strip() else '(空)'}")
-print(f"含 not found: {'not found' in txt[:400] and st != 200}")
-
-refs = sorted({m for m in re.findall(r'(?:href|src)="([^"]+)"', txt)
-               if not m.startswith(('http', '#', 'data:', 'mailto:', 'javascript:'))})
-print(f"\n文中相对引用 {len(refs)} 个, 逐条取:")
-bad = 0
-for r in refs:
-    u = urllib.parse.urljoin(URL, r)
-    s, c, b = get(q(u))
-    if s != 200:
-        bad += 1
-    print(f"  [{s}] {len(b):8d} 字节  {r}")
-print(f"\n结论: {'全部 200' if not bad else f'{bad} 个引用取不到'}")

+ 0 - 61
_修复记录_20260911/verify_guanlan_simsys.py

@@ -1,61 +0,0 @@
-#!/usr/bin/env python3
-# -*- coding: utf-8 -*-
-"""验证 /sim/sys/ (release/sim_sys_server.py) 恢复后: 页面能出 + 脱敏回扫干净。"""
-import sys
-import urllib.request
-import zipfile
-import pathlib
-
-ROOT = pathlib.Path(r"F:\temp\guanlan-rudong-v2_0.2.0")
-sys.path.insert(0, str(ROOT))
-from src.windscada import deid_public as DP
-
-
-def readable(name):
-    try:
-        return name.encode('cp437').decode('utf-8')
-    except UnicodeError:
-        return name
-
-
-zip_path = ROOT / 'release' / '如东SWT40_控制律仿真台_20260906.zip'
-with zipfile.ZipFile(zip_path) as z:
-    pages = [readable(n.rsplit('/', 1)[-1]) for n in z.namelist()
-             if n.endswith('.html') and readable(n.rsplit('/', 1)[-1])[:1] in {'0', '1', '2', '3', '4'}]
-
-print(f"页面 {len(pages)} 个")
-bad = 0
-for name in sorted(pages):
-    url = 'http://127.0.0.1:18792/' + urllib.parse.quote(name)
-    try:
-        with urllib.request.urlopen(url, timeout=20) as r:
-            body = r.read().decode('utf-8')
-            code = r.status
-    except Exception as ex:
-        print(f"  [X] {name}: {ex.__class__.__name__}: {ex}")
-        bad += 1
-        continue
-    leaks = DP.audit(body)
-    leak_txt = '干净' if not leaks else '; '.join(f'{n}×{c} {s}' for n, c, s in leaks)
-    print(f"  [{code}] {name:28s} {len(body):7d} 字节  回扫: {leak_txt}")
-    bad += bool(leaks)
-
-print("\n网关路由:")
-for path in ('/healthz', '/sim/sys/'):
-    try:
-        with urllib.request.urlopen('http://127.0.0.1:28084' + path, timeout=20) as r:
-            txt = r.read().decode('utf-8', 'replace')
-            if path == '/healthz':
-                import json
-                d = json.loads(txt)
-                print(f"  [{r.status}] {path}  status={d['status']} {d['mode']}")
-                for m in d['modules']:
-                    print(f"        {'ok  ' if m['ok'] else 'DOWN'} {m['path']} {m['name']}")
-            else:
-                print(f"  [{r.status}] {path}  {len(txt)} 字节, 含返回链接: {'guanlan-top' in txt}")
-    except Exception as ex:
-        print(f"  [X] {path}: {ex.__class__.__name__}: {ex}")
-        bad += 1
-
-print("\n结论:", "全部通过" if not bad else f"有 {bad} 项需要注意")
-sys.exit(0 if not bad else 1)

+ 238 - 0
docs/振动数据接入_v0.1.md

@@ -0,0 +1,238 @@
+# 振动数据接入(CMS / TCM)· v0.1
+
+> 日期: 2026-09-12 · 适用: `<安装目录>`(安装目录)
+> 本文回答: ①振动侧的现场件长什么样、放哪;②谁把它变成系统能用的东西;③改了哪些代码、怎么验收;
+> ④哪些还做不到(如实列,不假装)。
+> 相关: `docs\数据目录结构与落位约定_v0.2.md`(§2b 落位速查)· `docs\重算操作手册_v0.1.md`(一键重算)
+
+---
+
+## 1. 用户令与现场件盘点
+
+**用户令原文**(路径按本包文档约定写成占位符): 「数据层里的「CMS 振动评估报告」应遵循
+`<安装目录>\data\raw\如东\windcms`、「振动线 handoff」应遵循 `<安装目录>\data\raw\如东\m5_cms_tcm`
+存放; 要求 1. 修改 `<安装目录>` 下的观澜系统, 支持振动数据参与系统运行、重算等;
+2. 自 `<现场包目录>` 提取相应振动数据存放至上述目录」;随后用户把 `CMS_RuDong_CGN_202603-04.zip`
+放进现场包 —— **那正是此前缺的原始件**。
+
+现场包(`<现场包目录>`)里的振动相关件共三类:
+
+| # | 件 | 形态 | 能不能算 |
+|---|---|---|---|
+| ① | `CMS_RuDong_CGN_202603-04.zip`(38.8 GB 压缩 / 150 GB 解压 / 25,679 件) | Brande TCM Enterprise 导出: `measurement\<年>\<月>\<WTGxx>\<WTGxx>_<uuid>_decode.json`;每件是一次 API 响应 `{body:{body:{"<时间戳>":[{"Record":{…}}]}}}` | **能**(索引/谱/报告全部由它算) |
+| ② | 12 份月度「振动分析报告(用印版)」PDF(2025-06…2026-05) | **纯扫描件** | **不能** |
+| ③ | 上海电气 2026年07月报告 docx(4.87 MB)· 大生科技传动链报告 docx(20.81 MB) | Word, 有文本层 | 能(逐台判级转录) |
+
+**②为什么不能算(实测, 不是推测)**: 用 pypdf 打开 2026年5月那份 → `len(pages)=12`,
+每页 `len(images)=1`, `page.extract_text()` 长度 **0**(12 页全 0)。字节面也印证: `/Font` 0 处、
+`FlateDecode` 0 处、`/Image` 24 处、`DCTDecode` 12 处 = 每页一张 JPEG。
+⇒ 没有 OCR 就取不出任何数值。**处理方式: 只登记归档**(`_扫描件清单.json` 记 sha256/大小/无文本层),
+不参与判级、不假装读过。要它们的数值只有两条路: 现场给电子件(docx/xlsx),或上 OCR(本包不装)。
+
+**顺带核过的一条旧判断**: v0.2.0 初版把「`(8)…\振动分析报告\`」列在"故意不落"里,理由是
+"成品牌报告而非可再加工的测量数据, 且 windcms 要的测点索引包里没有"。**该判断在 2026-09-12 被推翻**
+(①到货后索引源件具备),故改为落位。现场包里的**同件副本**(根目录 `…2026年5月…(2).pdf`、
+嵌套包 `如东海上振动报告11份.zip`)不重复落 —— 落三遍会得到 3 份同名报告。
+
+---
+
+## 2. 落位约定
+
+```
+data\raw\如东\
+├─ windcms\                                   ← 数据层「CMS 振动评估报告」
+│    ├─ CMS_RuDong_CGN_202603-04\
+│    │    └─ measurement\2026\{03,04}\WTGxx\*_decode.json     25,679 件 / 150 GB
+│    └─ 厂家报告\上海电气_月度\
+│         ├─ 中广核如东海上风电场2025年06月…2026年05月振动分析报告用印版.pdf    12 件(扫描件)
+│         ├─ 中广核如东海上风电场2026年07月振动分析报告_上海电气.docx
+│         └─ _扫描件清单.json
+└─ m5_cms_tcm\                                ← 数据层「振动线 handoff」
+     └─ 厂家报告\
+          └─ 中广核如东海上风电场传动链振动分析报告_大生科技_20260311(1).docx
+```
+
+- 两个目录名已写进 `src\windscada\config.py` 的 `STATION_SUBDIRS`(扫描/落位/维护页从此认它们;
+  改目录名 = 换接口,要同步维护页与本文)。
+- 原始导出**保留包内 `measurement\` 这一层**(来源可追溯);摄入按 `rglob` 找 `*_decode.json`,
+  套不套这层都能吃。
+- 若现场给出 `handoff_vibration_v2.json` / `component_history.json` **正本**,放 `m5_cms_tcm\` 下即被优先采用
+  (包内那两份目前仍是随包快照 —— 振动线分支的产物没随包,见 §7)。
+- 落位命令(映射表可复核、可重跑,同尺寸文件自动跳过):
+  ```
+  .venv\Scripts\python.exe scripts\place_raw_data.py --src <现场包目录> --scope vib --dry-run
+  .venv\Scripts\python.exe scripts\place_raw_data.py --src <现场包目录> --scope vib
+  ```
+
+---
+
+## 3. 摄入链(谁把原始件变成系统能用的东西)
+
+```
+data/raw/如东/windcms/…/*_decode.json
+   │  scripts/rudong_tcm_index.py      → outputs/rudong/m5_cms_tcm/windows/<窗>/index.parquet   (54 列)
+   │  scripts/rudong_tcm_spectra.py    → outputs/rudong/m5_cms_tcm/windows/<窗>/spectra/*.npz
+   │                                     + <窗>/spectra_meta.parquet(谱库目录内另存一份, 兼顾两种读取路径)
+   │  [可选] scripts/windcms.py report / kb  → outputs/rudong/windcms/*(★默认不跑, 见 §3b)
+   ▼
+data/raw/如东/{windcms,m5_cms_tcm}/…/*.docx
+   │  scripts/vib_reports_build.py     → outputs/rudong/windcms/厂家报告提取_<报告期>.json
+   │                                     + 报告_CMS振动状态评估报告_<报告期>.md(与自产报告同构)
+   │                                     + outputs/rudong/m5_cms_tcm/报告_TCM传动链振动分析_<报告期>.md
+```
+
+**一键**: `scripts/vib_raw_build.py`(索引→谱;`--skip-spectra` 可关;`--with-report` 才跑报告/知识库)。
+它也写 `outputs/rudong/m5_cms_tcm/vib_raw_manifest.json`(源件路径、窗名、行数、谱数、时间窗、
+各步耗时/rc,以及 `missing_chain`)。
+
+### 3b. ★ 为什么"重生成 CMS 报告"默认不跑(2026-09-12 实测)
+
+跑一次 `scripts/windcms.py report` 在本包**会让产物变差**:
+
+| 件 | 随包快照 | 本包重生成后 | 变化 |
+|---|---|---|---|
+| `windcms/report.md` | 16,857 B(含逐台融合级表 + L4 过闸谱线) | **467 B**(融合级表空、L4 写"无") | **−97.2%** |
+| `windcms/overview.html` | 646,218 B | 548,466 B | −15.1% |
+| `windcms/index_eng.html` | 443,841 B | 408,112 B | −8.0% |
+| `windcms/index.html` | 389,997 B | 382,544 B | −1.9%(逐台页仍 38 个) |
+
+根因不在数据,在**缺件**:`src/windcms/data.py::load_model()` 要读 `m5/model_run_l6.parquet` 与
+`m5/fusion_38.csv`,而这两件属六层链的 `model_run` / `fusion` 两步 —— 那四步脚本没随包(§7),产物也不在。
+于是"重生成"= 用残缺输入覆盖完整快照。
+
+**处置**: 摄入默认只做索引/谱(加性、不覆盖任何随包件);报告/知识库改为 `--with-report` 显式开启,
+且要求先备份 `outputs/<场>/windcms/`。本次实测后已把 `outputs/rudong/windcms` **逐字节还原**为随包快照
+(56 件,哈希比对 0 差异),只保留厂家报告转录那两件。六层链补齐后这个默认值应当翻过来。
+
+**来源登记**: 摄入与转录产出的每一件都由构建脚本自登记进 `outputs/<场>/_derived_manifest.json`
+(`src/derived_manifest.py`),`_provenance.json` 生成时据此把它们记成 `raw-derived` ——
+避免"新造的件因不在随包快照里而整条不进台账"。
+
+**窗名口径**: `wMMDD` = **数据起始日**(与既有 `w0127`/`w0707`/`w0811` 同口径)。
+`src\windcms\pipeline.py::window_year_gate` 会拦"窗名无年份 + 数据过老"的误用(w1226 事故的护栏)。
+重复摄入同一批数据时,已存在的窗自动改名为 `<窗>_reimport_<时分>`,而 `data.EXCLUDE_DEFAULT`
+把带 `_reimport` 的窗排除在生产集外 —— **重复摄入不会污染分析集**。
+
+**消费者(无需改一行代码,窗是自动发现的)**:
+
+| 消费者 | 拿什么 |
+|---|---|
+| `src\windcms\data.py::windows()` | 扫 `m5\windows\w????\index.parquet` → 新窗进分析集 |
+| `src\windcms\data.py::load_scalars()` | 各窗标量(`ds_size==1` 行)拼接 → CMS 报告/页面 |
+| `src\windcms\data.py::spectrum()` | `spectra_meta.parquet` + `npz['values'][shard_row]` → **谱图能取到** |
+| `src\windscada\taxonomy.py` / `subsys\fusion.py::windcms_grades()` | 最新 `报告_CMS振动状态评估报告_*.md` 的 `## 附录 A` 表 → 设备状态转录 |
+
+> ★ `spectra` 这一环此前是**死的**: 包内 `m5\spectra\`、`m5\windows\` 两个目录整个不在,
+> `data.spectrum()` 只能返回 None。本次摄入让谱图第一次有数据可画。
+
+---
+
+## 4. 格式契约与对拍(改这块必看)
+
+摄入产物与**包内既有产物**是同构关系,基准就是包内那份 `outputs\rudong\m5_cms_tcm\tcm_index.parquet`
+(330,308 行 × 54 列,2026-01-27~02-03 窗,同一摄取逻辑的产物):
+
+| 检查 | 结果 |
+|---|---|
+| 列名与列序 | **完全一致**(54 列,逐字抄自基准,含列序) |
+| dtype | 差异 3 列: `rec_i`(float64/int64)、`ds_dim`(float64/str)、`parse_error`(str/object)—— 消费者按列名取数,不受影响 |
+| 取值(FFT 行) | `x_offset=0.0` `x_delta=0.9375` `x_unit='Hz'` `y_unit='m/s²'` 与基准同行**逐格一致** |
+| 内部一致性 | `ds_size == lines+1` 比例 **1.000** |
+
+**TCM 导出里踩过的两个坑(写在这里防复发)**:
+
+1. **X/Y 轴四件在 `Measurement.DataSets` 层,不在 `DataSet` 里**。`Size`/`Dimension`/`Values` 在
+   `DataSets.DataSet`,而 `X-axisOffset`/`X-axisDelta`/`X-axisUnit`/`Y-axisUnit` 在**上一层**。
+   第一版两层取错 → 四列整列为空,内部一致性检查算出 0.000(本该 1.000)才暴露。**先看一致性数再信列**。
+2. `DataSets.DataSet.Values` 是**空格分隔的数值字符串**(不是数组)。谱: `Size=Lines+1`;
+   `X-axisDelta = 带宽/Lines`(例: 6000/6400 = 0.9375)。标量: `Size==1` 时 `Values` 本身就是标量值。
+   另有少数测量(`FFT_250/2000/10000_Tr`, `Lines=400`)的 `x_delta = 带宽/800` = 半带宽口径 ——
+   **源数据自己的约定**,摄入原样转录,不做"修正"。
+
+---
+
+## 5. 厂商报告的转录纪律
+
+`scripts\vib_reports_build.py` 只做**转录**,不做判级:
+
+- **合并单元格必须按 `tc` 恒等还原**:python-docx 对合并区会把同一个 `tc` 在相邻列/行重复给出,
+  照抄会把 `优秀 | 优秀 | 优秀` 压成一格(第一版手工 dump 就吃了这个亏: 表里 `WTG-01` 行显示成
+  `WTG-01 | 优秀 | 无`,实际是 主轴承=优秀 齿轮箱=优秀 发电机=优秀 结论=无)。规则: 同行 `tc` 与左邻
+  相同 = 横向合并(沿用左值);与上行同列相同 = 纵向合并(沿用上值)。
+- **用厂家自己的总览句当校验和**:2026年07月报告原文"…优秀机组34台,良好3台,预警0台,无数据1台,
+  测点异常1台"。转录结果 `{优秀34, 良好3, 不可判1}` —— **逐项一致**("无数据"记 `不可判`: 没测到 ≠ 正常)。
+- **厂家没给的不猜**:报告按「主轴承/齿轮箱/发电机」给级(**不分前后**),转录表里主轴承前/后填同一值并注明;
+  「综合」= 三项**取严**(危险>报警>良好>优秀>不可判)的聚合规则结果,必须写明"这是我们的规则不是厂家判级";
+  「融合级/CMS 红黄/行动等级建议」厂家未给 → 写 `—`。
+- **两条轴不混**:厂商报告转录的 md 与观澜自产报告同构(消费端不必改代码),报告期用**文件名月份取月末**,
+  于是它不会顶掉更新的自产报告;随包自产报告缺失时它就是最新版(从零重算的机器上正好用得上)。
+  大生科技那份只写 `报告_TCM传动链振动分析_*.md`,**不冒充 handoff 判级**。
+- **口径不同别直接比**:厂家四级(优秀/良好/预警/报警)+「数据异常」≠ 观澜五级。实测两者对 38 台
+  有 33 台结论不同(厂家 2026-07 vs 观澜 2026-08)—— 来源、日期、判据都不同,属预期,不是谁错。
+
+---
+
+## 6. 验收(命令 + 实测数字)
+
+```
+.venv\Scripts\python.exe scripts\place_raw_data.py --src <现场包目录> --scope vib --dry-run   # 先看计划
+.venv\Scripts\python.exe scripts\place_raw_data.py --src <现场包目录> --scope vib             # 落位 (150 GB)
+.venv\Scripts\python.exe scripts\vib_raw_build.py --jobs 10                                   # 摄入 (索引+谱)
+.venv\Scripts\python.exe scripts\vib_reports_build.py --dump                                  # 只打印转录, 人工核对
+.venv\Scripts\python.exe scripts\vib_reports_build.py                                         # 写产物
+.venv\Scripts\python.exe scripts\scan_stations.py                                             # 两个新目录是否被认到
+```
+
+**落位**(实测): `windcms\CMS_RuDong_CGN_202603-04` 25,679 件 150.1 GB +
+`windcms\厂家报告\上海电气_月度` 13 件 51.4 MB + `m5_cms_tcm\厂家报告` 1 件 20.8 MB = **25,693 件 / 150.2 GB**
+(`rc=0`;余量要求: `F:` 需 ≥ 165 GB)。重跑幂等(同尺寸跳过;实测第二次跳过 2 件已存在的 docx)。
+
+**摄入**(实测, 10 并行): 索引步 **392.8 s** → `windows\w0316\index.parquet` **2,066,686 行 × 54 列**
+(标量行 861,917 · 谱行 1,204,769;38 台;时间 2026-03-16 17:27:22 → 2026-04-21 10:18:33);
+谱步 **477.2 s** → **420,742 条谱** / **1,712 个 npz 分片**(只转 `FFT_`,故小于"谱行"数;
+`Size>1` 行里另有 `Time_*` 波形 78 万行未转)。窗体合计 **3.27 GB**。全链 **875 s(15 分钟)**。
+
+**消费端**(实测, 改代码前后都是同一套消费者):
+
+| 环节 | 实测结果 |
+|---|---|
+| `data.windows()` | 发现 `['w0127', 'w0316']` —— 新窗自动进分析集 |
+| `data.load_scalars()` | 193,902 行 × 11 列(w0316 贡献 168,880;7 个关键标量齐) |
+| `data.spectra_meta()` | 420,742 行 × 18 列,8 个测点全 |
+| `data.spectrum()` | **✅ 取到真实谱**: `WTG01 / Gear_IMS / FFT_6000_Tr` → 6,401 点,x 0→6000 Hz,y 单位 m/s²(此前 `m5/spectra` 缺失 ⇒ 只能返回 None) |
+| 索引列契约 | `data._read_index()` 要的 10 列齐;无缺列 |
+
+**厂商报告转录**: 上海电气 2026-07 → 38 台,取严 `{优秀34, 良好3, 不可判1}` 与报告原文总览句
+`优秀34/良好3/预警0/无数据1/测点异常1` 逐项一致;大生科技 2026-03-11 → 38 台
+`{优秀29, 良好7, 预警1, 数据异常1}`(与其正文"WTG02 未取到数据"一致);
+消费端解析器(取 `## 附录 A` 后 `| WTG` 行第 6 列)读到 38 台不缺台。
+
+**来源台账**(`outputs/rudong/_provenance.json`): 改造前 **17 raw-derived / 572 shipped**;
+本次摄入+转录后 **1,740 / 571**(`m5_cms_tcm` 由 0/90 变 1,718/90、`windcms` 由 0/56 变 2/54)。
+逐件登记由 `src/derived_manifest.py` 承载(谁算的谁登记),`products_restore_missing.py` 读它;
+已存在的窗若漏登记,用 `python scripts/vib_raw_build.py --register-only` 幂等补登记。
+
+**页面侧**(`http://127.0.0.1:28084/detail/v2#tab=system` 的数据层表,接口 `/detail/api/maint_survey`):
+两行位置已改为 `<安装目录>\data\raw\如东\windcms\` 与 `<安装目录>\data\raw\如东\m5_cms_tcm\`。
+★ 组件服务**启动时会把这张表烘进缓存**,改完 `src/ontology/maintenance.py` 必须重启组件才生效
+(本次用运营台同一套脚本: `_ops_stop_keep_gateway.py` + `_ops_start_and_open.py --no-open`)。
+
+**其余闸门**: 全库 183 个 py 在 `-W error::SyntaxWarning` 下编译 0 失败;本体审计 `rc=0`;
+`check_transferable.py` 命中数 350 → **344**(`data\raw` 已排除出文本扫描;我的产出 0 处,
+manifest 里的源件路径已改为相对安装根)。
+
+## 7. 仍然做不到的(如实列)
+
+- **六层链四步未随包**: `rudong_tcm_oem_scan.py` / `rudong_line_energy_share.py` / `rudong_model_run.py` /
+  `rudong_fusion_run.py`(原在振动线分支 `claude/vibration-data-diagnosis-32b69e`)。
+  因此扫描线/能量占比/模型层/融合层那几类 parquet(`oem_frequency_scan`、`gear_freq_scan`、
+  `blade_1p_*`、`model_run_l6`、`fusion_38.csv` 等)**仍走包内随包快照**,不由 `data\raw` 重算;
+  连带的后果是 **CMS 报告/页面不能重生成**(§3b 实测 −97%)。`vib_raw_manifest.json` 的
+  `missing_chain` 字段如实列着这四项;缺口台账里对应 **B5**。
+- **扫描件 PDF 的数值**取不出(无文本层, 本包不带 OCR)。
+- `handoff_vibration_v2.json` / `component_history.json` 目前是**随包快照**:它们含大量人工裁决、
+  校准更新与开放项,不是能从测量数据直接算出来的东西 —— 现场给正本才改由现场件驱动。
+- 谱库的磁盘代价:整窗全转(`--meas ALL`)会显著吃盘(含 `Time_*` 波形, 单点 65,536/200,000);
+  默认只转 `FFT_` 谱(本轮 42 万条谱 = 3.2 GB)。
+- 一个可选的后续:把纯 Python 的 `pypdf` 轮子并入 `wheels\`,让 `vib_reports_build.py` 能**离线**
+  逐件核验扫描件页数/有无文本层(本轮是靠临时装的 pypdf 抽检一件得出的结论,清单里记了这一点)。

+ 69 - 9
docs/数据目录结构与落位约定_v0.2.md

@@ -16,7 +16,9 @@
 │         ├─ scada_1min\              (可选) 1min 导出
 │         ├─ 故障报警\                 报警事件导出 (SpreadsheetML *.xls)
 │         ├─ 风机故障记录\             检修工单台账 (*.xls/xlsx, 内按 {年}年故障记录\ 分年)
-│         └─ 油样报告\                 油液化验报告 (*.pdf)
+│         ├─ 油样报告\                 油液化验报告 (*.pdf)
+│         ├─ windcms\                 ★振动侧: CMS 原始测量导出 + 厂商月度评估报告 (2026-09-12 新增)
+│         └─ m5_cms_tcm\              ★振动侧: TCM 侧深度分析报告 + 现场给的 handoff 正本
 ├─ outputs\rudong\                   ★系统取数用的**产物仓** (页面 99% 读这里, 不直接读 raw)
 │    ├─ windscada\                   L0 标准仓: 37 个 parquet + 索引/日志
 │    ├─ ontology\                    本体对象库 objects.json + 检索索引 + release_r1/r2
@@ -38,9 +40,8 @@
 ├─ reference\rudong\                 契约与语言资产 (windscada_contract.yaml / 英文语言库 / 报警码表 …)
 ├─ wheels\win_amd64\                 离线轮子 (随包分发, 不入 git)
 ├─ vendor\                           便携 Python / Ollama 离线包 (不入 git)
-├─ docs\                             说明书与本文
-├─ logs\ run\                        运行日志 / pids.json (不入 git)
-└─ _修复记录_20260911\                现场修复记录与可重放脚本
+├─ docs\                             说明书、落位约定、振动接入与本文
+└─ (不随包: `_修复记录_*`/临时目录 —— 本包已纳入 git, 变更历史由版本库承载)
 ```
 
 ---
@@ -58,7 +59,7 @@
 | ③ | `data\raw\` 下**只有这一个**场站目录(单站部署) | 采用它, 但报告里标"凭单站唯一性", 不假装精确匹配 |
 | ④ | 多目录且都不匹配 | **不猜**: 报"未识别", 页面显示无数据并列出扫到的目录 |
 
-约定子目录(这五个名字是摄入接口, 改名等于换接口):
+约定子目录(这些名字是摄入接口, 改名等于换接口; 前五个是 SCADA/台账侧, 后两个是振动侧, 见 §2b):
 
 | 子目录 | 放什么 | 文件形态要求 |
 |---|---|---|
@@ -67,6 +68,8 @@
 | `故障报警\` | 报警事件导出 | `.xls` 实为 SpreadsheetML(XML), `<row>` 需带 `TimeOn` + `Alarmcode`;「全年/年至今」累计快照件会自动跳过 |
 | `风机故障记录\` | 检修工单台账 | 模板表(表头含 `机组编号`+`故障名称`+`故障代码`);表里 `风场名称` 必须属本场(集团导出件混着十来个场);按年分组不限层级 |
 | `油样报告\` | 油液化验报告 | `.pdf`, 文件名须含 `日期_台号_部件`(如 `17072025_10303681_…_1#_主轴后.pdf`) |
+| `windcms\` | CMS 原始测量导出 + 厂商月度评估报告 | 导出为 `测量根\<年>\<月>\<WTGxx>\*_decode.json`(Brande TCM 导出, 目录层级不限); 报告为 `.docx`(可解析) 或 `.pdf`(扫描件只归档) |
+| `m5_cms_tcm\` | TCM 侧深度分析报告 + handoff 正本 | `.docx` 报告; 若现场给出 `handoff_vibration_v2.json` / `component_history.json` 正本, 放这里即被优先采用 |
 
 机理层(厂商资料)按 A2 仍放 `data\raw\西门子4.0技术资料\`, **不在场站目录下**。
 
@@ -82,13 +85,61 @@
 |---|---|---|
 | `--scope a2`(默认) | 上表四类(`scada_10min`/`故障报警`/`风机故障记录`/`油样报告`) | A2 约定的四项数据层 |
 | `--scope mech` | `data\raw\西门子4.0技术资料\`(317 件 2.7 GB,含那份**对译表**)+ `<场站>\scada_1min\`(38 件 12.8 GB) | 本体层构建器 `kb_ingest.py` 的 `TECH` 与 `rd()` 四个文件名;`config.STATION_SUBDIRS` 声明的 `scada_1min` |
-| `--scope full` | 两组一起 | 从零重算要用的全集 |
+| `--scope vib` | `<场站>\windcms\`(CMS 原始测量导出 25,679 件 150 GB + 上海电气月度报告)+ `<场站>\m5_cms_tcm\`(大生科技 TCM 报告) | §2b |
+| `--scope full` | 三组一起 | 从零重算要用的全集 |
 
 > 为什么 `scada_1min` 也属"该落的": 它写在 `config.STATION_SUBDIRS` 里, `scan_stations.py` 会把它报成
 > "缺"。但**包内没有消费者** —— 落它是补数据层完整性, 不改任何页面数值。同理, 现场包里的
-> `scada数据(如东)\`(19 个月原始通道导出)、`fastlog数据\`、`(8)…\振动分析报告\` 三种**故意不落**:
-> 前者是 `scada_10min` 的上游且无人读, 中间那种全库只有 2 处注释提到, 后一种是振动线的成品牌报告
-> 而非可再加工的测量数据 —— 理由都逐条写在脚本的 `SKIPPED` 表里。
+> `scada数据(如东)\`(19 个月原始通道导出)、`fastlog数据\` 两种**故意不落**:
+> 前者是 `scada_10min` 的上游且无人读, 后者全库只有 2 处注释提到 —— 理由都逐条写在脚本的 `SKIPPED` 表里。
+>
+> ★ 关于 `(8)…\振动分析报告\`(12 份月度用印版 PDF): v0.2.0 初版把它列在"故意不落"里, 理由是
+> "成品牌报告而非可再加工的测量数据, 且 windcms 要的测点索引包里没有"。**那条判断在 2026-09-12 被推翻**:
+> 用户把 CMS 原始测量导出 (`CMS_RuDong_CGN_202603-04.zip`) 补进了现场包, 索引源件已具备, 于是
+> 报告作为数据层证据一并落位(见 §2b)。脚本的 `SKIPPED` 表里保留了这条记录并标注了推翻原因。
+
+---
+
+## 2b. 振动侧落位与摄入链(2026-09-12 新增)
+
+**用户令**: 「数据层里的「CMS 振动评估报告」应遵循 `<安装目录>\data\raw\如东\windcms`、
+「振动线 handoff」应遵循 `<安装目录>\data\raw\如东\m5_cms_tcm` 存放; 修改系统支持振动数据参与运行、重算;
+自现场包提取相应振动数据存放至上述目录」。
+
+落位后的实际形态(`--scope vib`):
+
+```
+data\raw\如东\
+├─ windcms\
+│    ├─ CMS_RuDong_CGN_202603-04\          ← 原始测量导出原样保留包内层级 (可追溯"哪来的")
+│    │    └─ measurement\2026\{03,04}\WTGxx\WTGxx_<uuid>_decode.json     25,679 件 / 150 GB
+│    └─ 厂家报告\上海电气_月度\
+│         ├─ 中广核如东海上风电场2025年06月…2026年05月振动分析报告用印版.pdf   12 件 (扫描件, 只归档)
+│         ├─ 中广核如东海上风电场2026年07月振动分析报告_上海电气.docx          (有文本层, 摄入)
+│         └─ _扫描件清单.json                                                (登记: sha256/大小/无文本层)
+└─ m5_cms_tcm\厂家报告\
+     └─ 中广核如东海上风电场传动链振动分析报告_大生科技_20260311(1).docx      (TCM M-system, 摄入)
+```
+
+**摄入链(两条, 分工不同, 别混)**:
+
+| 源 | 摄入命令 | 产出 | 谁消费 |
+|---|---|---|---|
+| CMS 原始测量导出(`*_decode.json`) | `python scripts\vib_raw_build.py`(= `rudong_tcm_index.py` → `rudong_tcm_spectra.py`;★**不**跑 `windcms.py report` —— 缺六层链产物时重生成会掉内容, 实测 `report.md` −97%, 见 `docs\振动数据接入_v0.1.md` §3b) | `outputs\rudong\m5_cms_tcm\windows\<窗>\index.parquet`(54 列)+ `…\<窗>\spectra\*.npz` + `spectra_meta.parquet` | `src\windcms\data.py`(窗/标量/谱图自动收录)、`scripts\windcms.py report/serve`、`src\windscada\taxonomy.py`(设备状态转录) |
+| 厂商 docx 评估报告 | `python scripts\vib_reports_build.py` | `outputs\rudong\windcms\厂家报告提取_<报告期>.json` + `报告_CMS振动状态评估报告_<报告期>.md`;`outputs\rudong\m5_cms_tcm\报告_TCM传动链振动分析_<报告期>.md` | 同上(报告 md 与自产报告**同构**, 消费端不必改代码) |
+
+窗名口径: `wMMDD` = **数据起始日**(与既有 `w0127`/`w0707`/`w0811` 同口径, `pipeline.window_year_gate`
+会拦"窗名无年份"的历史误用)。同一批数据重复摄入时, 已存在的窗会改名为 `<窗>_reimport_<时分>`,
+而 `data.EXCLUDE_DEFAULT` 会把带 `_reimport` 的窗排除在生产集外 —— 重复摄入不会污染分析集(2026-08-24 的教训)。
+
+**这一步之后"振动数据就真的参与了"的证据链**: `data.windows()` 自动发现新窗 → `load_scalars()` 把
+新窗标量并进 CMS 报告 → `data.spectrum()` 能取到谱(此前 `m5\spectra` 整个目录缺失, 谱图是死的)→
+`fusion.windcms_grades()` / `taxonomy.system_matrix()` 读最新报告 md 做设备状态转录。
+
+**仍然缺的(如实标注, 不冒充)**: 六层链的四步脚本 `rudong_tcm_oem_scan.py` / `rudong_line_energy_share.py` /
+`rudong_model_run.py` / `rudong_fusion_run.py` **未随包**(原在振动线分支 `claude/vibration-data-diagnosis-32b69e`)。
+`vib_raw_build.py` 产出的 `outputs\rudong\m5_cms_tcm\vib_raw_manifest.json` 里 `missing_chain` 字段列明这四项。
+因此扫描线/能量占比/模型层/融合层那几类 parquet 仍走"包内 shipped 快照", 只有**索引/谱/报告/知识库**是可重算的。
 
 ---
 
@@ -102,6 +153,8 @@
 | `故障报警\` | 否(只有摄入脚本读) | `scripts\windscada_alarms_ingest.py` | 数据层「报警事件」、故障分析、停机事件、限电绑定 |
 | `风机故障记录\` | 否 | `scripts\windscada_workorder_ingest.py` | 数据层「检修工单台账」、检修面、闭环验证 |
 | `油样报告\` | 否 | `scripts\windscada_watch_channels_build.py` | 数据层「油液化验」、油液时效胶囊、融合面油样轴 |
+| `windcms\`(CMS 原始导出) | 否(窗是**摄入时**落盘的) | `scripts\vib_raw_build.py`(索引→谱→报告/知识库) | CMS 系统的谱图/标量与报告(`windcms.py serve`);`taxonomy` 的设备状态转录 |
+| `windcms\`/`m5_cms_tcm\`(厂商报告) | 否 | `scripts\vib_reports_build.py` | 数据层的厂商评估报告转录(与自产报告同构的 md) |
 
 命令(在 `<安装目录>` 下跑):
 
@@ -111,6 +164,7 @@
 .venv\Scripts\python.exe scripts\rebuild_from_raw.py --scada # 再加 SCADA 侧 10 个构建器 (慢)
 .venv\Scripts\python.exe scripts\rebuild_from_raw.py --verify# 与随包基线逐值等价验收
 .venv\Scripts\python.exe scripts\place_raw_data.py           # 现场压缩包 → 约定子目录 (映射表)
+.venv\Scripts\python.exe scripts\vib_raw_build.py            # 振动侧: CMS 原始导出 → 窗索引/谱/报告
 ```
 
 ---
@@ -129,6 +183,9 @@
 | `powercurve_dev/bins` · `loss_monthly` · `curve_lenses/liveness` · `control_profile/schedule` · `stop_events` · `temp_bins` · `yaw_daily` · `hydraulic_accum` · `thermal_chain` · `system_aux` | scada_10min + alarms | `rebuild_from_raw.py --scada`(实测: `loss_monthly` 3729/3729、`powercurve_dev` 38/38、`powercurve_bins` 912/912 与随包基线**逐值完全一致**) |
 | `objects.json`(本体 2340 个对象)· `retrieval_index.json` | **厂商技术资料** `data\raw\西门子4.0技术资料`(不在场站目录下) | `python -m src.ontology.kb_ingest`;检索索引 `python -c "from src.ontology import retrieval as R; R.build(use_vec=False)"`(BM25 词法索引; 向量那半要 Ollama 的 `bge-m3`, 本机没装模型时会跳过并保持纯词法可用) |
 | `turbine_params.parquet`(1706 条整定值) | 同上, `Turbine+Parameters.*.xlsx` | `python -c "from src.ontology.maintenance import refresh_params as f; f()"` |
+| `m5_cms_tcm\windows\<窗>\index.parquet`(54 列标量索引) | **CMS 原始测量导出** `data\raw\如东\windcms\**\*_decode.json` | `scripts\rudong_tcm_index.py`(旧包缺这个脚本 → 2026-09-12 补齐; 列名列序 dtype 与包内 `tcm_index.parquet` 逐列对齐, 实测 FFT 行取值逐格一致) |
+| `m5_cms_tcm\windows\<窗>\spectra\*.npz` + `spectra_meta.parquet` | 同上 | `scripts\rudong_tcm_spectra.py`(`DataSets.DataSet.Values` 是空格分隔字符串: `Size=Lines+1`, `X-axisDelta=带宽/Lines`) |
+| `windcms\报告_CMS振动状态评估报告_<报告期>.md`(厂商报告转录) | 厂商 docx 评估报告 | `scripts\vib_reports_build.py`(合并单元格还原 + 厂家总览句作**校验和**: 实测 34优秀/3良好/1无数据逐项一致) |
 
 > 本体层这三件的前提是**技术资料在盘上**(`place_raw_data.py --scope mech` 会落, 见 §2);技术资料不在时
 > `kb_ingest` 会打印 `⚠ 源缺失` 并降级 —— 不静默。
@@ -300,3 +357,6 @@ P.venv_python()   # 跨平台探测 .venv/Scripts/python.exe 或 .venv/bin/pytho
 | 2026-09-11 | 验证闸门退出码修复:控制台 GBK 编不出 `✔` 时降级为 `?`(新增 `src\console.py`),修 `page_fingerprint` / `check_portability` / `rebuild_from_raw` / `scan_stations` / `portal_build` —— 此前"全部一致"会因 print 抛异常而**退出码 1(通过被报成失败)** |
 | 2026-09-11 | **从零重算实测** + 落位扩范围:`place_raw_data.py` 新增 `--scope a2\|mech\|full`(机理层技术资料 + 对译表 + `scada_1min`)、修"单文件规则恒不落件"与缺 `import os`(自 d90da05 起脚本跑不通)、同尺寸文件跳过;`src\ontology\` 全层 10 处文本 I/O 补 `encoding='utf-8'`(**中文 Windows 上本体层原本根本不可用**:默认 cp936 写 objects.json 崩在 `\u200b`);`check_portability.py` 新增 WARN 项 6「文本读写缺 encoding=」;实测数字见 §4 末 |
 | 2026-09-11 | 收尾这 18 处同类文本 I/O(`scripts\ontology_p2_verify.py`、`scripts\windscada_serve.py`、`src\ontology\scenario_29.py`、`src\windcms\{knowledge,pipeline,plugins}.py`、`src\windscada\subsys\fusion.py`、`src\windscada\taxonomy.py`),门禁 WARN 项 6 扫描面同时扩到 `open('w')` 与 `open(p)`(无 mode = 文本读);**WARN 归零** |
+| 2026-09-12 | **振动侧落位与摄入(本文 §2b)**:`data\raw\如东\` 新增 `windcms`(CMS 原始测量导出 25,679 件 150 GB + 上海电气月度报告) / `m5_cms_tcm`(大生科技 TCM 报告) 两个约定子目录(`config.STATION_SUBDIRS` 同步);补齐振动线分支缺失的摄入端 `scripts\rudong_tcm_index.py`(54 列窗索引, 与包内 `tcm_index.parquet` 同构)与 `scripts\rudong_tcm_spectra.py`(npz 谱库 + `spectra_meta.parquet`);新增一键 `scripts\vib_raw_build.py`(索引→谱→报告/知识库, 并行+窗体清单)与厂商报告摄入 `scripts\vib_reports_build.py`;`place_raw_data.py` 新增 `--scope vib`(含散装件规则与**每次开包一次**的写法修正: 旧写法对 2.5 万条目是平方复杂度);`rebuild_all.py` 新增 ④b 步;`check_transferable.py` 文本扫描排除 `data\raw\`(现场原始件不随包, 扫它只会把秒级闸门拖成小时级)。**细节见 `docs\振动数据接入_v0.1.md`** |
+| 2026-09-12 | **不再随包"修复记录/临时"类目录**(用户令: 本包已纳入 git 管理): `pack_dist.py` 的 `INCLUDE_DIRS` 去掉 `_修复记录_20260911`(其内容并入 `docs\振动数据接入_v0.1.md` 与本文 §2b), 相关文档/README 的引用一并改指正式文档 |
+| 2026-09-12 | 振动侧**实测复核**(见 `docs\振动数据接入_v0.1.md`): 落位 25,693 件/150.2 GB; 摄入出窗 `w0316`(2,066,686 行 × 54 列 + 420,742 条谱, 全链 875 s); `data.windows/load_scalars/spectrum` 三处消费实测通过(**谱图此前是死的**)。★ 同时逮到一条坑: 在本包跑 `windcms.py report` 会**用残缺输入覆盖完整快照**(`report.md` −97%)—— 因为六层链的 `model_run`/`fusion` 产物缺失, 故该步改为 `--with-report` 显式开启, `outputs\rudong\windcms` 已逐字节还原为随包快照(56 件/0 差异)。新增 `src\derived_manifest.py`(产物来源自登记)让台账认出新造件 |

+ 3 - 2
docs/移植与独立运行_v0.1.md

@@ -75,7 +75,7 @@ sh install.sh
 
 | 项 | 为什么 |
 |---|---|
-| `_products_off/` `_products_off_prev_*/` | 本机"清除产物"的暂存与存档(每份都是整份产物 ~220 MB); 目标机用不到 |
+| `_products_off/` `_products_off_prev_*/` | **旧设计**的"清除产物"暂存与存档(每份都是整份产物 ~220 MB)。2026-09-16 用户令"清除产物不留备份"后**不再产生**; 老机器上若还有, 确认不需要可手工删除 |
 | `logs/` `run/pids.json` | 本机运行日志与**旧 PID**; pids.json 到新机器上是无效引用(启动器会重写) |
 | `.git/` | 55 MB 仓库; 交付不需要(自己留版除外) |
 | `data/raw/`(视交付约定) | 30.9 GB 现场原始件; 按约定"原始件不随包分发"时不必带(缺它只影响"重算原始数据", 不影响页面) |
@@ -94,4 +94,5 @@ sh install.sh
 > 现状(2026-09-12 实测): 文本里仍有 350 处机器相关路径, **但都不在运行路径上** —— 分布是
 > `release/viewer/*`(三维单位的 `source_path` 出处注记, 186 处) · `outputs/**`(产物里记的历史路径/provenance, 95) ·
 > `src/sop/farm_paths.py`(盘符映射表本体, 门禁具名登记) · `scripts/build_oem_lexicon.py`(开发机 OEM 资料盘) ·
-> `_修复记录_20260911/fix_*.py`(一次性修复脚本留档) · 文档里的反面示例。**都不影响在别的机器上运行**。
+> 文档里的反面示例。**都不影响在别的机器上运行**。(`_修复记录_*` 这类目录已按用户令不再随包;
+> 文本扫描同时排除 `data/raw/` —— 现场原始件不随包分发, 而单那一个目录就是 2.5 万个文件/150 GB。)

+ 294 - 0
docs/系统设计说明.md

@@ -0,0 +1,294 @@
+# 观澜 · 如东样板 v2 · 系统设计说明
+
+> 版本 v0.2.0 · 2026-09-16 · 适用 `<安装目录>`(安装目录)
+> 本文是**系统设计正本**: ① 产物全景(类型/功用/输出路径/生成端/消费端);② 路径与进程的统一约定;
+> ③ 输入 ↔ 产物呼应关系;④ 重算的两种运行状态与"产物即时进页面"的实现;⑤ 启动与门户。
+> 与本文配套、可复跑的机器校验: `python scripts/inventory_products.py`(清点 + 呼应校验,`--check` 出退出码)。
+
+---
+
+## 1. 系统是什么(一页)
+
+```
+现场数据 data/raw/<场站>/  ──摄入──▶  outputs/<场>/(产物仓)  ──现读──▶  页面/接口
+     ▲  7 类约定目录                        ▲  9 个产物仓                  ▲
+     │  scada_10min / scada_1min /           │  windscada  L0 标准仓        │  网关 :28084(门户/运维控制台)
+     │  故障报警 / 风机故障记录 /             │  ontology   本体对象库        │  detail :18033(分析工作台)
+     │  油样报告 / windcms / m5_cms_tcm      │  windcms   CMS 振动           │  cms    :18020(振动诊断)
+     └─ 西门子4.0技术资料(机理层,不在场站目录下)│  m5_cms_tcm 振动线出件/窗     │  sim :18791 · sim_sys :18792 · viewer :64292
+                                            │  tcm_compatible_replay      │
+                                            │  sop / guanlan / pitch / paradigm_r1
+```
+
+四条设计铁律(贯穿全部代码,改代码时先读这四条):
+
+1. **路径唯一真源** `src/paths.py` —— 代码/配置里只写 ROOT 相对路径,运行期由助手解析成绝对路径(基准是安装根,**不是 cwd**)。
+2. **产物只落在 `outputs/<场>/`**;原始件只读 `data/raw/<场站>/`;`release/` 是交付静态层。三者不互相写。
+3. **判断在代码、模型只转述**;每个产物都要能说清"谁生成、谁消费、来源是 raw 重算还是随包补齐"(`_provenance.json`)。
+4. **窗口/进程统一口径** `src/proc.py` —— 子进程一律 `CREATE_NO_WINDOW`,日志落 `logs/`,不弹命令窗口。
+
+---
+
+## 2. 产物全景(自动生成,勿手改)
+
+> 由 `python scripts/inventory_products.py --write-doc` 生成;数据源是**盘上的实际文件** +
+> `outputs/<场>/_provenance.json`(逐件来源台账)+ `outputs/<场>/_derived_manifest.json`(构建脚本自登记)。
+
+<!-- INVENTORY:BEGIN (由 scripts/inventory_products.py --write-doc 生成, 勿手改) -->
+*自动生成于 2026-09-16 16:26;数据源: `outputs/rudong/_provenance.json` + `_derived_manifest.json` + 实际文件*
+
+| 产物仓 | 输出路径 (相对安装根) | 功用 | 生成端 | 消费端 | 件数 | 大小 | 来源(raw 重算/随包) |
+|---|---|---|---|---|---|---|---|
+| `windscada` | `outputs/rudong/windscada/` | L0 标准仓: SCADA/台账/派生分析的全部 parquet (页面主取数处) | rebuild_from_raw.py (三门台账) + --scada (10 个构建器) + windscada_monthly_build.py | scripts/windscada_serve.py 各视图 · src/windscada/taxonomy · subsys/fusion | 115 | 4.5 MB | 17 / 98 |
+| `ontology` | `outputs/rudong/ontology/` | 本体对象库: 码表/手册/工单展开/失效树 + 检索索引 + 实机参数 | python -m src.ontology.kb_ingest → populate → chain_ingest → trend_ingest → retrieval.build | 脚本 windscada_serve.py 本体页/问答 · scripts/guanlan_facts_contract.py | 25 | 82.9 MB | 3 / 22 |
+| `windcms` | `outputs/rudong/windcms/` | CMS 振动诊断产物: 状态评估报告/逐台页/工作台页/知识库 + 厂家报告转录 | scripts/windcms.py report/kb · scripts/vib_reports_build.py | 自服务 :18020 (每请求现读) · taxonomy.system_matrix (转录设备状态) | 56 | 30.1 MB | 2 / 54 |
+| `m5_cms_tcm` | `outputs/rudong/m5_cms_tcm/` | 振动线出件与窗级分析: handoff 接口 + 窗索引/谱库 + TCM 兼容件 | scripts/vib_raw_build.py (窗索引/谱) · 振动线出件 (handoff, 随包快照) | src/windscada/subsys/fusion.py · src/windcms/data.py · scripts/windcms.py report | 1808 | 3329.7 MB | 1718 / 90 |
+| `tcm_compatible_replay` | `outputs/rudong/tcm_compatible_replay/` | TCM 兼容链回放资产 (模型表/掩码阈值/裁决记录) | 随包快照 (无生成端) | src/windcms/report*.py · config.mask_thresholds | 59 | 32.7 MB | 0 / 59 |
+| `sop` | `outputs/rudong/sop/` | SOP 中间件/评审/台账与事实契约底稿 | 随包快照 (无生成端) | scripts/guanlan_facts_contract.py · 门户结论段 | 210 | 15.8 MB | 0 / 210 |
+| `guanlan` | `outputs/rudong/guanlan/` | 事实契约与对外派生 (可上云面孔) | scripts/guanlan_facts_contract.py | 门户 #findings · /api/facts | 5 | 0.5 MB | 0 / 5 |
+| `pitch` | `outputs/rudong/pitch/` | 变桨侧派生件 (零位/日粒度) | 随包快照 + rebuild_from_raw --scada | 脚本 windscada_serve.py 变桨面 | 2 | 0.3 MB | 0 / 2 |
+| `paradigm_r1` | `outputs/rudong/paradigm_r1/` | 范式实验件 (E3/E5/E8 底稿, 事实契约输入) | 随包快照 (无生成端) | scripts/guanlan_facts_contract.py | 29 | 0.5 MB | 0 / 29 |
+
+合计 2309 件 / 3496.9 MB;其中 raw 重算 1740 件、随包补齐 571 件。
+<!-- INVENTORY:END -->
+
+### 2.1 逐件来源口径(`_provenance.json`)
+
+| 来源 | 含义 | 判据 |
+|---|---|---|
+| `raw-derived` | 由 `data/raw` 重算出来的 | 该件在 `RAW_DERIVED` 表里,或由**构建脚本自登记**(`src/derived_manifest.py`) |
+| `shipped` | 包内没有生成端 / 规则未复现,用随包件补齐 | 既不在上表、也无自登记,且件在随包快照里存在 |
+
+*自登记(2026-09-16 新增)*: 振动侧的产物名随"窗名/分片序号"变化,写不进精确路径表,而按名字通配会误伤
+同名旧件(`报告_CMS振动状态评估报告_*.md` 既有随包/自产的、也有厂家报告转录的)。改成**谁算的谁登记**:
+`vib_raw_build.py` / `vib_reports_build.py` 落盘后把相对路径写进 `_derived_manifest.json`,
+`products_restore_missing.py` 生成台账时据此记为 `raw-derived`。漏登记可用
+`python scripts/vib_raw_build.py --register-only` 幂等补登记。
+
+---
+
+## 3. 路径约定(唯一真源 `src/paths.py`)
+
+### 3.1 全部路径助手(写新代码时只用这些)
+
+| 助手 | 返回 | 用途 |
+|---|---|---|
+| `P.ROOT` | `<安装目录>` | 一切解析的基准(`WINDSCADA_ROOT` 可覆盖,冻结构建时由启动器设) |
+| `P.RAW_ROOT` | `data/raw` | 现场原始件根(`WINDSCADA_RUDONG_SRC` / `serve.json.raw_dir` 可覆盖) |
+| `P.station_dir(name)` | `data/raw/<场站名称>` | 兜底约定位置(权威值来自 `config.raw_station_dir()` 的扫描辨识) |
+| `P.out_root(name)` | `outputs/<场>` | 产物仓根 |
+| `P.store(name)` | `outputs/<场>/windscada` | L0 标准仓(页面主取数处) |
+| `P.ont(name)` | `outputs/<场>/ontology` | 本体对象库 |
+| `P.objects_json(name)` | `…/ontology/objects.json` | 对象库文件 |
+| `P.cms(name)` | `outputs/<场>/windcms` | CMS 振动诊断产物 |
+| `P.m5(name)` | `outputs/<场>/m5_cms_tcm` | 振动线出件 / 窗索引 / 谱库 |
+| `P.tcm_replay(name)` | `outputs/<场>/tcm_compatible_replay` | TCM 兼容链回放资产 |
+| `P.sop(name)` | `outputs/<场>/sop` | SOP 中间件与评审落盘 |
+| `P.guanlan(name)` | `outputs/<场>/guanlan` | 事实契约与对外派生 |
+| `P.pitch(name)` | `outputs/<场>/pitch` | 变桨侧派生件 |
+| `P.paradigm(name)` | `outputs/<场>/paradigm_r1` | 范式实验件(E3/E5/E8 底稿)**2026-09-16 新增** |
+| `P.report_dir(name)` | `outputs/<场>/report` | 报告交付件(交接单/现场单)**2026-09-16 新增** |
+| `P.cloud(name)` | `outputs/<场>/guanlan/cloud` | 可上云面孔(脱敏后的契约/派生/页面)**2026-09-16 新增** |
+| `P.contract(name)` | `reference/<场>/windscada_contract.yaml` | 机型判据契约(属 reference 侧,不在 outputs) |
+| `P.rel(p)` / `P.disp(p)` / `P.disp_dir(p)` | 相对 POSIX 串 / 显示串 / 目录显示串 | 写进产物用 `rel()`;给人看(页面「位置」列)用 `disp()` |
+| `P.resolve(p)` | 绝对路径 | 把"可能是相对"的值按 **ROOT**(不是 cwd)解析 |
+| `P.venv_python()` / `P.python_exe()` | 解释器路径 | 跨平台探测 `.venv/Scripts/python.exe` 或 `.venv/bin/python` |
+
+### 3.2 2026-09-16 统一掉的路径问题(源码已改)
+
+| # | 位置 | 原样 | 改成 | 为什么要改 |
+|---|---|---|---|---|
+| 1 | `scripts/products_restore_missing.py` | `P.STORE if hasattr(P,'STORE') else ROOT/'outputs'/'rudong'` | `P.out_root()` | `P.STORE` **不存在** ⇒ 恒落到写死的 `rudong`;多场部署会把 rudong 的随包件补进别的场(静默串场) |
+| 2 | `scripts/ingest_ops_2025.py` | `pathlib.Path('outputs/rudong/ontology/objects.json')` + 裸 `write_text` | `_P.objects_json()` + `Store(…).save()` | cwd 相对 + 写死场名 + **绕过 store 的全库校闸与原子写**;且该脚本缺 `import os` 一直 NameError,属"上膛但没响"的凶器 |
+| 3 | `scripts/ontology_p2_verify.py` | `ROOT/'outputs/rudong/ontology/objects.json'`、`…/guanlan/facts_contract_v0.json` | `P.objects_json()`、`P.guanlan()` | 写死场名 |
+| 4 | `scripts/guanlan_facts_contract.py` | 4 条 `ROOT/"outputs/rudong/…"` | `P.guanlan()/P.sop()/P.paradigm()` | 同上(事实契约的输入/输出全在这里) |
+| 5 | `src/ontology/maintenance.py` | 自写 `ROOT=parents[2]`、`_RAW_STR`、`disp()` | `P.ROOT` / `P.RAW_ROOT` / `P.disp` | 同一文件里三处"影子真源"(显示规则两个实现) |
+| 6 | `src/sop/wrapup.py` | `…/cleaned/turbine.parquet`(硬编码) | 读 `clean_gate.json` 的 `section` → 否则取目录内唯一 parquet → 都取不到则**响亮打印** | 写侧落的是 `cleaned/<section>.parquet`(section 可变)⇒ 读侧硬编码时静默 `None`,柱2 悄悄降级 INSUFFICIENT |
+| 7 | `src/windscada/subsys/pitch.py` | `P.store()/'alarms.parquet'` | `cfg['store']`(`registry(cfg)` 透传) | `P.store()` 跟环境变量,`cfg['store']` 跟显式选定的场 ⇒ 多场下跨场串数据且不报错 |
+| 8 | `scripts/check_transferable.py` | 门户期望字节数写死 `20225828` | 运行时从主实例现量 | 门户外壳一改常量即过期,核验会打印永远不成立的"与主包不同" |
+| 9 | `src/windscada/taxonomy.py`、`subsys/fusion.py`、`scripts/windscada_serve.py` | 读侧硬绑 `P.cms()` | 新增 `windcms.config.cms_out()`,读侧与写侧同一解析口 | `WINDCMS_OUT` 只有写侧认 ⇒ 设了它就"**写到旁路、页面读生产**",表现为"重算完了页面还是旧数"且**不报错**(冻结构建自检正是这个组合) |
+| 10 | `src/windcms/pipeline.py::ingest` | 任何 `.zip/.rar/.7z` 都调 `rudong_tcm_ingest_raw.py`(**未随包**) | 先判包内容:含 `*_decode.json` → 按"已解码导出"走 `rudong_tcm_index/spectra`(支持 zip 当 root);否则报明确错误并给两条可行路径 | 把 CMS 导出 zip 直接指过来时原本必崩 `FileNotFoundError`,而其实不需要那个脚本 |
+
+### 3.3 仍然存在的不统一(**如实列出,未擅自大改**)
+
+| 类别 | 具体 | 影响 | 处置建议 |
+|---|---|---|---|
+| 一次性/历史脚本自拼路径 | `scripts/guanlan_cloud_{face,page,qa}.py`、`guanlan_portal_inject_claims.py`、`guanlan_baseline_manifest.py`、`llm_known_answer.py` 里仍有 `outputs/rudong/…` 字面量 | 只影响这些**一次性派生/云端面孔**脚本;多场时需手改 | 换场前把这几处改走 `P.guanlan()/P.cloud()/P.store()`(`P.paradigm()/P.report_dir()/P.cloud()` 已备好) |
+| 历史随包留档 | `scripts/sim_hub/*`(含 `/Users/yuanying/...` mac 路径)、`outputs/rudong/sop/*.py` | **非运行期**(装配脚本留档) | 已在 `check_portability.py` 具名登记;随包只为"可重放",不参与运行 |
+| 多场机制未接通 | `src/windscada/config.py::available()` glob `configs/farms/*.json`,而磁盘上 11 个场配置是 `*.yaml` | 现在只有内置 `rudong` 可用;`cfg['store']` 恒等于 `P.store('rudong')` | 二选一:把 `available()` 同时认 `*.yaml`,或把场配置转成 json |
+| `store` seam 三种写法 | `cfg['store']`(14 处写点)/ `P.store()`(1 处,已修)/ `ROOT/'outputs'/…`(4 处,已修 3 处) | 多场前无实际差异,属"潜在缺陷" | 新代码一律 `cfg['store']` 或 `P.store(name)` |
+| 配置/依赖指向不存在的件 | `configs/scenario_registry.yaml`、`configs/<场>/value_assumptions.yaml`、`configs/analysis_lock*.yaml`、`scripts/sop_check.py`、`scripts/analysis_lock_check.py`、`docs/审核规则_经验固化_v1.md`、`docs/振动诊断模型_六层_v1.md`、`.claude/skills/…`、`m5_cms_tcm/model_run_l6.parquet`、`m5_cms_tcm/fusion_38.csv` | SOP 场景模块读注册表必抛;锁闸恒空转;`wrapup` 电价只能走"假定锚";CMS 报告**不能重生成**(会掉内容) | 见 §7 缺口表;这批是"随包没发全",要么补件要么显式降级(`wrapup` 已改为响亮说明) |
+| 被引用但未随包的脚本(10 个) | `rudong_tcm_ingest_raw.py`(已加前置判断,见 §3.2-10)、`rudong_tcm_oem_scan.py`、`rudong_line_energy_share.py`、`rudong_model_run.py`、`rudong_fusion_run.py`、`rudong_build_baseline.py`、`dsh_learn_page.py`、`intake_scan.py`、`guanlan_cloud_sims.py`、`deploy_gate_check.py` | 前五个属振动六层链(见 §7);其余是开发/部署辅助脚本的引用 | 逐个二选一:补件,或在调用处加"缺失即明确报错 + 替代路径"(`pipeline.ingest` 已按此改) |
+| 有读无写的"孤儿产物"(12 件) | `genbearing_monthly` / `mblub_monthly` / `yaw_dynamic_monthly` / `yaw1min_liveness` / `yaw_err_clean` / `sector_power` / `duty_monthly` / `pc_monthly_bins` / `thermal_monthly` / `control_monthly` / `structure.parquet` / `watch_channels_monthly` | 趋势件、热链、扇区、偏航/润滑/控制面 | 维持随包件;要"自己算"须研发给口径(照 `temp_monthly` 的办法反推 + 逐值验证) |
+| **有意**重复落盘(不是缺陷,但要知道) | ① `spectra_meta.parquet` 同时写 `<窗>/spectra/` 与 `<窗>/`(约 4.6 MB/窗)—— 为兼容 `data.spectra_meta()` 的两条读取约定;② `报告_CMS振动状态评估报告_<期>.md` 有两个来源(`report_std.py` 自算 vs `vib_reports_build.py` 厂家转录),靠**日期口径**区分(自算=当天、转录=报告期月末) | 磁盘占用翻倍(仅 meta,非谱数据);命名空间共享 | 已在此登记;`data.py` 优先读 `<窗>/spectra_meta.parquet`,两份内容逐字节一致 |
+| `P.VIEWER` 常量全仓 0 引用 | `src/paths.py` 定义了 `VIEWER`,实际取 `viewer_dir` 走 `configs/serve.json` | 无功能影响 | 下次统一时二选一(用起来或删掉),避免"看着像真源其实没人用" |
+
+---
+
+## 4. 进程与命令窗口(统一口径 `src/proc.py`)
+
+**问题**: Windows 上"**无控制台的父进程** + 裸 spawn 一个控制台程序" = 系统给子进程**新建一个可见控制台窗口**。
+本系统里无控制台的父进程很多(网关、运维动作进程),于是"点页面按钮"会闪窗、跑重算会挂一个 20 分钟的黑窗、
+`/ops` 每 2 s 轮询 `tasklist` 会反复闪窗。
+
+**统一写法**: `src/proc.py`
+
+| API | 用途 |
+|---|---|
+| `proc.NO_WINDOW` / `proc.flags()` | `CREATE_NO_WINDOW`(Windows)/ 0(POSIX)。★与 `DETACHED_PROCESS` **互斥**,只认这一种 |
+| `proc.spawn(cmd, log=…, env=…, cwd=…, **popen_kw)` | 后台起无窗口子进程;`log=` 落日志,也可自己传 `stdout=`/`stderr=` 句柄(**透传**,不覆盖) |
+| `proc.run(cmd, **kw)` / `proc.run_text(cmd)` | 前台等待的无窗口子进程(`tasklist` / `git` / 短命令);`run_text` 默认 `text=True, errors='replace'`(控制台输出是 GBK,严格解码会得到 `None`) |
+
+已收敛的调用点(**凡"父进程无可见控制台"的入口都必须走这里**): `guanlan.py`(组件 spawn、`tasklist`、`taskkill`)、
+`scripts/_ops_launch.py`、`scripts/_ops_run.py`(整套重算的宿主进程)、`scripts/guanlan_ops.py`(启动器调用 + 每 2 s 的 `tasklist` 轮询)、
+`scripts/guanlan_gateway.py`(`git`,每次 `/api/version`)、`scripts/_ops_start_and_open.py`(起 serve)、
+`src/windcms/plugins.py`(CMS 页面「分析」按钮起的全链)、`scripts/windscada_serve.py`(详情工作台云端档问答)。
+
+**判定口径**(新加子进程前先问一句): 我的父进程**有没有可见控制台**?有(用户从终端跑的 CLI)→ 让它继承,用户能看见进度;
+没有(服务/网关/运维动作进程)→ **必须**走 `src/proc.py`,否则 Windows 会给子进程新建一个可见窗口。
+
+### 4.1 三个"静默故障"坑(都实炸过,已写进 `guanlan.py check` 作回归守卫)
+
+| # | 坑 | 现场表现 | 处置 |
+|---|---|---|---|
+| 1 | `CREATE_NO_WINDOW` 新建的控制台**会把子进程标准句柄吸走** | 从页面点"执行重算":`_ops_run` 自己那行进了日志,`rebuild_all.py` 及之后**一个字都没有**,任务随后挂住(0 CPU / 0 I/O / 1 线程);页面只剩"运行中" | `src/proc.py::_inherit_stdio()` 把父进程当前 `sys.stdout/stderr` **显式**交给子进程(STARTF_USESTDHANDLES) |
+| 2 | `capture_output=True` 与显式 `stdout=` **互斥**,同给抛 `ValueError` | 异常被 `guanlan_ops.job_running()` 的 `except` 吞成 `alive=False` ⇒ **任何在跑的任务都被立刻改写成"被强杀"**,页面显示完成、重算按钮重新可点(可能并发起两个重算) | `_inherit_stdio` 先让路(有 `capture_output` 就什么都不加);自检加一条"run_text(capture_output) 可用" |
+| 3 | 内嵌版 `/ops/recalc` 裁掉了「服务」卡,但 JS 仍 `$('#b_start').onclick=…` | 元素为 null → 抛 `TypeError` → **后面所有按钮绑定与 `refresh()` 全不执行** ⇒ 用户看到的正是"清除产物/执行重算点了没反应" | 所有元素访问与绑定走 `set()`/`on()` 容错包装;并加 try/catch 把脚本错误**显示在页面上**(不再静默) |
+
+**服务端配套**: `/ops*` 响应现在带 `Cache-Control: no-store, must-revalidate` —— 否则页面 HTML/脚本更新后浏览器仍用旧版,
+表现同样是"改了没反应"(`guanlan_gateway.py::_bytes(no_store=True)`)。
+
+### 4.2 用户怎么启动(不弹命令窗口)
+
+| 入口 | 行为 |
+|---|---|
+| `start_hidden.vbs`(推荐,可建桌面快捷方式) | `WScript.Shell.Run(..., 0, False)` 以 `pythonw.exe` 跑启动器 —— **完全不出现命令窗口** |
+| `start.bat`(双击) | 默认走上面的隐藏路径(自己立刻退出);`start.bat console` = 旧的前台模式(排障看实时日志) |
+| `scripts/guanlan_start_hidden.py` | 隐藏启动器本体: 端口已在 → 只开浏览器(幂等);否则无窗口起 `guanlan.py serve`、等 `/healthz`、开浏览器 |
+| 运维控制台/门户「数据重算」 | 动作进程同样无窗口;进度与日志尾巴在页面里看 |
+
+**日志去哪了**: `logs/serve.log`(隐藏模式下 serve 的输出)、`logs/start_hidden.log`(启动器轨迹:就绪耗时/失败原因/是否已开浏览器)、
+各组件 `logs/<name>.log`、运维动作 `logs/ops_<动作>_<时间>.log`。失败时隐藏启动器还会弹一个**消息框**(无窗口模式下唯一能让人看见的通道)。
+
+**实测**(2026-09-16): 全停后可见窗口 0;`start.bat` 2.3 s 返回;启动后可见窗口 **0**,6 个服务端口全开,
+7 s 内 `/healthz` 就绪。
+
+---
+
+## 5. 页面取数与"产物即时进页面"
+
+### 5.1 产品数据的读取口径
+
+| 服务 | 读法 | 产物改了要不要重启 |
+|---|---|---|
+| detail `:18033`(`scripts/windscada_serve.py`) | 首次请求把产物读进内存 `_CACHE` | **不需要**:每次取数前比**产物指纹**(`products_stamp()` = 相关产物文件的 `(路径, mtime_ns, size)` 摘要,TTL 2 s),指纹变了自动重载 |
+| cms `:18020`(`src/windcms/serve.py`) | 报告/页面/文件**每请求现读** | 不需要 |
+| cms 的谱元数据(`src/windcms/data.py::spectra_meta`) | 进程内缓存,缓存键含**各窗 index/spectra_meta 指纹** | 不需要(新摄入/重算后自动失效) |
+| 门户(网关 `:28084`) | `portal.html` 按 `(mtime, size)` 缓存改写结果 | 不需要(重建门户即生效) |
+| 网关的 `/ops` 页面模块 | `guanlan_ops` 模块级缓存 | **需要重启网关**(改了这个文件才需要) |
+
+> 实测证据: 只改一个产物文件的 mtime,`logs/detail.log` 立刻出现
+> `[reload] 产物指纹变化 b15e2652… → f0b457f7…, 重载`,页面随即用新数(无需重启)。
+> 另有显式入口 `GET /detail/api/reload`(清缓存并回显新旧指纹)。
+
+### 5.2 门户菜单
+
+顶部菜单: `总览 | 系统架构 | 方法 | 经验发现 | 案例·液压 | 振动·CMS | 仿真与回放 | 交付文档 | 系统状态 | 数据重算 | 登录`
+(2026-09-16 用户令: 「数据重算」落在「交付文档」与「登录」之间)。
+
+- 门外壳正本 `release/portal_src/shell.html` → 装配 `release/portal.html`(`scripts/portal_build.py`)。
+  改外壳后流程: `--check`(看漂移)→ `--rebaseline`(**有意**改动后重设基准)→ 装配 → `--check` 应"全部一致 ✔"。
+- `#recalc` 段落用 iframe **懒加载** `http://127.0.0.1:28084/ops/recalc`(同源;切到本页才挂 src,
+  否则隐藏时它还会每 2 s 轮询 `/ops/api/state`)。
+- `/ops/recalc` 是 `/ops` 的**内嵌版**(只保留 重算 + 产物 + 动作进度;服务卡与页头由门户承担)。
+  走独立路由而不是 `?embed=`:网关转给 `guanlan_ops.handle()` 的是 `u.path`,查询串会被丢掉 —— 用查询串
+  会变成"看起来支持、实际不生效"的静默坑。原 `/ops` 整页保留不变。
+
+---
+
+## 6. 输入 ↔ 产物呼应 与 重算
+
+### 6.1 呼应关系(机器可校验)
+
+`python scripts/inventory_products.py --check` 逐类输入算跨度/条数,与它喂出的产物对拍(退出码 5 = 有问题)。
+2026-09-16 实测**呼应正常**:
+
+| 输入类 | 输入跨度 | 件数 | 对拍产物 | 产物跨度/条数 |
+|---|---|---|---|---|
+| `scada_10min` | 2025-01-01 ~ 2026-07-07 | 38 | `temp_monthly.parquet` / `loss_monthly.parquet` / `powercurve_bins.parquet` | 2025-01 ~ 2026-07 / 19494 · 3729 行 |
+| `故障报警` | 2025 ~ 2026(文件名年粒度) | 16 | `alarms.parquet` | 2025-01-01 ~ 2026-07-15 / 39211 行 |
+| `风机故障记录` | 2021 ~ 2026(年粒度) | 135 | `workorders.parquet` | 2020-01-03 ~ 2026-07-08 / 5876 行 |
+| `油样报告` | 2024-11-19 ~ 2025-08-29 | 404 | `oil_samples_index.parquet` | 同跨度 / 404 行 |
+| `windcms`(CMS 原始导出) | 2026-03 ~ 2026-04 | 25693 | `m5_cms_tcm/windows/w0316/index.parquet` | 2026-03-16 ~ 2026-04-21 / 2,066,686 行 |
+
+★ 年粒度输入(报警/工单的"年度汇总表")**不参与"产物落后"判定** —— 用文件名年份推出来的跨度天然是年粒度,
+拿它当"输入到 2026-12"会造出假缺口(本脚本第一版就这么误报过 2 条)。
+
+### 6.2 重算的两种运行状态(都实测过)
+
+| 状态 | 入口 | 说明 |
+|---|---|---|
+| **系统运行中** | 门户「数据重算」或 `/ops` 的「执行重算」按钮;等价命令行 `python scripts/rebuild_all.py` | 动作经 `_ops_launch.py` 二次启动(不挂网关的父子树,避免 `taskkill /T` 把自己杀掉);跑完**不需要重启**就能在页面看到新数(§5.1 指纹重载)。重算期间按钮全灰、并发动作被后端拒(HTTP 409) |
+| **系统未运行** | 同一套命令行(先 `guanlan.py stop` 或本就关机状态) | 全部构建器都是普通 CLI,不依赖服务;跑完再 `start.bat`/`start_hidden.vbs` 起来,页面直接读新产物。**实测**: 全停后跑 `rebuild_from_raw.py` rc=0、`inventory_products.py --check` rc=0 |
+
+一键顺序(`rebuild_all.py`,14 步): ① 放数据(`--src` 才跑)/② 三门台账/③ SCADA 侧 10 构建器/④ 月度派生件/
+**④b 振动侧摄入**/⑤ 补齐随包件(★2026-09-16 用户令"清除产物不留备份"之后, 随包件不再有 `_products_off/` 暂存区
+⇒ 该步固定返回 6 跳过并说明; 要补齐须显式给交付包 `products_restore_missing.py --stash <交付包.zip>`, **不再打断整条链**)/
+⑥ 重启组件服务/
+⑦ 本体六步/⑧ 本体审计(+可选等价验收)。
+
+---
+
+## 7. 缺口与边界(如实列)
+
+| 缺口 | 影响的产物 | 依据/证据 | 补齐判据 |
+|---|---|---|---|
+| 振动六层链四步脚本未随包(`rudong_tcm_oem_scan` / `rudong_line_energy_share` / `rudong_model_run` / `rudong_fusion_run`) | `oem_frequency_scan`、`gear_freq_scan`、`blade_1p_*`、`model_run_l6.parquet`、`fusion_38.csv` 等 | `outputs/<场>/m5_cms_tcm/vib_raw_manifest.json` 的 `missing_chain` | 给脚本或口径;源件已在 `data/raw/<场>/windcms/` |
+| 厂商月度报告 12 份 PDF 是**纯扫描件** | 「CMS 振动评估报告」的厂商侧 | pypdf 实测 12 页 12 图、`extract_text()` 长度 0 | 现场给电子件(docx/xlsx),或上 OCR(本包不装) |
+| `handoff_vibration_v2.json` / `component_history.json` 是随包快照 | 融合面判级、`/cms/` | 该件含人工裁决/校准更新,不是测量数据的函数 | 现场给正本,放 `data/raw/<场>/m5_cms_tcm/` |
+| 随包件"暂存区"**按设计不再存在**(2026-09-16 用户令"清除产物不留备份") | 第⑤步"补齐随包件"固定返回 6 跳过 | 旧设计把清掉的产物挪到 `_products_off/` 以便还原;现口径 = 真删除 | 要补回"包内没有生成端"的随包件: `python scripts/products_restore_missing.py --stash <交付包.zip>`(从交付件按需补齐, **不在安装目录里留备份**)。`products_state.py --off` 需显式 `--yes`;`--on` 已移除 |
+| 台账等价验收基线 `outputs/<场>/windscada/_pre_rebuild_20260911/` 曾缺失 | 第④步与 `rebuild_from_raw.py --verify` 的逐值比对(缺基线时 ④ 返回 5) | 该目录在打包时被"清除产物"挪进了暂存区, 而暂存区随后被清掉 ⇒ 机器上无标准答案 | **2026-09-16 已重建**: 从当天 dist 包里取 4 件(`alarms/workorders/oil_samples_index/temp_monthly`)+ `_来源说明.json` 标明来历。★链已加固: ④ 容忍 rc=5 —— **缺基线只跳过"等价验收", 不再打断整条重算链**。注意: 「清除产物」会把这个目录一并删掉(它在 `windscada/` 里), 届时需按同样办法重建 |
+| 油样 2026-07 批 102 行(华标合并报告) | 数据层「油液化验」 | 源件 `BG-2026-07-YP013 ….pdf` 不在现场包 | 补那份 PDF 后 `rebuild_from_raw.py --verify` 无人工项 |
+| 无生成端的组级产物(`pc_monthly_bins` / `duty_monthly` / `thermal_monthly` / `sector_power` / `yaw_*` / `genbearing_monthly` / `mblub_monthly` / `pitch_daily` 等) | 趋势件、热链、扇区、偏航/润滑面 | 全库只有读取方、0 处写入方 | 研发补口径(照 `temp_monthly` 的办法反推 + 逐值验证) |
+| `configs/farms/*.yaml` 无代码读者;`available()` 只认 `*.json` | 多场部署 | `src/windscada/config.py::available()` | 二选一(见 §3.3) |
+
+---
+
+## 8. 变更记录
+
+| 日期 | 变更 |
+|---|---|
+| 2026-09-16 | **本文建立**(用户令 1): 产物全景(自动生成清单)+ 路径唯一真源与本次统一的 8 处 + 未统一项如实列表 + 进程无窗口口径 `src/proc.py` + 产物指纹热重载 + 输入↔产物呼应 + 重算两种状态 + 门户「数据重算」 |
+| 2026-09-16 | 用户令 2/3/4/5 的落地: `inventory_products.py`(清点+呼应校验,`--check` 出码);⑤步不再打断重算链(`StashMissing` rc=6);`start_hidden.vbs` + `start.bat` 默认无窗口;门户菜单新增「数据重算」+ `/ops/recalc` 内嵌版;`portal_build.py --rebaseline` |
+| 2026-09-16 | 振动侧接入(详见 `docs/振动数据接入_v0.1.md`): `data/raw/<场>/{windcms,m5_cms_tcm}` 两类源件、`rudong_tcm_index.py`/`rudong_tcm_spectra.py`/`vib_raw_build.py`/`vib_reports_build.py`、窗 `w0316` |
+| 2026-09-16 | 用户令"清除产物不留备份": `products_state.py --off --yes` 改为**真删除**(不再产生 `_products_off*/`)、`--on` 与门户「恢复产物」按钮移除;随包件的唯一来源改为**交付包 zip**(`products_restore_missing.py --stash <交付包.zip>`);`derived_manifest.prune()` 清掉陈旧自登记(`raw-derived` 台账 3444 → 1740 件,回到真实) |
+| 2026-09-16 | 用户令"打包不含 输入数据/产物/日志" → 交付包 **v0.4.0**(见 §9): `pack_dist.py` 增 `--no-products`、`VERSION='0.4.0'`、`dist-manifest.json` 记 `no_data/no_products/no_logs` 与逐条排除理由;`guanlan.py check` 读该清单,产物缺失显示 `[--] 待重算` 而非 FAIL |
+
+---
+
+## 9. 交付包组成(v0.4.0,2026-09-16 用户令:不含 输入数据 / 产物 / 日志)
+
+打包器:`scripts/pack_dist.py`。默认文件名 `<父目录>/guanlan-v<VERSION>_dist_<日期>.zip`,可用 `--out` 指定。
+
+| 进包 | 内容 |
+|---|---|
+| `src/` `scripts/` `guanlan.py` | 程序与全部构建/运维脚本 |
+| `configs/` `resources/` `reference/` | 配置(含 `serve.json` 端口真源、场站/通道台账) |
+| `release/` | 门户、仿真页、三维资产 `viewer/`、治理清单交付件 `release/如东/`(客户交付物,勿外传) |
+| `docs/` `README_先读我.txt` `测试须知.txt` | 交付文档与说明书 |
+| `wheels/win_amd64/`(42 件)、`vendor/python/`(3 平台便携运行时) | 离线安装件:Windows 完全离线可装;Linux/macOS 走联网安装(要离线就把轮子放进 `wheels/linux_x86_64` / `wheels/macos_arm64`) |
+| `install.bat/.ps1/.sh`、`check/start/stop.bat`、`requirements.txt` | 安装与起停入口 |
+
+| 不进包 | 理由(同时写进包内 `dist-manifest.json` 的 `excluded`) |
+|---|---|
+| `data/` | 现场原始输入件;约定"原始件不随包分发"(要带用 `--with-data`) |
+| `outputs/` | **用户令**:产物不进包(`--no-products`)。目标机放数据后 `rebuild_all.py` 重算;或缺"无生成端"的随包件时用 `products_restore_missing.py --stash <交付包.zip>` |
+| `logs/` `run/` | **用户令**:日志不进包(`run/pids.json` 里的 PID 到新机器上是无效引用) |
+| `.venv/` `.git/` `.github/` | venv 换机必失效(安装时重建);版本库不随交付件分发 |
+| `_products_off*/` | 旧设计的"清除产物"暂存档(2026-09-16 起清除=真删,不再产生;老机器上若有可手工删) |
+| `__pycache__/` `*.pyc`、顶层 `*.zip` | 解释器缓存、旧的交付压缩包(避免包中包) |
+
+开箱验证(一条命令给出"能不能装、能不能跑"的证据):`python scripts/pack_dist.py --verify <zip>` —— 解压到
+临时目录 → 离线安装 → 起服务核验 `/`·`/detail/`·`/cms/`·`/ops`·`/healthz` → 收尾清理,要求 ≥4 个页面可用。
+★ 核验要求组件端口(18033/18020/18791/18792/64292)空闲,**验证前先 `guanlan.py stop`**,否则整段核验被跳过(打印 `[!]`)。
+

+ 4 - 3
docs/说明书_观澜如东样板v2_v0.2.md

@@ -74,11 +74,12 @@ ollama serve &
 - 门户每页首屏一句话; "经验发现" 与 "案例" 页的每条结论旁有三枚标签: **证据级** (定论 / 准定论·预警 / 候选 / 参考 / 数据不足 / 撤回)、**审级**、**处置措施** (谁、什么时候、做什么)。候选不进业主正文。
 - 工作台 `/detail/`: 系统矩阵 (38 台 × 九系统判级) → 点台号看证据窗; 检修助手 (检修四链); **问答区**: 选模型 (默认 8B; 复杂问题会自动升档到 27B 并在尾注标明), 回答后自动附"结构化事实契约"引文块; 引不到契约条时会明写"未被契约背书"。汇报纸可导出。
 - `/cms/`: 每台机组六层判读, 机制未定不点零件。
-- **运维控制台 `/ops`** (2026-09-12 新增): 停/启服务 · 执行重算 · 清除产物 —— 一个按钮一件事。
+- **运维控制台** (2026-09-12 新增; 2026-09-16 起菜单入口为门户「**数据重算**」): 停/启服务 · 执行重算 · 清除产物 —— 一个按钮一件事。
   按钮按**真实状态**启用 (不能做的动作是灰的; 后端同样校验并返回 409, 直接打 API 也绕不过);
   "启动服务"起来后会自动打开门户; "停止组件服务"**保留控制台本身**(否则点完页面就没了);
-  "清除产物"是把产物挪到 `_products_off\` (可恢复), 不动 `data\raw\`。页面同时显示各服务端口状态、
-  产物件数、逐件来源台账、验收锚点, 以及当前任务的进度/退出码/日志尾巴。
+  "清除产物"是**直接删除、不留备份**(2026-09-16 用户令; 要随包件用
+  `scripts\products_restore_missing.py --stash <交付包.zip>` 从交付包补齐), 不动 `data\raw\`。
+  页面同时显示各服务端口状态、产物件数、逐件来源台账、验收锚点, 以及当前任务的进度/退出码/日志尾巴。
 
 ### 4.2 问答的可信边界
 - 台号与数字必须能回抓到源记录, 回抓不到的句子被校闸拦下并提示; 两档模型都拦则**不出文并给原因** (不会硬编)。

+ 48 - 41
docs/重算操作手册_v0.1.md

@@ -39,8 +39,9 @@ cd /d <安装目录>
 | ① | 放数据(给了 `--src` 才跑) | `place_raw_data.py --scope full` |
 | ② | 三门台账(报警/工单/油样) | `rebuild_from_raw.py` |
 | ③ | SCADA 侧 10 个构建器 | `rebuild_from_raw.py --scada` |
-| ④ | 月度派生件 | `windscada_monthly_build.py` |
-| ⑤ | 补齐"包内没有生成端"的产物 | `products_restore_missing.py` |
+| ④ | 月度派生件 | `windscada_monthly_build.py`(★无随包基线时返回 5 = "无法比对, 这不是通过" → 本步**容忍 5 并跳过等价验收**, 不再打断整条链; 要验收请先把随包件恢复成暂存区/基线) |
+| ④b | **振动侧摄入**(CMS 原始导出 → 窗索引/谱) | `vib_raw_build.py`(`--skip-vib` 可关;没振动原始件时空跑退出 0。★**不**重生成 CMS 报告/页面: 缺六层链产物时那会用残缺输入覆盖随包快照, 实测 `report.md` −97%, 见 `docs\振动数据接入_v0.1.md` §3b; 要跑用 `--vib-report`) |
+| ⑤ | 补齐"包内没有生成端"的产物 | `products_restore_missing.py`(★没有暂存区时返回 6 = "这一步没得做" → **跳过并说明**, 不打断链条; 想补齐先点"清除产物"生成暂存区) |
 | ⑥ | 重启服务(**必须**: ⑦ 要吃 `/api/fleet`) | `guanlan.py stop` / `serve` |
 | ⑦ | 本体层: 码表→铺开→决策链→趋势→检索→参数表 | `-m src.ontology.*` |
 | ⑧ | 本体审计 + (可选)等价验收 | `-m src.ontology.audit` 等 |
@@ -83,8 +84,7 @@ cd /d <安装目录>
 | **停止组件服务(保留控制台)** | 有组件服务在运行 | 只停 detail/cms/sim/sim_sys/viewer —— **保留网关**, 否则按钮点完页面就没了 |
 | **完整重启(含网关)** | 有组件服务在运行 | 分离进程先停再起; 页面断开约 15 s 后自动恢复 |
 | **执行重算** | 没有任务在跑 | `rebuild_all.py`(默认跳过 SCADA; 可勾选含 SCADA / 末尾加等价验收) |
-| **恢复产物** | 产物被清空过 | `products_state.py --on`(恢复后请点"启动服务") |
-| **清除产物(可恢复)** | 产物在位 且 没有任务在跑 | `products_state.py --off`, 暂存区有同名产物时自动加 `--archive-old` |
+| **清除产物(直接删除,不留备份)** | 产物在位 且 没有任务在跑 | `products_state.py --off --yes` —— ★2026-09-16 用户令: **不留备份、不可恢复**; 需要随包件时用 `products_restore_missing.py --stash <交付包.zip>` 从交付件补齐 |
 
 页面还实时显示: 各服务端口通不通 · 产物件数与分仓 · **来源台账**(raw 重算多少件 / 随包补齐多少件) ·
 **验收锚点**(报警 39211 行 · 工单 5876 · temp_monthly 19494 · 本体 9702 对象) · 最近一次动作的状态、
@@ -238,26 +238,33 @@ cd /d <安装目录>
 `--verify` 拿**随包基线**逐键逐值比对重算件, 把差异分成"旧链已知缺陷 / 本链新增覆盖 / 无法归类"三类,
 只有第三类才算不通过(退出码 4)。
 
-基线**不参与运行**, 平时是单独存着的(`_products_off\_baseline_kept\`), 验收时临时放回:
+基线是**标准答案**, 不参与运行; 它就在 `outputs\rudong\windscada\_pre_rebuild_20260911\`(4 件)。
+若该目录不在(例如刚清过产物), 先从交付包把它取回来:
 
 ```bat
-:: ① 临时放回基线(3 件 1.1 MB)
-xcopy /E /I /Y "_products_off\_baseline_kept\_pre_rebuild_20260911" "outputs\rudong\windscada\_pre_rebuild_20260911"
+:: ① 基线不在时: 从交付包取回台账件(或直接从同事的安装目录拷这个目录)
+.venv\Scripts\python.exe scripts\products_restore_missing.py --stash <交付包.zip>
 
 :: ② 验收(只比不算)
 .venv\Scripts\python.exe scripts\rebuild_from_raw.py --verify
-
-:: ③ 验完撤走, 让产物仓只留重算出来的东西
-rmdir /S /Q "outputs\rudong\windscada\_pre_rebuild_20260911"
 ```
 
 **期望结论**: `alarms` 39211/39211 逐值完全一致; `workorders` 8 行差异全部可归类(旧链的时间解析缺陷)
 + 69 行本链新增覆盖; **1 处"需人工看" = 油样 506→404**(那 102 行的源件不在包里)。
 
-## 5. 重启服务(**必做**)
+> ★ 基线目录位于 `windscada/` 内 ⇒ **「清除产物」会把它一并删掉**。要长期保留, 把它复制到包外
+> (例如交付包 zip 里那份), 或清完再用上面第 ① 步取回。
+
+## 5. 重启服务(通常**不必**)
+
+2026-09-16 起, 产物变了**不需要重启**: detail 服务每次取数前比"产物指纹", 变了自动重载
+(实测日志 `[reload] 产物指纹变化 …`); CMS 的报告/页面是**每请求现读**; 门户按 `(mtime,size)` 失效。
+所以重算完直接刷新页面即可。以下两种情况才需要重启:
 
-产物换了必须重启: CMS 等模块是**启动时**加载数据的, 不重启页面还是旧数; 反过来, 产物被挪走期间起的
-那次 CMS 会直接崩(表现为 `/cms/` 503)。
+- 改的是 `scripts/guanlan_ops.py`(运维控制台页面/接口) → **必须重启网关**(它被 `_ops_module()` 缓存);
+- 产物被清空期间起的 CMS 服务可能已崩(表现 `/cms/` 503) → 点一次"启动服务"。
+
+接口层的显式重载入口: `GET /detail/api/reload`(清缓存并回显新旧指纹)。
 
 ```bat
 .venv\Scripts\python.exe guanlan.py stop
@@ -266,36 +273,33 @@ rmdir /S /Q "outputs\rudong\windscada\_pre_rebuild_20260911"
 
 期望: `状态 degraded offline  模块 5/7 ok`, 并列出 DOWN 的项。
 
-## 5b. 清掉产物 / 再重算(完整操作, 2026-09-12 实测)
+## 5b. 清掉产物 / 再重算(完整操作, 2026-09-16 口径)
 
 ### 清掉产物(人工检查空状态用)
 
+> ★ 2026-09-16 用户令: **清除产物不留备份**。清掉就是**真删除**, 不再有 `_products_off/` 可 `--on` 还原;
+> 要随包件就从**交付包 zip** 按需补齐。演练实测(2026-09-16): 清 11 仓 / 2316 件 / 3499 MB 用时 <10 s;
+> 从交付包补回 589 件用时 <1 min。
+
 ```bat
-:: ① 清掉 (整体挪走, 不是删除 —— 一条命令能还原)
-.venv\Scripts\python.exe scripts\products_state.py --off --archive-old
-::    暂存区里已有上一代产物时**必须加 --archive-old**: 否则会被拦住(见下), 因为挪走的落点是
-::    _products_off\<场>\<目录>, 那里已有同名目录时 shutil.move 会嵌套成 windscada\windscada\,
-::    之后 --on 就把两代混在一起还原。加了它 = 自动把上一代存档到 _products_off_prev_<时间戳>\
-::    (随包基线 _baseline_kept\ 留在原地, rebuild_from_raw --verify 还要用)。
-
-:: ② 重启服务 (产物没了必须重启, CMS 等是启动时加载数据)
-.venv\Scripts\python.exe guanlan.py stop
-.venv\Scripts\python.exe guanlan.py serve
+:: ① 清掉 (真删除, 不可恢复; 必须显式加 --yes 防手滑)
+.venv\Scripts\python.exe scripts\products_state.py --off --yes
 
-:: ③ 看空状态: / 门户仍 200 (门户在 release\, 不是产物); /cms/ 变 503;
-::    /detail/api/fleet 回 err=no_products; 维护页「数据层」各项 条数=0 / 存在=False
+:: ② 看空状态: / 门户仍 200 (门户在 release\, 不是产物);
+::    /detail/api/fleet 回 err=no_products; /cms/ 变 503; 维护页「数据层」各项 条数=0 / 存在=False
+::    (服务不必重启: 产物指纹变了页面会自动重载; 见 §5c)
 
-:: ④ 还原
-.venv\Scripts\python.exe scripts\products_state.py --on
-.venv\Scripts\python.exe guanlan.py stop ; .venv\Scripts\python.exe guanlan.py serve
+:: ③ 恢复: 从**交付包 zip** 按需补回"包内没有生成端"的随包件
+.venv\Scripts\python.exe scripts\products_restore_missing.py --stash <交付包.zip>
+::    再把 raw 派生件重算出来 (或按需只跑某几环):
+.venv\Scripts\python.exe scripts\rebuild_all.py --skip-scada
 ```
 
-**不想用开关也行**(整目录搬走, 最直观):
+**不删产物、只想验证"没有产物会怎样"**: 把整个 `outputs\<场>` 目录改个名再改回来(手工), 效果等同清空且随时可逆:
 
 ```bat
-Move-Item outputs\rudong <暂存目录>       :: 清
-.venv\Scripts\python.exe guanlan.py stop ; .venv\Scripts\python.exe guanlan.py serve
-Move-Item <暂存目录> outputs\rudong       :: 还原
+Rename-Item outputs\rudong outputs\rudong_off      :: 相当于清空(页面立刻呈现无产物)
+Rename-Item outputs\rudong_off outputs\rudong      :: 还原
 ```
 
 > 清的是 `outputs\<场>\`(产物仓), **不动** `data\raw\`(现场数据)、`release\`(门户外壳与交付件)、
@@ -371,14 +375,17 @@ Move-Item <暂存目录> outputs\rudong       :: 还原
 | `populate` / `audit` 退出 1 | 卡在 `pitch\pitch_daily.parquet`(变桨产物无生成端) | 同上 |
 | `/local-ai/` 503 | 本机 Ollama 未运行或没有模型(环境项, 与数据无关) | `ollama pull` 见 `configs\models.json` |
 
-## 8. 两个必须知道的坑
+## 8. 三个必须知道的坑
 
-1. **`products_state.py --on` 会删掉重算产物**。它按清单把随包产物挪回来, 且**目标目录存在就先 `rmtree`**
-   (`scripts\products_state.py` 第 110-112 行)。当前清单覆盖 `outputs\rudong\{ontology, windscada, …}` ——
-   也就是说: **想让"原始重算"的结果留着, 就不要跑 `--on`**。要拿随包产物就先备份自己的重算结果。
-   只想取验收基线 → 用 §4 的 xcopy, 别用 `--on`。
-2. **顺序永远是: 停服务 → 挪/放数据与产物 → 算 → 起服务**。挪产物期间起的服务会带着"没有数据"的
-   状态跑起来(CMS 直接崩), 表现为页面 503, 而数据其实是好的。
+1. **「清除产物」= 真删除, 没有备份**(2026-09-16 用户令)。`products_state.py --off` 需显式 `--yes`;
+   旧的 `--on` 已移除。清完要恢复:
+   `products_restore_missing.py --stash <交付包.zip>`(补随包件) + `rebuild_all.py`(重算 raw 派生件)。
+   清之前如果想留一份回退路, 就把整个 `outputs\<场>` 改名(§5b 的手工做法) —— 那是**你自己的临时目录**, 不是系统备份。
+2. **产物变了通常不用重启服务**(2026-09-16 起): detail 按产物指纹自动重载、CMS 每请求现读、门户按 mtime 失效。
+   只有改 `scripts/guanlan_ops.py`(控制台本身)才必须重启**网关**。
+3. **重算中的按钮是灰的, 这是设计**: 有任务在跑时「执行重算/清除产物/停服务」全禁用, 防止并发;
+   此时页面显示"正在执行 … "与实时日志尾巴(2 s 刷新)。**日志真的一直空白**才是异常(2026-09-16 修过一次:
+   `CREATE_NO_WINDOW` 会吸走子进程句柄)。
 
 ---
 
@@ -406,4 +413,4 @@ Move-Item <暂存目录> outputs\rudong       :: 还原
 ```
 
 相关文档: `docs\数据目录结构与落位约定_v0.2.md`(§2 落位 · §4 重算边界表) ·
-`docs\重算缺口与补件清单_v0.1.md`(缺什么找谁补) · `_修复记录_20260911\README.md`(踩过的坑)
+`docs\重算缺口与补件清单_v0.1.md`(缺什么找谁补) · `docs\振动数据接入_v0.1.md`(振动侧 CMS/TCM 落位与摄入)

+ 24 - 9
docs/重算缺口与补件清单_v0.1.md

@@ -24,13 +24,14 @@
 | # | 缺什么 | 类型 | 卡住的页面/功能 | 找谁 | 补齐判据 |
 |---|---|---|---|---|---|
 | **A1** | 油样 2026-07 批的源件(1 份**合并报告**) | 源件 | 数据层「油液化验」**506→404**; 油液时效胶囊、油液判级 | 现场 / 化验机构 | 放进 `data\raw\如东\油样报告\` 重跑摄入后, `rebuild_from_raw.py --verify` **无人工项** |
-| **A2** | 振动线 **CMS 测点数据(handoff)**, 随包 90 件 56 MB | 源件 | ~~`/cms/` 503~~ **已由随包件补齐(页面已恢复)**; 但要"自己算"仍需测点件 | 振动线 / 研发 | `/cms/` 由重算件驱动而非随包件 |
+| **A2** | 振动线 **CMS 测点数据(handoff)**, 随包 90 件 56 MB | 源件 | ~~`/cms/` 503~~ **已由随包件补齐(页面已恢复)**; **2026-09-12 源件已到** → 索引/谱/报告可重算(见 §1 A2) | 振动线 / 研发 | 剩余缺口见 B5 |
 | **A3** | SOP 底稿 + 范式实验件(`sop` 210 件 16.5 MB, `paradigm_r1` 29 件) | 源件(内含脚本) | ~~门户契约段、`/api/facts`~~ **已由随包件补齐** | 研发 | `/api/facts` 由重算件驱动 |
 | **A4** | `temp_monthly.parquet` | 生成端 | ~~工作台全标签页无数据~~ **已解决**: 已反推出口径并 100% 复现 | — | ✅ 已闭合(见上) |
 | **B1** | 6 件组级/月度产物(见 §2) | 生成端 | 趋势件、热链、扇区、结构面 | 研发 | 提供口径后照 ① 反推+验证 |
 | **B2** | 6 件偏航/温度/润滑产物(见 §2) | 生成端 | 偏航面、润滑面 | 研发 | 同上。**源件 `scada_1min` 已落位 12.8 GB** |
 | **B3** | `pitch\pitch_daily.parquet` | 生成端 | 变桨面; 本体 `populate`/`audit`(已因随包件补齐而可跑) | 研发 | 液压四列口径未知(试过日界位移/越限穿越都不对) |
-| **B4** | CMS/TCM 兼容链(`windcms` 53 件 · `tcm_compatible_replay` 59 件) | 生成端 + 源件 | ~~`/cms/`、振动融合页~~ **已由随包件补齐** | 振动线 + 研发 | 由重算件驱动 |
+| **B4** | CMS/TCM 兼容链(`windcms` 53 件 · `tcm_compatible_replay` 59 件) | 生成端 + 源件 | ~~`/cms/`、振动融合页~~ **已由随包件补齐**; 其中**索引/谱/报告** 2026-09-12 起可重算 | 振动线 + 研发 | 由重算件驱动 |
+| **B5** | 振动六层链四步脚本(`rudong_tcm_oem_scan` / `rudong_line_energy_share` / `rudong_model_run` / `rudong_fusion_run`) | 生成端 | 扫描线/能量占比/模型层/融合层那批 parquet(`oem_frequency_scan`、`gear_freq_scan`、`blade_1p_*`、`model_run_l6`、`fusion_38.csv` 等) | 振动线 / 研发 | 给脚本或口径; 源件已在 `data\raw\如东\windcms\`(25,679 件 150 GB)。`vib_raw_manifest.json` 的 `missing_chain` 字段即此项 |
 
 ---
 
@@ -50,15 +51,29 @@
 
 ### A2 · 振动线 CMS 测点数据(handoff)
 
-- **缺的**: 随包里 `outputs\rudong\m5_cms_tcm\`(90 件 56 MB: handoff + TCM 兼容件 + 窗/谱/模型产物)
+> **2026-09-12 更新: 源件已到, 本条**部分闭合**。** 用户把 `CMS_RuDong_CGN_202603-04.zip`
+> (Brande TCM 导出 25,679 件 / 150 GB, 2026-03~04) 补进现场包, 落位到
+> `data\raw\如东\windcms\CMS_RuDong_CGN_202603-04\`, 并补齐了**摄入端**:
+> `scripts\rudong_tcm_index.py`(→ `windows\<窗>\index.parquet`, 54 列) +
+> `scripts\rudong_tcm_spectra.py`(→ `spectra\*.npz` + `spectra_meta.parquet`) +
+> 一键 `scripts\vib_raw_build.py`。**现在"索引/谱/报告/知识库"是可重算的了**
+> (此前连"含 `*_decode.json` 的目录"这条路径都跑不通 —— 两个摄入脚本随包缺失)。
+> **仍未闭合的**: 六层链四步(见新增 B5)、以及 handoff 接口件本身(它是振动线出件, 不是测量数据的函数)。
+
+- **缺的(原记录)**: 随包里 `outputs\rudong\m5_cms_tcm\`(90 件 56 MB: handoff + TCM 兼容件 + 窗/谱/模型产物)
   以及 `outputs\rudong\windcms\`(53 件 29 MB, 41 个 HTML 报告页)。
 - **证据**: `src\windscada\subsys\fusion.py` 的 `load_handoff()` 直接抛
   `FileNotFoundError: 振动 handoff 接口缺失 … 融合面不可静默降级`;
   `src\windcms\data.py` 的 `load_scalars()` 读的是 CMS **测点索引**(不是报告); 现场包里只有
-  `(8)…\振动分析报告\` 12 份**月度用印版 PDF** —— 那是成品牌报告, 不是可再加工的测量数据。
+  `(8)…\振动分析报告\` 12 份**月度用印版 PDF** —— 那是成品牌报告, 不是可再加工的测量数据,
+  **且实测为纯扫描件**(12 页 12 图, `extract_text()` 长度 0) ⇒ 连"从 PDF 反推"这条路也不通。
 - **影响面**: `/cms/` 503; 融合面判级; 台账里"振动级"这一轴; `trend_ingest`(在升/闭环证据)。
-- **怎么补**: 向振动线要 **CMS 测点导出/handoff 件**(按 `m5_cms_tcm` 的目录约定), 而不是 PDF 报告。
-- **判定**: `/cms/` 回 200 且能列出报告页。
+- **怎么补**: ~~向振动线要 CMS 测点导出/handoff 件~~ → **2026-09-12 已到**: CMS 测点导出落在
+  `data\raw\如东\windcms\`, TCM 侧报告落在 `data\raw\如东\m5_cms_tcm\`。
+  若现场还能给 `handoff_vibration_v2.json` **正本**, 放 `m5_cms_tcm\` 下即被优先采用。
+- **判定**: ① `/cms/` 回 200 且能列出报告页(已满足); ② `windcms` 的**谱图**能取到数据
+  (此前 `m5\spectra` 整个目录缺失, `data.spectrum()` 只能返回 None —— 摄入后可取);
+  ③ 新窗进入 `data.windows()` 分析集。细节与实测数字见 `docs\振动数据接入_v0.1.md`。
 
 ### A3 · SOP 底稿 + 范式实验件(事实契约的输入)
 
@@ -77,7 +92,7 @@
 
 - **缺的**: 随包 `outputs\rudong\windscada\` 里 118 件中的**非 16 件**(我们只重算得出 16 件)。
 - **证据**: 工作台取数面直接回结构化"无产物":
-  `{"err":"no_products","note":"temp_monthly.parquet 不存在 → outputs/rudong/windscada/temp_monthly.parquet; 先跑 scripts/rebuild_from_raw.py (SCADA 侧加 --scada) 生成产物, 或用 scripts/products_state.py --on 还原"}`。
+  `{"err":"no_products","note":"temp_monthly.parquet 不存在 → outputs/rudong/windscada/temp_monthly.parquet; 先跑 scripts/rebuild_from_raw.py (SCADA 侧加 --scada) 生成产物; "包内没有生成端"的件从交付包补齐: python scripts/products_restore_missing.py --stash <交付包.zip> (2026-09-16 起清除产物不留备份, 故不再有 --on 还原)"}`。
   注意那句提示在**这一件**上是误导的: `temp_monthly` 不在 `rebuild_from_raw.py` 的 10 个构建器里。
 - **怎么补**: 研发补 `temp_monthly` 的构建器(见 B1), 或明确它属于哪个上游产物。
 
@@ -133,7 +148,7 @@
 ## 4. 补齐之后怎么验(照这个顺序)
 
 > **手工操作的逐步手册见 `docs\重算操作手册_v0.1.md`** —— 含每步的期望输出、会踩的坑
-> (尤其:`products_state.py --on` 会删掉重算产物)、以及"算不出来"对照表。下面是压缩版命令。
+> (尤其: 2026-09-16 起「清除产物」= **真删除、不留备份**, 旧的 `--on` 已移除)、以及"算不出来"对照表。下面是压缩版命令。
 
 ```bat
 :: ① 源件落位 (现场包 → 约定目录; 同尺寸自动跳过, 可反复跑)
@@ -158,4 +173,4 @@
 - **累计快照必须归并**: 工单/报警目录里的"全年/年至今"件是季度件的累计快照, 摄入时自动跳过; 跨源内容重复自动归并 ——
   新加源件不会让台账虚高(本轮 `大部件维修记录` 就是被归并的)。
 - **本清单的维护**: 新增一项的条件是"能用一条命令重复验证它的缺失"; 每项都要写清证据路径与补齐判据。
-- 相关文档: `docs\数据目录结构与落位约定_v0.2.md`(§4 重算边界表) · `_修复记录_20260911\README.md`(从零重算一节)
+- 相关文档: `docs\数据目录结构与落位约定_v0.2.md`(§4 重算边界表) · `docs\振动数据接入_v0.1.md`(振动侧接入与缺口)

+ 47 - 3
guanlan.py

@@ -87,7 +87,12 @@ def services(c):
 def spawn(cmd, log: Path, e):
     f = open(log, "ab")
     kw = dict(cwd=str(ROOT), env=e, stdout=f, stderr=subprocess.STDOUT, stdin=subprocess.DEVNULL)
-    if WIN: kw["creationflags"] = subprocess.CREATE_NEW_PROCESS_GROUP | getattr(subprocess, "DETACHED_PROCESS", 0)
+    if WIN:
+        # 2026-09-16: 统一到 CREATE_NO_WINDOW (src/proc.py) —— 原先用 DETACHED_PROCESS, 子进程没有控制台,
+        # 而 .venv 的 python.exe 是**转发器**, 它会再拉一个真解释器; 落地实测那一步会冒出可见黑窗。
+        # CREATE_NO_WINDOW 给的是"没有窗口的控制台", stdio 照旧重定向到日志, 且与 DETACHED 互斥故只用前者。
+        from src import proc as _proc
+        kw["creationflags"] = _proc.flags()
     else: kw["start_new_session"] = True
     return subprocess.Popen(cmd, **kw).pid
 
@@ -99,14 +104,18 @@ def alive(pid):
         # errors="replace" 必须有: tasklist 的输出是**控制台代码页**(中文 Windows = GBK), 而本进程
         # 可能是 PYTHONUTF8=1 起的(默认文本编码变 UTF-8) → 解码失败会让 r.stdout 变成 None,
         # 下一行 `str(pid) in None` 直接 TypeError。2026-09-12 由运维控制台的动作进程实测逮到。
-        r = subprocess.run(["tasklist", "/FI", f"PID eq {pid}"], capture_output=True, text=True, errors="replace")
+        # ★ 走 src.proc.run_text: 本函数也会被**无控制台**的进程调用 (网关/运维任务), 裸 spawn tasklist
+        #   会给它新建可见控制台 → 反复闪窗 (2026-09-16)。
+        from src import proc as _proc
+        r = _proc.run_text(["tasklist", "/FI", f"PID eq {pid}"])
         return r.stdout is not None and str(pid) in r.stdout
     try: os.kill(pid, 0); return True
     except OSError: return False
 
 
 def kill(pid):
-    if WIN: subprocess.run(["taskkill", "/PID", str(pid), "/T", "/F"], capture_output=True)
+    from src import proc as _proc
+    if WIN: _proc.run(["taskkill", "/PID", str(pid), "/T", "/F"], capture_output=True)
     else:
         try: os.killpg(os.getpgid(pid), signal.SIGTERM)
         except OSError:
@@ -132,13 +141,48 @@ def cmd_check(c):
         for it in w:
             if issubclass(it.category, SyntaxWarning): bad.append(f"{Path(str(it.filename)).name}:{it.lineno} {it.message}")
     row(f"源码可编译 ({len(srcs)} 个文件, 无非法转义/语法错)", not bad, "; ".join(bad[:3]))
+    # 子进程口径自检 (2026-09-16 加): 这三条是**回归守卫** —— 它们各自都真炸过一次, 而炸法都是"静默":
+    #   ① capture_output=True 与显式 stdout 互斥 → 异常被 except 吞成 alive=False ⇒ 在跑的任务被
+    #      误判成"被强杀", 页面显示完成、重算按钮可再点(可能并发重算);
+    #   ② CREATE_NO_WINDOW 新建的控制台会把子进程标准句柄吸走 ⇒ 重算日志一个字都没有、任务挂住;
+    #   ③ 不弹窗 (NO_WINDOW 位必须真的设上, 否则"点按钮闪黑窗"回来)。
+    try:
+        from src import proc as _proc
+        _r = _proc.run_text([sys.executable, "-c", "print('proc-ok')"])
+        row("子进程口径: run_text(capture_output) 可用", _r.returncode == 0 and "proc-ok" in (_r.stdout or ""),
+            "src/proc.py (capture_output 与显式 stdout 互斥, 冲突会让 job_running 误判任务已死)")
+        import tempfile as _tf
+        _lg = Path(_tf.gettempdir()) / "_guanlan_proc_check.log"
+        _p = _proc.spawn([sys.executable, "-c", "print('spawn-ok')"], log=_lg); _p.wait(timeout=60)
+        _txt = _lg.read_text(encoding="utf-8", errors="replace") if _lg.exists() else ""
+        try: _lg.unlink()
+        except OSError: pass
+        row("子进程口径: spawn(log=) 输出进日志", "spawn-ok" in _txt,
+            "CREATE_NO_WINDOW 会吸走句柄, 不显式继承则日志空白 (重算看起来卡住)")
+        row("子进程口径: 无窗口位已设 (CREATE_NO_WINDOW)", bool(_proc.flags() & 0x08000000) or os.name != "nt",
+            f"flags={hex(_proc.flags())}")
+    except Exception as _e:
+        row("子进程口径自检", False, f"{type(_e).__name__}: {_e}")
     for m in ("numpy", "pandas", "pyarrow", "polars", "yaml", "matplotlib", "plotly", "jinja2", "docx"):
         try: __import__(m); row(f"依赖 {m}", True)
         except Exception as ex: row(f"依赖 {m}", False, f"未安装: {ex.__class__.__name__} (运行 install 脚本)")
+    # ★2026-09-16 用户令: 打包"不含 输入数据/产物/日志"。产物缺失时, 若包内 dist-manifest.json 就声明了
+    #   no_products/no_data, 这里显示 `--`(待重算) 而**不是 FAIL** —— 否则开箱自检会把"本来就该重算"的
+    #   状态报成失败, install 脚本收尾的那次 check 也会失败 (把正常状态说成问题)。
+    _man = {}
+    try:
+        _mp = ROOT / "dist-manifest.json"
+        _man = jload(_mp) if _mp.is_file() else {}
+    except Exception:
+        _man = {}
+    _no_products = bool(_man.get("no_products"))
     for rel, n in (("outputs/rudong/windscada", "L0/L1 产物 (parquet)"), ("outputs/rudong/ontology/objects.json", "本体对象库"), ("outputs/rudong/sop/findings.json", "findings"), ("outputs/rudong/guanlan/facts_contract_v0.json", "事实契约"),
                    (c["release_dir"] + "/portal.html", "门户"), (c["release_dir"] + "/sim_sys_server.py", "仿真合页服务"), (c["release_dir"] + "/如东SWT40_控制律仿真台_20260906.zip", "仿真合页资料包"), (c["viewer_dir"], "三维 viewer 资产"), (c["sim_dir"], "仿真回放资产"), (c["release_dir"] + "/如东", "治理清单交付件 (可选)")):
         p = ROOT / rel; exists = p.exists() and (any(p.iterdir()) if p.is_dir() else p.stat().st_size > 0)
         if "可选" in n: print(f"  [{'OK' if exists else '--'}] {n}: {rel}")
+        elif _no_products and rel.startswith("outputs/"):
+            print(f"  [--] {n}: {rel}  (本包按用户令**未随产物** —— 把现场包放好后跑 "
+                  f"scripts/place_raw_data.py --src <现场包目录> --scope full 再 scripts/rebuild_all.py)")
         else: row(n, exists, rel)
     raw = Path(c.get("raw_dir", "data/raw")); raw = raw if raw.is_absolute() else (ROOT / raw)
     print(f"  [{'OK' if raw.exists() else '--'}] 原始件目录 (补数据放这里): {raw}{'' if raw.exists() else ' (不存在; 原始件不随包分发, 只影响摄入命令, 不影响页面)'}")

+ 1 - 1
release/portal_src/README.md

@@ -54,5 +54,5 @@ release/portal.html   20,226,052 B   sha256 9b6aabeb6ca18d15…
 ## 相关
 
 - 契约与产物地图:`docs/数据目录结构与落位约定_v0.2.md` §6
-- 产物进出 git 的约定与那次误提交的复盘:仓根 `.gitignore` 注释 + `_修复记录_20260911/README.md`
+- 产物进出 git 的约定与那次误提交的复盘:仓根 `.gitignore` 注释 + git 提交历史(`_修复记录_*` 类目录已按 2026-09-12 用户令不再随包/不再保留)
 - 人工检查空状态:`scripts/products_state.py --off/--on/--status`

+ 5 - 4
release/portal_src/manifest.json

@@ -5,8 +5,8 @@
  },
  "shell": {
   "file": "shell.html",
-  "bytes": 138544,
-  "sha256": "0a42ef684034be2f78cb8d38b16039d18a903d116d44154800d6abbaef8c3a71"
+  "bytes": 140103,
+  "sha256": "957e43d5e85ad7fe41a6e8c7e7c0e31921f4755dbba4bf7938825bd5b98f5329"
  },
  "governance_sources": {
   "file": "governance_sources.json",
@@ -159,13 +159,14 @@
   }
  },
  "claims_section_stripped": 1,
- "expected_portal_sha256": "9b6aabeb6ca18d15c3c0cd874dbb3dc1f820db050bdef7944e661d93d4ead843",
+ "expected_portal_sha256": "67ce6c881e5ffa09d6ae4e2099d8e78ce6f618349a43911164c288061d57e447",
  "line_ending": "LF",
  "notes": [
   "shell.html = 门户外壳 (CSS/JS/导航/面板结构), 含 <!--@TEMPLATE:id--> 与 <!--@GOVERNANCE_SOURCES--> 标记;",
   "templates/ 与 governance_sources.json 是**产物/交付件内容** (按\"产物不进 git\"的约定不入库); manifest 里留 sha256 以便漂移检测 (--check);",
   "contract-claims 段由 scripts/guanlan_portal_inject_claims.py 在装配最后注入 (内容来自 outputs/<场>/guanlan/…);",
   "锚点修复 scripts/guanlan_portal_fix_anchors.py 的输出是**代码**, 已固化在 shell.html 里;",
-  "行尾必须是 LF: 三个脚本 (本器 + 两个注入器) 都已用字节级/newline=\"\" 写文件, 装配后有 CRLF 自检。"
+  "行尾必须是 LF: 三个脚本 (本器 + 两个注入器) 都已用字节级/newline=\"\" 写文件, 装配后有 CRLF 自检。",
+  "重基线 2026-09-16 16:20: shell 0a42ef684034be2f→957e43d5e85ad7fe, portal 9b6aabeb6ca18d15→67ce6c881e5ffa09 (有意改动门户外壳后重设漂移基准)"
  ]
 }

+ 4 - 4
release/portal_src/shell.html

@@ -1,4 +1,4 @@
-<!DOCTYPE html><html lang="zh"><head><meta charset="utf-8"><meta name="viewport" content="width=device-width,initial-scale=1"><title>观澜中文系统 · 详细分析</title><style>:root{--ground:#070C14;--ground-2:#0B1220;--panel:#101A2B;--panel-2:#152238;--line:#1F2E44;--hair:#2A3B55;--ink:#E6EDF5;--ink-2:#C2CEDB;--muted:#8393A8;--accent:#3FC1D3;--accent-deep:#1F8FA0;--accent-soft:#0F2A35;--critical:#E4645C;--warning:#E0A43A;--good:#4CC39B;--excellent:#9AD8C4;--sans:"PingFang SC","Hiragino Sans GB","Microsoft YaHei",Inter,ui-sans-serif,system-ui,sans-serif;--mono:"SF Mono",Menlo,Consolas,monospace;--r:8px}
+<!DOCTYPE html><html lang="zh"><head><meta charset="utf-8"><meta name="viewport" content="width=device-width,initial-scale=1"><title>观澜中文系统 · 详细分析</title><style>:root{--ground:#070C14;--ground-2:#0B1220;--panel:#101A2B;--panel-2:#152238;--line:#1F2E44;--hair:#2A3B55;--ink:#E6EDF5;--ink-2:#C2CEDB;--muted:#8393A8;--accent:#3FC1D3;--accent-deep:#1F8FA0;--accent-soft:#0F2A35;--critical:#E4645C;--warning:#E0A43A;--good:#4CC39B;--excellent:#9AD8C4;--sans:"PingFang SC","Hiragino Sans GB","Microsoft YaHei",Inter,ui-sans-serif,system-ui,sans-serif;--mono:"SF Mono",Menlo,Consolas,monospace;--r:8px}
 *{box-sizing:border-box;margin:0} html{background:var(--ground)} body{background:var(--ground);color:var(--ink);font:14px/1.6 var(--sans)} a{color:var(--accent);text-decoration:none} a:hover{text-decoration:underline}
 .top{height:56px;display:flex;align-items:center;gap:18px;padding:0 24px;border-bottom:1px solid var(--line);background:#0A1120;position:sticky;top:0;z-index:5} .logo{width:30px;height:30px;border-radius:7px;background:linear-gradient(135deg,#1F8FA0,#0F2A35);display:grid;place-items:center;color:#fff;font-weight:700} .brand{font-weight:600} .brand small{display:block;color:var(--muted);font-size:11px;font-weight:400;letter-spacing:.08em}
 .tabs{display:flex;gap:2px;margin-left:12px} .tabs a{padding:8px 12px;color:var(--ink-2);border-bottom:2px solid transparent} .tabs a.on{color:var(--ink);border-color:var(--accent)} .tabs a:hover{text-decoration:none;color:var(--ink)} .live{margin-left:auto;border:1px solid var(--accent-deep);color:var(--accent);border-radius:6px;padding:6px 12px;font-size:13px}
@@ -43,7 +43,7 @@ body{background:linear-gradient(120deg,#f1f7f9,#edf5f7);}.hero{background:transp
 @media(max-width:1050px) and (min-width:761px){#architecture>.hero>.wrap,#case_hydraulic>.hero>.wrap{padding-left:360px}#architecture>.hero>.wrap:after,#case_hydraulic>.hero>.wrap:after{width:290px}}
 @media(max-width:760px){#index>.hero>.wrap{padding:148px 20px 0}#index>.hero>.wrap:after{position:absolute;left:20px;right:20px;top:0;height:125px;margin:0;width:auto;background-size:auto 180px,cover}#index>.hero h1{font-size:29px}#architecture>.hero>.wrap,#case_hydraulic>.hero>.wrap,#method>.hero>.wrap,#admin>.hero>.wrap,#login>.hero>.wrap{padding:0 20px}#architecture>.hero>.wrap:after,#case_hydraulic>.hero>.wrap:after,#login>.hero>.wrap:after{position:relative;left:auto;right:auto;width:100%;height:160px;border-radius:12px}#method>.hero>.wrap:after{height:130px;background-size:auto 230px,cover}#cms>.hero>.wrap{padding:0 20px 150px}#cms>.hero>.wrap:after{position:absolute;left:20px;right:20px;bottom:0;width:auto;height:130px;margin:0}#documents>.hero>.wrap{padding:0 20px;min-height:0}#documents>.hero>.wrap:after{display:none}#admin>.hero>.wrap:after{height:115px}}
 </style><style id="guanlan-touch-targets">@media(max-width:760px){.tabs a,.return-dock a,.chips a,.btn,.x{min-height:44px;display:inline-flex;align-items:center;justify-content:center}.return-dock{flex-wrap:wrap}.foot{padding-bottom:115px}}</style></head><body>
-<div class="top"><div class="logo">≋</div><div class="brand">观澜<small>WIND ASSET INTELLIGENCE</small></div><nav class="tabs"><a href="#index" data-nav="index">总览</a><a href="#architecture" data-nav="architecture">系统架构</a><a href="#method" data-nav="method">方法</a><a href="#findings" data-nav="findings">经验发现</a><a href="#case_hydraulic" data-nav="case_hydraulic">案例·液压</a><a href="#cms" data-nav="cms">振动·CMS</a><a href="#sim" data-nav="sim">仿真与回放</a><a href="#documents" data-nav="documents">交付文档</a><a href="#admin" data-nav="admin">系统状态</a><a href="http://127.0.0.1:18033/v2">登录</a></nav></div>
+<div class="top"><div class="logo">≋</div><div class="brand">观澜<small>WIND ASSET INTELLIGENCE</small></div><nav class="tabs"><a href="#index" data-nav="index">总览</a><a href="#architecture" data-nav="architecture">系统架构</a><a href="#method" data-nav="method">方法</a><a href="#findings" data-nav="findings">经验发现</a><a href="#case_hydraulic" data-nav="case_hydraulic">案例·液压</a><a href="#cms" data-nav="cms">振动·CMS</a><a href="#sim" data-nav="sim">仿真与回放</a><a href="#documents" data-nav="documents">交付文档</a><a href="#admin" data-nav="admin">系统状态</a><a href="#recalc" data-nav="recalc">数据重算</a><a href="http://127.0.0.1:18033/v2">登录</a></nav></div>
 <div class="pg" id="index" data-title="总览"><section class="hero"><div class="wrap"><div class="kicker">Wind asset intelligence · 机群诊断与分析服务</div><h1>从运行证据到检修决策</h1><p class="lead">观澜把运行数据、状态监测与检修记录落到正确的部件上,然后给出一个带证据与边界的决定。</p>
 <div class="chips"><a class="chip ev" href="#sim" data-nav="sim">进入仿真与回放 →</a><a class="chip ev" href="#cms" data-nav="cms">振动 · CMS 在线系统 →</a><a class="chip" href="#method" data-nav="method">读方法</a><a class="chip" href="#case_hydraulic" data-nav="case_hydraulic">看一个发现案例:液压系统</a><a class="chip" href="#documents" data-nav="documents">看交付文档</a></div></div></section>
 <section><div class="wrap"><div class="kicker">证据到决定的系统</div><h2>多源证据分析与检修后验证</h2><p class="sub">十分钟与秒级 SCADA · 状态监测 · 报警 · 工单 · 油样与检查 · 技术文档</p>
@@ -878,7 +878,7 @@ requestAnimationFrame(frame);
 <section><div class="wrap"><div class="kicker">数据层</div><h2>如东 · 五层原生 WPS 已装配</h2><div class="grid g4"><div class="card"><div class="k">10-min / 1-min</div><h3>1min 2025-01-01 → 2026-06-30 (99.87%)</h3><p>cnt 层至 2026-07-07;交叠 546 天;136 号文断点 2025 旧机制 / 2026 新机制。</p></div><div class="card"><div class="k">温度 / 电网 / 标志 / 日汇总</div><h3>native_scturtemp · scturgrid · scturflag · dailysummary</h3><p>parquet 7.1 GB,指纹文件同目录。</p></div><div class="card"><div class="k">告警 / 工单 / 油样</div><h3>alarm.tsv 2016→ · 工单 4,298 · 油样 506</h3><p>本体对象:告警码 555、工单 4,298、油样 506。</p></div><div class="card"><div class="k">CMS · TCM</div><h3>2026-01-27 单窗 + 06-29→08-11 五窗</h3><p>320 万记录 / 160 万谱 / 7,688 条波形。</p></div></div></div></section>
 <section><div class="wrap"><div class="kicker">模型档</div><h2>分析规则与模型配置</h2><div class="grid g4"><div class="card"><div class="k">本地 · 轻量</div><h3>qwen3:8b</h3><p>本体问答 8–10 s;接地闸通过才上屏。</p></div><div class="card"><div class="k">本地 · 标准</div><h3>qwen3.8:27b / qwen3:32b</h3><p>交叉审与长文;本地审核票只作初筛(已知硬伤命中 0/4)。</p></div><div class="card"><div class="k">本地 · 推理</div><h3>deepseek-r1:14b / 32b</h3><p>已下载,未标定。</p></div><div class="card"><div class="k">云端</div><h3>DeepSeek · 千问</h3><p>训练与提案侧;密钥从环境读,不落盘。</p></div></div></div></section>
 <section><div class="wrap"><div class="kicker">审级与队列</div><h2>本体 19,321 个对象 · 写入闸 C1–C17</h2><div class="grid g4"><div class="card"><div class="k">审级</div><h3>云端跨厂商双票 = 最低审级</h3><p>上云 / 含经济数字 / 确诊 / 跨机系统性 触发。</p></div><div class="card"><div class="k">月度滚动</div><h3>新数据 → 重跑 → 本体增量 → 只报变化</h3><p>验收状态机:29# 修后首例。</p></div><div class="card"><div class="k">审计六项</div><h3>悬空引用 / 孤儿 / 链接密度 / 字段覆盖 / …</h3><p>第七项(引用未审定)待加。</p></div><div class="card"><div class="k">回归</div><h3>paradigm 14/14 · windscada 4/5 · ontology 25/28</h3><p>失败项 = 原始 CSV 硬路径缺 / 测试与对象库状态耦合,非功能故障。</p></div></div></div></section>
-</div><div class="pg" id="login" data-title="登录"><section class="hero" style="padding-bottom:64px"><div class="wrap"><div class="kicker">Wind asset intelligence</div><h1>欢迎来到观澜。</h1><p class="lead">面向复杂运营的 AI 辅助风电资产智能,为清晰决策而建。</p><div class="chips"><span class="chip ev">● 风电智能</span><span class="chip ev">● AI 辅助</span><span class="chip ev">● 访问受控</span></div></div></section>
+</div><div class="pg" id="recalc" data-title="数据重算"><section class="hero"><div class="wrap"><div class="kicker">运维 · 数据重算与产物</div><h1>数据重算</h1><p class="lead">从 <code>data/raw</code> 的现场数据重算全部产物:执行重算、查看产物在位状态、清除/恢复产物。这一页就是运维控制台的重算与产物两块(同一后端、同一把锁),页面顶部菜单直接到这里,不必再记 <code>/ops</code>。</p><div class="chips"><span class="chip ev">● 重算期间按钮全灰</span><span class="chip ev">● 后端拒绝 = 真生效(HTTP 409)</span><span class="chip ev">● 重算产物即时进页面</span></div></div></section><section><div class="wrap"><div class="opsframe" style="border:1px solid #dfe6ef;border-radius:10px;overflow:hidden;background:#fff"><iframe id="rcf" title="数据重算与产物" style="width:100%;height:1150px;border:0;display:block" loading="lazy"></iframe></div><p class="small muted" style="margin-top:10px">看不到内容?说明网关还没起来(本页由网关提供,组件服务停了也能用)。直接打开:<a href="http://127.0.0.1:28084/ops" target="_blank" rel="noopener">/ops 控制台</a>。</p></div></section></div><div class="pg" id="login" data-title="登录"><section class="hero" style="padding-bottom:64px"><div class="wrap"><div class="kicker">Wind asset intelligence</div><h1>欢迎来到观澜。</h1><p class="lead">面向复杂运营的 AI 辅助风电资产智能,为清晰决策而建。</p><div class="chips"><span class="chip ev">● 风电智能</span><span class="chip ev">● AI 辅助</span><span class="chip ev">● 访问受控</span></div></div></section>
 <section><div class="wrap" style="max-width:560px"><div class="form"><div class="kicker">受保护工作区</div><h2 style="margin-top:8px">欢迎回来</h2><p class="sub">登录以进入受保护的观澜工作区(本机演示:直接进入在线系统)。</p><label>用户名</label><input placeholder="username"><label>密码</label><input type="password" placeholder="••••••••"><a class="btn" href="http://127.0.0.1:18033/v2">进入详细分析 →</a></div></div></section>
 </div>
 <div class="foot"><span>观澜中文系统 · 详细分析</span><span>判断在证据 · 模型只转述</span><span>本机交互模块按对应入口启用</span></div>
@@ -886,7 +886,7 @@ requestAnimationFrame(frame);
 <!--@TEMPLATE:tpl-U6_变桨液压仿真台.html--><!--@TEMPLATE:tpl-standard_panel_zh.html--><!--@TEMPLATE:tpl-coverage_ch0_zh.html--><!--@TEMPLATE:tpl-如东传动链实际运行诊断_单文件版.html--><!--@TEMPLATE:tpl-U6_液压公共站健康报告_客户版.html--><!--@TEMPLATE:tpl-如东海上风电场_整机综合诊断与风险评估_主轴冲击深挖修订版V2.1.html-->
 <script>
 (function(){
-  function show(k){document.querySelectorAll('.pg').forEach(p=>p.classList.toggle('on',p.id===k));document.querySelectorAll('.tabs a').forEach(a=>a.classList.toggle('on',a.dataset.nav===k));window.scrollTo(0,0)}
+  function show(k){document.querySelectorAll('.pg').forEach(p=>p.classList.toggle('on',p.id===k));document.querySelectorAll('.tabs a').forEach(a=>a.classList.toggle('on',a.dataset.nav===k));/*数据重算: iframe 懒加载 —— 页面隐藏时不去轮询 /ops (它每 2 s 一次), 切到本页才挂 src*/if(k==='recalc'){var f=document.getElementById('rcf');if(f&&!f.getAttribute('src'))f.setAttribute('src','http://127.0.0.1:28084/ops/recalc')}window.scrollTo(0,0)}
   function route(){var k=(location.hash||'#index').slice(1);if(!document.getElementById(k))k='index';show(k)}
   window.addEventListener('hashchange',route);route();
   document.addEventListener('click',function(e){var a=e.target.closest('a[data-open]');if(!a)return;e.preventDefault();var t=document.getElementById('tpl-'+a.dataset.open);if(!t)return;document.getElementById('ovt').textContent=a.dataset.open;document.getElementById('ovf').srcdoc=t.content.textContent;/*guanlan-anchor-fix-v2*/(function(){var f=document.getElementById('ovf');if(!f||f._afix)return;f._afix=1;f.addEventListener('load',function(){try{var d=f.contentDocument;if(!d||d._afix)return;d._afix=1;d.addEventListener('click',function(ev){var a=ev.target&&ev.target.closest?ev.target.closest('a[href^="#"]'):null;if(!a)return;ev.preventDefault();var id=decodeURIComponent((a.getAttribute('href')||'#').slice(1));if(!id){d.defaultView.scrollTo(0,0);return}var el=d.getElementById(id)||(d.getElementsByName(id)||[])[0];if(el&&el.scrollIntoView)el.scrollIntoView({block:'start'})},true)}catch(e){}})})();document.getElementById('ov').classList.add('on')});

+ 7 - 7
scripts/_ops_launch.py

@@ -24,6 +24,8 @@ import subprocess
 import sys
 
 ROOT = pathlib.Path(__file__).resolve().parents[1]
+sys.path.insert(0, str(ROOT))
+from src import proc as _proc            # 无窗口子进程 (2026-09-16 用户令: 不弹命令窗口)
 
 
 def main() -> int:
@@ -38,13 +40,11 @@ def main() -> int:
     log = pathlib.Path(a.log)
     log.parent.mkdir(parents=True, exist_ok=True)
     env = dict(os.environ, PYTHONUTF8='1', PYTHONIOENCODING='utf-8', PYTHONUNBUFFERED='1')
-    flags = 0
-    if os.name == 'nt':
-        flags = subprocess.CREATE_NEW_PROCESS_GROUP | getattr(subprocess, 'DETACHED_PROCESS', 0)
-    with open(log, 'ab') as f:
-        p = subprocess.Popen([sys.executable, 'scripts/_ops_run.py', '--tag', a.tag, '--'] + cmd,
-                             cwd=str(ROOT), env=env, stdout=f, stderr=subprocess.STDOUT,
-                             stdin=subprocess.DEVNULL, creationflags=flags, close_fds=True)
+    # ★ 用 src.proc.spawn (CREATE_NO_WINDOW + 新进程组), **不要**再加 DETACHED_PROCESS:
+    #   两者互斥 (前者=给一个没有窗口的控制台, 后者=不给控制台); 原先这里用 DETACHED,
+    #   而本启动器自己可能是从"无控制台"的网关侧被拉起的, 落地实测会看到黑窗 (2026-09-16)。
+    p = _proc.spawn([sys.executable, 'scripts/_ops_run.py', '--tag', a.tag, '--'] + cmd,
+                    log=log, env=env, cwd=ROOT)
     print(p.pid)          # ← 最后一行是执行器 pid, 父进程读它
     return 0
 

+ 9 - 1
scripts/_ops_run.py

@@ -23,6 +23,8 @@ import threading
 import time
 
 ROOT = pathlib.Path(__file__).resolve().parents[1]
+sys.path.insert(0, str(ROOT))
+from src import proc as _proc            # 无窗口子进程 (2026-09-16 用户令: 不弹命令窗口)
 JOB = ROOT / 'run' / 'ops_job.json'
 HEARTBEAT_S = 15
 
@@ -64,7 +66,13 @@ def main() -> int:
     threading.Thread(target=_heartbeat, args=(stop_evt,), daemon=True).start()
     t0 = time.time()
     try:
-        rc = subprocess.run([sys.executable] + cmd, cwd=str(ROOT), env=dict(os.environ)).returncode
+        # 走 src.proc.run: 本进程是被无窗口起的 (没有可见控制台), 裸 spawn 会让 Windows
+        # 给子进程**新建一个可见控制台窗口** —— 从页面点"重算"就是这个窗口在跑 20 分钟 (2026-09-16 实测)。
+        # ★stdout/stderr **显式**传本进程的句柄: CREATE_NO_WINDOW 会新建控制台并把标准句柄重指过去,
+        #   不显式传的话, 子进程 (以及它的子进程) 的输出会掉进那个隐形控制台 —— 2026-09-16 实逮:
+        #   页面上"重算中"但 logs/ops_rebuild_*.log 里除了命令行一个字都没有。
+        rc = _proc.run([sys.executable] + cmd, cwd=str(ROOT), env=dict(os.environ),
+                       stdout=sys.stdout, stderr=sys.stderr).returncode
     except Exception as e:
         print(f'[X] 命令起不来: {type(e).__name__}: {e}', flush=True)
         rc = 99

+ 4 - 1
scripts/_ops_start_and_open.py

@@ -20,6 +20,7 @@ ROOT = pathlib.Path(__file__).resolve().parents[1]
 sys.path.insert(0, str(ROOT))
 
 import guanlan as G                                            # noqa: E402
+from src import proc as _proc                                  # noqa: E402 无窗口子进程
 
 
 def main() -> int:
@@ -27,7 +28,9 @@ def main() -> int:
     c = G.cfg()
     url = f"http://{c['host']}:{c['gateway']}/"
     print(f'① 启动服务 (guanlan.py serve) …', flush=True)
-    rc = subprocess.run([sys.executable, 'guanlan.py', 'serve'], cwd=str(ROOT)).returncode
+    # src.proc.run: 本脚本常由**无控制台**的运维任务/网关侧拉起, 裸 spawn 会让 serve 拿到一个
+    # 可见控制台 (2026-09-16 实测: 现场会攒下标题为 .venv\Scripts\python.exe 的黑窗)。
+    rc = _proc.run([sys.executable, 'guanlan.py', 'serve'], cwd=str(ROOT)).returncode
     if rc == 2:                                   # serve 的退出码 2 = 网关 60 s 没就绪
         print(f'[X] 网关未就绪 (serve 退出码 {rc}) —— 看 logs/gateway.log')
         return rc

+ 20 - 1
scripts/check_transferable.py

@@ -42,6 +42,11 @@ RE_ABS = re.compile(r"""(?<![\w:/])(?:[A-Za-z]:[\\/](?![\\/])|/Users/[A-Za-z0-9_
 TEXT_EXT = {'.py', '.bat', '.sh', '.ps1', '.cmd', '.json', '.md', '.txt', '.html', '.htm', '.yaml', '.yml', '.cfg', '.ini', '.toml'}
 SKIP_DIRS = {'.venv', '.git', '__pycache__', 'node_modules', '_products_off', 'logs', 'run'}
 SKIP_PATH_HINT = ('_products_off_prev',)
+# 扫描范围里要**排除现场原始件**: data/raw/** 是现场给的测量导出 (SCADA csv、Brande TCM JSON ——
+# 2026-09-12 落位后单这一个目录就是 2.5 万个 .json / 150 GB), 且**不随包分发** (pack_dist.py 的
+# EXCLUDE_GLOBS 含 'data')。文本闸门扫它: ①对"换机可用"零信息量 (那些文件里就算有绝对路径也不是我们的
+# 接线) ②每次验收要多读几十 GB 文本 (8 MB 以下的一律会读进内存), 把闸门从秒级拖成小时级。
+SKIP_RAW_PREFIX = ('data/raw/', 'data/_incoming/')
 # 允许的例外: 文档里明确在讲"路径约定"或演示用的占位
 ALLOW_LINE = ('portability-allow', '<安装目录>', '<现场包目录>', '<场站名称>', 'D:\\guanlan', '/opt/guanlan',
               'D:/guanlan', '示例', '例:', '例如')
@@ -55,6 +60,10 @@ def walk_files():
         parts = set(p.relative_to(ROOT).parts)
         if parts & SKIP_DIRS or any(h in rel for h in SKIP_PATH_HINT):
             continue
+        if rel.startswith(SKIP_RAW_PREFIX):
+            continue
+        if rel.startswith(SKIP_RAW_PREFIX):
+            continue
         if p.stat().st_size > 8 * 1024 * 1024:      # 超大文本(如 20MB 门户)另算, 见下
             continue
         yield p
@@ -248,7 +257,17 @@ def run_and_probe(dst: pathlib.Path, py: pathlib.Path, base_port=28084) -> int:
     print(f'   serve 退出码 {rs.returncode} (1 = 有模块未就绪, 属正常), 耗时 {time.time() - t0:.0f}s')
     ok = 0
     try:
-        urls = [('/', 20225828), ('/detail/', 229643), ('/cms/', None), ('/ops', None), ('/healthz', None)]
+        # ★门户期望字节数**不写死** (2026-09-16): 原来常量 20225828 是某一版门户"服务出来"的字节数,
+        #   门户外壳一改 (例如按用户令在菜单里加「数据重算」) 它就过期, 于是核验会打印一条永远不成立的
+        #   "★与主包不同", 看起来像移植出了问题。改为**从正在跑的主实例现量**: 主实例没起就不比这一项。
+        exp_portal = None
+        try:
+            with urllib.request.urlopen(f'http://127.0.0.1:{base_port}/', timeout=10) as rr:
+                exp_portal = len(rr.read())
+                print(f'   参照: 主实例门户 {exp_portal:,} B (端口 {base_port})')
+        except Exception as e:
+            print(f'   参照: 主实例 ({base_port}) 未响应 → 门户字节数不做同字节比对 ({type(e).__name__})')
+        urls = [('/', exp_portal), ('/detail/', None), ('/cms/', None), ('/ops', None), ('/healthz', None)]
         for path, exp_size in urls:
             try:
                 with urllib.request.urlopen(f'http://127.0.0.1:{gw}{path}', timeout=30) as rr:

+ 8 - 3
scripts/guanlan_facts_contract.py

@@ -9,10 +9,15 @@
 import datetime as dt, hashlib, json, os, re, shutil, subprocess, sys, tempfile
 from pathlib import Path
 
-ROOT = Path(__file__).resolve().parents[1]; OUT = ROOT / "outputs/rudong/guanlan"; DERIVED = OUT / "derived"
+ROOT = Path(__file__).resolve().parents[1]
+sys.path.insert(0, str(ROOT))
+from src import paths as P          # 路径唯一真源 (2026-09-16 统一: 原先写死 outputs/rudong/…)
+OUT = P.guanlan(); DERIVED = OUT / "derived"
 CONTRACT = OUT / "facts_contract_v0.json"
-SRC = {"findings": ROOT / "outputs/rudong/sop/findings.json", "e5": ROOT / "outputs/rudong/paradigm_r1/experiments/E5_coverage_table/底稿.json",
-       "e3": ROOT / "outputs/rudong/paradigm_r1/experiments/E3_candidate_closure/底稿.json", "s0": ROOT / "outputs/rudong/paradigm_r1/experiments/E8_rudong_rebuild/底稿_s0.json"}
+SRC = {"findings": P.sop() / "findings.json",
+       "e5": P.paradigm() / "experiments/E5_coverage_table/底稿.json",
+       "e3": P.paradigm() / "experiments/E3_candidate_closure/底稿.json",
+       "s0": P.paradigm() / "experiments/E8_rudong_rebuild/底稿_s0.json"}
 sha = lambda b: hashlib.sha256(b).hexdigest(); sha16f = lambda p: sha(Path(p).read_bytes())[:16]; J = lambda p: json.loads(Path(p).read_text(encoding="utf-8"))
 SIX = ("定论", "准定论·预警", "候选", "参考", "INSUFFICIENT", "撤回")
 BANNED = re.compile(r"中广核|华电|国电|大唐|华能|三峡|龙源|/Users/|[A-Za-z0-9._%+-]+@[A-Za-z0-9.-]+\.[a-z]{2,}|1[3-9]\d{9}")   # 业主/集团名 · 本机路径 · 邮箱 · 手机 (对外面孔闸)

+ 14 - 4
scripts/guanlan_gateway.py

@@ -12,6 +12,7 @@
 import sys as _sys, pathlib as _plb
 _sys.path.insert(0, str(_plb.Path(__file__).resolve().parents[1]))
 from src import paths as P
+from src import proc as _proc           # 无窗口子进程 (2026-09-16 用户令: 不弹命令窗口)
 import argparse, datetime as dt, hashlib, http.client, json, os, pathlib, re, socket, subprocess, sys, time, urllib.parse
 from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer
 from concurrent.futures import ThreadPoolExecutor
@@ -168,7 +169,9 @@ def healthz():
 def version():
     man = json.loads(MANIFEST.read_text(encoding="utf-8")) if MANIFEST.exists() else {}
     b = man.get("detail_style_baseline", {}); c = man.get("content_contract", {})
-    try: head = subprocess.run(["git", "-C", str(ROOT), "rev-parse", "--short", "HEAD"], capture_output=True, text=True, timeout=5).stdout.strip()
+    # src.proc.run_text: 网关无控制台, 裸 spawn git 会给它新建一个可见控制台窗口 ——
+    # 而 /api/version 会被页面反复请求 ⇒ 反复闪窗 (2026-09-16 实测)。
+    try: head = _proc.run_text(["git", "-C", str(ROOT), "rev-parse", "--short", "HEAD"], timeout=5).stdout.strip()
     except Exception: head = None
     con = json.loads(CONTRACT.read_text(encoding="utf-8")) if CONTRACT.exists() else {}
     try: psha = _portal_sha()                       # 缓存版: 同一份门户不重复读 20 MB
@@ -198,8 +201,15 @@ class H(BaseHTTPRequestHandler):
         b = json.dumps(obj, ensure_ascii=False, indent=1).encode("utf-8"); self.send_response(code); self.send_header("Content-Type", "application/json; charset=utf-8"); self.send_header("Content-Length", str(len(b))); self.send_header("Cache-Control", "no-store"); self.end_headers()
         if self.command != "HEAD": self.wfile.write(b)
 
-    def _bytes(self, b, ctype, code=200):
-        self.send_response(code); self.send_header("Content-Type", ctype); self.send_header("Content-Length", str(len(b))); self.end_headers()
+    def _bytes(self, b, ctype, code=200, no_store=False):
+        self.send_response(code); self.send_header("Content-Type", ctype); self.send_header("Content-Length", str(len(b)))
+        if no_store:
+            # 运维控制台/数据重算页**必须**不缓存 (2026-09-16 实逮): 页面 HTML/脚本是随包更新的
+            # (例如"内嵌版按钮点不动"这次修复), 没有 no-store 时浏览器会用旧页面 —— 用户看到的是
+            # "改了还是没反应", 排查方向会被带偏。数据面同样: 状态必须每次现取。
+            self.send_header("Cache-Control", "no-store, must-revalidate")
+            self.send_header("Pragma", "no-cache")
+        self.end_headers()
         if self.command != "HEAD": self.wfile.write(b)
 
     def _route(self, path):
@@ -226,7 +236,7 @@ class H(BaseHTTPRequestHandler):
             ops = _ops_module()
             body = self._read_body() if self.command == "POST" else b""
             code, mime, data = ops.handle(self.command, path, body)
-            return self._bytes(data, mime, code)
+            return self._bytes(data, mime, code, no_store=True)
         if path == "/local-ai/status":
             run, models = ollama_models(); return self._json(dict(running=run, models=models, endpoint="/local-ai/", note=None if run else "本机模型未启动"), 200 if run else 503)
         if path.startswith("/release/"):   # E9/E10 裁定: Release 层不合并主库, 由网关只读暴露: /release/ 列表; /release/r1/… (E9 契约层) /release/r2/… (E10 分层覆盖)

+ 86 - 63
scripts/guanlan_ops.py

@@ -8,14 +8,14 @@ r"""观澜运维控制台的后端 (2026-09-12) —— 把"停服务 / 起服务
 本模块给网关加三个东西:
   · `GET  /ops`                  一个自适应控制台页面(单文件, 无外部依赖);
   · `GET  /ops/api/state`        真实状态: 各服务端口通不通 · 产物在不在 · 有没有任务在跑 · 上次结果;
-  · `POST /ops/api/<动作>`       停服务 · 起服务 · 重算 · 清产物 · 恢复产物。
+  · `POST /ops/api/<动作>`       停服务 · 起服务 · 重算 · 清产物。
 
 ## 三条设计纪律
 
 1. **页面活着的服务不能被自己停掉。** 控制台由网关(28084)提供, 所以"停服务"默认**保留网关**,
    否则按钮刚点完页面就没了、再也没法"启动服务"。要连网关一起停, 用 `include_gateway=True`,
    这时用**分离进程**先停再起(页面会断开十几秒, 之后自动重连)。
-2. **一次只允许一个动作。** 四个动作互相冲突(重算要重启服务、清产物要挪走产物), 所以有一把锁:
+2. **一次只允许一个动作。** 各动作互相冲突(重算要重启服务、清产物要删产物), 所以有一把锁:
    `run/ops_job.json` 里记着当前任务; 任务在跑时, 所有会冲突的按钮在后端**与前端都会被禁掉**
    (后端拒绝 = 真生效, 前端禁用只是提示)。判断以后端为准。
 3. **日志与状态落盘。** 动作全部 `Popen` 到 `logs/ops_<动作>_<时间>.log`, 页面轮询 job 状态并把日志尾巴显示出来 ——
@@ -26,8 +26,7 @@ r"""观澜运维控制台的后端 (2026-09-12) —— 把"停服务 / 起服务
     停服务    : 至少有一个组件服务在监听
     启动服务  : 至少有一个组件服务没在监听
     执行重算  : 没有任务在跑
-    清除产物  : 没有任务在跑 且 产物在位 且 暂存区没有同名目录(或允许 --archive-old)
-    恢复产物  : 没有任务在跑 且 产物被清掉过(暂存区清单存在)
+    清除产物  : 没有任务在跑 且 产物在位 (**直接删除, 不留备份** —— 2026-09-16 用户令; 无"恢复产物"按钮)
 """
 from __future__ import annotations
 
@@ -43,6 +42,7 @@ import time
 ROOT = pathlib.Path(__file__).resolve().parents[1]
 sys.path.insert(0, str(ROOT))
 from src import paths as P                                        # noqa: E402
+from src import proc as _proc                                     # noqa: E402 无窗口子进程
 
 RUN = ROOT / 'run'
 LOGS = ROOT / 'logs'
@@ -112,8 +112,9 @@ def job_running() -> tuple[bool, dict]:
             if os.name == 'nt':
                 # errors='replace': tasklist 输出是控制台代码页(GBK), 而本进程可能是 PYTHONUTF8=1
                 # → 严格解码会失败并使 stdout 为 None (guanlan.py 的 alive() 踩过同一个坑)
-                out = subprocess.run(['tasklist', '/FI', f'PID eq {pid}'], capture_output=True,
-                                     text=True, errors='replace').stdout
+                # ★ 走 src.proc.run_text: 本模块跑在网关进程里 (无控制台), 裸 spawn tasklist 会让
+                #   Windows 新建可见控制台 —— 而 /ops 页面每 2 s 轮询一次 ⇒ **反复闪窗** (2026-09-16 实测)。
+                out = _proc.run_text(['tasklist', '/FI', f'PID eq {pid}']).stdout
                 alive = bool(out) and str(pid) in out
             else:
                 os.kill(pid, 0); alive = True
@@ -148,8 +149,9 @@ def spawn(args: list, tag: str, detached=False) -> dict:
     RUN.mkdir(exist_ok=True); LOGS.mkdir(exist_ok=True)
     log = LOGS / f'ops_{tag}_{time.strftime("%Y%m%d_%H%M%S")}.log'
     env = dict(os.environ, PYTHONUTF8='1', PYTHONIOENCODING='utf-8', PYTHONUNBUFFERED='1')
-    r = subprocess.run([PY, 'scripts/_ops_launch.py', '--tag', tag, '--log', str(log), '--'] + list(args),
-                       cwd=str(ROOT), env=env, capture_output=True, text=True, errors='replace', timeout=60)
+    # 走 src.proc.run_text: 网关进程无控制台, 裸 spawn 会让这次"拉启动器"也闪一下窗
+    r = _proc.run_text([PY, 'scripts/_ops_launch.py', '--tag', tag, '--log', str(log), '--'] + list(args),
+                       cwd=str(ROOT), env=env, timeout=60)
     pid = None
     for line in reversed((r.stdout or '').strip().splitlines()):
         if line.strip().isdigit():
@@ -173,18 +175,15 @@ def products_state() -> dict:
             stores[d.name] = sum(1 for _ in d.rglob('*') if _.is_file())
     prov = fam / '_provenance.json'
     prov_d = json.loads(prov.read_text(encoding='utf-8')) if prov.is_file() else None
-    stash_dirs = {}
-    if OFF.is_dir():
-        for d in sorted(x for x in OFF.iterdir() if x.is_dir() and x.name != '_baseline_kept'):
-            stash_dirs[d.name] = sum(1 for _ in (d / store.name).rglob('*') if _.is_file()) if (d / store.name).is_dir() else 0
-    kept = list((OFF / '_baseline_kept').rglob('*')) if (OFF / '_baseline_kept').is_dir() else []
+    # ★2026-09-16 用户令"不要备份清除的产物": 清除 = 真删除, 没有暂存区可列。
+    #   这里只报"有没有旧设计留下的 _products_off*"(有就提示手工删掉), 不再提供"恢复产物"。
+    legacy = sorted(x.name for x in ROOT.glob('_products_off*') if x.is_dir())
     return dict(
         fam=str(fam.relative_to(ROOT)).replace('\\', '/'), store=str(store.relative_to(ROOT)).replace('\\', '/'),
         files=n, in_place=n > 0, stores=stores,
         cleared=n == 0,
-        stash_has_same=any(v > 0 for v in stash_dirs.values()), stash=stash_dirs,
-        can_restore=OFF_MANIFEST.is_file() and n == 0,
-        baseline_kept=sum(1 for x in kept if x.is_file()),
+        no_backup=True,                                  # 现口径: 清除不留备份
+        legacy_backups=legacy,                           # 旧设计残留 (按用户令不自动删, 只提示)
         provenance=(dict(counts=prov_d.get('counts'), at=prov_d.get('at')) if prov_d else None),
     )
 
@@ -214,7 +213,6 @@ def state() -> dict:
             start_services=any_comp_down and not running,
             rebuild=not running,
             products_off=pr['in_place'] and not running,
-            products_on=(not pr['in_place']) and not running,
         ),
         anchors=dict(alarms=_rows(f'{pr["fam"]}/windscada/alarms.parquet'),
                      workorders=_rows(f'{pr["fam"]}/windscada/workorders.parquet'),
@@ -309,38 +307,47 @@ def act_rebuild(body: dict) -> tuple[int, dict]:
 
 
 def act_products_off(body: dict) -> tuple[int, dict]:
+    """清除产物 —— 2026-09-16 用户令: **直接删除, 不留备份**。
+
+    原实现是 `products_state.py --off`(移动到 `_products_off/`, 可 `--on` 还原)。现口径: 真删,
+    故这里传 `--yes`(脚本对不可恢复操作要求显式确认; 前端已 confirm 过一次)。
+    恢复缺失的**随包件**改用交付包补齐:
+    `python scripts/products_restore_missing.py --stash <交付包.zip>`。
+    """
     if (e := _guard('products_off')):
         return 409, dict(err=e)
     pr = products_state()
     if not pr['in_place']:
         return 409, dict(err='产物已经是清空状态, 无需再清')
-    args = ['scripts/products_state.py', '--off']
-    if pr['stash_has_same'] or body.get('archive_old'):
-        args.append('--archive-old')      # 暂存区已有同名产物时必须加, 否则会嵌套(见该脚本注释)
-    j = spawn(args, 'products_off')
-    return 200, dict(ok=True, job=j, note='正在把产物挪到 _products_off/ (挪完请点"启动服务"让页面呈现空状态)')
-
-
-def act_products_on(body: dict) -> tuple[int, dict]:
-    if (e := _guard('products_on')):
-        return 409, dict(err=e)
-    pr = products_state()
-    if pr['in_place']:
-        return 409, dict(err='产物已在位, 无需恢复')
-    if not OFF_MANIFEST.is_file():
-        return 409, dict(err='没有 _products_off/manifest.json (没清过或清单已删), 无法恢复')
-    j = spawn(['scripts/products_state.py', '--on'], 'products_on')
-    return 200, dict(ok=True, job=j, note='正在恢复产物; 恢复后请点"启动服务"')
+    j = spawn(['scripts/products_state.py', '--off', '--yes'], 'products_off')
+    return 200, dict(ok=True, job=j, note='正在**删除**产物(不留备份, 不可恢复); 清完请点"启动服务"让页面呈现空状态')
 
 
 ACTIONS = dict(stop_services=act_stop, start_services=act_start, rebuild=act_rebuild,
-               products_off=act_products_off, products_on=act_products_on)
+               products_off=act_products_off)
 
 
 def handle(method: str, path: str, body: bytes) -> tuple[int, str, bytes]:
-    """网关调用入口: → (http code, content-type, body bytes)。"""
+    """网关调用入口: → (http code, content-type, body bytes)。
+
+    路由:
+      `/ops`           运维控制台整页 (服务 + 重算 + 产物 + 动作进度)
+      `/ops/recalc`    **内嵌版**: 只保留 重算 + 产物 (+ 动作进度), 给门户菜单「数据重算」用
+                       (2026-09-16 用户令: 把 /ops 的重算与产物搬到门户菜单里)。
+                       走独立路由而不是 `?embed=…`: 网关转给本模块的是 `u.path` (查询串被丢掉),
+                       用查询串会变成"看起来支持、实际不生效"的静默坑。
+      `/ops/api/...`   运维动作 JSON API (GET state / POST 各动作)
+    """
     if path in ('/ops', '/ops/'):
         return 200, 'text/html; charset=utf-8', PAGE.encode('utf-8')
+    if path in ('/ops/recalc', '/ops/recalc/'):
+        import re as _re
+        html = _re.sub(r'<!--HIDE_IN_EMBED-->.*?<!--/HIDE_IN_EMBED-->', '', PAGE, flags=_re.S)
+        html = (html.replace('<title>观澜 · 运维控制台</title>', '<title>观澜 · 数据重算</title>')
+                    .replace('</style>', '.wrap{max-width:100%;padding:8px 10px 24px}'
+                                        'body{background:transparent}'
+                                        '.card{margin:0 0 12px}</style>', 1))
+        return 200, 'text/html; charset=utf-8', html.encode('utf-8')
     if path == '/ops/api/state':
         return 200, 'application/json; charset=utf-8', json.dumps(state(), ensure_ascii=False).encode('utf-8')
     if path.startswith('/ops/api/'):
@@ -385,10 +392,13 @@ th{color:var(--mut);font-weight:600}
 .busy{background:#FFF7E6;border:1px solid #F0D9A8;color:#7A5600;padding:8px 12px;border-radius:8px;margin:0 0 12px}
 label{font-size:13.5px;color:var(--mut);display:flex;gap:6px;align-items:center}
 </style></head><body><div class="wrap">
+<!--HIDE_IN_EMBED-->
 <h1>观澜 · 运维控制台</h1>
 <p class="sub">停/启服务 · 执行重算 · 清除产物 —— 按钮按真实状态启用; 不可用的动作后端也会拒绝。本页由网关(端口 <span id="gw"></span>)提供。</p>
+<!--/HIDE_IN_EMBED-->
 <div id="busy"></div>
 
+<!--HIDE_IN_EMBED-->
 <div class="card"><h2>服务</h2><div id="svc"></div>
   <div class="row" style="margin-top:12px">
     <button id="b_start" class="ghost">启动服务(并打开门户)</button>
@@ -397,6 +407,7 @@ label{font-size:13.5px;color:var(--mut);display:flex;gap:6px;align-items:center}
   </div>
   <p class="hint" id="svc_hint"></p>
 </div>
+<!--/HIDE_IN_EMBED-->
 
 <div class="card"><h2>重算</h2>
   <div class="row">
@@ -412,14 +423,14 @@ label{font-size:13.5px;color:var(--mut);display:flex;gap:6px;align-items:center}
 
 <div class="card"><h2>产物</h2><div id="prod"></div>
   <div class="row" style="margin-top:12px">
-    <button id="b_on" class="ghost">恢复产物</button>
-    <button id="b_off" class="danger">清除产物(挪到 _products_off,可恢复)</button>
+    <button id="b_off" class="danger">清除产物(直接删除,不留备份)</button>
   </div>
-  <p class="hint">清除后门户首页仍可打开(它是静态交付页), 但工作台会显示"无产物"; 恢复后请点"启动服务"。</p>
+  <p class="hint">按用户令(2026-09-16)**清除不留备份、不可恢复**;清完请点"启动服务"让页面呈现空状态。
+     要补回"包内没有生成端"的随包件:<code>python scripts/products_restore_missing.py --stash &lt;交付包.zip&gt;</code>(从交付包按需补齐)。</p>
 </div>
 
 <div class="card"><h2>最近一次动作</h2><div id="job">(无)</div><pre id="log"></pre></div>
-<p class="hint">命令行等价物: <code>guanlan.py stop/serve</code> · <code>scripts/rebuild_all.py</code> · <code>scripts/products_state.py --off/--on</code>(手册 §0/§5b)。</p>
+<p class="hint">命令行等价物: <code>guanlan.py stop/serve</code> · <code>scripts/rebuild_all.py</code> · <code>scripts/products_state.py --off --yes/--status</code>(手册 §0/§5b)。清除产物**不留备份**(用户令 2026-09-16);要补回缺失的随包件用 <code>scripts/products_restore_missing.py --stash &lt;交付包.zip&gt;</code>。</p>
 </div><script>
 const $ = s => document.querySelector(s);
 let S = null, busyTimer = null;
@@ -428,35 +439,37 @@ async function api(path, body){
   return [r.status, await r.json().catch(()=>({err:'非 JSON 响应'}))];
 }
 function pill(up){ return `<span class="pill ${up?'up':'down'}">${up?'运行中':'未运行'}</span>`; }
+/* set: 元素可能不存在 —— 内嵌版 (/ops/recalc) 会去掉"服务"卡与页头, 直接 $() 取值再赋值会抛错,
+   而这里一抛整个 render() 就断, 页面看着像"卡住不动"(2026-09-16 加内嵌版时踩过)。 */
+function set(sel, fn){ const e = $(sel); if(e) fn(e); }
 function render(){
   const s = S; if(!s) return;
-  $('#gw').textContent = s.services.gateway ? s.services.gateway.port : '?';
+  set('#gw', e=>e.textContent = s.services.gateway ? s.services.gateway.port : '?');
   const rows = Object.entries(s.services).map(([n,v]) =>
      `<tr><td>${n}${v.is_gateway?' <span class="mut">(控制台本体)</span>':''}</td><td>${v.port}</td><td>${pill(v.up)}</td></tr>`).join('');
-  $('#svc').innerHTML = `<table><tr><th>服务</th><th>端口</th><th>状态</th></tr>${rows}</table>`;
+  set('#svc', e=>e.innerHTML = `<table><tr><th>服务</th><th>端口</th><th>状态</th></tr>${rows}</table>`);
   const p = s.products;
   const st = Object.entries(p.stores||{}).map(([k,v])=>`${k} ${v}`).join(' · ') || '(空)';
-  $('#prod').innerHTML = `<table>
+  set('#prod', e=>e.innerHTML = `<table>
     <tr><th>产物仓</th><td>${p.store}</td></tr>
     <tr><th>件数</th><td>${p.files} 件 ${p.in_place?'<span class="pill up">在位</span>':'<span class="pill down">已清空</span>'}</td></tr>
     <tr><th>分布</th><td class="mut">${st}</td></tr>
     <tr><th>来源台账</th><td class="mut">${p.provenance?`raw 重算 ${p.provenance.counts['raw-derived']} 件 · 随包补齐 ${p.provenance.counts.shipped} 件 (${p.provenance.at})`:'(无 _provenance.json)'}</td></tr>
-    <tr><th>暂存区</th><td class="mut">${Object.keys(p.stash||{}).length?Object.entries(p.stash).map(([k,v])=>`${k}:${v} 件`).join(' · '):'(空)'}${p.stash_has_same?' <b>← 有同名产物, 清除时会自动存档上一代</b>':''}</td></tr>
+    <tr><th>备份</th><td class="mut">${p.no_backup?'**不留备份**(用户令 2026-09-16: 清除 = 直接删除)':''}${(p.legacy_backups&&p.legacy_backups.length)?` <b>← 旧设计残留 ${p.legacy_backups.join(' · ')} —— 确认不需要后手工删掉</b>`:''}</td></tr>
     <tr><th>验收锚点</th><td class="mut">报警 ${s.anchors.alarms??'—'} 行 · 工单 ${s.anchors.workorders??'—'} · temp_monthly ${s.anchors.temp_monthly??'—'} · 本体 ${s.anchors.objects??'—'} 对象</td></tr>
-  </table>`;
+  </table>`);
   const b = s.buttons, run = s.job && s.job.status==='running';
-  $('#b_start').disabled   = !b.start_services;
-  $('#b_stop').disabled    = !b.stop_services;
-  $('#b_restart').disabled = !b.stop_services;
-  $('#b_rebuild').disabled = !b.rebuild;
-  $('#b_off').disabled     = !b.products_off;
-  $('#b_on').disabled      = !b.products_on;
-  $('#svc_hint').textContent = b.start_services ? '有服务未运行 → 可启动。' : '全部组件服务已在运行 → 启动按钮已禁用。';
-  $('#rebuild_hint').textContent = b.rebuild ? '' : '(有任务在跑, 重算按钮已禁用)';
-  $('#busy').innerHTML = run ? `<div class="busy">正在执行: <b>${s.job.kind}</b>(起于 ${s.job.started})${s.job.note?' — '+s.job.note:''} · 完成后本页自动刷新</div>` : '';
+  set('#b_start',   e=>e.disabled = !b.start_services);
+  set('#b_stop',    e=>e.disabled = !b.stop_services);
+  set('#b_restart', e=>e.disabled = !b.stop_services);
+  set('#b_rebuild', e=>e.disabled = !b.rebuild);
+  set('#b_off',     e=>e.disabled = !b.products_off);
+  set('#svc_hint', e=>e.textContent = b.start_services ? '有服务未运行 → 可启动。' : '全部组件服务已在运行 → 启动按钮已禁用。');
+  set('#rebuild_hint', e=>e.textContent = b.rebuild ? '' : '(有任务在跑, 重算按钮已禁用)');
+  set('#busy', e=>e.innerHTML = run ? `<div class="busy">正在执行: <b>${s.job.kind}</b>(起于 ${s.job.started})${s.job.note?' — '+s.job.note:''} · 完成后本页自动刷新</div>` : '');
   const j = s.job;
-  $('#job').innerHTML = j ? `<div>${j.kind} · <b>${j.status==='running'?'执行中':(j.rc===0?'完成 (退出码 0)':'结束 (退出码 '+j.rc+')')}</b> · ${j.started||''}${j.seconds?' · 耗时 '+j.seconds+'s':''}</div><div class="mut">${j.cmd||''}</div>${j.note?'<div class="mut">'+j.note+'</div>':''}` : '(无)';
-  $('#log').textContent = (j && j.log_tail && j.log_tail.length) ? j.log_tail.join('\n') : '';
+  set('#job', e=>e.innerHTML = j ? `<div>${j.kind} · <b>${j.status==='running'?'执行中':(j.rc===0?'完成 (退出码 0)':'结束 (退出码 '+j.rc+')')}</b> · ${j.started||''}${j.seconds?' · 耗时 '+j.seconds+'s':''}</div><div class="mut">${j.cmd||''}</div>${j.note?'<div class="mut">'+j.note+'</div>':''}` : '(无)');
+  set('#log', e=>e.textContent = (j && j.log_tail && j.log_tail.length) ? j.log_tail.join('\n') : '');
 }
 async function refresh(){ const [c,d] = await api('/ops/api/state'); if(c===200){ S=d; render(); } }
 async function fire(name, body){
@@ -465,12 +478,22 @@ async function fire(name, body){
   else { if(d.note) console.log(d.note); }
   await refresh();
 }
-$('#b_start').onclick   = ()=>fire('start_services', {open_browser:true});
-$('#b_stop').onclick    = ()=>fire('stop_services', {include_gateway:false});
-$('#b_restart').onclick = ()=>{ if(confirm('完整重启会连网关一起停, 本页会断开约 15 秒后自动恢复。继续?')) fire('stop_services', {include_gateway:true}); };
-$('#b_rebuild').onclick = ()=>{ if(confirm('开始重算? 期间服务会被重启, 页面可能短暂打不开。')) fire('rebuild', {skip_scada: !$('#f_scada').checked, with_verify: $('#f_verify').checked}); };
-$('#b_off').onclick     = ()=>{ if(confirm('清除产物(挪到 _products_off,可恢复)?')) fire('products_off', {}); };
-$('#b_on').onclick      = ()=>{ if(confirm('恢复产物?')) fire('products_on', {}); };
-refresh(); setInterval(refresh, 2000);
+/* on: 与 set 同理 —— 内嵌版没有"服务"卡, 直接 $('#b_start').onclick=… 会抛 TypeError,
+   而**这一抛会让后面所有绑定与 refresh() 全不执行** (2026-09-16 用户实测: 门户 #recalc 页里
+   "执行重算 / 清除产物"点了没反应)。容错绑定 + 下面 try/catch 兜底 = 同类问题不再静默。 */
+function on(sel, fn){ const e = $(sel); if(e) e.onclick = fn; }
+try {
+  on('#b_start',   ()=>fire('start_services', {open_browser:true}));
+  on('#b_stop',    ()=>fire('stop_services', {include_gateway:false}));
+  on('#b_restart', ()=>{ if(confirm('完整重启会连网关一起停, 本页会断开约 15 秒后自动恢复。继续?')) fire('stop_services', {include_gateway:true}); });
+  on('#b_rebuild', ()=>{ if(confirm('开始重算? 期间服务会被重启, 页面可能短暂打不开。')) fire('rebuild', {skip_scada: !$('#f_scada').checked, with_verify: $('#f_verify').checked}); });
+  on('#b_off',     ()=>{ if(confirm('清除产物 = **直接删除, 不留备份, 不可恢复**。继续?')) fire('products_off', {}); });
+  refresh(); setInterval(refresh, 2000);
+} catch (err) {
+  /* 界面脚本自己坏了也要看得见 (否则表现就是"按钮点了没反应", 排查全靠猜) */
+  const b = document.getElementById('busy');
+  if (b) b.innerHTML = '<div class="busy">⚠ 本页脚本出错, 按钮可能不可用: ' + err + ' (请反馈该行)</div>';
+  else alert('本页脚本出错: ' + err);
+}
 </script></body></html>
 """

+ 150 - 0
scripts/guanlan_start_hidden.py

@@ -0,0 +1,150 @@
+#!/usr/bin/env python3
+# -*- coding: utf-8 -*-
+"""**无窗口**启动观澜 (2026-09-16 用户令: "启动观澜系统, 弹出的命令窗口, 改为不弹出方式")。
+
+为什么单开一个脚本: `guanlan.py serve` 是**前台阻塞**命令 (它在控制台里打启动横幅、等 healthz),
+双击 start.bat 就会一直挂着一个黑窗; 而各组件服务其实早已是 `DETACHED_PROCESS` 起的 (不弹窗)。
+所以只需要一个"把 serve 藏到后台、等就绪、打开浏览器、自己退出"的入口。
+
+做法:
+  1. 网关已在 → 只打开浏览器 (重复点不重复起, 幂等);
+  2. 否则用 `CREATE_NO_WINDOW | DETACHED_PROCESS` 起 `guanlan.py serve`, stdout/stderr 落
+     `logs/serve.log` (有窗口才有地方看日志, 藏起来就必须落文件);
+  3. 轮询 `/healthz` 直到就绪 (最多 `--wait` 秒), 就绪后按 `--open/--no-open` 决定是否开浏览器;
+  4. 全过程写 `logs/start_hidden.log` (含退出码与失败原因) —— **不静默**:
+     失败时不但写日志, 还会弹一个消息框 (无窗口模式下唯一能让人看见的方式)。
+
+由 `start_hidden.vbs` 以 `pythonw.exe` 调用 (窗口样式 0 = 完全不显示)。
+
+用法:
+    .venv\\Scripts\\pythonw.exe scripts\\guanlan_start_hidden.py          # 隐藏启动 + 开浏览器
+    .venv\\Scripts\\python.exe  scripts\\guanlan_start_hidden.py --no-open  # 隐藏启动, 不开浏览器
+    .venv\\Scripts\\python.exe  scripts\\guanlan_start_hidden.py --wait 180 # 加长等待
+"""
+from __future__ import annotations
+
+import argparse
+import datetime as dt
+import json
+import os
+import pathlib
+import socket
+import subprocess
+import sys
+import time
+import urllib.request
+
+ROOT = pathlib.Path(__file__).resolve().parents[1]
+sys.path.insert(0, str(ROOT))
+from src import paths as P          # noqa: E402
+from src.console import soft        # noqa: E402
+
+soft()
+LOG = P.LOGS / 'start_hidden.log'
+
+
+def log(msg: str) -> None:
+    P.LOGS.mkdir(parents=True, exist_ok=True)
+    line = f'{dt.datetime.now():%Y-%m-%d %H:%M:%S} {msg}'
+    with open(LOG, 'a', encoding='utf-8') as f:
+        f.write(line + '\n')
+
+
+def port_up(host: str, port: int, t: float = 0.4) -> bool:
+    try:
+        with socket.create_connection((host, port), timeout=t):
+            return True
+    except OSError:
+        return False
+
+
+def healthz(url: str, timeout=6.0) -> bool:
+    try:
+        with urllib.request.urlopen(url, timeout=timeout) as r:
+            return r.status == 200
+    except Exception:
+        return False
+
+
+def notify(title: str, text: str) -> None:
+    """无窗口模式下唯一能让人看见失败的通道 (Windows 消息框; 非 Windows 走 stderr)。"""
+    if os.name != 'nt':
+        print(f'{title}: {text}', file=sys.stderr)
+        return
+    try:
+        import ctypes
+        ctypes.windll.user32.MessageBoxW(0, text, title, 0x30)      # MB_ICONWARNING
+    except Exception:
+        pass
+
+
+def main() -> int:
+    ap = argparse.ArgumentParser()
+    ap.add_argument('--host', default='127.0.0.1')
+    ap.add_argument('--gateway', type=int, default=None, help='网关端口 (默认读 configs/serve.json)')
+    ap.add_argument('--wait', type=int, default=120, help='等 healthz 就绪的最长秒数')
+    ap.add_argument('--no-open', action='store_true', help='就绪后不开浏览器')
+    ap.add_argument('--quiet-fail', action='store_true', help='失败时不弹消息框 (只写日志)')
+    a = ap.parse_args()
+
+    gw = a.gateway
+    if gw is None:
+        try:
+            cfg = json.loads(P.SERVE_JSON.read_text(encoding='utf-8-sig'))
+            gw = int(cfg.get('gateway') or 28084)
+        except Exception:
+            gw = 28084
+    url = f'http://{a.host}:{gw}/'
+    log(f'--- 无窗口启动请求 (端口 {gw}, wait={a.wait}s) ---')
+
+    if port_up(a.host, gw):
+        log(f'网关 {gw} 已在运行 → 只打开页面')
+        if not a.no_open:
+            import webbrowser
+            webbrowser.open(url)
+        return 0
+
+    py = P.venv_python() or pathlib.Path(sys.executable)
+    cmd = [str(py), 'guanlan.py', 'serve']
+    flags = 0
+    if os.name == 'nt':
+        flags = (getattr(subprocess, 'CREATE_NO_WINDOW', 0x08000000)
+                 | getattr(subprocess, 'CREATE_NEW_PROCESS_GROUP', 0)
+                 | getattr(subprocess, 'DETACHED_PROCESS', 0))
+    P.LOGS.mkdir(parents=True, exist_ok=True)
+    lf = open(P.LOGS / 'serve.log', 'ab')
+    kw = dict(cwd=str(ROOT), stdout=lf, stderr=subprocess.STDOUT, stdin=subprocess.DEVNULL)
+    if os.name == 'nt':
+        kw['creationflags'] = flags
+    else:
+        kw['start_new_session'] = True
+    try:
+        pid = subprocess.Popen(cmd, **kw).pid
+    except Exception as e:
+        log(f'✘ 起 serve 失败: {type(e).__name__}: {e}')
+        if not a.quiet_fail:
+            notify('观澜启动失败', f'无法启动服务: {e}\n详见 logs\\start_hidden.log')
+        return 3
+    log(f'起 serve pid={pid} (无窗口; 日志 logs/serve.log)')
+
+    t0 = time.time()
+    while time.time() - t0 < a.wait:
+        if healthz(f'{url}healthz'):
+            log(f'✔ 就绪 ({time.time() - t0:.0f}s) → {url}')
+            if not a.no_open:
+                try:
+                    import webbrowser
+                    webbrowser.open(url)
+                except Exception as e:
+                    log(f'⚠ 打开浏览器失败: {e}')
+            return 0
+        time.sleep(1.5)
+    log(f'✘ {a.wait}s 内 healthz 未就绪 (serve pid={pid} 可能已退出)')
+    if not a.quiet_fail:
+        notify('观澜启动未就绪', f'{a.wait} 秒内网关未就绪, 请查看:\n'
+                                  f'logs\\serve.log\nlogs\\start_hidden.log')
+    return 2
+
+
+if __name__ == '__main__':
+    sys.exit(main())

+ 17 - 4
scripts/ingest_ops_2025.py

@@ -21,8 +21,10 @@
   重算成熟度时必须按实际非空率判, 不能因为"字段有了"就宣布该维可评估 —
   那会把 19/818 的覆盖说成能力达标。
 """
+import pathlib as _pl, sys as _sys
+_sys.path.insert(0, str(_pl.Path(__file__).resolve().parents[1]))   # 2026-09-16: 与其它摄入脚本一致
 from src import paths as _P
-import argparse, json, pathlib, sys, warnings
+import argparse, json, os, pathlib, sys, warnings
 warnings.filterwarnings('ignore')
 import pandas as pd
 
@@ -44,7 +46,12 @@ def _ledger_range(raw):
         return None, None
 
 
-OBJ = pathlib.Path('outputs/rudong/ontology/objects.json')
+OBJ = _P.objects_json()
+# ★2026-09-16 修: 原为 `pathlib.Path('outputs/rudong/ontology/objects.json')` —— **cwd 相对** + 写死场名。
+#   本文件在模块级用它, 于是: ① 从别处调用/换工作目录就会写到别处; ② 它还是全仓唯一没有
+#   `sys.path.insert(0, str(ROOT))` 的摄入脚本, 干净环境下连 `src.paths` 都导不进来。
+#   注意: 该脚本此前因为缺 `import os` (上面那行用 os.environ) 在模块级就 NameError, 从未真正跑过 ——
+#   所以"补 import os"必须和"改掉 cwd 相对路径 + 走 Store 落盘"同一次做完, 否则只是把一把静默的凶器上膛。
 OPS = pathlib.Path(os.environ.get('WINDSCADA_OPS_LIB') or (_P.RAW_ROOT / '工作库' / '20_structured' / 'ops'))
                           # 原为 mac 盘路径 (本机不存在): 改为 env 可配 + 安装目录相对默认位置
 
@@ -297,8 +304,14 @@ def main():
     if a.dry_run:
         print('  (dry-run, 未写盘)')
         return 0
-    OBJ.write_text(json.dumps(db, ensure_ascii=False, indent=1), encoding='utf-8')
-    print('  已写入')
+    # 走本体 store 落盘 (2026-09-16): 原来直接 `OBJ.write_text(...)` —— 绕过了 store 的
+    # ① 全库 validate 闸 ② F12 原子写 (tmp + os.replace)。对象库里混进一个坏对象时, 直写会
+    # 让**整库**在下次读取时才炸, 而 store.save() 会当场报出是哪个对象不合法。
+    from src.ontology.store import Store
+    st = Store(OBJ)
+    st.objects = db
+    st.save()
+    print(f'  已写入 {_P.rel(OBJ)} (经 Store 闸)')
     return 0
 
 

+ 349 - 0
scripts/inventory_products.py

@@ -0,0 +1,349 @@
+#!/usr/bin/env python3
+# -*- coding: utf-8 -*-
+r"""产物清点与「输入 ↔ 产物」呼应校验 (2026-09-16 用户令 1、2)。
+
+用户令:
+  1. 基于观澜源码, 认真检查、梳理**所有产物**(产物类型/功用/输出路径)确保无遗漏; 输出路径不统一的改源码统一;
+     完善到 `<安装目录>\docs\系统设计说明.md`。
+  2. 确保 `data/raw` 下的输入数据**可用于重算**, 且**产物与输入数据呼应**。
+
+本脚本是三件事的**单一实现** (别在文档里手抄第二份, 会飘):
+  · 清点: 每个产物仓的件数/大小/类型, 以及**逐件来源**(raw-derived / shipped, 读 `_provenance.json`
+    与 `_derived_manifest.json`);
+  · 呼应: 对每类输入算它的**数据跨度/条数**, 与它喂出来的产物**逐项对拍** (产物跨度必须落在输入跨度内,
+    且输入更新时产物不该明显落后);
+  · 写档: `--write-doc` 把清单块写进 `docs\系统设计说明.md` 的
+    `<!-- INVENTORY:BEGIN --> … <!-- INVENTORY:END -->` 之间 (文档其余部分手工维护)。
+
+用法:
+    python scripts/inventory_products.py                # 打清单 + 呼应结论 (人看)
+    python scripts/inventory_products.py --check        # 只做呼应校验 (有问题 → 退出码 5)
+    python scripts/inventory_products.py --write-doc    # 把清单块写进 docs/系统设计说明.md
+    python scripts/inventory_products.py --json         # 机器可读
+"""
+from __future__ import annotations
+
+import argparse
+import datetime as dt
+import json
+import pathlib
+import re
+import sys
+from collections import OrderedDict
+
+ROOT = pathlib.Path(__file__).resolve().parents[1]
+sys.path.insert(0, str(ROOT))
+from src.console import soft  # noqa: E402
+from src import paths as P    # noqa: E402
+
+soft()
+DOC = P.DOCS / '系统设计说明.md'
+MARK_B, MARK_E = '<!-- INVENTORY:BEGIN', '<!-- INVENTORY:END -->'
+
+# ── 产物仓: 路径助手 → (仓名, 功用, 谁生成, 谁消费) ───────────────────────────────
+# ★路径一律用 src/paths.py 的助手 (唯一真源); 这里写的是**相对安装根**的形式, 便于跨平台显示。
+STORES = OrderedDict([
+    ('windscada',  dict(helper='P.store()',            purpose='L0 标准仓: SCADA/台账/派生分析的全部 parquet (页面主取数处)',
+                        producer='rebuild_from_raw.py (三门台账) + --scada (10 个构建器) + windscada_monthly_build.py',
+                        consumer='scripts/windscada_serve.py 各视图 · src/windscada/taxonomy · subsys/fusion')),
+    ('ontology',   dict(helper='P.ont()',              purpose='本体对象库: 码表/手册/工单展开/失效树 + 检索索引 + 实机参数',
+                        producer='python -m src.ontology.kb_ingest → populate → chain_ingest → trend_ingest → retrieval.build',
+                        consumer='脚本 windscada_serve.py 本体页/问答 · scripts/guanlan_facts_contract.py')),
+    ('windcms',    dict(helper='P.cms()',              purpose='CMS 振动诊断产物: 状态评估报告/逐台页/工作台页/知识库 + 厂家报告转录',
+                        producer='scripts/windcms.py report/kb · scripts/vib_reports_build.py',
+                        consumer='自服务 :18020 (每请求现读) · taxonomy.system_matrix (转录设备状态)')),
+    ('m5_cms_tcm', dict(helper='P.m5()',               purpose='振动线出件与窗级分析: handoff 接口 + 窗索引/谱库 + TCM 兼容件',
+                        producer='scripts/vib_raw_build.py (窗索引/谱) · 振动线出件 (handoff, 随包快照)',
+                        consumer='src/windscada/subsys/fusion.py · src/windcms/data.py · scripts/windcms.py report')),
+    ('tcm_compatible_replay', dict(helper='P.tcm_replay()', purpose='TCM 兼容链回放资产 (模型表/掩码阈值/裁决记录)',
+                        producer='随包快照 (无生成端)', consumer='src/windcms/report*.py · config.mask_thresholds')),
+    ('sop',        dict(helper='P.sop()',              purpose='SOP 中间件/评审/台账与事实契约底稿',
+                        producer='随包快照 (无生成端)', consumer='scripts/guanlan_facts_contract.py · 门户结论段')),
+    ('guanlan',    dict(helper='P.guanlan()',          purpose='事实契约与对外派生 (可上云面孔)',
+                        producer='scripts/guanlan_facts_contract.py', consumer='门户 #findings · /api/facts')),
+    ('pitch',      dict(helper='P.pitch()',            purpose='变桨侧派生件 (零位/日粒度)',
+                        producer='随包快照 + rebuild_from_raw --scada', consumer='脚本 windscada_serve.py 变桨面')),
+    ('paradigm_r1', dict(helper='(无助手: P.out_root()/paradigm_r1)', purpose='范式实验件 (E3/E5/E8 底稿, 事实契约输入)',
+                        producer='随包快照 (无生成端)', consumer='scripts/guanlan_facts_contract.py')),
+])
+
+# ── 输入类 → (放什么, 喂哪些产物, 跨度判据) ────────────────────────────────────────
+# span: 'csv_col' = 逐台 CSV 的首/末数据行; 'fname' = 从文件名年份/日期推; 'dir' = 目录名 (年/月)
+INPUTS = OrderedDict([
+    ('scada_10min', dict(what='SCADA 10min 导出 (逐台 WTG01.csv…WTG38.csv)', span='csv_col', time_col='Time',
+                         feeds=['windscada/powercurve*.parquet', 'windscada/loss_monthly.parquet', 'windscada/temp_monthly.parquet',
+                                'windscada/temp_bins.parquet', 'windscada/stop_events.parquet', 'windscada/yaw_daily.parquet',
+                                'windscada/hydraulic_accum.parquet', 'windscada/thermal_chain.parquet',
+                                'windscada/curve_*.parquet', 'windscada/control_*.parquet', 'windscada/system_aux.parquet'])),
+    ('scada_1min', dict(what='1min 导出 (逐台)', span='csv_col', time_col='occur_time',
+                        feeds=['windscada/watch_channels_monthly.parquet (watch 月度)'])),
+    ('故障报警',   dict(what='报警事件导出 (SpreadsheetML *.xls)', span='fname_year',
+                        feeds=['windscada/alarms.parquet'])),
+    ('风机故障记录', dict(what='检修工单台账 (*.xls/xlsx)', span='fname_year',
+                        feeds=['windscada/workorders.parquet', 'windscada/mblub_monthly.parquet'])),
+    ('油样报告',   dict(what='油液化验报告 (*.pdf)', span='fname_date',
+                        feeds=['windscada/oil_samples_index.parquet'])),
+    ('windcms',    dict(what='CMS 原始测量导出 (Brande TCM *_decode.json) + 厂家评估报告', span='dir',
+                        feeds=['m5_cms_tcm/windows/<窗>/{index.parquet,spectra/*}', 'windcms/厂家报告提取_*.json'])),
+    ('m5_cms_tcm', dict(what='振动线出件与 TCM 侧报告', span='none',
+                        feeds=['m5_cms_tcm/handoff_vibration_v2.json (现场正本优先)', 'windcms/报告_TCM传动链振动分析_*.md'])),
+])
+# 机理层 (不在场站目录下)
+TECH_INPUT = ('西门子4.0技术资料', '厂商技术资料 (故障处理手册/维护 WI/图纸/对译表)',
+              ['ontology/objects.json', 'ontology/turbine_params.parquet', 'ontology/retrieval_index.json'])
+
+
+def _rows_of(path: pathlib.Path, probe: int = 1):
+    """只读首/末若干行, 不整表载入 (scada csv 单文件几百 MB)。"""
+    try:
+        with open(path, 'rb') as f:
+            head = f.readline()
+            f.seek(max(0, path.stat().st_size - 65536))
+            tail = f.read().splitlines()
+        return head, (tail[-1] if tail else b'')
+    except Exception:
+        return b'', b''
+
+
+def input_span(station: pathlib.Path, sub: str, spec: dict):
+    """→ (起, 止, 件数, 说明, 粒度)。粒度 ∈ {'日','月','年'}: 只有'日/月'才参与"产物落后"判定 ——
+    按文件名年份推出来的跨度天然是**年粒度**, 拿它当"输入到 2026-12"会造出假缺口 (本脚本第一版就这么
+    误报过 2 条: 报警/工单"落后 5 个月"), 故显式带粒度、按粒度决定能不能比。"""
+    d = station / sub
+    if not d.is_dir():
+        return None, None, 0, '目录不存在', '年'
+    files = [p for p in d.rglob('*') if p.is_file()]
+    kind = spec.get('span')
+    if kind == 'csv_col':
+        col = spec.get('time_col')
+        starts, ends = [], []
+        for p in files[:60]:
+            try:
+                with open(p, 'rb') as f:
+                    first = f.readline()          # 表头
+                    first = f.readline()          # 第一条数据行
+                    f.seek(max(0, p.stat().st_size - 65536))
+                    last = f.read().splitlines()[-1]
+                hdr = first.decode('utf-8', 'replace').strip().split(',')
+                i = hdr.index(col) if col in hdr else 0
+                starts.append(hdr[i])
+                ends.append(last.decode('utf-8', 'replace').split(',')[0])
+            except Exception:
+                continue
+        if starts and ends:
+            return min(starts), max(ends), len(files), f'{len(files)} 件 (抽样 {len(starts)} 件首/末数据行)', '日'
+        return None, None, len(files), 'CSV 首/末行解析失败', '年'
+    if kind == 'fname_year':
+        ys = sorted({int(m.group()) for p in files for m in [re.search(r'(20\d\d)', p.name)] if m})
+        return (f'{ys[0]}-01' if ys else None), (f'{ys[-1]}-12' if ys else None), len(files), \
+               f'{len(files)} 件 (文件名年份, 年粒度)', '年'
+    if kind == 'fname_date':
+        ds = []
+        for p in files:
+            m = re.search(r'(\d{2})(\d{2})(20\d\d)', p.name)
+            if m:
+                ds.append(f'{m.group(3)}-{m.group(2)}-{m.group(1)}')
+        return (min(ds) if ds else None), (max(ds) if ds else None), len(files), \
+               f'{len(files)} 件 (文件名日期)', '日'
+    if kind == 'dir':
+        months = sorted({f'{m.group(1)}-{m.group(2)}' for p in d.rglob('*')
+                         for m in [re.search(r'[/\\](20\d\d)[/\\](\d{2})[/\\]', str(p))] if m})
+        return (months[0] if months else None), (months[-1] if months else None), len(files), \
+               f'{len(files)} 件 (目录年月 {", ".join(months) or "未识别"})', '月'
+    return None, None, len(files), '按件数清点 (无跨度判据)', '年'
+
+
+def product_stats(farm: str):
+    """→ {仓: dict(n, bytes, files: [(名, 类型, 大小)], kinds)}。"""
+    out = {}
+    for name in STORES:
+        d = P.out_root(farm) / name
+        if not d.is_dir():
+            out[name] = dict(n=0, bytes=0, kinds={}, newest=None)
+            continue
+        files = [p for p in d.rglob('*') if p.is_file()]
+        kinds = {}
+        for p in files:
+            k = p.suffix.lower() or '(无扩展名)'
+            kinds[k] = kinds.get(k, 0) + 1
+        newest = max((p.stat().st_mtime for p in files), default=None)
+        out[name] = dict(n=len(files), bytes=sum(p.stat().st_size for p in files), kinds=kinds,
+                         newest=(dt.datetime.fromtimestamp(newest).strftime('%Y-%m-%d %H:%M') if newest else None))
+    return out
+
+
+def provenance(farm: str):
+    root = P.out_root(farm)
+    prov = {}
+    for fn in ('_provenance.json', '_derived_manifest.json'):
+        f = root / fn
+        if not f.exists():
+            continue
+        try:
+            d = json.loads(f.read_text(encoding='utf-8'))
+        except Exception:
+            continue
+        for rel, v in (d.get('files') or {}).items():
+            src = v.get('source') if isinstance(v, dict) else None
+            prov[rel] = dict(source=src or 'raw-derived',
+                             builder=(v.get('builder') if isinstance(v, dict) else str(v)) or '')
+    return prov
+
+
+def span_check(farm: str, station: pathlib.Path, verbose=True):
+    """逐类输入: 算输入跨度 + 对应产物的跨度/条数, 报联动结论。→ (rows, problems)"""
+    import pandas as pd
+    rows, problems = [], []
+    ST = P.store(farm)
+
+    def pspan(p: pathlib.Path, col=None, path_cols=()):
+        if p is None or not p.exists():
+            return None, None, 0
+        try:
+            if p.suffix == '.parquet':
+                d = pd.read_parquet(p)
+                n = len(d)
+                c = col if col in d.columns else next((c for c in path_cols if c in d.columns), None)
+                if c is None:
+                    return None, None, n
+                s = pd.to_datetime(d[c], errors='coerce')
+                return ((str(s.min())[:10] if s.notna().any() else None),
+                        (str(s.max())[:10] if s.notna().any() else None), n)
+        except Exception:
+            return None, None, 0
+        return None, None, 0
+
+    # 输入类 → 它喂出来的产物 (路径, 时间列)
+    PAIRS = {
+        'scada_10min': [('windscada/temp_monthly.parquet', ST / 'temp_monthly.parquet', 'month'),
+                        ('windscada/loss_monthly.parquet', ST / 'loss_monthly.parquet', 'month'),
+                        ('windscada/powercurve_bins.parquet', ST / 'powercurve_bins.parquet', None)],
+        '故障报警': [('windscada/alarms.parquet', ST / 'alarms.parquet', 't_on')],
+        '风机故障记录': [('windscada/workorders.parquet', ST / 'workorders.parquet', 't_report')],
+        '油样报告': [('windscada/oil_samples_index.parquet', ST / 'oil_samples_index.parquet', 'date')],
+    }
+    for sub, spec in INPUTS.items():
+        if sub == 'm5_cms_tcm':
+            continue                                     # 出件不是"测量输入", 无跨度判据
+        i_s, i_e, i_n, note, gran = input_span(station, sub, spec)
+        comparable = gran in ('日', '月')                 # 年粒度不参与"落后"判定 (见 input_span 注释)
+        if sub == 'windcms':
+            # ★glob 必须写 `w[0-9][0-9][0-9][0-9]`: `w[????]` 是"字符类里含 ? 之一", 匹配不到任何窗名
+            #  (本脚本第一版就这么写, windcms 那行因此整行不出现 —— 静默漏项, 不是报错)
+            wins = [w for w in sorted((P.m5(farm) / 'windows').glob('w[0-9][0-9][0-9][0-9]'))
+                    if (w / 'index.parquet').exists()]
+            if wins:
+                w = wins[-1]
+                p_s, p_e, p_n = pspan(w / 'index.parquet', 'trigger_time')
+                rows.append(dict(input=sub, note=note, in_span=f'{i_s} ~ {i_e}', files=i_n,
+                                 product=f'm5_cms_tcm/windows/{w.name}/index.parquet',
+                                 prod_span=f'{p_s} ~ {p_e}', rows_n=p_n))
+                if p_s and i_s and p_s[:7] < i_s[:7]:
+                    problems.append(f'{sub}: 窗 {w.name} 索引起点 {p_s} 早于输入最早月 {i_s} (跨窗混入?)')
+            continue
+        for label, path, col in PAIRS.get(sub, []):
+            p_s, p_e, p_n = pspan(path, col)
+            rows.append(dict(input=sub, note=note, in_span=f'{i_s} ~ {i_e}', files=i_n, gran=gran,
+                             product=label, prod_span=f'{p_s} ~ {p_e}', rows_n=p_n))
+            if comparable and i_e and p_e and _months_between(p_e[:7], i_e[:7]) >= 2:
+                problems.append(f'{sub}: 产物 {label} 只到 {p_e}, 输入到 {i_e} '
+                                f'(落后 {_months_between(p_e[:7], i_e[:7])} 个月) → 未重算 或 该产物无生成端')
+            elif not comparable and p_e:
+                pass                                     # 年粒度: 跨度只作参考, 不判落后 (避免假缺口)
+    if verbose:
+        for r in rows:
+            print(f"  {r['input']:12s} 输入 {r['in_span']:26s} {r['files']:6d} 件 | "
+                  f"{r['product']:52s} {r['rows_n']:8d} 行 {r['prod_span']}")
+    return rows, problems
+
+
+def _months_between(a: str, b: str) -> int:
+    try:
+        ya, ma = int(a[:4]), int(a[5:7])
+        yb, mb = int(b[:4]), int(b[5:7])
+        return (yb - ya) * 12 + (mb - ma)
+    except Exception:
+        return 0
+
+
+def doc_block(farm: str) -> str:
+    """生成写进 docs/系统设计说明.md 的清单块 (Markdown)。"""
+    ps = product_stats(farm)
+    prov = provenance(farm)
+    from collections import Counter
+    by_store = Counter()
+    for rel, v in prov.items():
+        by_store[(rel.split('/')[0], v['source'])] += 1
+    L = [f'<!-- INVENTORY:BEGIN (由 scripts/inventory_products.py --write-doc 生成, 勿手改) -->',
+         f'*自动生成于 {dt.datetime.now():%Y-%m-%d %H:%M};数据源: `outputs/{farm}/_provenance.json` + `_derived_manifest.json` + 实际文件*', '',
+         '| 产物仓 | 输出路径 (相对安装根) | 功用 | 生成端 | 消费端 | 件数 | 大小 | 来源(raw 重算/随包) |',
+         '|---|---|---|---|---|---|---|---|']
+    for name, meta in STORES.items():
+        st = ps[name]
+        rd = by_store.get((name, 'raw-derived'), 0)
+        sh = by_store.get((name, 'shipped'), 0)
+        path = f'outputs/{farm}/{name}/'
+        L.append(f"| `{name}` | `{path}` | {meta['purpose']} | {meta['producer']} | {meta['consumer']} | "
+                 f"{st['n']} | {st['bytes'] / 1048576:.1f} MB | {rd} / {sh} |")
+    L += ['', f'合计 {sum(s["n"] for s in ps.values())} 件 / {sum(s["bytes"] for s in ps.values()) / 1048576:.1f} MB;'
+              f'其中 raw 重算 {sum(v for (_, k), v in by_store.items() if k == "raw-derived")} 件、'
+              f'随包补齐 {sum(v for (_, k), v in by_store.items() if k == "shipped")} 件。', MARK_E]
+    return '\n'.join(L)
+
+
+def write_doc(farm: str) -> int:
+    block = doc_block(farm)
+    txt = DOC.read_text(encoding='utf-8') if DOC.exists() else ''
+    if MARK_B in txt and MARK_E in txt:
+        head = txt.split(MARK_B)[0]
+        tail = txt.split(MARK_E)[1]
+        new = head + block + tail
+    else:
+        new = (txt + '\n\n## 附: 产物清单 (自动生成)\n\n' + block + '\n') if txt else block + '\n'
+    DOC.parent.mkdir(parents=True, exist_ok=True)
+    DOC.write_text(new, encoding='utf-8')
+    print(f'已写产物清单 → {P.rel(DOC)}')
+    return 0
+
+
+def main() -> int:
+    ap = argparse.ArgumentParser()
+    ap.add_argument('--farm', default=os_farm())
+    ap.add_argument('--check', action='store_true', help='只做呼应校验 (有问题退出码 5)')
+    ap.add_argument('--write-doc', action='store_true', help='把清单块写进 docs/系统设计说明.md')
+    ap.add_argument('--json', action='store_true')
+    a = ap.parse_args()
+    from src.windscada.config import farm as _farm, raw_station_dir
+    station = pathlib.Path(raw_station_dir(a.farm))
+    ps = product_stats(a.farm)
+
+    if a.write_doc:
+        return write_doc(a.farm)
+
+    print(f'=== 产物仓 ({len(STORES)} 个) @ outputs/{a.farm}/ ===')
+    for name, meta in STORES.items():
+        st = ps[name]
+        kinds = ' '.join(f'{k}×{v}' for k, v in sorted(st['kinds'].items(), key=lambda kv: -kv[1])[:5])
+        print(f"  {name:24s} {st['n']:6d} 件 {st['bytes'] / 1048576:9.1f} MB  {meta['helper']:28s} {kinds}")
+        print(f"     功用: {meta['purpose']}")
+
+    print(f'\n=== 输入 ↔ 产物 呼应 (输入根 {P.rel(station)}) ===')
+    rows, problems = span_check(a.farm, station)
+    if not rows:
+        print('  (无可对拍项)')
+    print(f'\n=== 结论: {"呼应正常" if not problems else str(len(problems)) + " 项要处理"} ===')
+    for p in problems:
+        print(f'  [!] {p}')
+    if a.json:
+        print(json.dumps(dict(products={k: {kk: vv for kk, vv in v.items() if kk != "kinds"} for k, v in ps.items()},
+                              spans=rows, problems=problems), ensure_ascii=False, indent=1))
+    return 5 if (a.check and problems) else 0
+
+
+def os_farm() -> str:
+    import os
+    return os.environ.get('WINDSCADA_FARM') or 'rudong'
+
+
+if __name__ == '__main__':
+    sys.exit(main())

+ 61 - 21
scripts/pack_dist.py

@@ -12,25 +12,33 @@ r"""打一个"拷到别的电脑能装、能跑"的分发包 (2026-09-12)。
 ## 包含 / 排除 (规则即文档)
 
 **包含**: 程序(`src/` `scripts/` `guanlan.py`) · 配置(`configs/`) · 离线安装件(`wheels/win_amd64/` `vendor/python/`)
-· 页面与交付(`release/` `resources/` `reference/`) · **产物**(`outputs/`, 页面取数靠它) · 文档(`docs/` `README_先读我.txt`
-`测试须知.txt` `_修复记录_20260911/`) · 安装/起停脚本(`install.bat` `install.sh` `install.ps1` `check.bat` `start.bat`
-`stop.bat`) · `requirements.txt`。
+· 页面与交付(`release/` `resources/` `reference/`) · **产物**(`outputs/`, 页面取数靠它; 加 `--no-products`
+则不含) · 文档(`docs/` `README_先读我.txt` `测试须知.txt`) · 安装/起停脚本(`install.bat` `install.sh`
+`install.ps1` `check.bat` `start.bat` `stop.bat`) · `requirements.txt`。
+
+**不含**修复记录/临时目录(`_修复记录_*` 之类): 本包已纳入 git 管理, 变更历史由版本库承载 (2026-09-12 用户令)。
 
 **排除**(每条都写了理由):
     .git/ .venv/ .github/            版本库与虚拟环境 —— venv 换机必失效, 目标机安装时重建
-    _products_off*/                  本机"清除产物"的暂存与存档(每份含整份产物); 目标机用不到
+    _products_off*/                  历史遗留的"清除产物"暂存档(2026-09-16 起清除=真删, 不再产生此类目录)
     logs/ run/                       本机日志与 run/pids.json(旧 PID, 到新机器上是无效引用)
-    data/raw/                        31 GB 现场原始件(约定"原始件不随包分发"); 要一起交付用 --with-data
+    data/raw/                        现场原始件(约定"原始件不随包分发"); 要一起交付用 --with-data
     __pycache__/ *.pyc               解释器缓存
     *.zip(顶层)                      旧的交付压缩包; 避免包中包
 
 ## 用法
 
     python scripts/pack_dist.py --dry-run                  # 只报会打什么/多大, 不写文件
-    python scripts/pack_dist.py                            # 打包 → <父目录>/guanlan-rudong-v2_0.2.0_dist_<日期>.zip
+    python scripts/pack_dist.py                            # 打包 → <父目录>/guanlan-v<VERSION>_dist_<日期>.zip
+    python scripts/pack_dist.py --no-products              # **不含产物**(用户令 2026-09-16): 开箱页面显示"无产物"
     python scripts/pack_dist.py --with-data                # 连 data/raw 一起打 (30 GB, 慎用)
     python scripts/pack_dist.py --out D:\out\x.zip         # 指定输出
     python scripts/pack_dist.py --verify <zip>             # **开箱验证**: 解压 → 离线安装 → 起服务核验页面 → 清理
+
+## 排除结果写进包
+
+包里 `dist-manifest.json` 记下 version / no_data / no_products / no_logs 与逐条排除理由;
+`guanlan.py check` 读它 —— 打了"不含产物"的包, 产物那几行检查按 `[--]` 说明而不判 FAIL。
 """
 from __future__ import annotations
 
@@ -48,8 +56,13 @@ import zipfile
 ROOT = pathlib.Path(__file__).resolve().parents[1]
 sys.path.insert(0, str(ROOT))
 
-INCLUDE_DIRS = ['src', 'scripts', 'configs', 'release', 'resources', 'reference', 'outputs', 'docs',
-                'wheels', 'vendor', '_修复记录_20260911']
+INCLUDE_DIRS = ['src', 'scripts', 'configs', 'release', 'resources', 'reference', 'docs',
+                'wheels', 'vendor']
+PRODUCTS_DIR = 'outputs'          # 产物仓: 默认打进包 (页面开箱有数); --no-products 时排除
+# ★ 不打包"修复记录/临时/脚手架"类目录 (2026-09-12 用户令): 本包已纳入 git 管理, 变更历史由版本库承载,
+#   不必在交付物里再带一份 `_修复记录_<日期>\`。原先这里含 `_修复记录_20260911` (该目录的**内容**已并入
+#   `docs\振动数据接入_v0.1.md` 与 `docs\数据目录结构与落位约定_v0.2.md`, 目录本身按用户令删除)。
+VERSION = '0.4.0'                 # 包版本 (2026-09-16): 写进 dist-manifest.json 与默认文件名
 INCLUDE_FILES = ['guanlan.py', 'install.bat', 'install.ps1', 'install.sh', 'check.bat', 'start.bat', 'stop.bat',
                  'requirements.txt', 'README_先读我.txt', '测试须知.txt']
 EXCLUDE_DIRS = {'.git', '.venv', '.github', 'logs', 'run', '__pycache__'}
@@ -69,13 +82,25 @@ def _xfer():
     return m
 
 
-def plan(with_data=False):
-    """→ (要打包的顶层条目 list, 排除说明 list)"""
+def plan(with_data=False, no_products=False):
+    """→ (要打包的顶层条目 list, 排除说明 list)
+
+    2026-09-16 用户令: 打包**不含 输入数据 / 产物 / 日志**。
+      · 输入数据 `data/`   —— 一直默认排除 (要用 --with-data 才带);
+      · 产物     `outputs/` —— 新增 --no-products 排除 (默认仍带: 页面开箱即有数);
+      · 日志     `logs/`    —— 一直在 EXCLUDE_DIRS 里 (连同 run/ 旧 PID)。
+    排除产物时, 包内 `dist-manifest.json` 会记 `no_products: true`,
+    目标机 `guanlan.py check` 据此把"产物缺失"显示为 `--`(待重算) 而不是 FAIL。
+    """
     inc = []
     for d in INCLUDE_DIRS:
         p = ROOT / d
         if p.is_dir():
             inc.append(p)
+    if not no_products:
+        p = ROOT / PRODUCTS_DIR
+        if p.is_dir():
+            inc.append(p)
     if with_data:
         inc.append(ROOT / 'data')
     for f in INCLUDE_FILES:
@@ -83,8 +108,15 @@ def plan(with_data=False):
         if p.is_file():
             inc.append(p)
     skipped = [('.git/', '版本库'), ('.venv/', '虚拟环境(换机必失效, 目标机安装时重建)')]
-    skipped += [(g + '/', '本机痕迹/超大体量, 见脚本头部说明') for g in EXCLUDE_GLOBS]
-    skipped += [('logs/, run/', '本机日志与旧 PID')]
+    _glob_why = {
+        '_products_off*': '历史遗留的"清除产物"暂存档 (2026-09-16 起清除=真删, 不再产生; 老机器上若有可手工删)',
+        'data': '现场原始输入件 (约定"原始件不随包分发"; 要一起交付用 --with-data) —— 用户令: 输入数据不进包',
+    }
+    skipped += [(g + '/', _glob_why.get(g, '本机痕迹/超大体量, 见脚本头部说明')) for g in EXCLUDE_GLOBS]
+    skipped += [('logs/, run/', '本机日志与旧 PID —— 用户令: 日志不进包')]
+    if no_products and (ROOT / PRODUCTS_DIR).is_dir():
+        skipped += [('outputs/', '用户令: 产物不进包 (目标机放数据后 rebuild_all.py 重算; '
+                                 '或在目标机也用 --no-products 的同一口径)')]
     return inc, skipped
 
 
@@ -98,8 +130,8 @@ def size_of(p: pathlib.Path):
     return n, s
 
 
-def build(out: pathlib.Path, with_data=False, progress=True) -> int:
-    inc, skipped = plan(with_data)
+def build(out: pathlib.Path, with_data=False, progress=True, no_products=False) -> int:
+    inc, skipped = plan(with_data, no_products)
     tot_n = tot_s = 0
     print('== 要打进包的内容 ==')
     for p in inc:
@@ -112,10 +144,15 @@ def build(out: pathlib.Path, with_data=False, progress=True) -> int:
         print(f'   {name:26s} {why}')
 
     out.parent.mkdir(parents=True, exist_ok=True)
-    man = dict(built=time.strftime('%Y-%m-%d %H:%M:%S'), root=str(ROOT), includes=[p.relative_to(ROOT).as_posix() for p in inc],
+    man = dict(version=VERSION, built=time.strftime('%Y-%m-%d %H:%M:%S'), root=str(ROOT),
+               includes=[p.relative_to(ROOT).as_posix() for p in inc],
                excluded={n: w for n, w in skipped}, files=tot_n, bytes=tot_s,
-               note='本包由 scripts/pack_dist.py 生成; 目标机解压后跑 install.bat(Windows) 或 sh install.sh(Linux/macOS), '
-                    '再 check → start。.venv 不在包内(换机必失效, 安装时重建)。')
+               no_data=not with_data, no_products=bool(no_products), no_logs=True,
+               note='本包由 scripts/pack_dist.py 生成; **不含 输入数据(data/) · 产物(outputs/) · 日志(logs/, run/)** '
+                    '(2026-09-16 用户令)。目标机解压后跑 install.bat(Windows) 或 sh install.sh(Linux/macOS), '
+                    '再 check → start。要让页面有数: 把现场包放好后跑 '
+                    '`scripts/place_raw_data.py --src <现场包目录> --scope full` + `scripts/rebuild_all.py`。'
+                    '.venv 不在包内(换机必失效, 安装时重建)。')
     t0 = time.time()
     with zipfile.ZipFile(out, 'w', zipfile.ZIP_DEFLATED, compresslevel=6) as z:
         z.writestr('dist-manifest.json', json.dumps(man, ensure_ascii=False, indent=1))
@@ -175,17 +212,20 @@ def main() -> int:
     ap = argparse.ArgumentParser()
     ap.add_argument('--out', default=None)
     ap.add_argument('--with-data', action='store_true', help='连 data/raw 一起打(约 31 GB, 慎用)')
+    ap.add_argument('--no-products', action='store_true',
+                    help='**不含产物** (outputs/): 用户令 2026-09-16 "打包不含 输入数据/产物/日志"')
     ap.add_argument('--dry-run', action='store_true')
     ap.add_argument('--verify', default=None, help='对已打好的 zip 做开箱验证(解压→安装→起服务核验)')
     ap.add_argument('--keep', action='store_true', help='--verify 后保留解压目录')
     a = ap.parse_args()
     if a.verify:
         return verify(pathlib.Path(a.verify), a.keep)
-    out = pathlib.Path(a.out) if a.out else ROOT.parent / f'guanlan-rudong-v2_0.2.0_dist_{time.strftime("%Y%m%d")}.zip'
+    out = (pathlib.Path(a.out) if a.out
+           else ROOT.parent / f'guanlan-v{VERSION}_dist_{time.strftime("%Y%m%d")}.zip')
     if a.dry_run:
-        inc, skipped = plan(a.with_data)
+        inc, skipped = plan(a.with_data, a.no_products)
         tot_n = tot_s = 0
-        print('== 会打进包 ==')
+        print(f'== 会打进包 (version {VERSION}) ==')
         for p in inc:
             n, s = size_of(p); tot_n += n; tot_s += s
             print(f'   {p.relative_to(ROOT).as_posix():26s} {n:6d} 件  {s/1e6:9.1f} MB')
@@ -195,7 +235,7 @@ def main() -> int:
             print(f'   {name:26s} {why}')
         print('\n(dry-run, 未写文件)')
         return 0
-    return build(out, a.with_data)
+    return build(out, a.with_data, no_products=a.no_products)
 
 
 if __name__ == '__main__':

+ 133 - 33
scripts/place_raw_data.py

@@ -45,20 +45,55 @@ A2 定了四项数据源的落位: data/raw/<场站名称>/{scada_10min, 故障
             `src_farm_names=['如海','如东']` 同源(如海=如东项目)。**包内没有消费者**: 它补的是
             数据层完整性, 不改变任何页面数值 —— 落它是因为"扫描辨识"会把它报成缺口。
 
+## 振动侧 (--scope vib, 2026-09-12 用户令)
+
+用户令: 「数据层里的「CMS 振动评估报告」应遵循 `<安装目录>\\data\\raw\\如东\\windcms`、
+「振动线 handoff」应遵循 `<安装目录>\\data\\raw\\如东\\m5_cms_tcm` 存放; 修改系统支持振动数据参与
+运行、重算; 自现场包提取相应振动数据存放至上述目录」。
+
+  <场站>/windcms/CMS_RuDong_CGN_202603-04/measurement/<年>/<月>/<WTGxx>/*_decode.json
+      ← CMS_RuDong_CGN_202603-04.zip 的 measurement/** 全部成员 (25,679 件, 解压 153.7 GB)
+      依据: 这就是 `src/windcms/pipeline.py` 的 `detect_input()` 认的 **`tcm_decoded_json`**
+            (Brande TCM Enterprise 导出, SiteName=CGN Rudong, LocationName=WTGxx)。原样保留
+            包内 `measurement/` 这一层, 于是"哪个包来的哪一层"仍可追溯; 摄入按 rglob 找
+            `*_decode.json`, 套不套这层都能吃。**这是振动侧唯一能重算的源**: 六层链全部
+            输入都是它 (索引/谱 → 扫线 → 能量占比 → 模型 → 融合 → 报告)。
+            ★ 2026-09-11 版的 SKIPPED 里写着"振动分析报告是成品牌报告, windcms 要的测点索引
+              包里没有" —— 那句在当时成立; 2026-09-12 用户把 CMS 原始导出补进现场包后不再成立,
+              故本范围落地, 并在下面改注。
+
+  <场站>/windcms/厂家报告/上海电气_月度/     ← 如东风场数据.zip:
+      `数据收集/更新/(8)风机振动数据、油液分析记录(2025.1-至今)/振动分析报告/**`
+      (12 份月度用印版 PDF, 2025-06…2026-05) + 现场目录散装件
+      `中广核如东海上风电场2026年07月振动分析报告_上海电气.docx`
+      依据: 数据层「CMS 振动评估报告」的**厂商评估报告**侧 (VDI3834 / NB/T 31129-2018 判级)。
+            12 份 PDF 是纯扫描件 (无文本层, 实测 /Font=0、pdftotext 类抽取为空) → 只作归档证据,
+            数值不参与判级; docx 有文本层, 由 scripts/vib_reports_build.py 摄入。
+
+  <场站>/m5_cms_tcm/厂家报告/                ← 现场目录散装件
+      `中广核如东海上风电场传动链振动分析报告_大生科技_20260311(1).docx`
+      依据: 数据层「振动线 handoff」的**TCM 侧深度分析报告** (依托机组自带 TCM M-system 数据)。
+            若干现场给了 handoff 正本 (`handoff_vibration_v2.json` / `component_history.json`),
+            也放本目录 → 摄入优先采用现场正本, 不再用包内 shipped 快照。
+
 用法:
     python scripts/place_raw_data.py --dry-run                # 只报要落什么, 不写盘 (默认 a2)
     python scripts/place_raw_data.py                          # 真落位 (已存在则覆盖)
     python scripts/place_raw_data.py --scope mech --dry-run    # 看机理层/1min 会落什么
-    python scripts/place_raw_data.py --scope full              # 两组一起
+    python scripts/place_raw_data.py --scope vib               # 落振动侧 (CMS 原始导出 153.7 GB + 厂家报告)
+    python scripts/place_raw_data.py --scope vib --limit 300   # 冒烟: 只落 300 件 (试跑/校验用)
+    python scripts/place_raw_data.py --scope full              # 三组一起
     python scripts/place_raw_data.py --src D:\\别的现场数据目录
 
-## 故意不做的事 (两组范围共有的判断)
+## 故意不做的事 (各范围共有的判断)
 
   · 不落 `(3)现场检修记录/2025年检修记录` 与 `2026年检修记录`: 与 工作/风机故障记录/2025年故障记录、
     2026年故障记录 **逐件同名同大小**(19/19 件), 是同一批月度汇总表的副本。落两遍会让"按年目录"
     摄入看到两份重复台账 (2026-08-31 长停台账虚高 35 倍那类事故的同款成因: 快照重复必须归并, 不能叠加)。
-  · 不落 `(8)…/振动分析报告/`(12 份月度用印版 PDF): 是振动线 (CMS) 的**成品牌报告**, 不是可再加工的
-    测量数据; 而 `windcms`/`m5_cms_tcm` 要的是 CMS 测点索引(handoff), 包里没有。
+  · 振动报告只从 `(8)…/振动分析报告/` 落一遍 (不在 a2/mech 范围): 包内另有两处**同件副本** ——
+    根目录 `中广核如东海上风电场2026年5月振动分析报告用印版(2).pdf`(与 2026年05月 件同尺寸)
+    与嵌套包 `如东海上振动报告11份.zip`(其 11 份与 `振动分析报告/` 的 11 份逐件同尺寸)。
+    落三遍会得到 3 份同名报告, 摄入时会互相覆盖或重复计数 —— 故只在 vib 范围落目录里那一份。
   · 不落 `scada数据(如东)/**`(19 个月 × 12 个通道组 zip, 2.4 GB): 它是 `scada_10min/*.csv` 的**上游**
     原始通道导出, 包内没有任何脚本读它(构建器读的是已经平铺好的 10min CSV) —— 落了也不参与重算。
   · 不落 `fastlog数据/`(4 件 WTG0x.xls): 全库搜 `fastlog` 只有 2 处**注释**提到它("运行态见证"),
@@ -105,6 +140,27 @@ RULES_MECH = [
     ('1分钟数据.zip', '', 'scada_1min', True, 'station'),
 ]
 
+# 振动侧 (--scope vib): 数据层两条振动行的源件 —— CMS 原始测量导出 + 厂商评估报告
+ZIP_CMS = 'CMS_RuDong_CGN_202603-04.zip'
+DIR_SA_REPORTS = '如东风场数据/数据收集/更新/(8)风机振动数据、油液分析记录(2025.1-至今)/振动分析报告/'
+RULES_VIB = [
+    # ① CMS 原始测量导出 (Brande TCM Enterprise, SiteName=CGN Rudong): 25,679 件 *_decode.json。
+    #    strip=False 保留包内 measurement/ 层 → <包名>/measurement/2026/03/WTGxx/*.json, 来源可追溯。
+    (ZIP_CMS, 'measurement/', f'windcms/{ZIP_CMS[:-4]}', False, 'station'),
+    # ② 上海电气月度振动分析报告 (12 份扫描件 PDF): 数据层「CMS 振动评估报告」的厂商报告侧
+    ('如东风场数据.zip', DIR_SA_REPORTS, 'windcms/厂家报告/上海电气_月度', True, 'station'),
+]
+
+# 现场目录下的散装件 (不是压缩包成员): (源文件名, 目标相对 data/raw/<场站>/ 的路径)
+RULES_VIB_LOOSE = [
+    # 上海电气 2026年07月报告: 有文本层, 由 scripts/vib_reports_build.py 摄入成评估报告
+    ('中广核如东海上风电场2026年07月振动分析报告_上海电气.docx',
+     'windcms/厂家报告/上海电气_月度'),
+    # 大生科技传动链振动分析报告 (TCM M-system 数据): 数据层「振动线 handoff」的 TCM 侧
+    ('中广核如东海上风电场传动链振动分析报告_大生科技_20260311(1).docx',
+     'm5_cms_tcm/厂家报告'),
+]
+
 # 说明性的"故意不落", 只在报告里列出来
 SKIPPED = [
     ('如东风场数据.zip', '如东风场数据/数据收集/更新/(3)现场检修记录(2025.1-至今)/2025年检修记录/',
@@ -112,7 +168,13 @@ SKIPPED = [
     ('如东风场数据.zip', '如东风场数据/数据收集/更新/(3)现场检修记录(2025.1-至今)/2026年检修记录/',
      '与 工作/风机故障记录/2026年故障记录 逐件同名校验相同 (副本)'),
     ('如东风场数据.zip', '如东风场数据/数据收集/更新/(8)风机振动数据、油液分析记录(2025.1-至今)/振动分析报告/',
-     '振动线成品牌报告, 不是可再加工的测量数据; windcms 要的测点索引包里没有'),
+     '※ 2026-09-12 起**已改为落位** (vib 范围, → windcms/厂家报告/上海电气_月度)。原判断"成品牌报告+'
+     '包内没有测点索引"在当时成立; 用户补入 CMS_RuDong_CGN_202603-04.zip 后, 索引源件已具备, '
+     '报告作为数据层证据一并落位。此处保留记录以免后人以为漏了'),
+    ('如东风场数据.zip', '如东风场数据/中广核如东海上风电场2026年5月振动分析报告用印版(2).pdf',
+     '与 振动分析报告/…2026年05月…用印版.pdf 同尺寸 (副本, 只落目录里那一份)'),
+    ('如东风场数据.zip', '如东风场数据/如东海上振动报告11份.zip',
+     '其 11 份与 振动分析报告/ 的 11 份逐件同尺寸 (副本包, 只落目录里那一份)'),
     ('如东风场数据.zip', '如东风场数据/scada数据(如东)/',
      'scada_10min/*.csv 的上游原始通道导出 (2.4 GB), 包内无脚本读它 → 不参与重算'),
     ('如东风场数据.zip', '如东风场数据/fastlog数据/',
@@ -140,8 +202,11 @@ def human(n: float) -> str:
     return f'{n:.1f} GB'
 
 
-def plan(src: pathlib.Path, station: pathlib.Path, rules):
-    """→ [(zip, 目标文件, 包内条目名, 解压后大小, 目标根, 目标名)]; 只读压缩包目录, 不解压。"""
+def plan(src: pathlib.Path, station: pathlib.Path, rules, loose_rules=(), limit: int = 0):
+    """→ [(zip, 目标文件, 包内条目名, 解压后大小, 目标根, 目标名)]; 只读压缩包目录, 不解压。
+
+    zip=None 表示散装件 (现场目录下的单个文件, 如两份 docx 振动报告不在任何压缩包里)。
+    limit>0 时只取前 limit 件 (冒烟/试跑用; 大包 2.5 万件跑一遍要几十分钟, 得能小样验证)。"""
     out = []
     for zname, prefix, target, strip, root in rules:
         zp = src / zname
@@ -165,6 +230,13 @@ def plan(src: pathlib.Path, station: pathlib.Path, rules):
                         continue
                     rest = prefix.rstrip('/').rsplit('/', 1)[-1]
                 out.append((zname, base / target / rest, name, info.file_size, root, target))
+    for fname, target in loose_rules:
+        fp = src / fname
+        if not fp.exists():
+            raise SystemExit(f'缺散装源件: {fp}')
+        out.append((None, station / target / fname, fname, fp.stat().st_size, 'station', target))
+    if limit:
+        out = out[:limit]
     return out
 
 
@@ -172,8 +244,10 @@ def main() -> int:
     ap = argparse.ArgumentParser()
     ap.add_argument('--src', default=str(DEFAULT_SRC), help='现场数据目录 (默认 %(default)s)')
     ap.add_argument('--dry-run', action='store_true', help='只报计划, 不写盘')
-    ap.add_argument('--scope', choices=('a2', 'mech', 'full'), default='a2',
-                    help='a2=只落四项数据层(A3 语义, 默认) · mech=机理层源件+1min · full=两组一起')
+    ap.add_argument('--scope', choices=('a2', 'mech', 'vib', 'full'), default='a2',
+                    help='a2=四项数据层(A3 语义, 默认) · mech=机理层源件+1min · '
+                         'vib=振动侧(CMS 原始导出+厂家报告) · full=三组一起')
+    ap.add_argument('--limit', type=int, default=0, help='只落前 N 件 (冒烟/试跑; 默认 0=全落)')
     a = ap.parse_args()
 
     src = pathlib.Path(a.src)
@@ -186,8 +260,13 @@ def main() -> int:
     print(f'场站目录: {station}   (来自 src/windscada/config.py raw_station={cfg.get("raw_station")!r})')
     print(f'现场数据: {src}   范围: --scope {a.scope}\n')
 
-    rules = {'a2': RULES, 'mech': RULES_MECH, 'full': RULES + RULES_MECH}[a.scope]
-    items = plan(src, station, rules)
+    rules, loose = {
+        'a2': (RULES, ()),
+        'mech': (RULES_MECH, ()),
+        'vib': (RULES_VIB, RULES_VIB_LOOSE),
+        'full': (RULES + RULES_MECH + RULES_VIB, RULES_VIB_LOOSE),
+    }[a.scope]
+    items = plan(src, station, rules, loose, limit=a.limit)
     by_target = {}
     for zname, dst, entry, size, root, target in items:
         d = by_target.setdefault((root, target), [0, 0])
@@ -196,12 +275,14 @@ def main() -> int:
     print(f'== 计划落位 (scope={a.scope}) ==')
     for (root, target), (n, sz) in sorted(by_target.items(), key=lambda kv: kv[0][1]):
         where = (station / target) if root == 'station' else (station.parent / target)
-        print(f'  {target:18s} {n:5d} 件  {human(sz):>10s}   → {where}')
+        print(f'  {target:44s} {n:6d} 件  {human(sz):>10s}   → {where}')
     print(f'  合计 {len(items)} 件, {human(sum(i[3] for i in items))}')
 
-    print('\n== 故意不落 ==')
-    for zname, prefix, why in SKIPPED:
-        print(f'  {prefix}\n      理由: {why}')
+    if a.scope in ('vib', 'full'):
+        print('\n== 故意不落 (振动侧同件副本) ==')
+        for zname, prefix, why in SKIPPED:
+            if '振动' in prefix or '报告11份' in prefix:
+                print(f'  {prefix}\n      理由: {why}')
 
     if a.dry_run:
         print('\n(dry-run, 未写盘)')
@@ -210,24 +291,43 @@ def main() -> int:
     print('\n== 落位 ==')
     done = 0
     skipped = 0
-    for zname, dst, entry, size, root, target in items:
-        # 已有同尺寸文件 = 已经是最新 → 跳过。重跑一次不该把 15 GB 原样再抄一遍
-        # (mech 范围 15.5 GB, a2 范围 14.7 GB; 2026-09-11 修单文件规则时就是靠这个避免整盘重写)。
-        if dst.exists() and dst.stat().st_size == size:
-            skipped += 1
-            continue
-        dst.parent.mkdir(parents=True, exist_ok=True)
-        with zipfile.ZipFile(src / zname) as zf:
-            info = next(i for i in zf.infolist() if gbk_name(i).replace('\\', '/') == entry)
-            with zf.open(info) as fsrc, open(dst, 'wb') as fdst:
-                shutil.copyfileobj(fsrc, fdst, 1024 * 1024 * 4)
-        got = dst.stat().st_size
-        if got != size:
-            raise SystemExit(f'写出大小不符: {dst} 期望 {size} 实得 {got}')
-        done += 1
-        if done % 10 == 0 or size > 50 * 1024 * 1024:
-            print(f'  [{done}/{len(items)}] {human(size):>10s}  {dst.relative_to(station.parent)}', flush=True)
-    print(f'\n完成 (scope={a.scope}): 新写/更新 {done} 件, 已是最新跳过 {skipped} 件, 共 {len(items)} 件')
+    written = 0
+    # 按包分组 + **每包只开一次**。旧写法在循环体里 with ZipFile(...) 且用 next(...) 线性找条目 ——
+    # 对 2.5 万件的大包是 25,679 次开包 × 25,759 次名字比较 ≈ 6.6 亿次比较, 实测不可接受。
+    # 改成: 每包一次性建 {条目名: ZipInfo} 映射, 顺序流式写出。
+    by_zip = {}
+    for it in items:
+        by_zip.setdefault(it[0], []).append(it)
+    for zname, group in by_zip.items():
+        zf = zipfile.ZipFile(src / zname) if zname else None
+        try:
+            imap = {gbk_name(i).replace('\\', '/'): i for i in zf.infolist()} if zf else {}
+            for _, dst, entry, size, root, target in group:
+                # 已有同尺寸文件 = 已经是最新 → 跳过。重跑一次不该把 150 GB 原样再抄一遍
+                # (vib 范围解压 153.7 GB; a2 14.7 GB / mech 15.5 GB; 2026-09-11 修单文件规则时
+                #  就是靠这个避免整盘重写)。
+                if dst.exists() and dst.stat().st_size == size:
+                    skipped += 1
+                    continue
+                dst.parent.mkdir(parents=True, exist_ok=True)
+                if zf is None:
+                    with open(src / entry, 'rb') as fsrc, open(dst, 'wb') as fdst:
+                        shutil.copyfileobj(fsrc, fdst, 1024 * 1024 * 4)
+                else:
+                    with zf.open(imap[entry]) as fsrc, open(dst, 'wb') as fdst:
+                        shutil.copyfileobj(fsrc, fdst, 1024 * 1024 * 4)
+                got = dst.stat().st_size
+                if got != size:
+                    raise SystemExit(f'写出大小不符: {dst} 期望 {size} 实得 {got}')
+                done += 1
+                written += size
+                # 大包进度: 每 500 件报一次 (2.5 万件逐件打印会把日志刷爆); 单件 >200 MB 也报
+                if done % 500 == 0 or size > 200 * 1024 * 1024:
+                    print(f'  [{done}/{len(items)}] 已写 {human(written)}  {dst.name}', flush=True)
+        finally:
+            if zf is not None:
+                zf.close()
+    print(f'\n完成 (scope={a.scope}): 新写/更新 {done} 件 ({human(written)}), 已是最新跳过 {skipped} 件, 共 {len(items)} 件')
     return 0
 
 

+ 37 - 0
scripts/portal_build.py

@@ -46,6 +46,7 @@ import pathlib
 import re
 import sys
 import tempfile
+import time
 
 ROOT = pathlib.Path(__file__).resolve().parents[1]
 sys.path.insert(0, str(ROOT))
@@ -227,6 +228,39 @@ def do_verify() -> int:
     return 3
 
 
+def do_rebaseline() -> int:
+    """把当前 shell.html / portal.html 记为新的基线 (manifest 里的 sha256)。
+
+    为什么需要它: manifest 的 sha256 是"漂移检测"的基准。**有意**改了门户外壳 (例如 2026-09-16
+    按用户令在菜单里加「数据重算」) 之后, `--check` 会如实报漂移 —— 这是对的, 但不能让人只有
+    "手改 JSON" 或 "跑 --extract 重新拆包" 两条路: 后者会用 portal.html 反过来覆盖 shell.html
+    (而装配过程会注入契约结论段/内嵌模板), 一不小心就把手工维护的外壳冲掉。
+    本命令只改 manifest 里的两个基准值, 并记下是谁、什么时候、为什么重基线。
+    """
+    if not MANIFEST.is_file():
+        print(f'[X] 缺 {P.rel(MANIFEST)}; 先跑 --extract')
+        return 1
+    if not (SHELL.is_file() and PORTAL.is_file()):
+        print('[X] shell.html / portal.html 不全, 不能重基线')
+        return 1
+    man = json.loads(rd(MANIFEST))
+    old_shell = man.get('shell', {}).get('sha256', '?')[:16]
+    old_portal = (man.get('expected_portal_sha256') or '?')[:16]
+    shell_b = SHELL.read_bytes()
+    portal_b = PORTAL.read_bytes()
+    man['shell'] = dict(file='shell.html', bytes=len(shell_b), sha256=sha(shell_b))
+    man['expected_portal_sha256'] = sha(portal_b)
+    man['notes'] = list(man.get('notes') or []) + [
+        f'重基线 {time.strftime("%Y-%m-%d %H:%M")}: shell {old_shell}→{sha(shell_b)[:16]}, '
+        f'portal {old_portal}→{sha(portal_b)[:16]} (有意改动门户外壳后重设漂移基准)']
+    wr(MANIFEST, json.dumps(man, ensure_ascii=False, indent=1))
+    print(f'已重基线 {P.rel(MANIFEST)}')
+    print(f'  shell.html  {old_shell} → {sha(shell_b)[:16]}  ({len(shell_b):,} B)')
+    print(f'  portal.html {old_portal} → {sha(portal_b)[:16]}  ({len(portal_b):,} B)')
+    print('  下一步: --check 应显示"全部与 manifest 一致 ✔"')
+    return 0
+
+
 def do_check() -> int:
     if not MANIFEST.is_file():
         print(f'[X] 缺 {P.rel(MANIFEST)}; 先跑 --extract')
@@ -268,7 +302,10 @@ if __name__ == '__main__':
     g.add_argument('--extract', action='store_true', help='从现有 portal.html 拆出 portal_src/ (一次性)')
     g.add_argument('--verify', action='store_true', help='装配到临时文件并与现有门户逐字节比对')
     g.add_argument('--check', action='store_true', help='源件与 manifest 的 sha256 漂移检查')
+    g.add_argument('--rebaseline', action='store_true',
+                   help='**有意**改了 shell.html/portal.html 后, 把当前状态记为新的漂移基准 (只改 manifest 的两个 sha, 不碰文件)')
     ap.add_argument('--no-claims', action='store_true', help='装配时不注入契约结论段')
     a = ap.parse_args()
     sys.exit(do_extract() if a.extract else do_verify() if a.verify else
+             do_rebaseline() if a.rebaseline else
              do_check() if a.check else do_build(claims=not a.no_claims))

+ 78 - 13
scripts/products_restore_missing.py

@@ -81,7 +81,8 @@ def stash_dir(explicit=None) -> pathlib.Path:
     含它的目录才是"标准答案"来源。原先只看第一个候选, 于是 --on 还原后暂存区空了 → 台账被写成 0/0。
     """
     if explicit:
-        return pathlib.Path(explicit)
+        # 必须 resolve: 传相对路径时下面 `stash.relative_to(ROOT)` 会 ValueError (2026-09-12 实逮)
+        return pathlib.Path(explicit).resolve()
     marker = 'windscada/temp_monthly.parquet'
     cands = []
     for pat in ('_products_off/*', '_products_off_prev_*/*'):
@@ -93,7 +94,21 @@ def stash_dir(explicit=None) -> pathlib.Path:
     scored.sort(key=lambda x: (-x[0], -x[1]))
     if scored:
         return scored[0][2]
-    raise SystemExit('找不到随包产物原件 (_products_off/** 或 _products_off_prev_**); 用 --stash 指定')
+    raise StashMissing('找不到随包产物原件 (_products_off/** 或 _products_off_prev_**); 用 --stash 指定')
+
+
+class StashMissing(SystemExit):
+    """暂存区不存在 —— **不是崩溃, 是"这一步没得做"**。
+
+    2026-09-16 用户令"系统运行中/未运行都能重算"时实逮: `rebuild_all.py` 的第⑤步 (补齐"包内没有生成端"
+    的产物) 在**没有暂存区**的检出上必抛 SystemExit → 重建链当场中断, 后面的 ⑦本体链 与 ⑧审计**全不执行**
+    (脚本的中断语义是"后面的步骤依赖它")。而暂存区只有"清除产物"动作会产生 —— 首次全量重算、或
+    清完又还原过的机器上就都没有, 于是"从零重算"在交付包上跑不完。
+    处置: 本类带退出码 6, 由 rebuild_all.py 显式容忍并打印"跳过随包补齐" —— 缺的是"随包件"这一路,
+    不是重算本身; 如实报出来比假装通过好, 也比把整条链打断好。
+    """
+
+    exit_code = 6
 
 
 def main() -> int:
@@ -101,16 +116,25 @@ def main() -> int:
     ap.add_argument('--dry-run', action='store_true')
     ap.add_argument('--stash', default=None)
     a = ap.parse_args()
-    stash = stash_dir(a.stash)
-    dest_root = P.STORE if hasattr(P, 'STORE') else ROOT / 'outputs' / 'rudong'
-    print(f'随包件: {stash.relative_to(ROOT)}')
+    try:
+        stash = stash_dir(a.stash)
+    except StashMissing as e:
+        # 没有暂存区 → 这一步没得做。**不许静默、也不许把整条重算链打断**(见 StashMissing 的说明)。
+        print(f'[跳过] {e}')
+        print('       处置: ① 从**交付包 zip** 按需补齐: --stash <交付包.zip> (推荐; 2026-09-16 用户令'
+              '"清除产物不留备份"之后, 交付包就是随包件的来源); '
+              '② --stash <随包产物目录> 显式指定一个目录; ③ 只想让页面有数 → 产物就在位, 无需本步。')
+        return StashMissing.exit_code
+    # ★2026-09-16: 原写法 `P.STORE if hasattr(P,'STORE') else ROOT/'outputs'/'rudong'` —— P.STORE 根本
+    #   不存在, 于是**恒**落到硬编码的 rudong, 多场部署下会把 rudong 的随包件补进别的场 (静默串场)。
+    dest_root = P.out_root()
+    print(f'随包件: {P.rel(stash)}')
     print(f'产物仓: {dest_root.relative_to(ROOT)}\n')
 
     prov = {}
-    for src in sorted(stash.rglob('*')):
-        if not src.is_file():
-            continue
-        rel = src.relative_to(stash).as_posix()
+
+    def take(rel: str, src_path=None, src_bytes=None):
+        """补齐一件 (或在位则只记账)。src_path=目录来源, src_bytes=zip 来源。"""
         dst = dest_root / rel
         if dst.exists():
             # 已存在的件分两类: 我们自己重算的(raw-derived) 与 上轮已补齐的(shipped)。
@@ -120,17 +144,58 @@ def main() -> int:
                 prov[rel] = dict(source='raw-derived', builder=RAW_DERIVED[rel])
             else:
                 prov[rel] = dict(source='shipped', why='包内无生成端 / 规则未复现 → 随包件补齐 (本轮已在位)')
-            continue
+            return
         prov[rel] = dict(source='shipped', why='包内无生成端 / 规则未复现 → 用随包件补齐')
-        if not a.dry_run:
-            dst.parent.mkdir(parents=True, exist_ok=True)
-            shutil.copy2(src, dst)
+        if a.dry_run:
+            return
+        dst.parent.mkdir(parents=True, exist_ok=True)
+        if src_bytes is not None:
+            dst.write_bytes(src_bytes)
+        else:
+            shutil.copy2(src_path, dst)
+
+    if stash.is_file():
+        # 交付包 zip 形态 (2026-09-16): 按需从包里取 outputs/<场>/** 补齐 ——
+        # 这是"清除不留备份"之后**唯一**的随包件来源, 所以把它做成一等输入。
+        import zipfile
+        with zipfile.ZipFile(stash) as zf:
+            pref = f'outputs/{P.farm()}/'
+            members = [n for n in zf.namelist() if n.startswith(pref) and not n.endswith('/')]
+            print(f'  从交付包读取 {len(members)} 件 (前缀 {pref})')
+            for name in sorted(members):
+                if a.dry_run:
+                    take(name[len(pref):])
+                else:
+                    take(name[len(pref):], src_bytes=zf.read(name))
+    else:
+        for src in sorted(stash.rglob('*')):
+            if not src.is_file():
+                continue
+            take(src.relative_to(stash).as_posix(), src_path=src)
 
     # 重算件里有些**不在随包件里**(如我们新造的 turbine_params.parquet), 也要记进台账
     for rel, builder in RAW_DERIVED.items():
         if rel not in prov and (dest_root / rel).exists():
             prov[rel] = dict(source='raw-derived', builder=builder)
 
+    # 构建脚本的**自登记** (src/derived_manifest.py): 振动侧的产物名随窗/分片变, 写不进上面的精确表,
+    # 按名字通配又会误伤同名旧件 → 由"谁算的谁登记", 这里只认登记。缺失不影响其余记账。
+    try:
+        from src.derived_manifest import load as _load_manifest, prune as _prune_manifest
+        _gone = _prune_manifest(dest_root)      # 先清掉"登记了但盘上已删"的条目 (如被删的 _reimport 窗)
+        if _gone:
+            print(f'  自登记清理: {_gone} 条指向已删文件的条目')
+        _man = (_load_manifest(dest_root).get('files') or {})
+    except Exception:
+        _man = {}
+    for rel, info in _man.items():
+        if (dest_root / rel).exists():
+            prov[rel] = dict(source='raw-derived',
+                             builder=f"{info.get('builder', '?')} (自登记 by {info.get('by', '?')})")
+        else:
+            prov[rel] = dict(source='raw-derived',
+                             builder=f"{info.get('builder', '?')} (自登记, 但该件当前不在盘上)")
+
     by_store = {}
     for rel, m in prov.items():
         store = rel.split('/')[0]

+ 55 - 130
scripts/products_state.py

@@ -1,40 +1,42 @@
 #!/usr/bin/env python3
 # -*- coding: utf-8 -*-
-r"""产物开关 —— 把 outputs\<场>\ 下的产物整体挪走 / 挪回 (人工检查页面空状态用)。
+r"""产物开关 —— 把 outputs\<场>\ 下的产物**直接清掉** / 看当前状态 (人工检查页面空状态用)。
 
-## 为什么做成开关而不是删除
+## 2026-09-16 用户令: 清除产物**不留备份**
 
-产物是"上一次算出来的结果", 人工检查"没数据时页面长什么样"必须把它们清掉; 但清掉之后一定要能**一条命令还原**,
-否则就得重跑整条摄入链(而且像油样那 102 行华创合并报告根本重算不出来)。所以: **移动**(同盘瞬间完成) + 清单留痕。
+原设计是"移动到 `_products_off/<场>/` + 清单留痕, 一条命令还原"。用户令改为**不备份**:
+清掉就是清掉。于是本脚本:
+  · `--off --yes`  真删除 (rmtree), 打印删了什么; **不再产生 `_products_off*` 目录**;
+  · `--status`     看产物在位/已清 (清掉后就是"已清", 没有备份清单可列);
+  · `--on`         已**移除**: 没有备份就没有"挪回"。需要恢复缺失的**随包件**时, 用
+                   `python scripts/products_restore_missing.py --stash <交付包.zip>`
+                   (从交付包 zip 里按需补齐 —— 那是"从交付件恢复", 不是"留一份备份")。
 
 ## 保留不动的
 
 - `data\raw\` 现场数据 (检查的是"产物没了会怎样", 不是"数据没了会怎样")
-- `outputs\<场>\windscada\_pre_rebuild_20260911\` 随包基线备份 (它是备份, 不是产物; `rebuild_from_raw.py --verify` 要用)
 - `release\` 门户与三维/交付件 (那是外壳, 不是数据产物)
 - `logs\` `run\` 运行期文件
+- ★ `outputs\<场>\windscada\_pre_rebuild_20260911\` **不再特殊保留**: 它曾是"随包基线备份"而在清除时豁免,
+  现在按"不留备份"一并删掉 —— 需要等价验收时, 用 `products_restore_missing.py --stash <交付包.zip>`
+  把这个目录从交付包里补回来即可。
 
 用法:
-    python scripts/products_state.py --off      # 挪走全部产物, 打印清单
-    python scripts/products_state.py --status   # 看当前是开还是关
-    python scripts/products_state.py --on       # 全部挪回
+    python scripts/products_state.py --status            # 看当前产物状态
+    python scripts/products_state.py --off --yes         # 清掉全部产物 (真删除, 不可恢复; --yes 是防手滑)
 """
 from __future__ import annotations
 
 import argparse
-import json
 import pathlib
 import shutil
 import sys
-import time
 
 ROOT = pathlib.Path(__file__).resolve().parents[1]
 sys.path.insert(0, str(ROOT))
 from src import paths as P                                        # noqa: E402
 
-OFF = ROOT / '_products_off'                                       # 挪走后的存放处 (仓库外无需, 放安装目录下便于还原)
-MANIFEST = OFF / 'manifest.json'
-KEEP = ('_pre_rebuild_20260911',)                                  # 随包基线备份, 不算产物
+LEGACY_OFF = ROOT / '_products_off'                                # 旧版暂存区 (只用于提示, 不再写入)
 
 
 def farms() -> list[pathlib.Path]:
@@ -42,15 +44,9 @@ def farms() -> list[pathlib.Path]:
     return sorted(p for p in out.iterdir() if p.is_dir()) if out.is_dir() else []
 
 
-def plan() -> list[tuple[pathlib.Path, pathlib.Path]]:
-    """→ [(原位置, 挪走后的位置)]"""
-    items = []
-    for fam in farms():
-        for d in sorted(fam.iterdir()):
-            if d.name in KEEP:
-                continue
-            items.append((d, OFF / fam.name / d.name))
-    return items
+def plan() -> list[pathlib.Path]:
+    """→ 待清除的产物目录列表 (产物仓本身)。"""
+    return [d for fam in farms() for d in sorted(fam.iterdir())]
 
 
 def count(p: pathlib.Path) -> tuple[int, float]:
@@ -60,131 +56,60 @@ def count(p: pathlib.Path) -> tuple[int, float]:
     return len(fs), sum(f.stat().st_size for f in fs) / 1024 / 1024
 
 
-def archive_old() -> bool:
-    """把**上一代**暂存区改名存档, 给这一次挪走腾出干净的落点。
-
-    为什么必须做: 挪走的落点是 `_products_off/<场>/<目录>`; 若该目录已存在(shutil.move 的语义是
-    "目的地是已存在目录 → 把源挪进去"), 会嵌套成 `_products_off/rudong/windscada/windscada/…`,
-    随后 `--on` 就会把新旧两代混在一起还原 —— 2026-09-12 实逮这个隐患。
-    存档时**保留** `_baseline_kept/`(随包基线, `rebuild_from_raw.py --verify` 要用), 只动场站目录与清单。
-    """
-    arch = ROOT / f'_products_off_prev_{time.strftime("%Y%m%d_%H%M%S")}'
-    moved = []
-    for fam in farms():
-        src = OFF / fam.name
-        if src.exists():
-            arch.mkdir(parents=True, exist_ok=True)
-            shutil.move(str(src), str(arch / fam.name))
-            moved.append(fam.name)
-    if MANIFEST.exists():
-        arch.mkdir(parents=True, exist_ok=True)
-        shutil.move(str(MANIFEST), str(arch / 'manifest.json'))
-    if not moved and not arch.exists():
-        return False
-    print(f'  上一代暂存区已存档 → {P.rel(arch)}/  (场站: {", ".join(moved) or "无"}; 基线 _baseline_kept 留在原地)')
-    return True
-
-
-def do_off(archive=False) -> int:
-    items = plan()
+def do_off(assume_yes=False) -> int:
+    items = [p for p in plan() if p.exists()]
     if not items:
-        print('产物已经全部挪走了 (用 --status 看清单)')
+        print('产物已经是空的 (用 --status 看)')
         return 0
-    OFF.mkdir(parents=True, exist_ok=True)
-    conflicts = [dst for _, dst in items if dst.exists()]
-    if conflicts:
-        if not archive:
-            print(f'[X] 暂存区里已经有同名产物 ({len(conflicts)} 个, 例: {P.rel(conflicts[0])}) —— 直接挪会嵌套成'
-                  f' <目录>/<目录>/, 之后 --on 会把两代混在一起。\n'
-                  f'    二选一: ① 加 --archive-old 把上一代存档后重跑 (推荐);\n'
-                  f'           ② 手工把整个 {P.rel(OFF)} 改名后再跑 (例: Move-Item _products_off _products_off_旧)。')
-            return 1
-        archive_old()
-    man = []
-    # 随包基线备份藏在 windscada 里面: 先摘出来, 它不是产物 (rebuild_from_raw.py --verify 要用)
-    baseline = P.store() / KEEP[0]
-    baseline_kept = False
-    kept_at = OFF / '_baseline_kept' / KEEP[0]
-    if baseline.exists():
-        if not kept_at.exists():
-            kept_at.parent.mkdir(parents=True, exist_ok=True)
-            n, mb = count(baseline)
-            shutil.move(str(baseline), str(kept_at))
-            print(f'  保留 {P.rel(baseline):34s} {n:5d} 件 {mb:8.1f} MB  (随包基线, 不算产物)')
-    kept_at = OFF / '_baseline_kept' / KEEP[0]
-    baseline_kept = kept_at.exists()          # ★必须在 move **之后**判 (2026-09-11 曾写成之前 → 恒 False, --on 就不还原基线)
-    for src, dst in items:
-        if not src.exists():
-            continue
+    total_n = total_mb = 0
+    rows = []
+    for src in items:
         n, mb = count(src)
-        dst.parent.mkdir(parents=True, exist_ok=True)
-        shutil.move(str(src), str(dst))
-        man.append(dict(rel=P.rel(src), to=P.rel(dst), files=n, mb=round(mb, 1)))
-        print(f'  挪走 {P.rel(src):34s} {n:5d} 件 {mb:8.1f} MB')
-    MANIFEST.write_text(json.dumps(dict(at=time.strftime('%Y-%m-%d %H:%M:%S'),
-                                        baseline_kept=baseline_kept, items=man),
-                                   ensure_ascii=False, indent=1), encoding='utf-8')
-    print(f'\n共 {len(man)} 项已挪到 {P.rel(OFF)} (清单: {P.rel(MANIFEST)})')
-    print(f'还原: python scripts/products_state.py --on')
-    return 0
-
-
-def do_on() -> int:
-    if not MANIFEST.exists():
-        print('没有清单 (没挪过, 或清单已删)')
-        return 1
-    man = json.loads(MANIFEST.read_text(encoding='utf-8'))
-    n = 0
-    for it in man['items']:
-        src, dst = P.resolve(it['to']), P.resolve(it['rel'])
-        if not src.exists():
-            print(f'  [!] 备份里没有 {it["to"]}, 跳过')
-            continue
-        dst.parent.mkdir(parents=True, exist_ok=True)
-        if dst.exists():
-            shutil.rmtree(dst) if dst.is_dir() else dst.unlink()
-        shutil.move(str(src), str(dst))
-        n += 1
-        print(f'  还原 {it["rel"]:34s} {it["files"]:5d} 件')
-    if man.get('baseline_kept') or (OFF / '_baseline_kept' / KEEP[0]).exists():
-        b = OFF / '_baseline_kept' / KEEP[0]        # 不看清单标记, 直接看备份在不在 (更稳)
-        if b.exists():
-            target = P.store() / KEEP[0]
-            target.parent.mkdir(parents=True, exist_ok=True)
-            if target.exists():
-                shutil.rmtree(target)
-            shutil.move(str(b), str(target))
-            print(f'  还原 {P.rel(target):34s} (随包基线)')
-    if OFF.exists() and not any(OFF.rglob('*')):
-        shutil.rmtree(OFF, ignore_errors=True)
-    print(f'\n共还原 {n} 项')
+        rows.append((P.rel(src), n, mb))
+        total_n += n
+        total_mb += mb
+    print('即将**删除**以下产物 (按用户令: 不留备份, 删除后不可恢复):')
+    for rel, n, mb in rows:
+        print(f'  {rel:34s} {n:5d} 件 {mb:8.1f} MB')
+    print(f'  合计 {len(rows)} 项 / {total_n} 件 / {total_mb:.1f} MB')
+    print('  恢复办法: python scripts/products_restore_missing.py --stash <交付包.zip> (从交付包补缺失件)')
+    if not assume_yes:
+        print('\n[X] 未执行: 这是不可恢复操作, 请显式加 --yes (运维控制台里点按钮时已由前端确认过一次)')
+        return 2
+    for src in items:
+        shutil.rmtree(src, ignore_errors=True) if src.is_dir() else src.unlink(missing_ok=True)
+    if LEGACY_OFF.exists():
+        print(f'  [!] 检测到旧版暂存区 {P.rel(LEGACY_OFF)} —— 按"不留备份"的口径, 本次**未删除**它; '
+              f'确认不需要后请手工删掉 (它是旧设计留下的产物备份)')
+    print(f'\n已删除 {len(rows)} 项 / {total_n} 件 / {total_mb:.1f} MB (无备份)')
     return 0
 
 
 def do_status() -> int:
-    items = plan()
+    items = [p for p in plan() if p.exists()]
     if not items:
-        print('当前: 产物**已全部挪走** (页面应呈空状态)')
-        if MANIFEST.exists():
-            man = json.loads(MANIFEST.read_text(encoding='utf-8'))
-            print(f'  挪走时间 {man["at"]}, 共 {len(man["items"])} 项')
-            for it in man['items']:
-                print(f'    {it["rel"]:34s} {it["files"]:5d} 件 {it["mb"]:8.1f} MB')
+        print('当前: 产物**已清空** (页面应呈空状态)')
+        print('  恢复: python scripts/products_restore_missing.py --stash <交付包.zip>')
         return 0
     print(f'当前: 产物**在位** ({len(items)} 项)')
-    for src, _ in items:
+    for src in items:
         n, mb = count(src)
         print(f'    {P.rel(src):34s} {n:5d} 件 {mb:8.1f} MB')
+    print('  说明: 按 2026-09-16 用户令, 清除产物不留备份 (没有 _products_off/ 可还原)')
     return 0
 
 
 if __name__ == '__main__':
     ap = argparse.ArgumentParser()
     g = ap.add_mutually_exclusive_group(required=True)
-    g.add_argument('--off', action='store_true', help='挪走全部产物')
-    g.add_argument('--on', action='store_true', help='全部挪回')
+    g.add_argument('--off', action='store_true', help='清掉全部产物 (**直接删除, 不留备份**)')
     g.add_argument('--status', action='store_true', help='看当前状态')
-    ap.add_argument('--archive-old', action='store_true',
-                    help='--off 前把上一代暂存区存档到 _products_off_prev_<时间戳>/ (暂存区已有同名产物时必须加)')
+    g.add_argument('--on', action='store_true', help='[已移除] 旧版的"挪回" —— 见文件头说明')
+    ap.add_argument('--yes', action='store_true', help='确认执行不可恢复的清除 (--off 必带)')
     a = ap.parse_args()
-    sys.exit(do_off(archive=a.archive_old) if a.off else do_on() if a.on else do_status())
+    if a.on:
+        print('[X] --on 已移除: 按用户令"清除产物不留备份", 没有备份可挪回。\n'
+              '    要从交付包补回缺失的随包件: python scripts/products_restore_missing.py --stash <交付包.zip>\n'
+              '    要重新算出数据面产物: python scripts/rebuild_all.py (SCADA 侧加 --scada)')
+        sys.exit(3)
+    sys.exit(do_off(assume_yes=a.yes) if a.off else do_status())

+ 33 - 2
scripts/rebuild_all.py

@@ -53,8 +53,35 @@ def build_plan(a) -> list:
     plan.append(step_cmd('② 三门台账', [PY, 'scripts/rebuild_from_raw.py']))
     if not a.skip_scada:
         plan.append(step_cmd('③ SCADA 侧 10 个构建器 (约 15 分钟)', [PY, 'scripts/rebuild_from_raw.py', '--scada']))
-    plan.append(step_cmd('④ 月度派生件', [PY, 'scripts/windscada_monthly_build.py']))
-    plan.append(step_cmd('⑤ 补齐"包内没有生成端"的产物', [PY, 'scripts/products_restore_missing.py']))
+    # ④ 月度派生件: `windscada_monthly_build.py` 会拿随包基线**逐值比对**, 没有基线时它返回 5
+    #    ("无法比对 —— 这不是通过", 这是它的诚实口径, 不要改它的退出码)。
+    #    ★但"重算链"与"等价验收"是两件事: 基线只是**验收标准答案**, 缺它不该让整条链停在第④步
+    #    (2026-09-16 实逮: 基线目录被清理后, 从页面点"执行重算" 387 s 后 rc=1 断在这里, ⑦⑧ 全不执行)。
+    #    故这里容忍 5 并写清"跳过了什么"; 要真做验收, 先把随包件恢复成暂存区/基线(见 products_state.py)。
+    plan.append(step_cmd('④ 月度派生件', [PY, 'scripts/windscada_monthly_build.py'],
+                         tolerate=(5,),
+                         note='rc=5 = 找不到随包基线可比对 ⇒ **跳过等价验收**(不是通过)。'
+                              '要恢复验收能力: 把随包产物放回 _products_off*/<场>/ 或 '
+                              'outputs/<场>/windscada/_pre_rebuild_20260911/'))
+    if not a.skip_vib:
+        # 振动侧 (2026-09-12 用户令"振动数据参与运行、重算"): 原始导出的 Brande TCM JSON
+        # (data/raw/<场站>/windcms/**/*_decode.json) → 窗索引 + 谱库。消费者 (src/windcms/data.py)
+        # 自动发现新窗, 所以这一步之后 CMS 的谱图/标量/页面就带上了新数据。
+        # ★**不跑** `windcms.py report` —— 本包缺六层链的 model_run/fusion 产物, 重生成会用残缺输入
+        #   覆盖随包快照 (report.md 16,857 B → 467 B, overview.html −15%; 实测见 vib_raw_build.py 文件头)。
+        #   六层链补齐后应把该步改为带 --with-report。
+        # 现场没有振动原始件时本步**空跑退出 0** (不是每台机器都有振动导出, 不该报失败)。
+        plan.append(step_cmd('④b 振动侧摄入 (原始导出 → 窗索引/谱)',
+                             [PY, 'scripts/vib_raw_build.py'] + (['--with-report'] if a.vib_report else []),
+                             note='没有振动原始件时空跑属正常; 若解析出错会以非零退出 (不静默)'))
+    # ⑤ 补齐"包内没有生成端"的产物: 源件是**暂存区**(清除产物时挪出来的那份随包件)。
+    #    暂存区不存在时该步返回 6 (不是崩溃) —— 首次全量重算/清完又还原过的机器上都没有暂存区,
+    #    若不容忍就会把后面的 ⑦本体链 与 ⑧审计 一起打断 ("从零重算"整条链跑不完; 2026-09-16 实逮)。
+    plan.append(step_cmd('⑤ 补齐"包内没有生成端"的产物', [PY, 'scripts/products_restore_missing.py'],
+                         tolerate=(6,),
+                         note='rc=6 = 没有随包件来源, 该步跳过。★2026-09-16 用户令"清除产物不留备份"之后, '
+                              '随包件不再保留 _products_off/ 暂存区; 要补齐请显式给交付包: '
+                              'scripts/products_restore_missing.py --stash <交付包.zip> (或在 --src 的现场包场景下先补齐)'))
     if not a.no_restart:
         # 只重启**组件服务**, 网关留着 —— 两个理由:
         #  ① 控制台页面是网关提供的, 连它一起停的话, 从页面点的重算会把自己的界面(甚至自己)弄没;
@@ -84,6 +111,10 @@ def main() -> int:
     ap = argparse.ArgumentParser()
     ap.add_argument('--src', default=None, help='现场数据包目录 (给了就先跑 place_raw_data --scope full)')
     ap.add_argument('--skip-scada', action='store_true', help='跳过 SCADA 侧 10 个构建器 (没换 10min 数据时用)')
+    ap.add_argument('--skip-vib', action='store_true', help='跳过振动侧摄入 (没换 CMS 原始导出时用)')
+    ap.add_argument('--vib-report', action='store_true',
+                    help='振动侧额外重生成 CMS 报告/知识库 —— ★默认关: 本包缺六层链的 model_run/fusion '
+                         '产物, 重生成会掉内容 (report.md −97%), 详见 scripts/vib_raw_build.py 文件头')
     ap.add_argument('--no-restart', action='store_true', help='不自动重启服务')
     ap.add_argument('--with-verify', action='store_true', help='末尾加等价验收')
     ap.add_argument('--dry-run', action='store_true', help='只打印计划')

+ 383 - 0
scripts/rudong_tcm_index.py

@@ -0,0 +1,383 @@
+#!/usr/bin/env python3
+# -*- coding: utf-8 -*-
+"""TCM 解码 JSON → 窗索引 index.parquet (2026-09-12 补; `src/windcms/pipeline.py` 的 tcm_decoded_json 第一步).
+
+## 它是谁、为什么在这里
+
+`src/windcms/pipeline.py::ingest()` 对"含 `*_decode.json` 的目录"跑两步:
+    ① rudong_tcm_index.py   --root <dir> --out <m5>/windows/<w>/index.parquet
+    ② rudong_tcm_spectra.py --root <dir> --out <m5>/windows/<w>/spectra
+这两个脚本在 v0.2.0 里**缺失**(振动线分支 claude/vibration-data-diagnosis-32b69e 上的东西没随包),
+于是"振动数据参与重算"这条路是断的: 只有产物 `tcm_index.parquet`(w0127 汇总) 在, 没有生成端。
+2026-09-12 用户补入 CMS 原始导出 (`CMS_RuDong_CGN_202603-04.zip`, 25,679 件 Brande TCM 导出) 后,
+按消费端契约把这两步补齐。
+
+## 输入
+
+Brande TCM Enterprise 导出 (西门子机组自带 M-system), 每个文件是一次 API 响应:
+    {"expiry", "buildId", "method", "controller", "serial", "requestInfo",
+     "body": {"body": {"<ISO时间戳>": [{"Record": {...}}, ...]}}}
+`Record` 下才有真数据: `Site` / `Location` / `ConfigurationSettings` / `Sensor` / `Measurement`
+(后者再带 `Conditions` 与 `DataSets`)。文件布局实测两种, 本脚本都认:
+    measurement/<YYYY>/<MM>/<WTGxx>/<WTGxx>_<uuid>_decode.json     (本包; --root 指到 measurement 或更上层都行)
+    <WTGxx>/<YYYY>/<MM>/<WTGxx>_<uuid>_decode.json                 (w0127 那批的布局)
+
+## 输出契约 (消费者 = src/windcms/data.py)
+
+`_read_index()` 只挑 `turbine, sensor_name, meas_name, trigger_time, rpm, condition_key, alarm_type,
+ds_size, scalar_value, overload`; `load_alarms` 读 `turbine, alarm_type`; 其余列是谱/工况/报警阈值元数据,
+供风电场页面与后续六层链用。**列名与列义必须与包内 `outputs/<场>/m5_cms_tcm/tcm_index.parquet`
+(330,308 行 × 54 列, 2026-01-27~02-03 窗) 逐列对齐** —— 那份是同一摄取逻辑的产物, 是唯一的格式基准;
+本脚本的 54 列与它同名同义 (见 COLSPEC), 因此新旧窗可以拼接进入同一分析集。
+
+## 并行与可重入
+
+150 GB / 2.5 万件的单线程 json 解析要 ~1 小时, 机器 14 核 → 按机组切片并行 (`--jobs`, 默认 8)。
+每个子进程写自己的分片 parquet (`<out>.parts/p<i>.parquet`), 父进程合并后删除分片。
+单文件/单记录异常不中断整窗: 记进 `parse_error` 列 (响亮留痕), 文件级失败计数并在末尾汇总报告。
+
+用法:
+    python scripts/rudong_tcm_index.py --root data/raw/如东/windcms/CMS_RuDong_CGN_202603-04/measurement \
+        --out outputs/rudong/m5_cms_tcm/windows/w0317/index.parquet
+    python scripts/rudong_tcm_index.py --root <dir> --out <index.parquet> --jobs 12 --turbines WTG01,WTG02
+    python scripts/rudong_tcm_index.py --root <dir> --out <index.parquet> --limit 200 --dry-run
+"""
+from __future__ import annotations
+
+import argparse
+import json
+import math
+import os
+import pathlib
+import shutil
+import sys
+import time
+import zipfile
+
+ROOT = pathlib.Path(__file__).resolve().parents[1]
+sys.path.insert(0, str(ROOT))
+from src.console import soft  # noqa: E402
+
+soft()
+
+# 厂内 JSON 层 → 索引列 (缺键给 None; 表里写的就是 JSON 里的键名, 便于逐列对拍)
+REC_MAP = {
+    'serial': ('Location', 'SerialNumber'),
+    'turbine': ('Location', 'LocationName'),
+    'config_name': ('ConfigurationSettings', 'ConfigurationName'),
+    'config_time': ('ConfigurationSettings', 'ConfigurationTime'),
+    'turbine_type': ('ConfigurationSettings', 'TurbineType'),
+    'recording_time_s': ('ConfigurationSettings', 'RecordingTime_s'),
+    'monitoring_cycle_min': ('ConfigurationSettings', 'MonitoringCycle-Minutes'),
+    'sensor_name': ('Sensor', 'SensorName'),
+    'sensor_type': ('Sensor', 'SensorType'),
+    'sensor_addr': ('Sensor', 'SensorAddress'),
+    'sensor_sn': ('Sensor', 'Sensor_Sn'),
+    'sensitivity_mvpg': ('Sensor', 'Sensitivity_mVpG'),
+    'sensor_unit': ('Sensor', 'SensorUnit'),
+    'hw_serial': ('Sensor', 'HwSerial'),
+    'meas_type': ('Measurement', 'MeasurementType'),
+    'meas_name': ('Measurement', 'MeasurementName'),
+    'meas_source': ('Measurement', 'Source'),
+    'category': ('Measurement', 'Category'),
+    'meas_key': ('Measurement', 'MeasurementKey'),
+    'trigger_time': ('Measurement', 'Trigger_Time'),
+    'trigger_ms': ('Measurement', 'Trigger_Time_ms'),
+    'duration_s': ('Measurement', 'Measurement_Duration-s'),
+    'rpm': ('Measurement', 'RPM'),
+    'nominal_freq_hz': ('Measurement', 'NominalFrequency-Hz'),
+    'min_freq_hz': ('Measurement', 'MinimumFrequency-Hz'),
+    'max_freq_hz': ('Measurement', 'MaximumFrequency-Hz'),
+    'lower_freq_hz': ('Measurement', 'LowerFrequency-Hz'),
+    'upper_freq_hz': ('Measurement', 'UpperFrequency-Hz'),
+    'overload': ('Measurement', 'Overload'),
+    'n_averages': ('Measurement', 'NumberOfAverages'),
+    'bandwidth_hz': ('Measurement', 'Bandwidth-Hz'),
+    'lines': ('Measurement', 'Lines'),
+    'integration': ('Measurement', 'Integration'),
+    'condition_name': ('Measurement', 'Conditions', 'ConditionsName'),
+    'condition_key': ('Measurement', 'ConditionKey'),
+    'alarm_type': ('Measurement', 'AlarmType'),
+    'red_alarm': ('Measurement', 'RedAlarm'),
+    'yellow_alarm': ('Measurement', 'YellowAlarm'),
+    'blue_alarm': ('Measurement', 'BlueAlarm'),
+    'red_hys': ('Measurement', 'RedHysAlarm'),
+    'yellow_hys': ('Measurement', 'YellowHysAlarm'),
+    'trend_hys': ('Measurement', 'TrendHysAlarm'),
+    'fault_freq': ('Measurement', 'FaultFreq'),
+}
+# DataSets 下的列。★层级别踩错 (2026-09-12 对拍逮到): `Size`/`Dimension`/`Values` 在
+# `Measurement.DataSets.DataSet` 里, 而 X/Y 轴四件在**上一层** `Measurement.DataSets` 里 ——
+# 一开始全按 DataSet 取, 结果 x_offset/x_delta/x_unit/y_unit 四列整列为空,
+# 而 "x_delta == 带宽/lines" 的内部一致性检查比例是 0.000 (本该 1.000), 就是这里露的马脚。
+DS_MAP = {
+    'ds_dim': 'Dimension',
+    'ds_size': 'Size',
+}
+AXIS_MAP = {
+    'x_offset': 'X-axisOffset',
+    'x_delta': 'X-axisDelta',
+    'x_unit': 'X-axisUnit',
+    'y_unit': 'Y-axisUnit',
+}
+TXT_COLS = ('file', 'serial', 'turbine', 'ts_key', 'config_name', 'config_time', 'turbine_type',
+            'sensor_name', 'sensor_type', 'sensor_addr', 'sensor_sn', 'sensor_unit', 'hw_serial',
+            'meas_type', 'meas_name', 'meas_source', 'category', 'meas_key', 'trigger_time',
+            'integration', 'condition_name', 'condition_key', 'alarm_type', 'red_hys', 'yellow_hys',
+            'trend_hys', 'x_unit', 'y_unit', 'parse_error')
+NUM_COLS = ('rec_i', 'recording_time_s', 'monitoring_cycle_min', 'sensitivity_mvpg', 'trigger_ms',
+            'duration_s', 'rpm', 'nominal_freq_hz', 'min_freq_hz', 'max_freq_hz', 'lower_freq_hz',
+            'upper_freq_hz', 'overload', 'n_averages', 'bandwidth_hz', 'lines', 'fault_freq',
+            'red_alarm', 'yellow_alarm', 'blue_alarm', 'ds_dim', 'ds_size', 'x_offset', 'x_delta',
+            'scalar_value')
+# 54 列的**列序**逐字抄自包内 `outputs/<场>/m5_cms_tcm/tcm_index.parquet` (振动线 v0.2.0 产物,
+# 330,308 行) —— 消费者按列名取数, 但"列序也要一致"是为了让新旧窗在人工对拍/并排打印时能直接比。
+# (文本列与数值列是交错的, 所以不能用 TXT_COLS + NUM_COLS 拼。)
+COLSPEC = ['file', 'serial', 'turbine', 'ts_key', 'rec_i', 'config_name', 'config_time', 'turbine_type',
+           'recording_time_s', 'monitoring_cycle_min', 'sensor_name', 'sensor_type', 'sensor_addr',
+           'sensor_sn', 'sensitivity_mvpg', 'sensor_unit', 'hw_serial', 'meas_type', 'meas_name',
+           'meas_source', 'category', 'meas_key', 'trigger_time', 'trigger_ms', 'duration_s', 'rpm',
+           'nominal_freq_hz', 'min_freq_hz', 'max_freq_hz', 'lower_freq_hz', 'upper_freq_hz', 'overload',
+           'n_averages', 'bandwidth_hz', 'lines', 'integration', 'condition_name', 'condition_key',
+           'alarm_type', 'red_alarm', 'yellow_alarm', 'blue_alarm', 'red_hys', 'yellow_hys', 'trend_hys',
+           'fault_freq', 'ds_dim', 'ds_size', 'x_offset', 'x_delta', 'x_unit', 'y_unit', 'scalar_value',
+           'parse_error']
+
+
+def _dig(d, path):
+    """按键路径取 (任一环缺失返回 None, 不抛)。"""
+    for k in path:
+        if not isinstance(d, dict):
+            return None
+        d = d.get(k)
+    return d
+
+
+def _f(v):
+    """→ float; 空/非数 → NaN (索引列必须能进 pandas float 列, 字符串混进去会让整列变 object)。"""
+    if v is None or v == '':
+        return math.nan
+    try:
+        return float(v)
+    except (TypeError, ValueError):
+        return math.nan
+
+
+def _s(v):
+    return None if v is None else str(v)
+
+
+def _datasets(m):
+    """Measurement.DataSets → (DataSets 层, DataSet 记录)。DataSet 可能是 dict 也可能是 list。"""
+    dss = m.get('DataSets') or {}
+    if not isinstance(dss, dict):
+        return {}, {}
+    ds = dss.get('DataSet')
+    if isinstance(ds, list):
+        ds = ds[0] if ds else {}
+    return dss, (ds if isinstance(ds, dict) else {})
+
+
+def rows_of_doc(doc, relpath, ts_key, recs, rows):
+    """一个 (文件, 时间戳) 下的一批记录 → 追加进 rows。"""
+    for rec_i, item in enumerate(recs):
+        rec = item.get('Record') if isinstance(item, dict) else None
+        if not isinstance(rec, dict):
+            rows.append(dict(file=relpath, ts_key=ts_key, rec_i=rec_i,
+                             parse_error='记录非 Record 结构: ' + str(type(item).__name__)))
+            continue
+        try:
+            m = rec.get('Measurement') or {}
+            dss, ds = _datasets(m)
+            vals = ds.get('Values')
+            size = _f(ds.get('Size'))
+            row = dict(file=relpath, ts_key=ts_key, rec_i=rec_i)
+            for col, path in REC_MAP.items():
+                row[col] = _dig(rec, path)
+            for col, key in DS_MAP.items():
+                row[col] = ds.get(key)
+            for col, key in AXIS_MAP.items():          # X/Y 轴在 DataSets 层, 不在 DataSet 里
+                row[col] = dss.get(key)
+            # 标量 (Size==1): Values 字符串本身就是标量值; 谱/波形 (Size>1) 的 Values 是长串, 值不进索引
+            row['scalar_value'] = _f(vals) if (size == 1 and isinstance(vals, (str, int, float))) else None
+            row['parse_error'] = None
+            rows.append(row)
+        except Exception as exc:                      # 单记录异常不许断整窗
+            rows.append(dict(file=relpath, ts_key=ts_key, rec_i=rec_i,
+                             parse_error=f'{type(exc).__name__}: {exc}'))
+
+
+def rows_of_file(raw: bytes, relpath: str, rows: list):
+    doc = json.loads(raw.decode('utf-8', 'replace'))
+    body = ((doc.get('body') or {}).get('body')) or {}
+    if not isinstance(body, dict):
+        rows.append(dict(file=relpath, parse_error='body.body 非 dict'))
+        return
+    for ts_key, recs in body.items():
+        if isinstance(recs, list):
+            rows_of_doc(doc, relpath, ts_key, recs, rows)
+        else:
+            rows.append(dict(file=relpath, ts_key=ts_key, parse_error='记录集非 list'))
+
+
+def relpath_of(name: str) -> str:
+    """包内成员名 → `file` 列。去掉 `measurement/` 这一层 (包本身的组织层, 不是数据层)。"""
+    n = name.replace('\\', '/')
+    if n.startswith('measurement/'):
+        n = n[len('measurement/'):]
+    return n
+
+
+def discover(root: pathlib.Path, turbines, limit):
+    """→ [(成员名/相对路径, 完整路径|None, zip 路径|None)]。目录与 zip 都支持。"""
+    out = []
+    if root.is_file() and root.suffix.lower() == '.zip':
+        with zipfile.ZipFile(root) as zf:
+            for i in zf.infolist():
+                if i.is_dir() or not i.filename.endswith('_decode.json'):
+                    continue
+                out.append((i.filename, None, str(root)))
+    else:
+        for p in sorted(root.rglob('*_decode.json')):
+            out.append((str(p.relative_to(root)).replace('\\', '/'), str(p), None))
+    if turbines:
+        want = {t.upper() for t in turbines}
+        out = [x for x in out if any(t in x[0].upper() for t in want)]
+    if limit:
+        out = out[:limit]
+    return out
+
+
+def flatten(items, rows):
+    """逐文件解析 (目录: 直接读; zip: 按需打开一次)。"""
+    cur_zip = None
+    cur_path = None
+    try:
+        for name, full, zpath in items:
+            rel = relpath_of(name)
+            try:
+                if zpath:
+                    if cur_zip is None or cur_path != zpath:
+                        if cur_zip:
+                            cur_zip.close()
+                        cur_zip = zipfile.ZipFile(zpath)
+                        cur_path = zpath
+                    raw = cur_zip.read(name)
+                else:
+                    raw = pathlib.Path(full).read_bytes()
+                rows_of_file(raw, rel, rows)
+            except Exception as exc:
+                rows.append(dict(file=rel, parse_error=f'文件级失败 {type(exc).__name__}: {exc}'))
+    finally:
+        if cur_zip:
+            cur_zip.close()
+
+
+def worker(payload):
+    """子进程: 解析分片 → 写分片 parquet → 回 (分片路径, 行数, 文件数, 失败数)。"""
+    part_path, items = payload
+    import pandas as pd
+    rows = []
+    flatten(items, rows)
+    df = pd.DataFrame(rows)
+    for c in COLSPEC:
+        if c not in df.columns:
+            df[c] = None
+    df = df[list(COLSPEC)]
+    for c in NUM_COLS:
+        # 与 shipped 一致的 float64 (整列同型; 有 NaN 的列本就会升为 float, 没有 NaN 的列不升则出现 int64)
+        df[c] = pd.to_numeric(df[c], errors='coerce').astype('float64')
+    for c in TXT_COLS:
+        df[c] = df[c].astype('str')          # pandas 3 的 'str' dtype (shipped 同款)
+    bad = int(df['parse_error'].notna().sum()) if 'parse_error' in df else 0
+    df.to_parquet(part_path, index=False)
+    return part_path, len(df), len(items), bad
+
+
+def main() -> int:
+    ap = argparse.ArgumentParser()
+    ap.add_argument('--root', required=True, help='含 *_decode.json 的目录, 或导出 zip')
+    ap.add_argument('--out', required=True, help='index.parquet 输出路径')
+    ap.add_argument('--jobs', type=int, default=min(8, (os.cpu_count() or 4)),
+                    help='并行进程数 (按机组切片; 默认 %(default)s)')
+    ap.add_argument('--turbines', default=None, help='只处理这些机组 (逗号分隔, 如 WTG01,WTG02)')
+    ap.add_argument('--limit', type=int, default=0, help='只处理前 N 个文件 (冒烟; 默认 0=全部)')
+    ap.add_argument('--dry-run', action='store_true', help='只报将处理多少文件, 不解析')
+    a = ap.parse_args()
+
+    root = pathlib.Path(a.root)
+    if not root.exists():
+        raise SystemExit(f'输入不存在: {root}')
+    out = pathlib.Path(a.out)
+    turbines = [x.strip() for x in a.turbines.split(',')] if a.turbines else None
+    t0 = time.time()
+    items = discover(root, turbines, a.limit)
+    if not items:
+        raise SystemExit(f'没找到 *_decode.json: {root}')
+    size = sum((pathlib.Path(f).stat().st_size if f else 0) for _, f, _ in items)
+    print(f'输入: {root}')
+    print(f'发现 {len(items)} 个 *_decode.json  目录内合计 {size / 1073741824:.2f} GB(仅目录模式可量)')
+    if a.dry_run:
+        print('(dry-run, 未解析)')
+        return 0
+
+    import pandas as pd
+    import concurrent.futures as cf
+
+    # 切片: 优先按机组 (一个子进程只碰自己那几台 → 负载均匀且写入互不干扰)
+    groups = {}
+    for it in items:
+        key = next((seg for seg in it[0].replace('\\', '/').split('/') if seg.upper().startswith('WTG')), '_other')
+        groups.setdefault(key, []).append(it)
+    chunks = []
+    n = max(1, a.jobs)
+    per = math.ceil(len(items) / n)
+    cur = []
+    for k in sorted(groups):
+        cur.extend(groups[k])
+        if len(cur) >= per:
+            chunks.append(cur)
+            cur = []
+    if cur:
+        chunks.append(cur)
+
+    parts_dir = out.parent / (out.stem + '.parts')
+    if parts_dir.exists():
+        shutil.rmtree(parts_dir)
+    parts_dir.mkdir(parents=True, exist_ok=True)
+    payloads = [(str(parts_dir / f'p{i:02d}.parquet'), ch) for i, ch in enumerate(chunks)]
+    print(f'并行 {min(n, len(payloads))} 进程 × {len(payloads)} 片  →  {out}')
+
+    done_files = 0
+    total_rows = 0
+    total_bad = 0
+    part_files = []
+    with cf.ProcessPoolExecutor(max_workers=min(n, len(payloads))) as ex:
+        for part, nrow, nfile, nbad in ex.map(worker, payloads):
+            part_files.append(part)
+            done_files += nfile
+            total_rows += nrow
+            total_bad += nbad
+            print(f'  [{done_files}/{len(items)} 文件] 累计 {total_rows} 行, 异常 {total_bad}  '
+                  f'({time.time() - t0:.0f}s, {part})', flush=True)
+
+    df = pd.concat([pd.read_parquet(p) for p in sorted(part_files)], ignore_index=True)
+    out.parent.mkdir(parents=True, exist_ok=True)
+    df.to_parquet(out, index=False)
+    shutil.rmtree(parts_dir, ignore_errors=True)
+
+    print(f'\n完成: {len(df)} 行 × {df.shape[1]} 列 → {out}')
+    print(f'  机组 {df.turbine.nunique()} 台; 传感器 {df.sensor_name.nunique()} 种; 测量名 {df.meas_name.nunique()} 种')
+    if 'trigger_time' in df:
+        tt = pd.to_datetime(df.trigger_time, errors='coerce')
+        print(f'  时间窗 {tt.min()} → {tt.max()}')
+    print(f'  标量行(Size=1) {int((df.ds_size == 1).sum())}; 谱/波形行(Size>1) {int((df.ds_size > 1).sum())}; '
+          f'解析异常 {int(df.parse_error.notna().sum())}')
+    if total_bad:
+        print(f'  ⚠ 有 {total_bad} 条记录带 parse_error (见该列), 未静默丢弃')
+    print(f'  用时 {time.time() - t0:.0f}s')
+    return 0
+
+
+if __name__ == '__main__':
+    sys.exit(main())

+ 334 - 0
scripts/rudong_tcm_spectra.py

@@ -0,0 +1,334 @@
+#!/usr/bin/env python3
+# -*- coding: utf-8 -*-
+"""TCM 解码 JSON → 谱库 (npz 分片 + spectra_meta.parquet) —— pipeline.py 的 tcm_decoded_json 第二步.
+
+## 契约 (消费者 = src/windcms/data.py::spectra_meta / spectrum)
+
+    spectra_meta.parquet 每行一条谱: turbine, sensor, meas_name, trigger_time, shard, shard_row,
+                                    x_offset, x_delta (+ rpm/y_unit/alarm_type/n_points 等附列)
+    <store>/<shard>.npz  key='values' = 二维数组 (n 条 × n 点); 第 shard_row 行是该条谱
+    取数: x = x_offset + arange(len(v)) * x_delta,  v = npz['values'][shard_row]
+
+`data.spectra_meta()` 找 meta 的两条路径 (都写, 免得换 `--out` 就失联):
+    <window>/spectra/spectra_meta.parquet   ← 单窗 (本脚本 --out 指 windows/<w>/spectra)
+    <store>/spectra_meta.parquet            ← 合并库 (--out 指 m5/spectra 时)
+
+## 输入
+
+同上一步 (Brande TCM 导出 JSON)。**谱在 `Record.Measurement.DataSets.DataSet.Values` 里, 是空格
+分隔的数值字符串** (不是数组! 实测 6401 点 ~ 60 KB 字符串/条), Size = Lines + 1, X-axisDelta =
+带宽/Lines, X-axisUnit='Hz'。标量记录 Size==1 (其 Values 就是标量值, 归索引脚本处理, 这里跳过)。
+
+## 只转需要的测量
+
+原始导出里绝大部分字节是 `Time_*` 波形 (65536/200000 点), 而页面/判据用的是 `FFT_*` 谱。
+默认 `--meas FFT_` 只转谱; 要全转用 `--meas ALL` (磁盘会显著变大: 实测单个窗的谱 npz 约 11 GB 量级)。
+
+## 落盘与并行
+
+按 (测量名, 点数) 分组装分片, 每片 `--shard-rows` 条 (默认 256) 写一个 npz (float32, 见 --dtype)。
+并行按机组切片, 子进程写自己的 `p<NN>/` 命名空间 (shard 是相对 store 的路径, 带子目录合法),
+末尾父进程合并各分片的 meta。分片原子落盘 (先 .tmp 再改名), 中断不会留半截 npz。
+
+用法:
+    python scripts/rudong_tcm_spectra.py --root <measurement 目录> --out outputs/rudong/m5_cms_tcm/windows/w0317/spectra
+    python scripts/rudong_tcm_spectra.py --root <dir> --out <store> --meas FFT_ --jobs 8 --turbines WTG09
+"""
+from __future__ import annotations
+
+import argparse
+import json
+import math
+import os
+import pathlib
+import shutil
+import sys
+import time
+import zipfile
+
+ROOT = pathlib.Path(__file__).resolve().parents[1]
+sys.path.insert(0, str(ROOT))
+from src.console import soft  # noqa: E402
+
+soft()
+
+
+def _f(v):
+    if v is None or v == '':
+        return math.nan
+    try:
+        return float(v)
+    except (TypeError, ValueError):
+        return math.nan
+
+
+def _datasets(m):
+    """Measurement.DataSets → (DataSets 层, DataSet 记录)。X/Y 轴在 DataSets 层, Size/Values 在 DataSet 里。"""
+    dss = m.get('DataSets') or {}
+    if not isinstance(dss, dict):
+        return {}, {}
+    ds = dss.get('DataSet')
+    if isinstance(ds, list):
+        ds = ds[0] if ds else {}
+    return dss, (ds if isinstance(ds, dict) else {})
+
+
+def values_to_array(vals):
+    """Values (空格分隔字符串 / list) → float 数组; 不合法返回 None。"""
+    import numpy as np
+    if isinstance(vals, list):
+        try:
+            return np.asarray(vals, dtype=float)
+        except (TypeError, ValueError):
+            return None
+    if not isinstance(vals, str):
+        return None
+    try:
+        arr = np.fromstring(vals, sep=' ')          # C 级解析, 6401 点 ~ 0.2 ms
+    except Exception:
+        return None
+    if arr.size == 0:
+        try:
+            arr = np.asarray(vals.split(), dtype=float)
+        except (TypeError, ValueError):
+            return None
+    return arr if arr.size else None
+
+
+class ShardWriter:
+    """按 (meas_name, 点数) 攒批写 npz。"""
+
+    def __init__(self, store: pathlib.Path, ns: str, shard_rows: int, dtype):
+        self.store = store
+        self.ns = ns
+        self.shard_rows = shard_rows
+        self.dtype = dtype
+        self.buf = {}          # key → list[np.ndarray]
+        self.idx = {}          # key → 该 key 已写出片数
+        self.meta = []
+        self.shards = 0
+
+    def add(self, key, turbine, sensor, meas, trigger, x_off, x_delta, extra, arr):
+        rows = self.buf.setdefault(key, [])
+        rows.append(arr)
+        row_in_shard = len(rows) - 1
+        shard = f'{self.ns}/{self._stem(key, self.idx.get(key, 0))}'
+        self.meta.append(dict(turbine=turbine, sensor=sensor, meas_name=meas, trigger_time=trigger,
+                              shard=shard, shard_row=row_in_shard, x_offset=x_off, x_delta=x_delta,
+                              n_points=int(arr.size), **extra))
+        if len(rows) >= self.shard_rows:
+            self.flush_key(key)
+
+    def _stem(self, key, i):
+        meas, npts = key
+        safe = ''.join(c if (c.isalnum() or c in '-_.') else '_' for c in meas)
+        return f'{safe}_n{npts}_{i:04d}.npz'
+
+    def flush_key(self, key):
+        import numpy as np
+        rows = self.buf.get(key)
+        if not rows:
+            return
+        i = self.idx.get(key, 0)
+        rel = f'{self.ns}/{self._stem(key, i)}'
+        dst = self.store / rel
+        dst.parent.mkdir(parents=True, exist_ok=True)
+        # 不同测量/点数的分片点数一致 (同 key 同长度), 可直接 stack
+        arr = np.stack(rows).astype(self.dtype, copy=False)
+        tmp = dst.with_suffix('.npz.tmp')
+        with open(tmp, 'wb') as fh:
+            np.savez_compressed(fh, values=arr)
+        os.replace(tmp, dst)
+        self.shards += 1
+        self.idx[key] = i + 1
+        self.buf[key] = []
+
+    def close(self):
+        for key in list(self.buf):
+            self.flush_key(key)
+        return self.meta, self.shards
+
+
+def rows_of_file(raw: bytes, meas_prefix: str, writer: ShardWriter, stats: dict):
+    doc = json.loads(raw.decode('utf-8', 'replace'))
+    body = ((doc.get('body') or {}).get('body')) or {}
+    if not isinstance(body, dict):
+        stats['bad'] += 1
+        return
+    for _ts, recs in body.items():
+        if not isinstance(recs, list):
+            continue
+        for item in recs:
+            rec = item.get('Record') if isinstance(item, dict) else None
+            if not isinstance(rec, dict):
+                stats['bad'] += 1
+                continue
+            try:
+                m = rec.get('Measurement') or {}
+                meas = str(m.get('MeasurementName') or '?')
+                if meas_prefix != 'ALL' and not meas.startswith(meas_prefix):
+                    stats['skipped'] += 1
+                    continue
+                dss, ds = _datasets(m)
+                size = _f(ds.get('Size'))
+                if not (size and size > 1):
+                    stats['skipped'] += 1          # 标量行归索引脚本
+                    continue
+                arr = values_to_array(ds.get('Values'))
+                if arr is None or arr.size <= 1:
+                    stats['bad'] += 1
+                    continue
+                loc = (rec.get('Location') or {}).get('LocationName') or '?'
+                sens = (rec.get('Sensor') or {}).get('SensorName') or '?'
+                extra = dict(rpm=_f(m.get('RPM')), y_unit=dss.get('Y-axisUnit'),
+                             x_unit=dss.get('X-axisUnit'), alarm_type=m.get('AlarmType'),
+                             condition_key=m.get('ConditionKey'), ds_size=size,
+                             meas_key=m.get('MeasurementKey'))
+                writer.add((meas, int(arr.size)), loc, sens, meas, m.get('Trigger_Time'),
+                           _f(dss.get('X-axisOffset')), _f(dss.get('X-axisDelta')), extra, arr)
+                stats['rows'] += 1
+            except Exception:
+                stats['bad'] += 1
+
+
+def worker(payload):
+    part_path, items, store, ns, shard_rows, meas, dtype = payload
+    import pandas as pd
+    writer = ShardWriter(pathlib.Path(store), ns, shard_rows, dtype)
+    stats = dict(rows=0, bad=0, skipped=0)
+    cur_zip = None
+    cur_path = None
+    try:
+        for name, full, zpath in items:
+            try:
+                if zpath:
+                    if cur_zip is None or cur_path != zpath:
+                        if cur_zip:
+                            cur_zip.close()
+                        cur_zip = zipfile.ZipFile(zpath)
+                        cur_path = zpath
+                    raw = cur_zip.read(name)
+                else:
+                    raw = pathlib.Path(full).read_bytes()
+                rows_of_file(raw, meas, writer, stats)
+            except Exception:
+                stats['bad'] += 1
+    finally:
+        if cur_zip:
+            cur_zip.close()
+    meta, shards = writer.close()
+    pd.DataFrame(meta).to_parquet(part_path, index=False)
+    row = dict(part=part_path, rows=stats['rows'], bad=stats['bad'], skipped=stats['skipped'],
+               shards=shards, files=len(items))
+    return row
+
+
+def discover(root: pathlib.Path, turbines, limit):
+    out = []
+    if root.is_file() and root.suffix.lower() == '.zip':
+        with zipfile.ZipFile(root) as zf:
+            for i in zf.infolist():
+                if not i.is_dir() and i.filename.endswith('_decode.json'):
+                    out.append((i.filename, None, str(root)))
+    else:
+        for p in sorted(root.rglob('*_decode.json')):
+            out.append((str(p.relative_to(root)).replace('\\', '/'), str(p), None))
+    if turbines:
+        want = {t.upper() for t in turbines}
+        out = [x for x in out if any(t in x[0].upper() for t in want)]
+    return out[:limit] if limit else out
+
+
+def main() -> int:
+    ap = argparse.ArgumentParser()
+    ap.add_argument('--root', required=True, help='含 *_decode.json 的目录, 或导出 zip')
+    ap.add_argument('--out', required=True, help='谱库目录 (通常 <window>/spectra)')
+    ap.add_argument('--meas', default='FFT_', help="只转这些测量 (前缀匹配); 'ALL' = 全转 (默认 %(default)s)")
+    ap.add_argument('--jobs', type=int, default=min(8, (os.cpu_count() or 4)), help='并行进程数')
+    ap.add_argument('--shard-rows', type=int, default=256, help='每个 npz 装多少条谱 (默认 %(default)s)')
+    ap.add_argument('--dtype', default='float32', choices=('float32', 'float64'),
+                    help='谱点存储精度 (默认 float32; 判据标量不从这里算, 谱图显示足够)')
+    ap.add_argument('--turbines', default=None, help='只处理这些机组 (逗号分隔)')
+    ap.add_argument('--limit', type=int, default=0, help='只处理前 N 个文件 (冒烟)')
+    ap.add_argument('--dry-run', action='store_true')
+    a = ap.parse_args()
+
+    root = pathlib.Path(a.root)
+    if not root.exists():
+        raise SystemExit(f'输入不存在: {root}')
+    out = pathlib.Path(a.out)
+    turbines = [x.strip() for x in a.turbines.split(',')] if a.turbines else None
+    import numpy as np
+    dtype = np.float32 if a.dtype == 'float32' else np.float64
+
+    t0 = time.time()
+    items = discover(root, turbines, a.limit)
+    if not items:
+        raise SystemExit(f'没找到 *_decode.json: {root}')
+    print(f'输入: {root}\n发现 {len(items)} 个 *_decode.json; 测量过滤 {a.meas}; 谱库 → {out}')
+    if a.dry_run:
+        print('(dry-run, 未解析)')
+        return 0
+
+    import pandas as pd
+    import concurrent.futures as cf
+
+    groups = {}
+    for it in items:
+        key = next((seg for seg in it[0].replace('\\', '/').split('/') if seg.upper().startswith('WTG')), '_other')
+        groups.setdefault(key, []).append(it)
+    n = max(1, a.jobs)
+    per = math.ceil(len(items) / n)
+    chunks, cur = [], []
+    for k in sorted(groups):
+        cur.extend(groups[k])
+        if len(cur) >= per:
+            chunks.append(cur)
+            cur = []
+    if cur:
+        chunks.append(cur)
+
+    parts_dir = out.parent / (out.name + '.meta.parts')
+    if parts_dir.exists():
+        shutil.rmtree(parts_dir)
+    parts_dir.mkdir(parents=True, exist_ok=True)
+    payloads = [(str(parts_dir / f'p{i:02d}.parquet'), ch, str(out), f'p{i:02d}', a.shard_rows, a.meas, dtype)
+                for i, ch in enumerate(chunks)]
+    print(f'并行 {min(n, len(payloads))} 进程 × {len(payloads)} 片')
+
+    rows_total = bad = skipped = shards = files_done = 0
+    parts = []
+    with cf.ProcessPoolExecutor(max_workers=min(n, len(payloads))) as ex:
+        for r in ex.map(worker, payloads):
+            parts.append(r['part'])
+            rows_total += r['rows']
+            bad += r['bad']
+            skipped += r['skipped']
+            shards += r['shards']
+            files_done += r['files']
+            print(f"  [{files_done}/{len(items)} 文件] 谱 {rows_total} 条, 分片 {shards}, "
+                  f"跳过(非目标测量/标量) {skipped}, 异常 {bad}  ({time.time() - t0:.0f}s)", flush=True)
+
+    metas = [pd.read_parquet(p) for p in sorted(parts) if pathlib.Path(p).stat().st_size > 0]
+    meta = pd.concat(metas, ignore_index=True) if metas else pd.DataFrame(
+        columns=['turbine', 'sensor', 'meas_name', 'trigger_time', 'shard', 'shard_row', 'x_offset', 'x_delta'])
+    for p in parts:
+        pathlib.Path(p).unlink(missing_ok=True)
+    parts_dir.rmdir()
+
+    out.mkdir(parents=True, exist_ok=True)
+    meta.to_parquet(out / 'spectra_meta.parquet', index=False)
+    # 单窗约定: data.spectra_meta() 对 <window>/spectra 优先读 <window>/spectra_meta.parquet
+    if out.name == 'spectra':
+        meta.to_parquet(out.parent / 'spectra_meta.parquet', index=False)
+    print(f'\n完成: {len(meta)} 条谱 → {out} (+ {out.parent / "spectra_meta.parquet" if out.name == "spectra" else "—"})')
+    print(f'  分片文件 {shards} 个; 机组 {meta.turbine.nunique() if len(meta) else 0} 台; '
+          f'测量 {meta.meas_name.nunique() if len(meta) else 0} 种')
+    if len(meta):
+        print('  逐测量条数:\n' + meta.meas_name.value_counts().head(20).to_string())
+    print(f'  用时 {time.time() - t0:.0f}s')
+    return 0
+
+
+if __name__ == '__main__':
+    sys.exit(main())

+ 271 - 0
scripts/vib_raw_build.py

@@ -0,0 +1,271 @@
+#!/usr/bin/env python3
+# -*- coding: utf-8 -*-
+"""振动侧一键摄入: data/raw/<场站>/{windcms,m5_cms_tcm} 的原始件 → 窗索引/谱库 → (可选)CMS 报告与知识库.
+
+## 为什么需要它 (2026-09-12 用户令)
+
+用户令: 「数据层里的「CMS 振动评估报告」应遵循 `<安装目录>\\data\\raw\\如东\\windcms`、
+「振动线 handoff」应遵循 `<安装目录>\\data\\raw\\如东\\m5_cms_tcm` 存放; 修改系统支持振动数据参与
+运行、重算」。此前振动侧只有**产物**(`outputs/<场>/m5_cms_tcm` 90 件 + `outputs/<场>/windcms` 53 件),
+源件不在 data/raw, 生成端脚本 (振动线分支的 rudong_tcm_*.py 七件) 也没随包 ⇒ 从零重算时振动侧是断的。
+
+两个事实把这件事限定得很清楚:
+  · `src/windcms/pipeline.py` 认的输入就是"含 `*_decode.json` 的目录"(Brande TCM Enterprise 导出),
+    这一步**必须能跑**, 否则数据进了 data/raw 也只是躺着;
+  · 同分支的六层链 (oem_scan / energy_share / model_run / fusion) 脚本仍未随包 —— 本脚本**不假装**
+    能把那几步跑出来: 它只做"索引 + 谱 + (可选)报告/知识库", 其余在报告里如实写"缺失"。
+
+## 这一步跑完, 振动数据就真的参与了运行
+
+    data/raw/<场站>/windcms/.../*_decode.json
+        → outputs/<场>/m5_cms_tcm/windows/<窗>/index.parquet     (54 列, 与包内 tcm_index.parquet 同构)
+        → outputs/<场>/m5_cms_tcm/windows/<窗>/spectra/*.npz + spectra_meta.parquet
+    消费者 (无需改一行代码, 窗是自动发现的):
+        src/windcms/data.py::windows()      → 新窗进分析集
+        src/windcms/data.py::load_scalars() → 标量进 CMS 报告/页面
+        src/windcms/data.py::spectrum()     → 谱图能取到 (此前 spectra 库整个缺失, 谱图是死的)
+        scripts/windcms.py report / kb      → 报告与知识库重生成
+
+## ★ 为什么"重生成 CMS 报告"默认**不跑** (2026-09-12 实测)
+
+跑一次 `scripts/windcms.py report` 在本包会**让产物变差**, 不是变好:
+
+| 件 | 随包快照 | 本包重生成后 | 变化 |
+|---|---|---|---|
+| `windcms/report.md` | 16,857 B (含逐台融合级表 + L4 过闸谱线) | **467 B** (融合级表空、L4 写"无") | **−97.2%** |
+| `windcms/overview.html` | 646,218 B | 548,466 B | −15.1% |
+| `windcms/index_eng.html` | 443,841 B | 408,112 B | −8.0% |
+
+根因不在数据, 在**缺件**: `src/windcms/data.py::load_model()` 要读
+`m5/model_run_l6.parquet` 与 `m5/fusion_38.csv`, 而这两件属六层链的 `model_run` / `fusion` 两步 ——
+**那四步脚本没随包**(见 `missing_chain`), 产物也不在。于是重生成 = 用残缺输入覆盖完整快照。
+
+所以: 摄入(索引/谱)默认跑, **报告/知识库默认不跑**; 要跑得显式 `--with-report`, 并且跑之前**先备份**
+`outputs/<场>/windcms/`。六层链补齐后这个默认值应当翻过来 (那时重生成才会≥快照)。
+
+## 用法
+
+    python scripts/vib_raw_build.py --farm rudong                     # 摄入(索引+谱), 不动报告
+    python scripts/vib_raw_build.py --window w0316                    # 指定窗名
+    python scripts/vib_raw_build.py --jobs 12 --meas FFT_             # 并行/测量过滤
+    python scripts/vib_raw_build.py --with-report                     # 额外重生成 CMS 报告/知识库(见上, 慎用)
+    python scripts/vib_raw_build.py --dry-run                         # 只报会做什么
+"""
+from __future__ import annotations
+
+import argparse
+import json
+import os
+import pathlib
+import shutil
+import subprocess
+import sys
+import time
+
+ROOT = pathlib.Path(__file__).resolve().parents[1]
+sys.path.insert(0, str(ROOT))
+from src.console import soft  # noqa: E402
+
+soft()
+PY = sys.executable
+
+
+def human(n):
+    for u in ('B', 'KB', 'MB', 'GB'):
+        if n < 1024 or u == 'GB':
+            return f'{n:.1f} {u}'
+        n /= 1024
+
+
+def find_roots(station: pathlib.Path):
+    """→ 振动侧原始件根列表 (含 *_decode.json 的目录或导出 zip)。"""
+    roots = []
+    for name in ('windcms', 'm5_cms_tcm'):
+        d = station / name
+        if not d.is_dir():
+            continue
+        for z in sorted(d.rglob('*.zip')):
+            if z.name.upper().startswith('CMS') or 'decode' in z.name.lower():
+                roots.append(z)
+        for sub in sorted({p.parent for p in d.rglob('*_decode.json')}):
+            # 取**最上层**那个含 decode 文件的目录 (measurement/), 避免每台机组一个 root
+            top = sub
+            while top.parent != d and any(top.parent.rglob('*_decode.json')):
+                top = top.parent
+            if top not in roots:
+                roots.append(top)
+    return roots
+
+
+def run(step, cmd, log):
+    t0 = time.time()
+    print(f'[{step}] {" ".join(str(c) for c in cmd)}', flush=True)
+    p = subprocess.run([str(c) for c in cmd], cwd=str(ROOT), capture_output=True, text=True,
+                       encoding='utf-8', errors='replace')
+    tail = (p.stdout or '') + (p.stderr or '')
+    rec = dict(step=step, cmd=' '.join(str(c) for c in cmd), rc=p.returncode,
+               seconds=round(time.time() - t0, 1), tail=tail[-1200:])
+    log.append(rec)
+    print(f'    rc={p.returncode}  {rec["seconds"]}s', flush=True)
+    if p.returncode != 0:
+        print(tail[-2000:], flush=True)
+        raise SystemExit(f'步骤 {step} 失败 (rc={p.returncode})')
+    return rec
+
+
+def _rels_of_window(win: pathlib.Path, m5: pathlib.Path) -> dict:
+    """某窗的产物 → {相对产物仓的路径: 构建器说明}; 供正常摄入与 --register-only 共用 (单一实现)。"""
+    rels = {f'm5_cms_tcm/windows/{win.name}/index.parquet':
+            'scripts/rudong_tcm_index.py (54 列, 与包内 tcm_index.parquet 同列名列序)',
+            f'm5_cms_tcm/windows/{win.name}/spectra_meta.parquet':
+            'scripts/rudong_tcm_spectra.py',
+            'm5_cms_tcm/vib_raw_manifest.json': 'scripts/vib_raw_build.py'}
+    sp = win / 'spectra'
+    if sp.is_dir():
+        for p in sp.rglob('*.npz'):
+            rels[f'm5_cms_tcm/windows/{win.name}/spectra/{p.relative_to(sp).as_posix()}'] = \
+                'scripts/rudong_tcm_spectra.py (npz 分片: DataSets.DataSet.Values 空格串 → float 数组)'
+        # 谱库目录里那份 meta 是同一内容的两条读取路径之一, 也登记
+        rels[f'm5_cms_tcm/windows/{win.name}/spectra/spectra_meta.parquet'] = \
+            'scripts/rudong_tcm_spectra.py (与窗根同名件同内容: 兼顾两种读取约定)'
+    return rels
+
+
+def main() -> int:
+    ap = argparse.ArgumentParser()
+    ap.add_argument('--farm', default='rudong')
+    ap.add_argument('--window', default=None, help='窗名 (默认按数据起始日推 wMMDD)')
+    ap.add_argument('--jobs', type=int, default=min(8, (os.cpu_count() or 4)))
+    ap.add_argument('--meas', default='FFT_', help="谱摄入的测量过滤 (默认 FFT_; ALL=全转)")
+    ap.add_argument('--skip-spectra', action='store_true', help='只做索引 (谱很占盘)')
+    ap.add_argument('--with-report', action='store_true',
+                    help='额外跑 windcms.py report/kb —— ★默认关: 本包缺六层链的 model_run/fusion 产物, '
+                         '重生成会让 report.md/overview.html 掉内容 (见文件头实测表); 跑前先备份 windcms\\')
+    ap.add_argument('--register-only', action='store_true',
+                    help='不摄入, 只把**已存在**的窗体件补进来源自登记 (幂等修复: 例如先前的摄入跑在自登记'
+                         '功能之前, 台账就漏记了那些件)')
+    ap.add_argument('--limit', type=int, default=0, help='冒烟: 只摄入前 N 个文件')
+    ap.add_argument('--dry-run', action='store_true')
+    a = ap.parse_args()
+
+    from src.windscada.config import farm, raw_station_dir
+    from src import paths as P
+    cfg = farm(a.farm)
+    station = pathlib.Path(raw_station_dir(a.farm))
+    m5 = P.m5(a.farm)
+    roots = find_roots(station)
+    print(f'场站原始件目录: {station}')
+    print(f'振动侧产物目录: {m5}')
+
+    if a.register_only:
+        # 只补登记: 找已存在的窗 (--window 指定, 否则取名字最大的那个), 逐件登记。
+        # 放在"源件存在性检查"之前 —— 补登记不需要重新读原始件。
+        cands = sorted(p for p in (m5 / 'windows').glob('w[0-9][0-9][0-9][0-9]')
+                       if (p / 'index.parquet').exists())
+        if a.window:
+            cands = [p for p in cands if p.name == a.window]
+        if not cands:
+            print('没有已存在的窗可登记 (先正常跑一次摄入)')
+            return 0
+        win = cands[-1]
+        rels = _rels_of_window(win, m5)
+        from src.derived_manifest import record as _rec
+        p = _rec(P.out_root(a.farm), rels, by='scripts/vib_raw_build.py')
+        print(f'已登记 {len(rels)} 件 (窗 {win.name}) → {p}')
+        return 0
+
+    if not roots:
+        print('没找到振动原始件 (*_decode.json 目录或 CMS 导出 zip)。'
+              '先用 scripts/place_raw_data.py --scope vib 从现场包落位。')
+        return 0
+    for r in roots:
+        n = len(list(r.rglob('*_decode.json'))) if r.is_dir() else '(zip)'
+        sz = sum(p.stat().st_size for p in r.rglob('*_decode.json')) if r.is_dir() else r.stat().st_size
+        print(f'  源件根: {r}   {n} 个 decode 文件, {human(sz)}')
+    if a.dry_run:
+        print('(dry-run, 未摄入)')
+        return 0
+
+    windows_dir = m5 / 'windows'
+    staging = windows_dir / '_staging_ingest'          # 'test'/'_reimport' 之外的名字, 会被 data.windows() 看见 → 立即改名
+    if staging.exists():
+        shutil.rmtree(staging)
+    staging.mkdir(parents=True, exist_ok=True)
+
+    log = []
+    t0 = time.time()
+    root = roots[0] if len(roots) == 1 else station / 'windcms'   # 多个根时交给各自目录 (index 支持 rglob)
+    idx = staging / 'index.parquet'
+    run('index', [PY, str(ROOT / 'scripts/rudong_tcm_index.py'), '--root', str(root), '--out', str(idx),
+                  '--jobs', str(a.jobs)] + (['--limit', str(a.limit)] if a.limit else []), log)
+
+    n_spectra = 0
+    if not a.skip_spectra:
+        run('spectra', [PY, str(ROOT / 'scripts/rudong_tcm_spectra.py'), '--root', str(root),
+                        '--out', str(staging / 'spectra'), '--meas', a.meas, '--jobs', str(a.jobs)]
+            + (['--limit', str(a.limit)] if a.limit else []), log)
+        sm = staging / 'spectra' / 'spectra_meta.parquet'
+        if sm.exists():
+            import pandas as pd
+            n_spectra = len(pd.read_parquet(sm))
+
+    # 窗名: 用数据自身的起始日 (wMMDD), 与既有 w0127/w0707/w0811 同口径
+    import pandas as pd
+    d = pd.read_parquet(idx)
+    tt = pd.to_datetime(d.trigger_time, errors='coerce')
+    tmin, tmax = tt.min(), tt.max()
+    win = a.window or f'w{tmin:%m%d}'
+    final = windows_dir / win
+    if final.exists():
+        final = windows_dir / f'{win}_reimport_{time.strftime("%m%d%H%M")}'
+        print(f'⚠ 窗 {win} 已存在 → 本次摄入落到 {final.name} (data.EXCLUDE_DEFAULT 会排除 _reimport 窗, 不污染生产集)')
+    staging.rename(final)
+    print(f'\n窗: {final.name}   数据窗 {tmin} → {tmax}   行 {len(d)}   谱 {n_spectra}')
+
+    if a.with_report:
+        run('report', [PY, str(ROOT / 'scripts/windcms.py'), 'report', '--farm', a.farm], log)
+        run('kb', [PY, str(ROOT / 'scripts/windcms.py'), 'kb', '--farm', a.farm], log)
+    else:
+        print('(跳过 CMS 报告/知识库重生成 —— 本包缺 model_run/fusion 产物, 重生成会掉内容; '
+              '要跑用 --with-report, 且先备份 windcms\\)')
+
+    man = dict(farm=a.farm, window=final.name, window_dir=str(final.relative_to(ROOT)),
+               # 源件路径写**相对安装根**的形态: 这文件会随包分发, 绝对路径换机后就是死链
+               # (check_transferable.py 会把 outputs 下的绝对路径算作"机器相关路径")
+               sources=[str(r.relative_to(ROOT)) if str(r).startswith(str(ROOT)) else str(r) for r in roots],
+               sources_note='路径相对<安装目录>', rows=int(len(d)), spectra=int(n_spectra),
+               time_min=str(tmin), time_max=str(tmax),
+               turbines=int(d.turbine.nunique()), sensors=int(d.sensor_name.nunique()),
+               meas_names=int(d.meas_name.nunique()),
+               steps=[{k: v for k, v in r.items() if k != 'tail'} for r in log],
+               finished=time.strftime('%Y-%m-%d %H:%M:%S'), seconds=round(time.time() - t0, 1),
+               missing_chain=['rudong_tcm_oem_scan', 'rudong_line_energy_share', 'rudong_model_run',
+                              'rudong_fusion_run'],
+               missing_note='六层链的四步脚本未随包 (振动线分支 claude/vibration-data-diagnosis-32b69e); '
+                            '本脚本只做索引/谱/报告/知识库, 不冒充跑过那四步')
+    (m5 / 'vib_raw_manifest.json').write_text(json.dumps(man, ensure_ascii=False, indent=1), encoding='utf-8')
+
+    # 来源自登记: 让 _provenance.json 把这些件记成 raw-derived (见 src/derived_manifest.py 的说明)
+    try:
+        from src.derived_manifest import record as _rec
+        rels = _rels_of_window(final, m5)
+        if a.with_report:
+            for p in (P.cms(a.farm)).glob('报告_CMS振动状态评估报告_*.md'):
+                # 只登记"今天生成的"这一件 (随包/厂家转录的那些不冒充自算)
+                if time.strftime('%Y-%m-%d') in p.name:
+                    rels[f'windcms/{p.name}'] = 'scripts/windcms.py report (数据窗: ' + final.name + ')'
+        _rec(P.out_root(a.farm), rels, by='scripts/vib_raw_build.py')
+        print(f'  自登记 {len(rels)} 件产物 → {P.out_root(a.farm) / "_derived_manifest.json"}')
+    except Exception as _e:
+        print(f'  ⚠ 来源自登记失败 ({type(_e).__name__}: {_e}) —— 台账会漏记这些件, 但不影响产物本身')
+
+    print(f'\n完成: 用时 {time.time() - t0:.0f}s; 清单 → {m5 / "vib_raw_manifest.json"}')
+    print(f'  新窗 {final.name} 已进分析集: windcms 的谱图/标量/报告页会自动收录它 '
+          f'(消费者读取时现扫 m5\\windows\\w????\\index.parquet, 无需改代码)')
+    print('  windscada 融合面读 handoff (m5_cms_tcm/handoff_vibration_v2.json) —— 那是振动线出件, 本次不动它')
+    print('  看效果: python scripts/windcms.py serve --port 8020  → http://127.0.0.1:8020/')
+    return 0
+
+
+if __name__ == '__main__':
+    sys.exit(main())

+ 357 - 0
scripts/vib_reports_build.py

@@ -0,0 +1,357 @@
+#!/usr/bin/env python3
+# -*- coding: utf-8 -*-
+"""厂家振动报告 (docx) → 结构化提取 + 消费端格式的评估报告 (2026-09-12 用户令).
+
+## 范围与诚实边界 (先读这段, 再看代码)
+
+数据层「CMS 振动评估报告」的源件有两类, 处理方式**不同**:
+
+  · CMS 原始测量导出 (Brande TCM `*_decode.json`)  → 可重算, 走 scripts/vib_raw_build.py
+    (索引/谱/六层链), 这一路是本目录下 `report_CMS振动状态评估报告_*.md` 的**正路**。
+  · 厂家月度评估报告 (上海电气 12 份用印版 PDF + 2026年07月 docx; 大生科技传动链 docx)
+    → **PDF 全部是扫描件** (2026-09-12 用 pypdf 抽检: 12 页 12 图, extract_text 长度 0),
+      没有 OCR 就取不出数值 ⇒ 只登记归档, 不进判级;
+      **docx 有文本层**, 本脚本把里面的逐台判级**逐字转录**成结构化件与一份 report md。
+
+**转录纪律** (照抄本包既有做法, 防"看起来像算出来的"): 每个产物的 meta 里写清源文件+sha256+表头+行数;
+取值一律来自报告原文, 不补不猜; 报告没给的列写 '—' 并在说明里点明"厂家未给"; 聚合量(如综合级)
+必须标注它是**取严规则**得出的, 不是厂家判的。
+
+## 输出的两份东西 (都在 outputs/<场>/ 下, 与既有消费者兼容)
+
+  windcms/厂家报告提取_<报告期>.json    逐台原文 + provenance (给机器读/复核)
+  windcms/报告_CMS振动状态评估报告_<报告期>.md
+       与 windcms 自产报告**同构**: `## 附录 A` 下 `| WTGxx | 主轴承前 | 主轴承后 | 齿轮箱 | 发电机 | 综合 | …`
+       (消费者: src/windscada/subsys/fusion.py::windcms_grades 取第 6 列=综合,
+        src/windscada/taxonomy.py 取第 2..5 列)。日期用**报告期**, 于是它不会顶掉更新的自产报告,
+       但随包自产报告缺失时它就是最新版 (从零重算的机器上正好用得上)。
+  m5_cms_tcm/厂家报告提取_<报告期>.json + 报告_TCM传动链振动分析_<报告期>.md
+       大生科技 (TCM M-system 数据) 逐台状态等级/分析/建议, 原文透传; **不**冒充 handoff 判级。
+
+用法:
+    python scripts/vib_reports_build.py --dump     # 只打印解析结果 (人工核对用, 不写盘)
+    python scripts/vib_reports_build.py            # 摄入并写产物
+"""
+from __future__ import annotations
+
+import argparse
+import hashlib
+import json
+import pathlib
+import re
+import sys
+import time
+
+ROOT = pathlib.Path(__file__).resolve().parents[1]
+sys.path.insert(0, str(ROOT))
+from src.console import soft  # noqa: E402
+
+soft()
+
+STATE_ORDER = ['危险', '报警', '良好', '优秀', '不可判']      # 取严: 越靠前越严重 (报告用四级 + 不可判)
+NO_DATA_MARKS = ('--', '—', '无数据', '通讯中断', '无通讯')
+
+
+def sha256(p: pathlib.Path) -> str:
+    h = hashlib.sha256()
+    with open(p, 'rb') as f:
+        for chunk in iter(lambda: f.read(1 << 20), b''):
+            h.update(chunk)
+    return h.hexdigest()[:16]
+
+
+def grid_of(table):
+    """docx 表 → 二维网格, **正确还原合并单元格**。
+
+    为什么不能直接用 `row.cells`: python-docx 对合并区会把同一个 `tc` 在相邻列/相邻行重复给出 ——
+    照抄等于把 "优秀 | 优秀 | 优秀" 压成 "优秀", 列就对不齐了 (第一版手工 dump 就吃了这个亏:
+    WTG-01 行显示成 `WTG-01 | 优秀 | 无`, 而实际是 主轴承=优秀 齿轮箱=优秀 发电机=优秀 结论=无)。
+    规则: 同一行里 tc 与左邻相同 = 横向合并 (沿用左值); tc 与上方同行同列相同 = 纵向合并 (沿用上值)。
+    """
+    ncol = len(table.columns)
+    out, tc_above = [], [None] * ncol
+    for row in table.rows:
+        vals, tcs, prev = [], [], None
+        for ci, cell in enumerate(row.cells):
+            tc = cell._tc
+            if tc is prev:                      # 横向合并的续格
+                vals.append(vals[-1] if vals else '')
+                tcs.append(tc)
+                continue
+            if tc is tc_above[ci] and ci < len(tc_above):   # 纵向合并的续格
+                vals.append(out[-1][ci] if out else '')
+            else:
+                vals.append(cell.text.strip().replace('\n', ' '))
+            tcs.append(tc)
+            prev = tc
+        tc_above = tcs
+        out.append(vals)
+    return out
+
+
+def find_table(doc, must_have, header_hint=None):
+    """按表头关键字找表 → (index, grid)。must_have: 表头行必须都含这些词。"""
+    import docx
+    for i, tb in enumerate(doc.tables):
+        g = grid_of(tb)
+        if g and all(any(k in c for c in g[0]) for k in must_have):
+            return i, g
+    return None, None
+
+
+def norm_turbine(s):
+    m = re.search(r'WTG[\s_-]*0*(\d{1,2})', str(s).upper())
+    return f'WTG{int(m.group(1)):02d}' if m else None
+
+
+def state_of(s):
+    """文本 → 五级之一; 认不出给 '不可判' (宁可说不可判, 不冒充优秀)。"""
+    t = str(s or '').strip()
+    if any(k in t for k in NO_DATA_MARKS) or t == '':
+        return '不可判'
+    for v in ('危险', '报警', '预警', '良好', '优秀'):
+        if v in t:
+            return {'预警': '报警'}.get(v, v)          # 上海电气四级里没有"预警", 大生科技有; 此处按四级归一
+    if '异常' in t:
+        return '不可判'
+    return '不可判'
+
+
+def worst(*states):
+    """取严 (综合列的口径; 厂家未给综合列时必须写明这是我们的聚合规则)。"""
+    live = [s for s in states if s]
+    return min(live, key=lambda s: STATE_ORDER.index(s) if s in STATE_ORDER else 99) if live else '不可判'
+
+
+def parse_sa(path: pathlib.Path):
+    """上海电气月度振动分析报告: 表「机组号|主轴承|齿轮箱|发电机|诊断结论和维护建议」+ 总览句校验和。"""
+    import docx
+    d = docx.Document(str(path))
+    ti, g = find_table(d, ['机组号', '主轴承', '齿轮箱'])
+    if g is None:
+        return None
+    head = g[0]
+    idx = {name: next((i for i, c in enumerate(head) if name in c), None)
+           for name in ('机组号', '主轴承', '齿轮箱', '发电机', '诊断')}
+    rows = []
+    for r in g[1:]:
+        t = norm_turbine(r[idx['机组号']] if idx['机组号'] is not None else '')
+        if not t:
+            continue
+        rows.append(dict(turbine=t,
+                         主轴承=state_of(r[idx['主轴承']]) if idx['主轴承'] is not None else '不可判',
+                         齿轮箱=state_of(r[idx['齿轮箱']]) if idx['齿轮箱'] is not None else '不可判',
+                         发电机=state_of(r[idx['发电机']]) if idx['发电机'] is not None else '不可判',
+                         厂家结论=(r[idx['诊断']] if idx['诊断'] is not None and idx['诊断'] < len(r) else '')))
+    # 总览句 = 报告的**自带校验和** (2026-07 报告原文: 优秀34 良好3 预警0 无数据1 测点异常1)。
+    # ★它在**表格单元格**里, 不是段落 (第一版只扫 paragraphs → 拿到空串, 于是"校验和"形同虚设)。
+    overview, sums = '', {}
+    texts = [p.text for p in d.paragraphs]
+    for tb in d.tables:
+        for row in tb.rows:
+            for c in row.cells:
+                texts.append(c.text)
+    for t in texts:
+        if '检测结果' in t and '机组' in t and '台' in t:
+            overview = t.strip()
+            break
+    for k in ('优秀', '良好', '预警', '无数据', '测点异常'):
+        # 原话是"运行状态优秀机组34台" —— 等级词与数字之间夹着"机组"二字, 只写 `{k}\s*(\d+)` 会漏掉优秀
+        m = re.search(rf'{k}[^0-9]{{0,6}}(\d+)\s*台', overview)
+        if m:
+            sums[k] = int(m.group(1))
+    period = re.search(r'(20\d\d)\s*年\s*(\d{1,2})\s*月', path.name)
+    # 报告期: 文件名 2026年07月 → 用**月末** (报告覆盖整个月; 用月初会让它比同类报告显得更新)
+    import calendar
+    if period:
+        y, mo = int(period.group(1)), int(period.group(2))
+        period = f'{y:04d}-{mo:02d}-{calendar.monthrange(y, mo)[1]:02d}'
+    else:
+        period = time.strftime('%Y-%m-%d')
+    return dict(kind='上海电气月度', file=path.name, sha256=sha256(path), table_index=ti,
+                header=[c for c in head], period=period, overview=overview, overview_counts=sums,
+                rows=rows)
+
+
+def parse_ds(path: pathlib.Path):
+    """大生科技传动链振动分析报告: 表「机组号|状态等级|分析与结论|建议」+ 数据窗 (报告正文)。"""
+    import docx
+    d = docx.Document(str(path))
+    ti, g = find_table(d, ['机组号', '状态等级', '分析'])
+    if g is None:
+        return None
+    head = g[0]
+    idx = {name: next((i for i, c in enumerate(head) if name in c), None)
+           for name in ('机组号', '状态等级', '分析', '建议')}
+    rows = []
+    for r in g[1:]:
+        t = norm_turbine(r[idx['机组号']] if idx['机组号'] is not None else '')
+        if not t:
+            continue
+        rows.append(dict(turbine=t,
+                         状态等级=str(r[idx['状态等级']]).strip() if idx['状态等级'] is not None else '',
+                         分析与结论=r[idx['分析']] if idx['分析'] is not None else '',
+                         建议=r[idx['建议']] if idx['建议'] is not None else ''))
+    txt = '\n'.join(p.text for p in d.paragraphs)
+    # 数据窗: 正文原话 "本次分析采集了 A 至 B 期间的振动监测数据" (两个时间戳都要带时刻)
+    win = re.search(r'(\d{4}年\d{1,2}月\d{1,2}日\s*\d{1,2}:\d{2}:\d{2})\s*至\s*'
+                    r'(\d{4}年\d{1,2}月\d{1,2}日\s*\d{1,2}:\d{2}:\d{2})', txt)
+    rep = re.search(r'报告日期[::]\s*([\d.]{8,10})', txt)
+    period = (rep.group(1).replace('.', '-') if rep else time.strftime('%Y-%m-%d'))
+    if re.fullmatch(r'\d{4}-\d{1,2}-\d{1,2}', period):
+        y, m, dd = period.split('-')
+        period = f'{y}-{int(m):02d}-{int(dd):02d}'
+    return dict(kind='大生科技传动链', file=path.name, sha256=sha256(path), table_index=ti,
+                header=[c for c in head], period=period,
+                data_window=(f'{win.group(1)} ~ {win.group(2)}' if win else '未标'),
+                rows=rows)
+
+
+def md_sa(ext, src_note):
+    """上海电气提取 → 与 windcms 自产报告同构的评估报告 (附录 A 表)。"""
+    L = [f'# CMS 振动状态评估报告 (厂家件转录) {ext["period"]}', '',
+         f'> 来源: `{ext["file"]}` (sha256 {ext["sha256"]}, 表 {ext["table_index"]}) —— {src_note}', '',
+         f'> 报告原文总览: {ext["overview"] or "(未找到总览句)"}', '',
+         '> ★ 本表是从**厂家月度振动分析报告**逐字转录的(上海电气, 依据 VDI3834 与 NB/T 31129-2018):',
+         '> ① 厂家按「主轴承 / 齿轮箱 / 发电机」给级, **主轴承不分前后** —— 本表两列填同一值, 不是我们分的;',
+         '> ② 厂家**未给综合列**, 本表「综合」= 三项**取严**(危险>报警>良好>优秀>不可判) 的聚合规则结果;',
+         '> ③ 厂家未给「融合级/CMS 红黄/行动等级建议」, 一律写 `—` (不猜);',
+         '> ④ 判级细节(谱/门槛)以厂家报告与 CMS 原始导出分析为准, 本表只承接"哪台什么级"。', '',
+         '## 附录 A 38 台状态等级表(厂家报告转录,不代表确诊数量)', '',
+         '| 机组号 | 主轴承前 | 主轴承后 | 齿轮箱 | 发电机 | 综合 | 融合级(模型) | 厂家结论与建议 |',
+         '|---|---|---|---|---|---|---|---|']
+    for r in sorted(ext['rows'], key=lambda x: x['turbine']):
+        zh = worst(r['主轴承'], r['齿轮箱'], r['发电机'])
+        note = (r['厂家结论'] or '—').replace('|', '/')
+        L.append(f"| {r['turbine']} | {r['主轴承']} | {r['主轴承']} | {r['齿轮箱']} | {r['发电机']} | "
+                 f"{zh} | — | {note} |")
+    L += ['', '## 说明事项', '',
+          '1. 本件是**厂家报告的转录**, 不是观澜自算结果; 观澜自算报告 (CMS 原始导出 → 六层链) 见同目录 `报告_CMS振动状态评估报告_*.md`。',
+          '2. 厂家报告里的「无数据/通讯中断/测点异常」在本表记为 `不可判` —— 没测到 ≠ 正常。',
+          f'3. 报告期为 {ext["period"]} (文件名月份, 取月末)。', '']
+    return '\n'.join(L)
+
+
+def md_ds(ext, src_note):
+    L = [f'# TCM 传动链振动分析报告 (厂家件转录) {ext["period"]}', '',
+         f'> 来源: `{ext["file"]}` (sha256 {ext["sha256"]}, 表 {ext["table_index"]}) —— {src_note}', '',
+         f'> 数据窗: {ext["data_window"]}    报告日期: {ext["period"]}', '',
+         '> ★ 原文透传: 状态等级/分析与结论/建议按厂家报告逐字记录。**本件不构成 handoff 判级** ——',
+         '> handoff 的定谳链判级 (候选以上/证据窗) 是振动线接口件, 本报告只作 TCM 侧证据与交叉参考。', '',
+         '| 机组 | 状态等级 | 分析与结论 | 建议 |', '|---|---|---|---|']
+    for r in sorted(ext['rows'], key=lambda x: x['turbine']):
+        cells = [r['turbine'], r['状态等级'], r['分析与结论'] or '—', r['建议'] or '—']
+        L.append('| ' + ' | '.join(str(c).replace('|', '/') for c in cells) + ' |')
+    L += ['', '## 说明事项', '',
+          '1. 判据/图谱在厂家报告原文 (本包只落可解析的 docx; 若同批有 PDF 扫描件, 扫描件只归档不解析)。',
+          '2. 状态等级为厂家四级口径 (优秀/良好/预警/报警) + "数据异常"(系统故障无通讯), 与观澜五级不是同一套, 消费时别直接比。', '']
+    return '\n'.join(L)
+
+
+def main() -> int:
+    ap = argparse.ArgumentParser()
+    ap.add_argument('--farm', default='rudong')
+    ap.add_argument('--dump', action='store_true', help='只打印, 不写盘')
+    a = ap.parse_args()
+    from src.windscada.config import raw_station_dir
+    from src import paths as P
+
+    station = pathlib.Path(raw_station_dir(a.farm))
+    sa_dir = station / 'windcms' / '厂家报告' / '上海电气_月度'
+    ds_dir = station / 'm5_cms_tcm' / '厂家报告'
+    out_cms = P.cms(a.farm)
+    out_m5 = P.m5(a.farm)
+
+    sa_files = sorted(sa_dir.glob('*.docx')) if sa_dir.is_dir() else []
+    ds_files = sorted(ds_dir.glob('*.docx')) if ds_dir.is_dir() else []
+    print(f'上海电气 docx: {len(sa_files)} 件   {sa_dir}')
+    for p in sa_files:
+        print(f'   {p.name}  {p.stat().st_size / 1048576:.2f} MB')
+    print(f'大生科技 docx: {len(ds_files)} 件   {ds_dir}')
+    for p in ds_files:
+        print(f'   {p.name}  {p.stat().st_size / 1048576:.2f} MB')
+    pdfs = sorted(sa_dir.glob('*.pdf')) if sa_dir.is_dir() else []
+    if pdfs:
+        print(f'扫描件 PDF: {len(pdfs)} 件 (登记归档, 不解析)')
+    if not sa_files and not ds_files:
+        print('没有可解析的 docx 厂家报告 (只有扫描件 PDF 属正常) → 空跑退出')
+        return 0
+
+    written = []
+    for p in sa_files:
+        ext = parse_sa(p)
+        if not ext:
+            print(f'⚠ {p.name}: 没找到"机组检测结果列表"表 (跳过, 不猜)')
+            continue
+        zh = [worst(r['主轴承'], r['齿轮箱'], r['发电机']) for r in ext['rows']]
+        print(f'\n== {p.name} ({ext["kind"]}, 表{ext["table_index"]}) ==')
+        print(f'   报告期 {ext["period"]}  台数 {len(ext["rows"])}  原文总览: {ext["overview"]}')
+        from collections import Counter
+        got = Counter(zh)
+        print(f'   转录结果取严分布: {dict(got)}')
+        print(f'   厂家总览句: {ext["overview_counts"]}')
+        if a.dump:
+            for r in ext['rows'][:6]:
+                print('   ', r)
+            continue
+        out_cms.mkdir(parents=True, exist_ok=True)
+        jp = out_cms / f'厂家报告提取_{ext["period"]}.json'
+        jp.write_text(json.dumps(ext, ensure_ascii=False, indent=1), encoding='utf-8')
+        mp = out_cms / f'报告_CMS振动状态评估报告_{ext["period"]}.md'
+        mp.write_text(md_sa(ext, '现场包 厂家月度评估报告'), encoding='utf-8')
+        written += [jp, mp]
+        print(f'   → {jp.name}\n   → {mp.name}')
+    for p in ds_files:
+        ext = parse_ds(p)
+        if not ext:
+            print(f'⚠ {p.name}: 没找到"信号分析与结论描述"表 (跳过, 不猜)')
+            continue
+        print(f'\n== {p.name} ({ext["kind"]}, 表{ext["table_index"]}) ==')
+        print(f'   报告期 {ext["period"]}  数据窗 {ext["data_window"]}  台数 {len(ext["rows"])}')
+        from collections import Counter
+        print(f'   状态等级分布: {dict(Counter(r["状态等级"] for r in ext["rows"]))}')
+        if a.dump:
+            for r in ext['rows'][:4]:
+                print('   ', {k: (v[:60] + '…' if isinstance(v, str) and len(v) > 60 else v) for k, v in r.items()})
+            continue
+        out_m5.mkdir(parents=True, exist_ok=True)
+        jp = out_m5 / f'厂家报告提取_{ext["period"]}.json'
+        jp.write_text(json.dumps(ext, ensure_ascii=False, indent=1), encoding='utf-8')
+        mp = out_m5 / f'报告_TCM传动链振动分析_{ext["period"]}.md'
+        mp.write_text(md_ds(ext, '现场包 大生科技 TCM 传动链分析报告 (docx)'), encoding='utf-8')
+        written += [jp, mp]
+        print(f'   → {jp.name}\n   → {mp.name}')
+
+    if not a.dump and pdfs:
+        reg = dict(note='扫描件登记: 无文本层 → 只归档, 数值不进系统; 需数值须 OCR 或取厂家电子件',
+                   verified='2026-09-12 用 pypdf 抽检一件: 12 页 / 每页 1 图 / extract_text() 长度 0',
+                   files=[dict(name=q.name, bytes=q.stat().st_size, sha256=sha256(q)) for q in pdfs])
+        rp = sa_dir / '_扫描件清单.json'
+        rp.write_text(json.dumps(reg, ensure_ascii=False, indent=1), encoding='utf-8')
+        written.append(rp)
+        print(f'\n扫描件登记 → {rp}')
+
+    # 来源自登记 (见 src/derived_manifest.py): 转录件也是"从 data/raw 来的", 台账要记 raw-derived
+    if not a.dump and written:
+        try:
+            from src.derived_manifest import record as _rec
+            from src import paths as _P
+            rels = {}
+            for p in written:
+                if p.suffix.lower() not in ('.json', '.md') or '_扫描件清单' in p.name:
+                    continue
+                if p.parent == out_cms:
+                    rels[f'windcms/{p.name}'] = 'scripts/vib_reports_build.py (厂家 docx 报告逐台转录)'
+                elif p.parent == out_m5:
+                    rels[f'm5_cms_tcm/{p.name}'] = 'scripts/vib_reports_build.py (厂家 docx 报告逐台转录)'
+            if rels:
+                _rec(_P.out_root(a.farm), rels, by='scripts/vib_reports_build.py')
+                print(f'自登记 {len(rels)} 件 → {_P.out_root(a.farm) / "_derived_manifest.json"}')
+        except Exception as _e:
+            print(f'⚠ 来源自登记失败 ({type(_e).__name__}: {_e})')
+    print(f'\n完成: 写出 {len(written)} 件')
+    return 0
+
+
+if __name__ == '__main__':
+    sys.exit(main())

+ 113 - 27
scripts/windscada_serve.py

@@ -9,6 +9,7 @@ import sys as _sys, pathlib as _plb
 _sys.path.insert(0, str(_plb.Path(__file__).resolve().parents[1]))
 from src import paths as P
 import sys, json, pathlib, re, threading, urllib.parse, collections
+import hashlib, time
 sys.path.insert(0, str(pathlib.Path(__file__).resolve().parents[1]))     # 安装根入 sys.path (先于 src.* 导入)
 from src import paths as _P                                              # 路径唯一真源 (与 cwd 无关)
 import numpy as np, pandas as pd
@@ -511,6 +512,56 @@ LOCK = threading.Lock()
 _CACHE = {}
 
 
+# ── 产物指纹 → **运行中自动重载** (2026-09-16 用户令: "重算后的产物可即时用于页面呈现") ──────
+# 背景 (实测): _load() 把产物**一次性**读进内存, 之后只在启动时读一次 ⇒ 重算完成后页面仍显示旧数,
+# 必须重启组件服务才更新。判定依据: 改产物后 /detail/api/fleet 等仍回旧值, 重启组件后才变。
+# 做法: 取"页面取数依赖的产物文件"的 (路径, mtime_ns, size) 摘要当指纹; _load() 前比一次
+# (带 TTL, 免得每请求都 stat 一遍), 指纹变了就重载 ⇒ 重算产物**不再需要重启服务**。
+# 也提供 /api/reload 显式重载 (运维控制台重算结束后可点一下)。
+_STAMP_TTL = 2.0          # 秒: 指纹有效期 (stat ~150 个文件 ≈ 1~3 ms, 不值得每请求都做)
+_STAMP = {'val': None, 'at': 0.0}
+
+
+def _product_files():
+    """页面取数依赖的产物文件清单 (顺序稳定: 供指纹; 只列**产物**, 不含 reference/ 随包契约)。
+
+    覆盖 _load() 直读的件 + 它经 taxonomy/temp_nbm/hydraulic/fusion 间接读的件:
+      windscada/*.parquet|csv|json  · ontology/*.json  · pitch/*.parquet
+      m5_cms_tcm/{handoff_vibration_v2,component_history,baseline_38}.json  · windcms/报告_CMS*.md
+    """
+    try:
+        from src.windcms.config import cms_out as _cms_out     # 与写侧同一解析口 (WINDCMS_OUT)
+        _cms_dir = _cms_out()
+    except Exception:
+        _cms_dir = _P.cms()
+    pats = ((ST, ('*.parquet', '*.csv', '*.json')),
+            (_P.ont(), ('*.json',)),
+            (_P.pitch(), ('*.parquet',)),
+            (_P.m5(), ('handoff_vibration_v2.json', 'component_history.json', 'baseline_38.json')),
+            (_cms_dir, ('报告_CMS振动状态评估报告_*.md',)))
+    files = []
+    for d, ps in pats:
+        for pat in ps:
+            files.extend(sorted(d.glob(pat)))
+    return files
+
+
+def products_stamp(force=False):
+    """产物指纹 (sha1 前 16 位)。force=True 时忽略 TTL 立即重算。"""
+    now = time.time()
+    if not force and _STAMP['val'] is not None and (now - _STAMP['at']) < _STAMP_TTL:
+        return _STAMP['val']
+    h = hashlib.sha1()
+    for p in _product_files():
+        try:
+            st = p.stat()
+            h.update(f'{_P.rel(p)}|{st.st_mtime_ns}|{st.st_size}\n'.encode('utf-8'))
+        except OSError:
+            h.update(f'{_P.rel(p)}|MISSING\n'.encode('utf-8'))
+    _STAMP.update(val=h.hexdigest()[:16], at=now)
+    return _STAMP['val']
+
+
 class ProductsMissing(RuntimeError):
     """产物缺失 (outputs/<场>/… 被清空或尚未生成)。
 
@@ -524,34 +575,55 @@ class ProductsMissing(RuntimeError):
         self.what, self.path = name, path
 
 
+def reload_products(reason=''):
+    """显式清缓存 (下一次请求重载)。返回清前的指纹, 供日志/接口回显。"""
+    with LOCK:
+        old = _CACHE.get('__stamp')
+        _CACHE.clear()
+    print(f'[reload] 清产物缓存 ({reason or "手动"}) 旧指纹={old}', flush=True)
+    return old
+
+
 def _load():
     with LOCK:
-        if _CACHE.get('__loaded'): return
+        stamp = products_stamp()
+        if _CACHE.get('__loaded'):
+            if _CACHE.get('__stamp') == stamp:
+                return
+            # 产物变了 (典型: 刚跑完重算) → 重载。**先建后换**: 下面任何一步抛异常都不会破坏旧缓存,
+            # 请求照旧能用旧数 (降级但不空白), 同时日志留痕。
+            print(f'[reload] 产物指纹变化 {_CACHE.get("__stamp")} → {stamp}, 重载', flush=True)
         tmp = {}
-        for key, name in (('tm', 'temp_monthly.parquet'), ('al', 'alarms.parquet'),
-                          ('lm', 'loss_monthly.parquet'), ('bins', 'powercurve_bins.parquet'),
-                          ('pcd', 'powercurve_dev.parquet')):
-            f = ST / name
-            if not f.exists():
-                raise ProductsMissing(name, f)
-            tmp[key] = pd.read_parquet(f)
-        tmp['al']['month'] = tmp['al']['t_on'].dt.to_period('M').astype(str)
-        tmp['pcd'] = tmp['pcd'].set_index('turbine')
-        for key, name in (('duty', 'duty_monthly.parquet'),):
-            f = ST / name
-            tmp[key] = pd.read_parquet(f) if f.exists() else None
         try:
-            tmp['sysmx'] = taxonomy.system_matrix()
-            tmp['treg'] = temp_nbm.registry()
-            tmp['treg'] = tmp['treg'][0] if isinstance(tmp['treg'], tuple) else tmp['treg']
-            tmp['hyd'], _ = hydraulic.registry()
-            tmp['hyd'] = tmp['hyd'].set_index('turbine')
-        except FileNotFoundError as e:            # 这些派生件同样在产物仓里; 缺了就按"无产物"处理
-            raise ProductsMissing(getattr(e, 'filename', '派生产物'), getattr(e, 'filename', ST))
-        zp = _P.pitch() / 'pitch_zero_monthly.parquet'
-        tmp['zero'] = pd.read_parquet(zp) if zp.exists() else None
-        tmp['__loaded'] = True
-        _CACHE.update(tmp)                        # 原子提交: 失败时不留下半截缓存
+            for key, name in (('tm', 'temp_monthly.parquet'), ('al', 'alarms.parquet'),
+                              ('lm', 'loss_monthly.parquet'), ('bins', 'powercurve_bins.parquet'),
+                              ('pcd', 'powercurve_dev.parquet')):
+                f = ST / name
+                if not f.exists():
+                    raise ProductsMissing(name, f)
+                tmp[key] = pd.read_parquet(f)
+            tmp['al']['month'] = tmp['al']['t_on'].dt.to_period('M').astype(str)
+            tmp['pcd'] = tmp['pcd'].set_index('turbine')
+            for key, name in (('duty', 'duty_monthly.parquet'),):
+                f = ST / name
+                tmp[key] = pd.read_parquet(f) if f.exists() else None
+            try:
+                tmp['sysmx'] = taxonomy.system_matrix()
+                tmp['treg'] = temp_nbm.registry()
+                tmp['treg'] = tmp['treg'][0] if isinstance(tmp['treg'], tuple) else tmp['treg']
+                tmp['hyd'], _ = hydraulic.registry()
+                tmp['hyd'] = tmp['hyd'].set_index('turbine')
+            except FileNotFoundError as e:            # 这些派生件同样在产物仓里; 缺了就按"无产物"处理
+                raise ProductsMissing(getattr(e, 'filename', '派生产物'), getattr(e, 'filename', ST))
+            zp = _P.pitch() / 'pitch_zero_monthly.parquet'
+            tmp['zero'] = pd.read_parquet(zp) if zp.exists() else None
+            tmp['__stamp'] = stamp
+            tmp['__loaded'] = True
+            _CACHE.update(tmp)                        # 原子提交: 失败时不留下半截缓存
+        except Exception:
+            if _CACHE.get('__loaded'):
+                print('[reload] 重载失败 → 继续用上一份缓存 (页面不会空白, 但数是旧的; 看上面的异常)', flush=True)
+            raise
 
 WINDOWS = ['近30日', '近90日', '2026年', '2026H1', '2025H2', '全程']
 def months_of(win):
@@ -1397,7 +1469,10 @@ def _ask_worker(q, model, lang='zh'):
             cmd.append(q + ' Answer in English only. End every conclusion sentence with [object/id]. Finally run claim_check.')
         else:
             cmd.append(q + ' 每个结论句末尾用[对象id]标注来源。最后调用claim_check。')
-        r = subprocess.run(cmd, capture_output=True, text=True, timeout=600)
+        # 2026-09-16: 走 src.proc.run —— 本函数跑在 detail 服务进程里 (无可见控制台), 裸 spawn
+        # 会让 Windows 给这条"云端档问答"新建可见控制台窗口 (本地档不走这里)。
+        from src import proc as _proc
+        r = _proc.run(cmd, capture_output=True, text=True, timeout=600)
         ans = r.stdout.strip() or ('(空输出) stderr: ' + r.stderr[-500:])
         out.write_text(ans, encoding='utf-8')
         sys.path.insert(0, str(pathlib.Path('scripts').resolve()))
@@ -5031,6 +5106,14 @@ class H(BaseHTTPRequestHandler):
                         self._send('404', code=404)
                 else:
                     self._send('403', code=403)
+            elif u.path == '/api/reload':
+                # 产物重载 (2026-09-16 用户令: 重算产物即时可用于页面呈现)。
+                # 平时**不需要**调它 —— _load() 每次请求都会比产物指纹, 变了自动重载;
+                # 这个端点给"重算结束想立刻生效"的显式动作, 也便于运维界面/脚本确认当前指纹。
+                _old = reload_products(urllib.parse.parse_qs(u.query).get('why', ['manual'])[0])
+                self._send(_jdump(dict(ok=True, old_stamp=_old, stamp=products_stamp(force=True),
+                                       files=len(_product_files())), ensure_ascii=False),
+                           'application/json; charset=utf-8')
             else:
                 self._send('404', code=404)
         except ProductsMissing as e:
@@ -5038,8 +5121,10 @@ class H(BaseHTTPRequestHandler):
             self._send(_jdump(dict(kind='multiline', title='无产物', unit='', months=[], series=[],
                                    err='no_products',
                                    note=f'{e.what} 不存在 → {_P.rel(e.path)}; 先跑 '
-                                        f'scripts/rebuild_from_raw.py (SCADA 侧加 --scada) 生成产物, '
-                                        f'或用 scripts/products_state.py --on 还原'), ensure_ascii=False),
+                                        f'scripts/rebuild_from_raw.py (SCADA 侧加 --scada) 生成产物; '
+                                        f'"包内没有生成端"的件从交付包补齐: '
+                                        f'python scripts/products_restore_missing.py --stash <交付包.zip> '
+                                        f'(2026-09-16 起清除产物不留备份, 故不再有 --on 还原)'), ensure_ascii=False),
                        'application/json; charset=utf-8')
         except Exception as e:
             import traceback
@@ -5075,6 +5160,7 @@ if __name__ == '__main__':
         _me.CFG = _cfg.farm()
         _me.ST = pathlib.Path(_me.CFG['store'])
         _CACHE.clear()
+        _STAMP['val'] = None            # 换场后产物指纹必须重算, 否则新场的缓存判为"未变化"
         print(f'场: {_me.CFG["name"]} ({a.farm}, {_me.CFG["n_turbines"]} 台, 仓 {_me.ST})')
     else:
         print(f'场: {CFG["name"]} (默认; 可用 {list(__import__("src.windscada.config", fromlist=["available"]).available())})')

+ 82 - 0
src/derived_manifest.py

@@ -0,0 +1,82 @@
+# -*- coding: utf-8 -*-
+"""产物来源**自登记**: 构建脚本落盘后把"这一件是我从 data/raw 算出来的"记进 `outputs/<场>/_derived_manifest.json`。
+
+## 为什么要有它
+
+`outputs/<场>/_provenance.json` 是**逐件来源台账**(raw-derived = 由 data/raw 重算 / shipped = 包内无生成端,
+用随包件补齐)。它由 `scripts/products_restore_missing.py` 生成, 而那个脚本是按**随包快照**逐件走一遍的 ——
+于是**新造的、快照里根本没有的产物**不会自动进台账 (既不算 raw-derived 也不算 shipped)。
+
+早期的做法是在 `products_restore_missing.py` 里维护一张 `RAW_DERIVED` 精确路径表。对"件数少、名字固定"
+的产物够用; 但振动侧的产物是 `<窗>/index.parquet` + `<窗>/spectra/*.npz`(分片名带序号) —— 窗名与分片数
+都随数据变, 写不进精确表, 而**按名字通配**又会误伤同名旧件 (例如 `报告_CMS振动状态评估报告_*.md`
+既有随包/自产的、也有厂家报告转录的, 名字形态一样)。
+
+所以改成**自登记**: 谁算的谁登记, 台账只认这份登记。名字对不上不是问题, 因为登记的是**相对路径本身**。
+
+用法 (构建脚本内):
+    from src.derived_manifest import record
+    record(P.out_root('rudong'), {rel: 'scripts/rudong_tcm_index.py (54 列, 与包内 tcm_index.parquet 同构)'},
+           by='scripts/vib_raw_build.py')
+"""
+from __future__ import annotations
+
+import json
+import pathlib
+import time
+
+FILENAME = '_derived_manifest.json'
+
+
+def path_of(store_root) -> pathlib.Path:
+    return pathlib.Path(store_root) / FILENAME
+
+
+def load(store_root) -> dict:
+    p = path_of(store_root)
+    if not p.exists():
+        return {}
+    try:
+        return json.loads(p.read_text(encoding='utf-8'))
+    except Exception:
+        return {}
+
+
+def prune(store_root) -> int:
+    """删掉**登记了但盘上已不存在**的条目, 返回删除数。
+
+    为什么需要 (2026-09-16 实逮): 振动摄入对同一批数据重跑时会落 `<窗>_reimport_<时分>` 窗
+    (设计如此, 该窗被 `data.EXCLUDE_DEFAULT` 排除在生产集外), 而登记是**追加式**的 ——
+    只补不删。重算几次后 `_derived_manifest.json` 里就攒了成百上千条指向已删目录的条目,
+    `_provenance.json` 的 raw-derived 计数随之虚增 (实测 1,740 → 3,444, 而盘上并没有多出这些件)。
+    台账是本包的"来源正本", 虚高等于说假话 ⇒ 每次生成台账前先 prune。
+    """
+    store_root = pathlib.Path(store_root)
+    cur = load(store_root)
+    files = cur.get('files') or {}
+    keep = {rel: v for rel, v in files.items() if (store_root / rel).exists()}
+    gone = len(files) - len(keep)
+    if gone:
+        cur['files'] = keep
+        cur['pruned'] = f'{time.strftime("%Y-%m-%d %H:%M")} 清理 {gone} 条不在盘的登记'
+        path_of(store_root).write_text(json.dumps(cur, ensure_ascii=False, indent=1), encoding='utf-8')
+    return gone
+
+
+def record(store_root, files: dict, by: str) -> pathlib.Path:
+    """把 {相对产物仓的路径: 构建器说明} 合并进登记 (幂等: 同路径后写覆盖先写)。
+
+    幂等很关键 —— 重跑摄入不该让登记无限膨胀; 同时**不删**别的构建器登记的条目
+    (振动摄入与厂家报告摄入是两个脚本, 各登各的)。"""
+    store_root = pathlib.Path(store_root)
+    cur = load(store_root)
+    entries = cur.get('files') or {}
+    for rel, builder in files.items():
+        entries[pathlib.Path(rel).as_posix()] = dict(builder=builder, by=by,
+                                                     at=time.strftime('%Y-%m-%d %H:%M:%S'))
+    cur = dict(note='产物来源自登记: 由构建脚本落盘后写入; _provenance.json 生成时把这些件记为 raw-derived',
+               at=time.strftime('%Y-%m-%d %H:%M:%S'), files=entries)
+    p = path_of(store_root)
+    p.parent.mkdir(parents=True, exist_ok=True)
+    p.write_text(json.dumps(cur, ensure_ascii=False, indent=1), encoding='utf-8')
+    return p

+ 21 - 16
src/ontology/maintenance.py

@@ -15,18 +15,15 @@ from .. import paths as P
 import re
 import json, os, pathlib, sys, datetime as _dt
 
-ROOT = pathlib.Path(__file__).resolve().parents[2]   # v2: 不再依赖 cwd (xzy 测试报告)
+ROOT = P.ROOT                     # 2026-09-16 统一: 原为 `parents[2]` 自推一份 (与本模块的 P.ROOT 同值,
+                                  # 但属"影子真源"; 同一文件里还并存自写的 disp()/_RAW_STR —— 一并归到 src/paths.py)
 ST = P.store()
 ONT = P.ont()
 CMS = P.cms()
-# 原始件目录: 由 src.windscada.config 统一给 (env WINDSCADA_RUDONG_SRC / configs/serve.json raw_dir 可覆盖;
-# 默认 <安装目录>/data/raw)。2026-09-08 xzy 测试逮: 原为开发机绝对路径 /Users/yuanying/rudong/…,
+# 原始件目录: 唯一真源就是 src/paths.py 的 RAW_ROOT (env WINDSCADA_RUDONG_SRC / configs/serve.json
+# 的 raw_dir 最终都经它生效)。2026-09-08 xzy 测试逮: 原为开发机绝对路径 /Users/yuanying/rudong/…,
 # 在 Windows 上既非法也不存在, 页面照着抄必然补不了数据。
-try:
-    from src.windscada.config import RUDONG_SRC as _RAW_STR
-except Exception:
-    _RAW_STR = os.environ.get('WINDSCADA_RUDONG_SRC') or str(ROOT / 'data' / 'raw')
-RAW = pathlib.Path(_RAW_STR)
+RAW = P.RAW_ROOT
 # ★2026-09-11 用户令 A2 + 场站扫描辨识: 数据层源在 data/raw/<场站名称>/ 下。场站目录名不再写死在
 #   代码里 —— config.farm() 扫 data/raw 的下一级目录辨识 (辨识依据见 station_note), 这里取结果用。
 try:
@@ -44,11 +41,11 @@ TECH = RAW / '西门子4.0技术资料'
 
 
 def disp(p) -> str:
-    """给人看的路径: 安装目录内的写成 <安装目录>/…, 其余给绝对路径; 分隔符随本机 (Windows 上是反斜杠).
+    """给人看的路径 (安装目录内 → `<安装目录>/…`) —— 2026-09-16 起直接转发 src/paths.py 的同名函数。
+
+    原来本模块自己写了一份 (语义相同但少一步 resolve), 与 P.disp 并存 = 同一份"显示规则"两个实现;
     2026-09-09 用户看页面反馈: 直接摆一串本机绝对路径, 客户读不出该往哪放。"""
-    p = pathlib.Path(p)
-    try: return str(pathlib.Path('<安装目录>') / p.relative_to(ROOT))
-    except ValueError: return str(p)
+    return P.disp(p)
 
 
 def _mtime(p):
@@ -114,11 +111,19 @@ SOURCES = [
          说明='SGS + 华标两家, 报告按台号/部件分目录; 文件名自带日期/台号/部件/sample_id; '
               '新一轮到货后重跑, 时效胶囊自动翻绿'),
     dict(层='数据', 名='CMS 振动评估报告', 产物=CMS / '报告_CMS振动状态评估报告_*.md', 时间列=None,
-         位置=disp(CMS), 摄入='windcms analyze (振动线会话)',
-         频率='月/事件驱动', 责任='振动线', 说明='设备状态五级来源; 观澜自动取最新版不写死日期'),
+         位置=P.disp_dir(STATION / 'windcms'),
+         摄入='python scripts/vib_raw_build.py  (索引→谱; 重生成 CMS 报告用 --with-report, 见 docs 振动数据接入 §3b)',
+         频率='月/事件驱动', 责任='振动线',
+         说明='本目录两类源件: ①CMS 原始测量导出 (Brande TCM *_decode.json → 索引/谱/六层链, 可重算); '
+              '②厂商月度评估报告 (上海电气 12 份用印版 PDF 实测为扫描件无文本层 ⇒ 只作归档证据, 不进判级)。'
+              '设备状态五级来源; 观澜自动取最新版不写死日期'),
     dict(层='数据', 名='振动线 handoff', 产物=P.m5() / 'handoff_vibration_v2.json', 时间列=None,
-         位置=disp(P.m5()), 摄入='振动线出件 → fusion 自动读',
-         频率='事件驱动', 责任='振动线', 说明='定谳链判级 (候选以上); 与 CMS 报告是两条轴, 取严合并'),
+         位置=P.disp_dir(STATION / 'm5_cms_tcm'),
+         摄入='振动线出件 → fusion 自动读 (现场正本落本目录则优先采用)',
+         频率='事件驱动', 责任='振动线',
+         说明='定谳链判级 (候选以上); 与 CMS 报告是两条轴, 取严合并。本目录放 TCM 侧深度分析报告 '
+              '(大生科技) 与现场给的接口正本 (handoff_vibration_v2.json / component_history.json); '
+              '包内那份 json 目前是 shipped 快照 (振动线分支的产物未随包), 现场给正本后即改由现场件驱动'),
     # ── 机理层 ──
     dict(层='机理', 名='报警码表与处置手册', 产物=ONT / 'objects.json', 时间列=None,
          位置=disp(TECH / '故障处理/故障处理手册.xlsx'), 摄入='python -m src.ontology.kb_ingest',

+ 18 - 0
src/paths.py

@@ -107,6 +107,24 @@ def pitch(name: str | None = None) -> pathlib.Path:
     return out_root(name) / 'pitch'
 
 
+def paradigm(name: str | None = None) -> pathlib.Path:
+    """范式实验件 (E3/E5/E8 底稿) —— 事实契约的输入之一 (2026-09-16 补: 原先直接用
+    `ROOT/'outputs'/'rudong'/'paradigm_r1'` 拼, 既写死场名又绕过了本模块)。"""
+    return out_root(name) / 'paradigm_r1'
+
+
+def report_dir(name: str | None = None) -> pathlib.Path:
+    """报告交付件目录 (`交接_振动→状态评估报告_*.md` / `现场单_*.md`) —— 由振动线出件,
+    并被 `src/windcms/config.py` 的 knowledge_docs 引用 (2026-09-16 补: 该目录在 v0.2.0 里
+    没有明确归属, 一直以 `out_root()/'report'` 的裸拼形式出现)。"""
+    return out_root(name) / 'report'
+
+
+def cloud(name: str | None = None) -> pathlib.Path:
+    """可上云面孔 (脱敏后的契约/派生件/页面) —— `scripts/guanlan_cloud_*.py` 的落点。"""
+    return guanlan(name) / 'cloud'
+
+
 def contract(name: str | None = None) -> pathlib.Path:
     """场契约 (机型判据参数), 属 reference 侧, 不在 outputs。"""
     return REFERENCE / farm(name) / 'windscada_contract.yaml'

+ 156 - 0
src/proc.py

@@ -0,0 +1,156 @@
+# -*- coding: utf-8 -*-
+"""子进程创建的统一口径 —— **不弹命令窗口** (2026-09-16 用户令: 启动/操作观澜时不弹命令窗口)。
+
+## 为什么要收敛到一处
+
+Windows 上"**无控制台的父进程** + 裸 spawn 一个控制台程序" = 系统给子进程**新建一个可见控制台窗口**。
+本项目里这类父进程很多:
+  · 网关 `guanlan_gateway.py` 自己是被 `DETACHED_PROCESS` 起来的 (无控制台) → 它调的 `tasklist` / `git` 会闪窗;
+  · 运维动作进程 (`scripts/_ops_launch.py` → `_ops_run.py`) 同样无控制台 → 从页面点"重算"时,
+    动作全过程都在一个可见窗口里跑 (最长 20 分钟), 页面每 2 s 轮询 `tasklist` 还会**反复闪窗**;
+  · 组件服务原先有的地方用 `DETACHED_PROCESS` (子进程干脆没有控制台), 有的地方什么都不加 (于是弹窗),
+    两种写法混用 —— 实测现场会攒下多个标题为 `.venv\\Scripts\\python.exe` 的黑窗。
+
+统一到本模块后: 只认 `NO_WINDOW` (`CREATE_NO_WINDOW`) 一种写法。
+★ `CREATE_NO_WINDOW` 与 `DETACHED_PROCESS` **互斥**, 不要叠加 —— 前者是"给一个没有窗口的控制台",
+  后者是"不给控制台"; 叠在一起行为依赖 Windows 版本。需要"子进程活过父进程"时用
+  `NEW_GROUP` (`CREATE_NEW_PROCESS_GROUP`) + 不共享控制台即可, 不需要 DETACHED。
+
+## 日志去哪了 (hide 窗口不等于看不见)
+
+窗口藏起来后, 子进程的 stdout/stderr 一律重定向到 `logs/<name>.log` (`spawn(log=…)`),
+`/ops` 页面也会显示任务日志尾巴 —— 排障路径不变, 只是不再靠一个黑窗。
+
+## 用法
+
+    from src.proc import spawn, run, NO_WINDOW
+    pid = spawn([py, 'scripts/x.py'], log=P.LOGS / 'x.log', env=e, cwd=ROOT)   # 后台, 不弹窗
+    r = run(['tasklist', '/FI', f'PID eq {pid}'], capture_output=True, text=True)  # 等待, 不弹窗
+"""
+from __future__ import annotations
+
+import os
+import pathlib
+import subprocess
+import sys
+
+WIN = os.name == 'nt'
+# CREATE_NO_WINDOW = 0x08000000: 给子进程一个**没有窗口**的控制台 (stdio 仍可重定向)
+NO_WINDOW = 0x08000000 if WIN else 0
+# CREATE_NEW_PROCESS_GROUP = 0x00000200: 子进程不受父进程 Ctrl-C 影响, 也不共享父的控制台事件
+NEW_GROUP = 0x00000200 if WIN else 0
+
+ROOT = pathlib.Path(__file__).resolve().parents[1]
+
+
+def flags(*, new_group: bool = True) -> int:
+    """本模块的统一 creationflags (非 Windows 返回 0, 调用方无需分支)。"""
+    f = NO_WINDOW
+    if new_group:
+        f |= NEW_GROUP
+    return f
+
+
+def _kw(kw: dict, *, new_group: bool = True) -> dict:
+    if WIN:
+        kw.setdefault('creationflags', flags(new_group=new_group))
+    else:
+        kw.setdefault('start_new_session', True)
+    return kw
+
+
+def _inherit_stdio(kw: dict) -> dict:
+    """把父进程**当前**的标准句柄显式交给子进程 (STARTF_USESTDHANDLES)。
+
+    为什么必须显式 (2026-09-16 实逮, 是本模块第一版引入的坑): Windows 上 `CREATE_NO_WINDOW` 会给
+    子进程新建一个"没有窗口的控制台", 而**新建控制台会把该进程的标准句柄重指到新控制台的缓冲区** ——
+    于是"父进程 stdout 已被重定向到日志文件"这件事, 在**孙子辈**就丢了。
+    现场表现: 从页面点"执行重算" → `_ops_launch`(显式 stdout=日志) → `_ops_run`(输出进了日志 ✔)
+    → `rebuild_all.py`(没显式句柄 → 输出掉进那个隐形控制台, 页面只剩"运行中"、日志里一个字都没有)。
+    显式传 `sys.stdout/stderr` 即可让整条链写进同一个日志; `sys.stdout is None` (pythonw) 时跳过。
+    """
+    if kw.get('capture_output'):
+        # ★ `capture_output=True` 与显式 `stdout=` **互斥**, 同给会抛 ValueError。
+        #   2026-09-16 实逮: 忘了这一步 → guanlan_ops.job_running() 里的 `tasklist` 每次都抛,
+        #   异常被 except 吞成 alive=False ⇒ **任何在跑的重算都被立刻改写成"被强杀"**,
+        #   页面显示完成、重算按钮重新可点(可能并发起两个重算)。
+        #   这里什么都不加就对了; 若调用方自己又传了 stdout, 让 Python 照旧抛错 (不替它吞)。
+        return kw
+    if 'stdout' not in kw:
+        s = sys.stdout
+        if s is not None and hasattr(s, 'fileno'):
+            try:
+                s.fileno()
+                kw['stdout'] = s
+            except Exception:
+                pass
+    if 'stderr' not in kw:
+        s = sys.stderr
+        if s is not None and hasattr(s, 'fileno'):
+            try:
+                s.fileno()
+                kw['stderr'] = s            # 不合并到 stdout: 让调用方自己决定 (spawn(log=) 才合并)
+            except Exception:
+                pass
+    return kw
+
+
+def spawn(cmd, log=None, env=None, cwd=None, *, new_group: bool = True, stdin_devnull: bool = True, **popen_kw):
+    """后台起一个**无窗口**子进程。`log` 给路径时把 stdout/stderr 追加进该文件。
+
+    `**popen_kw` 透传给 `Popen` —— 调用方要自己给 `stdout=`/`stderr=` 句柄时 (例如"日志由子进程
+    自己写、父进程只留 fd") 用得上; 给了就**不覆盖**它。没给则继承父进程当前的标准句柄 (见 _inherit_stdio)。
+    → Popen 对象 (拿 `.pid`)。
+    """
+    kw = dict(cwd=str(cwd or ROOT), env=env)
+    if stdin_devnull and 'stdin' not in popen_kw:
+        kw['stdin'] = subprocess.DEVNULL
+    fh = None
+    if log is not None and 'stdout' not in popen_kw:
+        log = pathlib.Path(log)
+        log.parent.mkdir(parents=True, exist_ok=True)
+        fh = open(log, 'ab')
+        kw['stdout'] = fh
+        kw['stderr'] = subprocess.STDOUT
+    kw.update(popen_kw)
+    _inherit_stdio(kw)
+    _kw(kw, new_group=new_group)
+    p = subprocess.Popen([str(c) for c in cmd], **kw)
+    if fh is not None:
+        # 父进程不一定等子进程结束; 句柄由子进程持有, 这里不要 close 掉它 —— 交给 GC/进程退出即可。
+        p._dsh_log = fh          # 仅作引用保存, 避免过早回收
+    return p
+
+
+def run(cmd, *, new_group: bool = True, **kw):
+    """前台等待的 `subprocess.run`, 但**不弹窗** (tasklist / git / 短命令都该走这里)。
+
+    ★ 文本解码请显式带 `errors='replace'`: `tasklist` 的输出是**控制台代码页**(中文 Windows = GBK),
+      而本进程可能是 PYTHONUTF8=1 起的 (默认文本编码 UTF-8) → 解码失败会让 `r.stdout` 变成 None,
+      调用方再 `str(pid) in r.stdout` 就 TypeError (2026-09-12 实逮过)。
+    ★ 未显式给 stdout/stderr 时继承父进程当前句柄 (见 _inherit_stdio) —— 否则 `CREATE_NO_WINDOW`
+      新建的控制台会把子进程输出"吸走", 日志里什么都看不到。
+    """
+    _inherit_stdio(kw)
+    _kw(kw, new_group=new_group)
+    return subprocess.run([str(c) for c in cmd], **kw)
+
+
+def run_text(cmd, **kw):
+    """`run` + text=True + errors='replace' (控制台输出来源的默认姿势)。"""
+    kw.setdefault('capture_output', True)
+    kw.setdefault('text', True)
+    kw.setdefault('errors', 'replace')
+    return run(cmd, **kw)
+
+
+def python_exe() -> str:
+    """跑脚本用的解释器 (venv 优先)。放这里是为了让调用方一处取值, 别再各写一遍。"""
+    try:
+        from src import paths as P
+        v = P.venv_python()
+        if v:
+            return str(v)
+    except Exception:
+        pass
+    return sys.executable

+ 29 - 2
src/sop/wrapup.py

@@ -154,9 +154,36 @@ if __name__ == '__main__':
 
     # 柱2 day_costs: cleaned parquet 场中位功率 × price.py 参数化价 (config 驱动 + per月季节 + provenance)
     day_costs, provenance = None, None
-    pq = P.out_root(a.farm) / 'cleaned' / 'turbine.parquet'
+    # ★ 写读要配对 (2026-09-16 修): 写侧 `src/sop/contract_gate.py:131` 落的是 `cleaned/<section>.parquet`
+    #   (section 来自契约配置, 可能是 turbine / turbine_10min_cnt / …), 而这里原先**硬编码** 'turbine.parquet'
+    #   ⇒ section 不是 turbine 时 pq.exists() 恒假 → day_costs 静默变 None、柱2 悄悄降级成 INSUFFICIENT,
+    #   一句报错都没有。改为: 优先读 clean_gate.json 里记的 section, 否则取该目录下唯一的 parquet;
+    #   仍然取不到就**响亮**说明 (而不是让下游以为"这场的经济性算不出来")。
+    _cdir = P.out_root(a.farm) / 'cleaned'
+    pq, _why = None, ''
+    _gate = _cdir / 'clean_gate.json'
+    try:
+        if _gate.exists():
+            import json as _j
+            _sec = (_j.loads(_gate.read_text(encoding='utf-8')) or {}).get('section')
+            if _sec and (_cdir / f'{_sec}.parquet').exists():
+                pq = _cdir / f'{_sec}.parquet'
+    except Exception:
+        pass
+    if pq is None:
+        _cands = sorted(_cdir.glob('*.parquet')) if _cdir.is_dir() else []
+        if len(_cands) == 1:
+            pq = _cands[0]
+        elif _cands:
+            pq, _why = _cands[0], f'目录里有 {len(_cands)} 个 parquet, 取第一个 {_cands[0].name}'
+        else:
+            _why = f'{P.rel(_cdir)} 下没有 cleaned parquet (清洗闸未落盘? persist=False?)'
+    if pq is None:
+        print(f'[wrapup] 柱2 无 cleaned parquet → day_costs 保持 None: {_why}', flush=True)
+    elif _why:
+        print(f'[wrapup] 柱2 cleaned parquet 选取说明: {_why}', flush=True)
     vcfg = ROOT / 'configs' / a.farm / 'value_assumptions.yaml'
-    if pq.exists():
+    if pq is not None and pq.exists():
         import pandas as pd
         import yaml
         from src.sop.price import build_price_by_hour

+ 11 - 0
src/windcms/config.py

@@ -95,6 +95,17 @@ FARMS = {
     }
 }
 
+def cms_out():
+    """CMS 产物目录的**唯一解析口** —— 读侧与写侧必须走同一个 (2026-09-16 统一)。
+
+    原状: 只有写侧认 `WINDCMS_OUT` (上面 FARMS 里的 `'out'`), 而读侧三处硬绑 `P.cms()` ——
+    `src/windscada/taxonomy.py` 的报告转录、`src/windscada/subsys/fusion.py::windcms_grades`、
+    `scripts/windscada_serve.py` 的产物指纹。设了该环境变量就会"**写到旁路、页面读生产**",
+    表现为"重算完了页面还是旧数", 且**一句报错都没有** (冻结构建自检就活在这个组合上)。
+    """
+    return Path(os.environ.get('WINDCMS_OUT') or P.cms())
+
+
 def farm(name='rudong'):
     if name not in FARMS:
         raise SystemExit(f'未知场 {name}; 可用: {list(FARMS)}')

+ 34 - 4
src/windcms/data.py

@@ -68,19 +68,46 @@ def load_alarm_counts(cfg):
 _META = {}
 
 
+def _mtime_stamp(paths) -> str:
+    """一组文件的 (路径, mtime_ns, size) 摘要 —— 用来判断"产物是否换过"。
+
+    2026-09-16 用户令"重算后的产物可即时用于页面呈现": 本模块的 _META 是**进程内**缓存,
+    长期跑着的服务 (windcms serve / 观澜组件) 在摄入新窗之后会一直用旧 meta (实测同类问题
+    在 scripts/windscada_serve.py 的 _CACHE 上出现过)。把文件指纹并进缓存键即可自动失效,
+    不必重启服务。"""
+    h = hashlib.sha1()
+    for p in paths:
+        try:
+            st = p.stat()
+            h.update(f'{p}|{st.st_mtime_ns}|{st.st_size}\n'.encode('utf-8'))
+        except OSError:
+            h.update(f'{p}|MISSING\n'.encode('utf-8'))
+    return h.hexdigest()[:12]
+
+
 def spectra_meta(cfg):
-    """合并谱元数据: 合并库 (cfg.spectra_dir, 覆盖 w0127) + 各窗 spectra 库; 加列 store (npz 所在目录)."""
-    key = str(cfg['out'])
+    """合并谱元数据: 合并库 (cfg.spectra_dir, 覆盖 w0127) + 各窗 spectra 库; 加列 store (npz 所在目录).
+
+    缓存键 = 产物目录 + **各窗 index/spectra_meta 的指纹** → 重算/新摄入后自动失效 (见 _mtime_stamp)。"""
+    base = cfg['spectra_dir']
+    wins = windows(cfg)
+    stamp_files = [base / 'spectra_meta.parquet']
+    for _w, info in wins.items():
+        stamp_files.append(info['index'])
+        if info['spectra_dir']:
+            sd = info['spectra_dir']
+            stamp_files.append(sd.parent / 'spectra_meta.parquet')
+            stamp_files.append(sd / 'spectra_meta.parquet')
+    key = f'{cfg["out"]}|{_mtime_stamp(stamp_files)}'
     if key in _META:
         return _META[key]
     parts = []
-    base = cfg['spectra_dir']
     if (base / 'spectra_meta.parquet').exists():
         m = pd.read_parquet(base / 'spectra_meta.parquet')
         m['store'] = str(base)
         m['window'] = 'w0127'
         parts.append(m)
-    for w, info in windows(cfg).items():
+    for w, info in wins.items():
         sd = info['spectra_dir']
         if not sd:
             continue
@@ -92,6 +119,9 @@ def spectra_meta(cfg):
             parts.append(m)
     M = pd.concat(parts, ignore_index=True) if parts else pd.DataFrame(columns=['turbine', 'sensor', 'meas_name', 'trigger_time', 'store'])
     M['trigger_time'] = M['trigger_time'].astype(str)
+    # 只留最近 4 个指纹键, 免得长期运行的服务把每代 meta 都攒在内存里
+    if len(_META) > 4:
+        _META.clear()
     _META[key] = M
     return M
 

+ 30 - 1
src/windcms/pipeline.py

@@ -59,10 +59,39 @@ def window_year_gate(m5_root, window, max_age_days=400):
     return tmax
 
 
+def _zip_has_decoded_json(path) -> bool:
+    """压缩包里是否已经是**解码后**的 `*_decode.json` (2026-09-16)。
+
+    为什么要有这个判断: `detect_input()` 把任何 .zip/.rar/.7z 都归为 `tcm_archive`, 而那条路要调
+    `scripts/rudong_tcm_ingest_raw.py` (解原始 base64+XML 导出) —— 该脚本**没随包** ⇒ 现场直接
+    把 CMS 导出的 zip 指过来必然崩在 `FileNotFoundError`。但其实"zip 里就是 `*_decode.json`"时
+    根本不需要它: 本包的 `rudong_tcm_index.py` / `rudong_tcm_spectra.py` **支持把 zip 当 --root**
+    (直接按成员读), 走它们即可。这里只做"是不是解码后包"的判定, 不假装支持原始导出。
+    """
+    import zipfile
+    try:
+        with zipfile.ZipFile(path) as zf:
+            return any(not i.is_dir() and i.filename.endswith('_decode.json') for i in zf.infolist())
+    except Exception:
+        return False
+
+
 def ingest(cfg, path, window, log, env):
     kind = detect_input(path)
+    if kind == 'tcm_archive' and _zip_has_decoded_json(path):
+        # 解码后的导出包 (zip 形态): 直接走索引/谱, 不必经缺失的 rudong_tcm_ingest_raw.py
+        print(f'[ingest] {Path(path).name}: 内含 *_decode.json → 按"已解码导出"处理 (zip 当 root)', flush=True)
+        kind = 'tcm_decoded_json'
     if kind == 'tcm_archive':
-        _run('ingest:tcm_archive', [PY, str(ROOT / 'scripts/rudong_tcm_ingest_raw.py'), '--rar', str(path), '--window', window], env, log)
+        raw = ROOT / 'scripts/rudong_tcm_ingest_raw.py'
+        if not raw.is_file():
+            raise RuntimeError(
+                f'输入 {path} 是**原始** TCM 导出包 (rar/7z, 需先解 base64+XML), 而解码脚本 '
+                f'{raw.name} 未随包 (振动线分支产物)。两条可走的路: '
+                f'① 现场若能直接给"含 *_decode.json 的导出"(zip 或目录), 把它放 data/raw/<场>/windcms/ 下, '
+                f'用 scripts/vib_raw_build.py 摄入 (本包支持 zip 当 root); '
+                f'② 向振动线索取 scripts/rudong_tcm_ingest_raw.py 后再跑本步。')
+        _run('ingest:tcm_archive', [PY, str(raw), '--rar', str(path), '--window', window], env, log)
     elif kind == 'tcm_decoded_json':
         out = cfg['m5'] / 'windows' / window
         out.mkdir(parents=True, exist_ok=True)

+ 5 - 1
src/windcms/plugins.py

@@ -369,8 +369,12 @@ def analyze(ctx, input='', window='', steps='', confirm=False):
     ctx['cfg']['out'].mkdir(parents=True, exist_ok=True)
     # stdout= 只用到这个文件的 fd (内容由子进程自己写), 所以这里不 with-close:
     # 句柄留着, 免得父进程提前关掉; 显式 encoding 是为了不留"缺 encoding"的门禁告警。
+    # ★2026-09-16: 本函数跑在 **CMS 服务进程**里 (由 guanlan.py 以无窗口方式拉起, 自己没有可见控制台),
+    #   裸 Popen 会让 Windows 给全链分析**新建一个可见控制台窗口** (从页面点"分析"就会看到黑窗)。
+    #   统一走 src.proc.spawn (CREATE_NO_WINDOW)。
+    from src import proc as _proc
     _logf = open(logp, 'w', encoding='utf-8')
-    p = subprocess.Popen(cmd, cwd=str(ROOT), stdout=_logf, stderr=subprocess.STDOUT)
+    p = _proc.spawn(cmd, cwd=ROOT, stdout=_logf, stderr=subprocess.STDOUT)
     lock.write_text(str(p.pid), encoding='utf-8')
     return dict(kind='text', data=dict(pid=p.pid, log=str(logp)), text=f'全链已在后台启动 (pid {p.pid}); 进度看 job_status; 日志 {logp}. 完成后产物/报告/知识库自动刷新 (界面需刷新页面).')
 

+ 8 - 3
src/windscada/config.py

@@ -46,10 +46,15 @@ REQUIRED = ('name', 'n_turbines', 'turbines', 'src_10min', 'src_alarm', 'store',
 RUDONG_SRC = str(P.RAW_ROOT)
 RAW_ROOT = P.RAW_ROOT
 
-# 场站目录下的**约定子目录名** — 这五个名字是摄入侧的接口, 改名等于换接口, 要同步 README 与维护页
-STATION_SUBDIRS = ('scada_10min', 'scada_1min', '故障报警', '风机故障记录', '油样报告')
+# 场站目录下的**约定子目录名** — 这些名字是摄入侧的接口, 改名等于换接口, 要同步 README 与维护页
+# 振动侧两目录 (2026-09-12 用户令: 「CMS 振动评估报告」遵循 data/raw/如东/windcms、
+# 「振动线 handoff」遵循 data/raw/如东/m5_cms_tcm) —— 它们此前只在 outputs/ 下以**产物**形式存在,
+# 源件不在 data/raw, 于是"从零重算"时振动侧无源可算。加进来后扫描/落位/数据层页都能认它们。
+STATION_SUBDIRS = ('scada_10min', 'scada_1min', '故障报警', '风机故障记录', '油样报告',
+                   'windcms', 'm5_cms_tcm')
 _SRC_KEYS = (('src_10min', 'scada_10min'), ('src_1min', 'scada_1min'), ('src_alarm', '故障报警'),
-             ('src_workorder', '风机故障记录'), ('src_oil', '油样报告'))
+             ('src_workorder', '风机故障记录'), ('src_oil', '油样报告'),
+             ('src_windcms', 'windcms'), ('src_m5', 'm5_cms_tcm'))
 
 _BUILTIN = {
     'rudong': {

+ 2 - 1
src/windscada/subsys/fusion.py

@@ -322,7 +322,8 @@ def windcms_grades():
         return _WCMS
     _WCMS = {}
     try:
-        rp = sorted((P.cms()).glob('报告_CMS振动状态评估报告_*.md'))[-1]
+        from src.windcms.config import cms_out     # 与写侧同一解析口 (WINDCMS_OUT; 见 windcms.config.cms_out)
+        rp = sorted(cms_out().glob('报告_CMS振动状态评估报告_*.md'))[-1]
         md = rp.read_text(encoding='utf-8')
         for l in md.split('## 附录 A')[1].splitlines():
             if not l.startswith('| WTG'):

+ 8 - 4
src/windscada/subsys/pitch.py

@@ -29,9 +29,12 @@ ALARM_FAM = {  # M4b 报警轴 (windscada alarms.parquet; 分册监测量对齐)
 }
 
 
-def _alarm_counts(win_start, win_end):
+def _alarm_counts(win_start, win_end, store=None):
     import re as _re
-    ap = P.store() / 'alarms.parquet'
+    # ★store 显式传入 (2026-09-16 修): 原写 `P.store()` —— 它跟的是环境变量 WINDSCADA_FARM,
+    #   而调用方 registry(cfg) 拿的是**显式选定的场**; 多场部署下这里会读到 rudong 的 alarms.parquet
+    #   而其余输入都来自当前场 ⇒ 跨场串数据, 且因为"文件存在、列名兼容"而完全不报错。
+    ap = (pathlib.Path(store) if store else P.store()) / 'alarms.parquet'
     if not ap.exists():
         return None
     al = pd.read_parquet(ap)
@@ -43,7 +46,8 @@ def _alarm_counts(win_start, win_end):
     return out
 
 
-def registry():
+def registry(cfg=None):
+    """变桨面登记表。`cfg` 给定时, 其 `store` 决定读哪个场的 alarms (见 _alarm_counts 的注释)。"""
     daily, zero = load()
     daily['date'] = pd.to_datetime(daily['date']).dt.date
     cur = _cur_win(daily).copy()
@@ -73,7 +77,7 @@ def registry():
             zsum[t] = dict(months=int(len(g)), dev_med=float(g['zero_dev'].median()), last=float(g['zero_dev'].iloc[-1]),
                            n_alarm=int((g['zero_dev'].abs() >= ZERO_ALARM).sum()), last_month=str(g['month'].iloc[-1]))
     win_end = daily['date'].max(); win_start = pd.Timestamp(win_end) - pd.Timedelta(days=60)
-    ac = _alarm_counts(win_start, win_end)
+    ac = _alarm_counts(win_start, win_end, store=(cfg or {}).get('store') if cfg else None)
     rows, detail = [], {}
     for t, r in fleet.iterrows():
         d = dict(维度={}, 依据={})

+ 2 - 1
src/windscada/taxonomy.py

@@ -211,7 +211,8 @@ def system_matrix(cfg=None):
             _WCOL = {'主轴承': ('主轴承前', '主轴承后'), '齿轮箱': ('齿轮箱',), '发电机': ('发电机',)}
             _ORD = ['危险', '报警', '良好', '优秀', '不可判']
             import pathlib as _pl, re as _re3
-            _wmd = sorted(P.cms().glob('报告_CMS振动状态评估报告_*.md'))
+            from src.windcms.config import cms_out as _cms_out     # 读写同一口 (WINDCMS_OUT 不能被读侧忽略)
+            _wmd = sorted(_cms_out().glob('报告_CMS振动状态评估报告_*.md'))
             _wdate = _wmd[-1].stem.rsplit('_', 1)[-1] if _wmd else '?'
             _wpart = {}
             if _wmd:

+ 45 - 1
start.bat

@@ -1,7 +1,38 @@
 @echo off
+rem ============================================================================
+rem  观澜 start.bat
+rem  2026-09-16 用户令: 启动观澜系统时"弹出的命令窗口"改为不弹出方式。
+rem
+rem  默认 (双击本文件)      : 走**无窗口**启动 —— 交给 start_hidden.vbs (窗口样式 0),
+rem                           后台起服务、等就绪、自动打开浏览器, 不留任何命令窗口。
+rem  start.bat console      : 前台模式 (旧行为) —— 在当前窗口里跑 serve 并实时打印日志,
+rem                           只在排障时用; 关掉这个窗口等于停掉 serve。
+rem  start.bat help         : 说明。
+rem
+rem  为什么分开: 各组件服务本来就是 DETACHED_PROCESS 起的 (不弹窗), 唯一会弹的就是
+rem  "在控制台里前台跑 guanlan.py serve" 这件事本身。排障又确实需要看见实时输出,
+rem  所以保留一个显式的 console 档, 而不是把日志彻底藏掉。
+rem ============================================================================
 chcp 65001 >nul
 set PYTHONUTF8=1
 cd /d "%~dp0"
+
+if /i "%~1"=="help" goto help
+if /i "%~1"=="console" goto console
+if /i "%~1"=="-c" goto console
+
+if not exist ".venv\Scripts\pythonw.exe" (
+  echo [X] Not installed yet: .venv\Scripts\pythonw.exe not found
+  echo     Run install.bat first.
+  pause
+  exit /b 2
+)
+
+rem 无窗口启动: wscript 以隐藏窗口方式调用 pythonw.exe 跑启动器, 本窗口立即退出。
+wscript.exe //nologo "%~dp0start_hidden.vbs"
+exit /b 0
+
+:console
 if not exist ".venv\Scripts\python.exe" (
   echo [X] Not installed yet: .venv\Scripts\python.exe not found
   echo     Run install.bat first. If install printed errors, send that screen back
@@ -9,6 +40,9 @@ if not exist ".venv\Scripts\python.exe" (
   pause
   exit /b 2
 )
+echo [i] Foreground mode - this window stays open and shows live logs.
+echo     Close it to stop the server; for no-window start just run start.bat with no argument.
+echo.
 ".venv\Scripts\python.exe" guanlan.py serve
 if errorlevel 1 (
   rem serve exits 1 in two different cases: the gateway never came up, or the gateway is
@@ -27,7 +61,7 @@ if errorlevel 1 (
   echo [!] Gateway is up, but at least one module is degraded ^(the [DOWN] lines above^).
   echo     Usual cause: local Ollama is not running or has no models pulled, which only
   echo     disables the Q^&A / local-review pages. Every other page works.
-  echo     Details: logs\ for the degraded module.
+  echo     Details: logs/ for the degraded module.
 )
 start "" http://127.0.0.1:28084/
 echo.
@@ -35,3 +69,13 @@ echo [i] Ops console - stop/start services, rebuild, clear products:
 echo     http://127.0.0.1:28084/ops
 echo     It stays available while component services are stopped (only the gateway must be up);
 echo     buttons follow real state and the backend refuses impossible actions (HTTP 409).
+exit /b 0
+
+:help
+echo Usage:
+echo   start.bat            no-window start (default, recommended): hidden services + browser
+echo   start.bat console    foreground start with live logs (troubleshooting)
+echo   stop.bat             stop everything
+echo   check.bat            self-check
+echo Logs: logs\serve.log (no-window mode), logs\start_hidden.log (launcher trace)
+exit /b 0

BIN
start_hidden.vbs