Kaynağa Gözat

从零重算: 修本体层 GBK 文本 I/O (中文 Windows 上根本跑不了) + 落位脚本两处静默丢件

按"不要随包产物, 基于 data/raw 从零重算"实测, 逮到三个真 bug (都不是参数问题, 是静默失效):

1. src/ontology/store.py 的 read_text()/write_text() 没写 encoding →
   Python 文本 I/O 默认用**系统 locale 编码**, 中文 Windows 上是 cp936(GBK), 而对象库是 UTF-8 JSON。
   读: 中文乱码/UnicodeDecodeError; 写: 内容里的零宽空格 \u200b 直接 UnicodeEncodeError 崩。
   实测: `python -m src.ontology.kb_ingest` 崩在 `\u200b` —— 也就是说**本体层在中文 Windows 上
   根本不可用**(开发机是 UTF-8 locale 所以从没暴露)。这是"兼容 windows/linux 部署"要防的那一类。
   一并修 `src/ontology/` 全层 10 处 (store/retrieval/monthly_roll/audit/mcp_server/seed_rudong29);
   修后 objects.json 一次写成 (AlarmCode 555 · WI 559 · 失效模式 150 · 维护任务 559 · Doc 303)。
   存量还有 12 处 (windscada/subsys/fusion.py、windscada_serve.py、windcms/* 等) 未改。

2. scripts/check_portability.py 新增 WARN 项 6: "文本读写缺 encoding=" —— 把上面这条变成可度量、
   不阻塞的门禁 (跨行调用要括号配对后再找 encoding= 才能不漏)。修完本体层后剩 12 处 / 6 文件。

3. scripts/place_raw_data.py:
   · 缺 `import os` —— 模块级 DEFAULT_SRC 就用 os.environ, 于是**脚本自 d90da05 起一直跑不通**
     (A3 那次落位不是它干的), 修后 --dry-run 正常;
   · **单文件规则从来没落过东西**: 包内前缀恰好等于条目本身时, 去前缀后 rest 为空, 旧代码当"目录条目"
     直接 continue。受害两例: `大部件维修记录.20240619143912557.xlsx` (A2 规则, 从 A3 起就少落) 与
     `广核如海…对译表.xlsx` (本轮新规则) —— 修后按前缀文件名落位;
   · 新增 `--scope a2|mech|full` 与 RULES_MECH: 把"机理层源件 + scada_1min"写成同一张可复核的表
     (依据: src/ontology/kb_ingest.py 的 TECH 与 rd() 四个文件名; config.STATION_SUBDIRS 的 scada_1min);
   · 落位改为"同尺寸即跳过": 重跑不再把 15 GB 原样重抄。

实测数据 (便携包 F:\temp\如东风场数据, --scope full 落 948 件):
  · 三门台账从零重算: 报警 0→39211 行 · 工单 0→5876 行 · 油样 0→404 行;
  · SCADA 侧 10 个构建器全跑通 (loss_monthly 3729 · temp_bins 24624 · yaw_daily/hydraulic_accum/
    thermal_chain/system_aux/curve_lenses/control_*/stop_events/powercurve_*), L0 仓 16 件;
  · 本体层从"包内没有源件"变为可重算: Doc 303 件 + turbine_params.parquet 1706 行;
  · 等价验收: alarms 39211/39211 **逐值完全一致**; workorders 8 行差异全部可归类 (旧链时间解析缺陷) +
    69 行本链新增覆盖; 唯一人工项 = 油样 506→404, 那 102 行的源件 (2026-07 合并报告) 不在包内。
zhouyang.xie 1 ay önce
ebeveyn
işleme
add485fb0c

+ 59 - 1
scripts/check_portability.py

@@ -14,7 +14,14 @@ r"""跨平台/相对路径 门禁 (2026-09-11 用户令) —— 换电脑、换
 ## 其余是 WARN (不阻塞, 但列出来慢慢收敛)
 
 4. 运行期仍写死 `ROOT / 'outputs/<场名>/…'` → 建议改 `src/paths.py` 的对应函数;
-5. 交付侧离线工具 (`scripts/guanlan_cloud_*` 等) 里的同类写法。
+5. 交付侧离线工具 (`scripts/guanlan_cloud_*` 等) 里的同类写法;
+6. **文本读写没写 `encoding=`** → 这是同一类"换机就坏"的坑: Python 文本 I/O 默认用**系统 locale 编码**,
+   中文 Windows 上是 cp936(GBK), 而本库的 JSON/HTML/索引全是 UTF-8。后果有两种, 都很隐蔽:
+     · 读: UTF-8 文件按 GBK 解 → 中文乱码, 或直接 UnicodeDecodeError;
+     · 写: 内容里出现 GBK 编不出的字符 (零宽空格 `\u200b`、emoji、罕见汉字) → UnicodeEncodeError 崩溃。
+   2026-09-11 从零重算时实逮: `src/ontology/store.py` 写 objects.json 崩在 `\u200b` 上 —— 也就是说
+   **本体层在中文 Windows 上根本不可用**(开发机是 UTF-8 locale, 所以一直没暴露)。
+   本项按 WARN 列出 (存量较多, 逐步收敛); 新写的代码请一律 `encoding='utf-8'`。
 
 ## 例外必须**具名**登记
 
@@ -27,6 +34,7 @@ r"""跨平台/相对路径 门禁 (2026-09-11 用户令) —— 换电脑、换
 from __future__ import annotations
 
 import argparse
+import collections
 import pathlib
 import re
 import sys
@@ -45,6 +53,43 @@ RE_ABS = re.compile(r"""(/Users/[A-Za-z0-9_.\-]+|/Volumes/[A-Za-z0-9_\- ]+|/home
 RE_CWD = re.compile(r"""(?:open|Path|pathlib\.Path)\(\s*['"](outputs|data|release|configs|src|scripts|reference|resources|logs|run|wheels|vendor)/""")
 # os.sep 出现在"拼字符串"里 (显示函数除外)
 RE_OSSEP = re.compile(r"""\+\s*os\.sep""")
+# 文本读写调用 (用于 WARN 项 6: 缺 encoding= 时按系统 locale 编解码, 中文 Windows 上是 GBK)
+RE_TXTIO = re.compile(r"""\.(read_text|write_text)\(""")   # portability-allow: 本门禁的规则源里必须写出这两个方法名
+
+
+def scan_encoding(globs, extra_files=()) -> list:
+    """WARN: 找出 `.read_text(...)` / `.write_text(...)` 没给 encoding= 的调用。   # portability-allow: 说明文字即规则本身
+
+    跨行调用 (如 retrieval.py 里分三行写的 write_text) 要把括号配对后再找 encoding,
+    所以不能只看一行。返回 [(相对路径, 行号, 片段)]。
+    """
+    out = []
+    files = []
+    for g in globs:
+        files += sorted(ROOT.glob(g))
+    files += [ROOT / f for f in extra_files if (ROOT / f).exists()]
+    for f in files:
+        rel = f.relative_to(ROOT).as_posix()
+        try:
+            text = f.read_text(encoding='utf-8')
+        except Exception:
+            continue
+        for m in RE_TXTIO.finditer(text):
+            i = m.end()
+            depth, j = 1, i
+            while j < len(text) and depth:
+                if text[j] == '(':
+                    depth += 1
+                elif text[j] == ')':
+                    depth -= 1
+                j += 1
+            call = text[m.start():j]
+            if 'encoding' in call:
+                continue
+            if 'portability-allow' in text[text.rfind('\n', 0, m.start()) + 1:text.find('\n', m.start())]:
+                continue
+            out.append((rel, text.count('\n', 0, m.start()) + 1, ' '.join(call.split())[:70]))
+    return out
 
 ALLOW = {
     'src/sop/farm_paths.py': '它本身就是"win 盘符 → POSIX 盘"的映射表 (自带用例), 路径即测试数据',
@@ -111,6 +156,19 @@ def main() -> int:
             print(f'  [ERROR] {kind:12s} {rel}:{i}  ← {found}')
             print(f'          {line}')
     print(f'\n例外登记 (ALLOW, 共 {len(ALLOW)} 项 + {len(DELIVERY_ALLOW)} 类): 见本文件注释, 每项都写了理由')
+
+    # ── WARN: 文本读写缺 encoding= (项 6; 不阻塞, 但中文 Windows 上会乱码/崩溃) ──
+    warns = scan_encoding(RUNTIME_GLOBS, RUNTIME)
+    if not a.runtime:
+        warns += scan_encoding(DELIVERY_GLOBS)
+    if warns:
+        by_file = collections.Counter(w[0] for w in warns)
+        print(f'\n[WARN] 文本读写缺 encoding= : {len(warns)} 处 / {len(by_file)} 个文件 '
+              f'—— 默认按系统 locale 编解码 (中文 Windows = GBK), 换机/换内容就可能乱码或崩溃;')
+        for rel, n in by_file.most_common(12):
+            print(f'         {n:3d} 处  {rel}')
+        print('       修法: read_text(encoding="utf-8") / write_text(..., encoding="utf-8")')
+
     if errs:
         print(f'\n结论: {len(errs)} 处 ERROR —— 换机部署前必须清零')
         return 1

+ 95 - 33
scripts/place_raw_data.py

@@ -24,23 +24,55 @@ A2 定了四项数据源的落位: data/raw/<场站名称>/{scada_10min, 故障
                                       (去掉"2025年油样"这一层, 直接落成 <台号>/<部件>/<pdf>,
                                        与维护页说明"按台号/部件分目录"一致)
 
-## 故意不做的事
+## 两组范围 (--scope)
 
-  · 不落 `(3)现场检修记录/2025年检修记录` 与 `2026年检修记录`: 与 工作/风机故障记录/2025年故障记录、
-    2026年故障记录 **逐件同名同大小**(19/19 件), 是同一批月度汇总表的副本。落两遍会让"按年目录"
-    摄入看到两份重复台账 (2026-08-31 长停台账虚高 35 倍那类事故的同款成因: 快照重复必须归并, 不能叠加)。
-  · 不落 `(8)…/振动分析报告/`: 属振动线 (CMS), 不在 A2 四项数据层内。
-  · 不落 1分钟数据.zip / scada数据(如东)/ / 西门子4.0技术资料/ 等: A3 本轮范围是 A2 那四类。
+`--scope a2`(默认) 只落上面那四类, 即 A3 当时的范围, 语义不变。
+
+`--scope mech` 落的是**"能重算的缺口"**(2026-09-11 用户令: 缺什么从现场包里抽)。它解决的是
+"从零重算时, 某些产物因为源件不在 data/raw 而算不出来"这件事, 因此只挑**有消费者**的源:
+
+  <data/raw>/西门子4.0技术资料/  ← 如东风场数据.zip: 西门子4.0技术资料/**           (317 件 2.8 GB)
+                                   + 如东风场数据/广核如海风电场西门子风机故障代码中英文对译表.xlsx
+      依据: `src/ontology/kb_ingest.py` 里 `TECH = P.RAW_ROOT/'西门子4.0技术资料'`, 且 `rd()` 要读
+            `故障处理/故障处理手册.xlsx`、`广核如海风电场西门子风机故障代码中英文对译表.xlsx`、
+            `如海故障代码表.xlsx`、`维护相关/维护作业指导书.xlsx` 四个文件。前三个里**对译表在
+            技术资料目录里没有**(包里它在 `scada数据(如东)/` 与包根各有一份), 所以单独补一条规则。
+      效果: 本体层 `objects.json` 从"包内没有源件(⛔)"变成**可重算**。
+
+  <场站>/scada_1min/           ← 1分钟数据.zip 全部 38 个 <台号>.csv  (01E.csv…37B.csv)
+      依据: `src/windscada/config.py` 的 STATION_SUBDIRS 声明了这个子目录, `scan_stations.py`
+            报它"缺"; 表头带 `source_file=如海风机1分钟数据/如海测点_2025-01.csv`, 与场配置
+            `src_farm_names=['如海','如东']` 同源(如海=如东项目)。**包内没有消费者**: 它补的是
+            数据层完整性, 不改变任何页面数值 —— 落它是因为"扫描辨识"会把它报成缺口。
 
 用法:
-    python scripts/place_raw_data.py --dry-run      # 只报要落什么, 不写盘
-    python scripts/place_raw_data.py                # 真落位 (已存在则覆盖)
+    python scripts/place_raw_data.py --dry-run                # 只报要落什么, 不写盘 (默认 a2)
+    python scripts/place_raw_data.py                          # 真落位 (已存在则覆盖)
+    python scripts/place_raw_data.py --scope mech --dry-run    # 看机理层/1min 会落什么
+    python scripts/place_raw_data.py --scope full              # 两组一起
     python scripts/place_raw_data.py --src D:\\别的现场数据目录
+
+## 故意不做的事 (两组范围共有的判断)
+
+  · 不落 `(3)现场检修记录/2025年检修记录` 与 `2026年检修记录`: 与 工作/风机故障记录/2025年故障记录、
+    2026年故障记录 **逐件同名同大小**(19/19 件), 是同一批月度汇总表的副本。落两遍会让"按年目录"
+    摄入看到两份重复台账 (2026-08-31 长停台账虚高 35 倍那类事故的同款成因: 快照重复必须归并, 不能叠加)。
+  · 不落 `(8)…/振动分析报告/`(12 份月度用印版 PDF): 是振动线 (CMS) 的**成品牌报告**, 不是可再加工的
+    测量数据; 而 `windcms`/`m5_cms_tcm` 要的是 CMS 测点索引(handoff), 包里没有。
+  · 不落 `scada数据(如东)/**`(19 个月 × 12 个通道组 zip, 2.4 GB): 它是 `scada_10min/*.csv` 的**上游**
+    原始通道导出, 包内没有任何脚本读它(构建器读的是已经平铺好的 10min CSV) —— 落了也不参与重算。
+  · 不落 `fastlog数据/`(4 件 WTG0x.xls): 全库搜 `fastlog` 只有 2 处**注释**提到它("运行态见证"),
+    没有读取代码。
+  · 结论: 上面三项都是"有源件、无生成端/无消费者"。`genbearing_monthly`、`mblub_monthly`、
+    `yaw_dynamic_monthly`、`yaw1min_liveness`、`sector_power`、`duty_monthly`、`pc_monthly_bins`、
+    `thermal_monthly`、`structure.parquet`、`watch_channels_monthly` 这批产物同理 ——
+    全库只有读取方、**0 处写入方**, 属于 v0.2.0 未附构建脚本(见 docs §4)。
 """
 from __future__ import annotations
 
 import argparse
 import pathlib
+import os
 import shutil
 import sys
 import zipfile
@@ -53,15 +85,26 @@ DEFAULT_SRC = pathlib.Path(os.environ.get('GUANLAN_PLACE_SRC') or (ROOT / 'data'
 ZIP_10MIN = '10分钟数据.zip'
 ZIP_FARM = '如东风场数据.zip'
 
-# (压缩包, 包内前缀, 目标子目录, 是否去掉前缀这一层)
+# (压缩包, 包内前缀, 目标, 是否去掉前缀这一层, 目标根: station= data/raw/<场站>/, raw= data/raw/)
 RULES = [
-    ('10分钟数据.zip', '', 'scada_10min', True),
-    ('如东风场数据.zip', '如东风场数据/报警数据/', '故障报警', True),
-    ('如东风场数据.zip', '如东风场数据/数据收集/更新/(2)故障记录(首发故障有标识)(2025.1-至今)/故障记录/', '故障报警', True),
-    ('如东风场数据.zip', '如东风场数据/工作/风机故障记录/', '风机故障记录', True),
-    ('如东风场数据.zip', '如东风场数据/数据收集/更新/(3)现场检修记录(2025.1-至今)/大部件维修记录.20240619143912557.xlsx', '风机故障记录', True),
-    ('如东风场数据.zip', '如东风场数据/数据收集/更新/(8)风机振动数据、油液分析记录(2025.1-至今)/2025年油样/', '油样报告', True),
+    ('10分钟数据.zip', '', 'scada_10min', True, 'station'),
+    ('如东风场数据.zip', '如东风场数据/报警数据/', '故障报警', True, 'station'),
+    ('如东风场数据.zip', '如东风场数据/数据收集/更新/(2)故障记录(首发故障有标识)(2025.1-至今)/故障记录/', '故障报警', True, 'station'),
+    ('如东风场数据.zip', '如东风场数据/工作/风机故障记录/', '风机故障记录', True, 'station'),
+    ('如东风场数据.zip', '如东风场数据/数据收集/更新/(3)现场检修记录(2025.1-至今)/大部件维修记录.20240619143912557.xlsx', '风机故障记录', True, 'station'),
+    ('如东风场数据.zip', '如东风场数据/数据收集/更新/(8)风机振动数据、油液分析记录(2025.1-至今)/2025年油样/', '油样报告', True, 'station'),
 ]
+
+# 机理层与可选层 (--scope mech): 只挑**有消费者**的源 —— 目的是让"从零重算"能覆盖到本体层。
+RULES_MECH = [
+    # 本体层源件: kb_ingest.py 的 TECH = data/raw/西门子4.0技术资料 (317 件, 含故障处理/维护相关/各类图纸)
+    ('如东风场数据.zip', '如东风场数据/西门子4.0技术资料/', '西门子4.0技术资料', True, 'raw'),
+    # …但它要读的对译表不在技术资料目录里, 包里在别处 → 单独补一条 (见本文件开头"两组范围")
+    ('如东风场数据.zip', '如东风场数据/广核如海风电场西门子风机故障代码中英文对译表.xlsx', '西门子4.0技术资料', True, 'raw'),
+    # 数据层声明里的 scada_1min (config.STATION_SUBDIRS); 包内无消费者, 补的是数据层完整性
+    ('1分钟数据.zip', '', 'scada_1min', True, 'station'),
+]
+
 # 说明性的"故意不落", 只在报告里列出来
 SKIPPED = [
     ('如东风场数据.zip', '如东风场数据/数据收集/更新/(3)现场检修记录(2025.1-至今)/2025年检修记录/',
@@ -69,7 +112,11 @@ SKIPPED = [
     ('如东风场数据.zip', '如东风场数据/数据收集/更新/(3)现场检修记录(2025.1-至今)/2026年检修记录/',
      '与 工作/风机故障记录/2026年故障记录 逐件同名校验相同 (副本)'),
     ('如东风场数据.zip', '如东风场数据/数据收集/更新/(8)风机振动数据、油液分析记录(2025.1-至今)/振动分析报告/',
-     '属振动线 (CMS), 不在 A2 四项数据层内'),
+     '振动线成品牌报告, 不是可再加工的测量数据; windcms 要的测点索引包里没有'),
+    ('如东风场数据.zip', '如东风场数据/scada数据(如东)/',
+     'scada_10min/*.csv 的上游原始通道导出 (2.4 GB), 包内无脚本读它 → 不参与重算'),
+    ('如东风场数据.zip', '如东风场数据/fastlog数据/',
+     '全库搜 fastlog 只有 2 处注释提到, 无读取代码'),
 ]
 
 
@@ -93,13 +140,14 @@ def human(n: float) -> str:
     return f'{n:.1f} GB'
 
 
-def plan(src: pathlib.Path, station: pathlib.Path):
-    """→ [(zip, 目标文件, 包内条目名, 解压后大小)]; 只读压缩包目录, 不解压。"""
+def plan(src: pathlib.Path, station: pathlib.Path, rules):
+    """→ [(zip, 目标文件, 包内条目名, 解压后大小, 目标根, 目标名)]; 只读压缩包目录, 不解压。"""
     out = []
-    for zname, prefix, target, strip in RULES:
+    for zname, prefix, target, strip, root in rules:
         zp = src / zname
         if not zp.exists():
             raise SystemExit(f'缺压缩包: {zp}')
+        base = station if root == 'station' else station.parent      # 'raw' = data/raw (机理层不在场站目录下)
         with zipfile.ZipFile(zp) as zf:
             for info in zf.infolist():
                 if info.is_dir():
@@ -109,8 +157,14 @@ def plan(src: pathlib.Path, station: pathlib.Path):
                     continue
                 rest = name[len(prefix):] if strip else name
                 if not rest:
-                    continue
-                out.append((zname, station / target / rest, name, info.file_size))
+                    # 前缀恰是条目本身 = 单文件规则 (如 '…/大部件维修记录.20240619143912557.xlsx')。
+                    # 旧写法在这里直接 continue, 于是**单文件规则从来没落过东西** —— 2026-09-11 逮到两例:
+                    # `大部件维修记录.20240619143912557.xlsx`(A2 规则, 从 A3 起就没落) 与
+                    # `广核如海…对译表.xlsx`(mech 规则)。只有"目录条目"(prefix 以 / 结尾) 才该跳过。
+                    if prefix.endswith('/'):
+                        continue
+                    rest = prefix.rstrip('/').rsplit('/', 1)[-1]
+                out.append((zname, base / target / rest, name, info.file_size, root, target))
     return out
 
 
@@ -118,6 +172,8 @@ def main() -> int:
     ap = argparse.ArgumentParser()
     ap.add_argument('--src', default=str(DEFAULT_SRC), help='现场数据目录 (默认 %(default)s)')
     ap.add_argument('--dry-run', action='store_true', help='只报计划, 不写盘')
+    ap.add_argument('--scope', choices=('a2', 'mech', 'full'), default='a2',
+                    help='a2=只落四项数据层(A3 语义, 默认) · mech=机理层源件+1min · full=两组一起')
     a = ap.parse_args()
 
     src = pathlib.Path(a.src)
@@ -128,19 +184,19 @@ def main() -> int:
     station = pathlib.Path(raw_station_dir())
     cfg = farm()
     print(f'场站目录: {station}   (来自 src/windscada/config.py raw_station={cfg.get("raw_station")!r})')
-    print(f'现场数据: {src}\n')
+    print(f'现场数据: {src}   范围: --scope {a.scope}\n')
 
-    items = plan(src, station)
+    rules = {'a2': RULES, 'mech': RULES_MECH, 'full': RULES + RULES_MECH}[a.scope]
+    items = plan(src, station, rules)
     by_target = {}
-    for zname, dst, entry, size in items:
-        t = dst.relative_to(station).parts[0]
-        d = by_target.setdefault(t, [0, 0])
+    for zname, dst, entry, size, root, target in items:
+        d = by_target.setdefault((root, target), [0, 0])
         d[0] += 1
         d[1] += size
-    print('== 计划落位 ==')
-    for t in ('scada_10min', '故障报警', '风机故障记录', '油样报告'):
-        n, sz = by_target.get(t, (0, 0))
-        print(f'  {t:12s} {n:5d} 件  {human(sz):>10s}   → {station / t}')
+    print(f'== 计划落位 (scope={a.scope}) ==')
+    for (root, target), (n, sz) in sorted(by_target.items(), key=lambda kv: kv[0][1]):
+        where = (station / target) if root == 'station' else (station.parent / target)
+        print(f'  {target:18s} {n:5d} 件  {human(sz):>10s}   → {where}')
     print(f'  合计 {len(items)} 件, {human(sum(i[3] for i in items))}')
 
     print('\n== 故意不落 ==')
@@ -153,7 +209,13 @@ def main() -> int:
 
     print('\n== 落位 ==')
     done = 0
-    for zname, dst, entry, size in items:
+    skipped = 0
+    for zname, dst, entry, size, root, target in items:
+        # 已有同尺寸文件 = 已经是最新 → 跳过。重跑一次不该把 15 GB 原样再抄一遍
+        # (mech 范围 15.5 GB, a2 范围 14.7 GB; 2026-09-11 修单文件规则时就是靠这个避免整盘重写)。
+        if dst.exists() and dst.stat().st_size == size:
+            skipped += 1
+            continue
         dst.parent.mkdir(parents=True, exist_ok=True)
         with zipfile.ZipFile(src / zname) as zf:
             info = next(i for i in zf.infolist() if gbk_name(i).replace('\\', '/') == entry)
@@ -164,8 +226,8 @@ def main() -> int:
             raise SystemExit(f'写出大小不符: {dst} 期望 {size} 实得 {got}')
         done += 1
         if done % 10 == 0 or size > 50 * 1024 * 1024:
-            print(f'  [{done}/{len(items)}] {human(size):>10s}  {dst.relative_to(station)}', flush=True)
-    print(f'\n完成: {done} 件写入 {station}')
+            print(f'  [{done}/{len(items)}] {human(size):>10s}  {dst.relative_to(station.parent)}', flush=True)
+    print(f'\n完成 (scope={a.scope}): 新写/更新 {done} 件, 已是最新跳过 {skipped} 件, 共 {len(items)} 件')
     return 0
 
 

+ 1 - 1
src/ontology/audit.py

@@ -12,7 +12,7 @@ ONT = P.objects_json()
 
 
 def audit(sample_turbines=3, seed=7):
-    d = json.loads(ONT.read_text())
+    d = json.loads(ONT.read_text(encoding='utf-8'))
     issues, stats = [], {}
 
     # ① 悬空引用

+ 1 - 1
src/ontology/mcp_server.py

@@ -27,7 +27,7 @@ def db():
         raise FileNotFoundError(f'本体库缺失: {OBJ_PATH} (先跑 populate) — 不可静默降级')
     mt = OBJ_PATH.stat().st_mtime
     if _DB is None or _DB_MTIME != mt:
-        _DB = json.loads(OBJ_PATH.read_text())
+        _DB = json.loads(OBJ_PATH.read_text(encoding='utf-8'))
         _DB_MTIME = mt
     return _DB
 

+ 3 - 3
src/ontology/monthly_roll.py

@@ -18,7 +18,7 @@ def data_month():
 
 
 def ont_month():
-    o = json.loads((ONT / 'objects.json').read_text())
+    o = json.loads((ONT / 'objects.json').read_text(encoding='utf-8'))
     mw = o.get('mwindow/rudong/climatology', {}).get('props', {})
     roll = o.get('meta/roll_state', {}).get('props', {})
     return roll.get('data_month')
@@ -46,7 +46,7 @@ def diff_summary(before, after):
 
 
 def roll(dry=False):
-    before = json.loads((ONT / 'objects.json').read_text())
+    before = json.loads((ONT / 'objects.json').read_text(encoding='utf-8'))
     dm, om = data_month(), ont_month()
     actions = []
     if dm == om:
@@ -61,7 +61,7 @@ def roll(dry=False):
         actions.append(f'数据月 {om} → {dm}: L1仓重建 + populate增量 + 沙盘刷新 + 验收状态机')
     else:
         actions.append(f'[dry] 检测到新数据月 {dm} (本体记 {om}), 未执行重跑')
-    after = json.loads((ONT / 'objects.json').read_text())
+    after = json.loads((ONT / 'objects.json').read_text(encoding='utf-8'))
     chg = diff_summary(before, after)
     if not dry:
         # roll 状态只在真跑时落库 — dry 落状态会把随后的真跑骗成空转 (2026-08-28 实逮自修)

+ 5 - 2
src/ontology/retrieval.py

@@ -95,7 +95,10 @@ def build(use_vec=True, verbose=False):
     dl = [sum(x.values()) for x in docs]
     IDX.write_text(json.dumps(dict(ids=ids, inv={k: v for k, v in inv.items()},
                                    idf=idf, dl=dl, avgdl=avgdl,
-                                   src_mtime=OBJ.stat().st_mtime), ensure_ascii=False))
+                                   src_mtime=OBJ.stat().st_mtime), ensure_ascii=False),
+                   # 索引里全是中文词条: 不给 encoding 会按系统 locale(cp936) 写, 中文 Windows 上直接抛
+                   # UnicodeEncodeError (与 src/ontology/store.py 同一类坑, 2026-09-11 逮)
+                   encoding='utf-8')
     n_vec = 0
     if use_vec:
         try:
@@ -133,7 +136,7 @@ _M = {}
 def _load():
     if _M.get('ids') and _M.get('mtime') == OBJ.stat().st_mtime:
         return _M
-    if not IDX.exists() or json.loads(IDX.read_text())['src_mtime'] != OBJ.stat().st_mtime:
+    if not IDX.exists() or json.loads(IDX.read_text(encoding='utf-8'))['src_mtime'] != OBJ.stat().st_mtime:
         build(verbose=False)
     j = json.loads(IDX.read_text(encoding='utf-8'))
     d = json.loads(OBJ.read_text(encoding='utf-8'))

+ 1 - 1
src/ontology/seed_rudong29.py

@@ -35,7 +35,7 @@ def build(out=None):
     n64 = len(a64)
     m64 = sorted(a64.t_on.dt.to_period('M').astype(str).unique())
 
-    h = json.loads(HANDOFF.read_text())
+    h = json.loads(HANDOFF.read_text(encoding='utf-8'))
     tcm16 = next(r for r in h['per_turbine']['new_2026_08_26'] if r.get('id') == 'TCM-16')
 
     s = Store(out or OUT)

+ 5 - 2
src/ontology/store.py

@@ -13,7 +13,10 @@ class Store:
         self.path = pathlib.Path(path)
         self.objects = {}
         if self.path.exists():
-            self.objects = json.loads(self.path.read_text())
+            # encoding 必须显式给: Python 文本 I/O 默认用**系统 locale 编码**, 中文 Windows 上是 cp936(GBK)。
+            # 对象库是 UTF-8 的 JSON, 不写 encoding 会在中文 Windows 上读出乱码/解码失败
+            # (开发机是 UTF-8 locale 所以一直没暴露; 2026-09-11 从零重算时逮到)。
+            self.objects = json.loads(self.path.read_text(encoding='utf-8'))
 
     _DECISION_KINDS = {'验收树', '情景沙盘', '用户裁决', '验收状态机', '滚动元数据', '排程计划'}
 
@@ -76,6 +79,6 @@ class Store:
         self.path.parent.mkdir(parents=True, exist_ok=True)
         # F12 原子写: tmp + os.replace (同目录保证同文件系统, replace 原子)
         tmp = self.path.with_suffix(self.path.suffix + '.tmp')
-        tmp.write_text(json.dumps(self.objects, ensure_ascii=False, indent=1))
+        tmp.write_text(json.dumps(self.objects, ensure_ascii=False, indent=1), encoding='utf-8')
         os.replace(tmp, self.path)
         return str(self.path)