Jelajahi Sumber

振动页"看不到 2026-03 数据": 数据在(占 32%), 是呈现没说清区间 + 趋势聚合粒度

用户报: 输入振动数据时间范围 2026-03 至 2026-04, 重算后从页面看振动相关产物存在**没有呈现 2026-03 数据**的情况。

先查数据(逐件扫源, 不猜):
- 扫 data/raw/如东/windcms/CMS_RuDong_CGN_202603-04 全部 25,679 个 *_decode.json 正文时间键:
  源件最早记录 **2026-03-16 17:03:56**(逐台最早 03-16~03-30, 中位 03-17) ⇒ 源里**没有** 03-01~03-15 的记录。
- 重算出的窗 w0316 = 2026-03-16 17:27 → 2026-04-21 10:18, 2,066,686 行 38 台; 其中 3 月
  索引 661,935 行(32%)、谱 139,135 条(33%) ⇒ **3 月数据在产物里**。
⇒ 不是数据没进来, 是**呈现**上两处没说清:

1) 窗区间从没写出来。窗名 w0316 只是"窗起始日"代号, 页面上只有它 + "证据窗末 2026-04-21" ⇒ 看不出区间。
   新增 src/windcms/data.py::window_spans/span_text(实算, 只读索引 trigger_time+turbine, 进程内缓存);
   接进: CMS 首页副标题、**逐台页副标题(含"本台数据 起→止")**、标准模版报告头部与 §1/§5 说明事项、
   report.md 头部、/api/vibcms(新增 时间范围/窗 字段) + v2 振动评估页显示(lang.py 增 fus.cms_window 双语条目)。

2) 趋势图把 3 月压没了。workbench._trend_block 原按 `groupby('window').median()` 聚合, 而本机只有一个窗
   ⇒ 整条趋势只画**一个点**(时间戳≈窗内中位≈4 月上旬)。实测 WTG29/主轴承前/Peak 窗内 9 条记录、
   跨 2026-03-18 → 04-16 共 9 天。改为按 (窗, 日) 聚合 ⇒ 单窗也画逐日序列, 3 月如实出现在轴上。

顺带(体检查出来的两处噪声):
- raw_data_check 不再把 Access 锁/临时文件(.laccdb/.ldb/~$*)当"源件后缀不在约定内"(MDB 重建期间会留下它);
- logs/ops 动作日志按既有保留策略清理 1 份。

复检:
- /detail/api/vibcms → 时间范围 = w0316 2026-03-16 → 2026-04-21(逐台起始不等: 最早 03-16 · 最晚 03-30);
- /cms/turbines/WTG01.html 副标题 = 本台数据 2026-03-30 → 2026-04-21 · w0316 …; 页内 2026-03 出现 0 → 4 次;
- 标准模版报告 §5 增一条"本报告覆盖区间"的说明(窗名不是区间本身; 更早数据需现场重导);
- guanlan.py check 只剩三条已知 FAIL(人裁底稿两件 + 本机 Ollama 未启动)。

docs/页面输出体检_逐页原因_v0.1.md 增 §6 记这次核查的证据链与改动。
zhouyang.xie 3 minggu lalu
induk
melakukan
5f9acbef6a

+ 41 - 0
docs/页面输出体检_逐页原因_v0.1.md

@@ -114,3 +114,44 @@ src/windcms/report.py:504   md.append(l6[cols].to_markdown(index=False) …)
 `/detail/api/facts` 503(人裁底稿)、`/release/` 与 `/release/r1|r2/manifest.json`(随包发布层,研发出件)、
 `/detail/api/ask_status` 与 `/local-ai/`(本机 Ollama 未启动)。
 
+## 6. 追加:振动页"看不到 2026-03 数据"的核查(用户令 2026-09-19)
+
+用户报:"输入的振动数据(时间范围 2026-03 至 2026-04),重算后从页面看振动相关产物存在没有呈现 2026-03 振动数据的情况。"
+
+### 6.1 先查数据本身(逐件扫源,不猜)
+
+扫 `data/raw/如东/windcms/CMS_RuDong_CGN_202603-04/measurement/2026/{03,04}` 全部 **25,679** 个
+`*_decode.json` 的正文时间键:
+
+| 项 | 实测 |
+|---|---|
+| 源件最早记录 | **2026-03-16 17:03:56**(3 月目录内,38 台全扫) |
+| 逐台最早记录分布 | 最早一台 03-16 · 中位 03-17 · **最晚一台 03-30** |
+| 4 月目录里也有 3 月记录 | 最早 2026-03-28(导出按"文件落哪个目录"分月,不严格按记录时间) |
+| 重算出的窗 `w0316` | 2026-03-16 17:27 → 2026-04-21 10:18 · 2,066,686 行 · 38 台 |
+| 窗内 3 月占比 | 索引 **661,935 行(32%)** · 谱库 139,135 条(33%) |
+
+⇒ **3 月的数据在产物里**(占了三分之一),源件本身也**只从 03-16 起**(不存在 03-01~03-15 的记录)。
+所以"没有呈现 3 月"不是数据没进来,而是**呈现上没把覆盖区间与聚合粒度说清** —— 两处都改了:
+
+### 6.2 改了什么
+
+1. **窗区间实算并到处写明**(新增 `src/windcms/data.py::window_spans/span_text`):窗名 `w0316` 只是
+   "窗起始日 03-16"的代号,原来页面上只有它和"证据窗末 2026-04-21" ⇒ 看不出区间。
+   现在 CMS 首页副标题、**逐台页副标题(含"本台数据 起→止")**、标准模版报告头部与 §1/§5、`report.md`、
+   `/api/vibcms`(新增 `时间范围`/`窗` 字段,v2 振动评估页会显示)都写实算区间;
+   逐台起始不等的还会写明"最早/最晚"。
+2. **趋势图从"逐窗中位"改成"逐(窗,日)中位"**(`src/windcms/workbench.py::_trend_block`):
+   本机只有一个窗时,原来 `groupby('window').median()` 把整段区间压成**一个点**(时间戳≈4 月上旬)
+   ⇒ 3 月在图上彻底不见。实测 WTG29/主轴承前/Peak 窗内 9 条记录、跨 03-18 → 04-16 共 9 天,改后逐日可见。
+3. 顺带修两处体检噪声:`raw_data_check` 不再把 Access 锁文件(`.laccdb`/`.ldb`/`~$*`)当"源件后缀不合约定"
+   (实测 MDB 重建期间会留下它);`logs/ops` 动作日志按既有保留策略清理了 1 份。
+
+### 6.3 复检
+
+`/detail/api/vibcms` 现回:`时间范围 = w0316 2026-03-16 → 2026-04-21(逐台起始不等: 最早 2026-03-16 · 最晚 2026-03-30)`;
+`/cms/turbines/WTG01.html` 副标题 = `本台数据 2026-03-30 → 2026-04-21 · w0316 2026-03-16 → 2026-04-21(…)`,
+页内 `2026-03` 出现次数从 **0 → 4**(趋势逐日点带上 3 月);`guanlan.py check` 只剩三条已知 FAIL
+(`sop/findings.json` 与人裁底稿派生的事实契约、本机 Ollama 未启动)。
+
+

+ 459 - 450
scripts/raw_data_check.py

@@ -1,450 +1,459 @@
-#!/usr/bin/env python3
-# -*- coding: utf-8 -*-
-r"""输入数据放置体检 + 增量放置指导 (2026-09-17 用户令 3)。
-
-用户要的是两件事:
-  ① **检查**放进去的输入数据(特别是**增量**放的)是否符合规则;
-  ② 给出**放置指导**(该放哪、命名怎么写、放完跑什么)。
-
-## 规则 (每条都能指到具体文件, 依据 = 数据目录约定 + 各摄入脚本的真实读法)
-
-    R1 顶层只许 场站目录 / 共享资料(西门子4.0技术资料) / 说明文件(*.txt|*.md);
-       常见错误: 把 10 个 csv 直接解到 data/raw 根下, 或把整包 zip 丢在根上。
-    R2 场站目录下只许约定子目录 (src/windscada/config.py::STATION_SUBDIRS);
-       常见错误: 解压时多套一层包名目录 (如 `10分钟数据/10分钟数据/WTG01.csv`)。
-    R3 各源类的结构与命名:
-       scada_10min/  scada_1min/  平铺 `<机组>.csv`, 机组集合要与场配置一致
-       故障报警/      .xls/.xlsx/.xml, 不许多套层
-       风机故障记录/  年目录 + .xls/.xlsx/.rar/.jpg
-       油样报告/      两级 `台号/部件/*.pdf` (维护页按此分目录)
-       windcms/       含 CMS 原始导出: `*_decode.json` 或 `<包名>/measurement/**`
-       m5_cms_tcm/    含 handoff 正本 (handoff_vibration_v2.json / component_history.json) + 厂家报告/
-    R4 可读性抽样: CSV 能按 utf-8/gbk 解出非空表头; TCM `*_decode.json` 能解析出场站/机组字段。
-        (`--deep` 才做, 默认只查结构与命名, 免得在 GB 级目录上耗时)
-    R5 **增量语义** (给了 --src 时): 逐件对比"源包会落到哪"与"现在有什么", 出三清单:
-       新增 / 相同(将跳过) / **冲突**(同名不同大小 ⇒ 会覆盖已摄入的数据, 必须人来定夺)。
-    R6 时间空洞: scada_* 的文件名里若带年月, 检查是否缺月/重月 (数据层页面的时间轴会缺一段)。
-    R7 体量提示: 单文件 > 2 GB、目录 > 60 GB、或混进 zip/rar (应先解压再放)。
-    R8 放置指导: 按"新增了哪类源"给出该跑的重算命令 (台账/SCADA/振动/全量)。
-
-## 退出码
-    0 合规(可能带提示) · 5 结构或命名违例 · 6 增量冲突(同名不同大小) · 7 缺源类/时间空洞 · 8 仅有提示
-
-## 用法
-    python scripts/raw_data_check.py                      # 体检现有 data/raw/<场站>
-    python scripts/raw_data_check.py --deep               # 加可读性抽样
-    python scripts/raw_data_check.py --src <现场包目录>     # 增量体检: 出新增/相同/冲突三清单
-    python scripts/raw_data_check.py --src <现场包> --json
-    python scripts/raw_data_check.py --write-doc          # 把规则表与结论写进放置指导文档
-"""
-from __future__ import annotations
-
-import argparse
-import json
-import os
-import pathlib
-import re
-import sys
-import zipfile
-
-ROOT = pathlib.Path(__file__).resolve().parents[1]
-sys.path.insert(0, str(ROOT))
-from src import paths as P                                          # noqa: E402
-
-RC_OK, RC_STRUCT, RC_CLASH, RC_GAP, RC_NOTE = 0, 5, 6, 7, 8
-
-TOP_OK_DIRS = ('西门子4.0技术资料',)          # data/raw 下与"场站目录"并列的共享资料
-TOP_OK_FILES = ('.txt', '.md', '.pdf')        # 说明类文件
-ZIP_MOD = {'.zip', '.rar', '.7z', '.tar', '.gz', '.zst'}
-BIG_FILE = 2 * 1024 ** 3
-BIG_DIR = 60 * 1024 ** 3
-
-# 各类源的期望 (后缀集合, 说明) —— 与摄入脚本的真实读法一致 (见 place_raw_data.py 的映射表注释)
-EXPECT = {
-    'scada_10min': ({'.csv'}, '平铺 <机组>.csv (每台一个文件, 不许再套层)'),
-    'scada_1min': ({'.csv'}, '平铺 <机组>.csv'),
-    '故障报警': ({'.xls', '.xlsx', '.xml'}, '年度/季度报警导出 (.xls) + XML 报警导出'),
-    '风机故障记录': ({'.xls', '.xlsx', '.rar', '.jpg', '.png'}, '年目录 + 月度汇总表/现场照片'),
-    '油样报告': ({'.pdf'}, '两级: <台号>/<部件>/*.pdf'),
-    'windcms': ({'.json', '.pdf', '.docx'}, 'CMS 原始导出 *_decode.json(可套 <包名>/measurement/ 层) + 厂家报告/'),
-    'm5_cms_tcm': ({'.json', '.docx', '.pdf', '.md'}, 'handoff 正本 + 厂家报告/'),
-    # 2026-09-17 用户令: 现场交来的年度 MDB 归档(25年.zip / 26年.zip)就地放在这里 —— 它就是
-    # 原始收资(Access 按月分类),取代此前的"已导出 CSV"两个 zip(10分钟数据.zip / 1分钟数据.zip)。
-    # 允许 .zip(原样归档)+ .mdb(已解出时);解出的目录层次 `<年>/<月>/` 由 R2 的子目录规则放行。
-    'scada_mdb': ({'.zip', '.mdb'}, '年度归档 <年>年.zip(内层 <年>/<月>/<年-月-类>.zip → .mdb)'),
-}
-DATE_RE = re.compile(r'(20\d{2})[-_年]?(0[1-9]|1[0-2])')
-
-
-def station_dir(farm: str | None = None) -> pathlib.Path:
-    """当前场的原始件目录: data/raw/<场站目录名>。"""
-    from src.windscada import config as C
-    try:
-        d = C.raw_station_dir(farm)
-    except BaseException:
-        d = P.RAW_ROOT / (farm or P.farm())
-    return pathlib.Path(d)
-
-
-def sizes_map(base: pathlib.Path) -> dict:
-    """{相对路径: 字节数} —— 只算文件。"""
-    out = {}
-    if base.is_dir():
-        for f in base.rglob('*'):
-            if f.is_file():
-                out[f.relative_to(base).as_posix()] = f.stat().st_size
-    return out
-
-
-def check_tree(st: pathlib.Path, deep=False) -> tuple[list, list, dict]:
-    """规则体检 → (问题列表[(level, 项, 说明, rc)], 提示列表, 统计)。"""
-    res, notes = [], []
-    stat = dict(exists=st.is_dir(), files=0, bytes=0, subdirs={}, units=[])
-
-    # R1 data/raw 顶层
-    if P.RAW_ROOT.is_dir():
-        for p in sorted(P.RAW_ROOT.iterdir()):
-            if p.is_dir():
-                if p.name != st.name and p.name not in TOP_OK_DIRS:
-                    res.append(('!', f'data/raw/{p.name}', f'顶层目录未在约定内 (场站目录应为 {st.name}/; '
-                                                           f'共享资料放 {TOP_OK_DIRS})', RC_STRUCT))
-            elif p.suffix.lower() not in TOP_OK_FILES:
-                res.append(('!', f'data/raw/{p.name}', 'data/raw 顶层只放场站目录与说明文件; '
-                                                       '源件要放进 <场站>/<源类>/', RC_STRUCT))
-    if not st.is_dir():
-        notes.append(('i', f'{P.rel(st)}', '尚未放置输入数据 —— 页面会显示"无产物/无数据" (属预期); '
-                                           '放置见 docs/输入数据放置指导_v0.1.md', RC_OK))
-        return res, notes, stat
-
-    # R2 场站内子目录 + R3 各源类
-    from src.windscada import config as C
-    allowed = set(C.STATION_SUBDIRS)
-    for d in sorted(st.iterdir()):
-        if d.is_file():
-            if d.suffix.lower() in ZIP_MOD:
-                res.append(('!', P.rel(d), '压缩包直接放在场站目录下 —— 先解压再按源类落位', RC_NOTE))
-            else:
-                notes.append(('i', P.rel(d), '场站目录下的散装文件 (多半是现场随手拷的; 建议归入对应源类子目录)', RC_OK))
-            continue
-        if d.name not in allowed:
-            res.append(('!', f'{P.rel(d)}', f'未约定的子目录 (约定: {sorted(allowed)})'
-                                           ' —— 常见成因: 解压时多套了一层包名', RC_STRUCT))
-            continue
-        files = [f for f in d.rglob('*') if f.is_file()]
-        nbytes = sum(f.stat().st_size for f in files)
-        stat['subdirs'][d.name] = dict(files=len(files), bytes=nbytes)
-        stat['files'] += len(files)
-        stat['bytes'] += nbytes
-        exts, bad = {}, []
-        for f in files:
-            e = f.suffix.lower()
-            exts[e] = exts.get(e, 0) + 1
-            want = EXPECT.get(d.name, (None, ''))[0]
-            if want and e not in want and e not in ZIP_MOD:
-                bad.append(f)
-        if bad:
-            sample = ', '.join(sorted({f.suffix.lower() for f in bad})[:4])
-            res.append(('!', P.rel(d), f'{len(bad)} 个文件后缀不在该源类的约定内 ({sample}); '
-                                       f'约定: {sorted(EXPECT[d.name][0])}', RC_STRUCT))
-        if not files:
-            notes.append(('i', P.rel(d), '目录是空的 (放了数据才有用)', RC_OK))
-        # 各类专项
-        if d.name in ('scada_10min', 'scada_1min'):
-            names = {f.name for f in files}
-            nested = [f for f in files if len(f.relative_to(d).parts) > 1]
-            if nested:
-                res.append(('!', P.rel(d), f'{len(nested)} 个文件在子目录里 —— 摄入按 <源类>/<机组>.csv 平铺读, '
-                                           f'多一层就读不到', RC_STRUCT))
-            try:
-                want = set(C.farm(None)['turbines'])
-                n_want = len(want)
-            except BaseException:
-                want, n_want = set(), 0
-            got = {f.stem for f in files if f.parent == d}
-            # ★ 两类命名都对: scada_10min 用显示机组号 (WTG01.csv); scada_1min 的现场导出用的是
-            #   **内部台号** (01E.csv…37B.csv, 见 place_raw_data.py 的映射说明) —— 后者按"台数一致"判完整,
-            #   不再按名字判 (2026-09-17 实测: 按 WTG 名字判会把 38/38 齐全的 1min 目录误报成"缺 38 台")。
-            if d.name == 'scada_10min':
-                miss = sorted(want - got) if want else []
-                extra = sorted(got - want) if want else []
-                stat['units'].append((d.name, len(got)))
-                if miss:
-                    res.append(('!', P.rel(d), f'缺 {len(miss)} 台机组的数据: {miss[:6]}'
-                                               f'{" …" if len(miss) > 6 else ""} (页面按台取数, 缺台就是空白)', RC_GAP))
-                if extra:
-                    res.append(('!', P.rel(d), f'{len(extra)} 个文件名不是本场机组: {extra[:6]}', RC_STRUCT))
-            else:
-                stat['units'].append((d.name, len(got)))
-                if n_want and len(got) != n_want:
-                    res.append(('!', P.rel(d), f'台数不符: 现有 {len(got)} 个 <台号>.csv, 场配置是 {n_want} 台'
-                                               f' (1min 导出用内部台号 01E/37B 这类, 按台数判完整)', RC_GAP))
-            # R6 时间空洞 (文件名带年月时才判)
-            months = sorted({m.group(0) for m in (DATE_RE.search(f.stem) for f in files) if m})
-            if months:
-                stat['months'] = months
-        if d.name == '油样报告':
-            flat = [f for f in files if len(f.relative_to(d).parts) < 2]
-            if flat:
-                res.append(('!', P.rel(d), f'{len(flat)} 个 PDF 直接躺在源类目录下 —— 约定是 <台号>/<部件>/*.pdf '
-                                           f'(维护页按此分组)', RC_STRUCT))
-        if d.name == 'windcms':
-            js = [f for f in files if f.suffix.lower() == '.json']
-            dec = [f for f in js if f.name.endswith('_decode.json')]           # TCM 导出本体
-            if not dec:
-                res.append(('!', P.rel(d), '没有 *_decode.json (CMS 原始导出) —— 振动侧就没有可重算的源件', RC_GAP))
-            if deep and dec:
-                # 抽样核对"摄入脚本真正要的结构" —— 依据 scripts/rudong_tcm_index.py:
-                #   doc['body']['body'][<ISO时间戳>] = [ {Record: {...}}, … ], 机组取 Record.Location / Record.LocationName
-                # (顶层还有 method/controller/buildId 等信封字段, 别拿它们当判据 —— 2026-09-17 实测踩过这个坑)
-                import random as _rnd
-                _rnd.seed(0)
-                sample = _rnd.sample(dec, min(12, len(dec)))
-                ok_n, bad_n, shapes = 0, 0, {}
-                for f in sample:
-                    try:
-                        o = json.loads(f.read_text(encoding='utf-8', errors='replace'))
-                        body = ((o or {}).get('body') or {}).get('body') if isinstance(o, dict) else None
-                        if not isinstance(body, dict) or not body:
-                            shapes[str(sorted(o)[:4]) if isinstance(o, dict) else type(o).__name__] = \
-                                shapes.get(str(sorted(o)[:4]) if isinstance(o, dict) else type(o).__name__, 0) + 1
-                            bad_n += 1
-                            continue
-                        recs = next((v for v in body.values() if isinstance(v, list) and v), None)
-                        keys = sorted((recs[0] or {}).get('Record', {}) or {}) if recs else []
-                        if recs and (('Location' in keys) or ('LocationName' in keys)):
-                            ok_n += 1
-                        else:
-                            shapes[f'Record 缺 Location/LocationName: {keys[:6]}'] = \
-                                shapes.get(f'Record 缺 Location/LocationName: {keys[:6]}', 0) + 1
-                            bad_n += 1
-                    except Exception as e:
-                        shapes[f'解析失败 {type(e).__name__}'] = shapes.get(f'解析失败 {type(e).__name__}', 0) + 1
-                        bad_n += 1
-                stat['tcm_sample'] = dict(ok=ok_n, bad=bad_n, shapes=shapes)
-                if bad_n:
-                    res.append(('!', P.rel(d), f'TCM 抽样 {len(sample)} 件里 {bad_n} 件结构对不上摄入脚本的取法 '
-                                               f'(body.body[时间戳]=[{{Record:…}}], 机组取 Record.Location/LocationName): '
-                                               f'{list(shapes.items())[:3]}', RC_STRUCT))
-                else:
-                    stat['tcm_ok'] = True
-        if d.name == 'm5_cms_tcm':
-            hand = [f for f in files if f.name in ('handoff_vibration_v2.json', 'component_history.json')]
-            if not hand:
-                notes.append(('i', P.rel(d), '缺 handoff 正本 (handoff_vibration_v2.json / component_history.json) —— '
-                                             '融合面判级与 /cms/ 只能用随包快照 (docs §7 的已知缺口)', RC_OK))
-
-    # R4 可读性抽样 (CSV 编码/表头)
-    if deep:
-        for name in ('scada_10min', 'scada_1min'):
-            d = st / name
-            if d.is_dir():
-                f = next(iter(sorted(x for x in d.glob('*.csv'))), None)
-                if f:
-                    head = None
-                    for enc in ('utf-8-sig', 'gbk'):
-                        try:
-                            with f.open(encoding=enc, errors='strict') as fh:
-                                head = fh.readline().strip()
-                            stat[f'{name}_enc'] = enc
-                            break
-                        except Exception:
-                            continue
-                    if not head:
-                        res.append(('!', P.rel(f), 'CSV 表头读不出来 (utf-8/gbk 都不行)', RC_STRUCT))
-                    elif len(head.split(',')) < 5:
-                        res.append(('!', P.rel(f), f'表头列数异常 ({len(head.split(","))} 列): {head[:60]}', RC_STRUCT))
-
-    # R7 体量
-    if stat['bytes'] > BIG_DIR:
-        notes.append(('?', P.rel(st), f'源件合计 {stat["bytes"]/2**30:.1f} GB —— 摄入与重算会很慢, 建议分批', RC_NOTE))
-    return res, notes, stat
-
-
-def increment_findings(plan: dict, src_name: str = '') -> list:
-    """把增量三清单里的**冲突**变成一条可拦人的问题 (供本模块与 place_raw_data 共用)。"""
-    out = []
-    if plan.get('clash'):
-        c = plan['clash'][0]
-        out.append(('!', f'--src {src_name}'.strip(), f'{len(plan["clash"])} 件**同名不同大小**: 直接放会覆盖已摄入的源件 '
-                                                        f'(示例 {c["path"]}: 现有 {c.get("now", 0):,} B → 源包 '
-                                                        f'{c["size"]:,} B) —— 先确认哪一份是对的, 再决定 --force 覆盖还是改名并存',
-                    RC_CLASH))
-    return out
-
-
-def increment_src(src: pathlib.Path, st: pathlib.Path) -> tuple[list, dict]:
-    """R5 增量三清单: 遍历源包 zip 成员/散件, 算出它们会落到哪, 与现有树对比。"""
-    res, plan = [], dict(new=[], same=[], clash=[], unmapped=[])
-    import importlib.util
-    spec = importlib.util.spec_from_file_location('_place', ROOT / 'scripts' / 'place_raw_data.py')
-    place = importlib.util.module_from_spec(spec)
-    try:
-        spec.loader.exec_module(place)
-    except SystemExit:
-        pass
-    rules = list(getattr(place, 'RULES', [])) + list(getattr(place, 'RULES_MECH', [])) + list(getattr(place, 'RULES_VIB', []))
-    loose = list(getattr(place, 'RULES_VIB_LOOSE', []))
-    have = sizes_map(st)
-    have_root = sizes_map(P.RAW_ROOT)
-
-    def classify(rel_path: str, size: int, base: pathlib.Path):
-        target = base / rel_path
-        cur = have.get(rel_path) if base == st else have_root.get(rel_path)
-        item = dict(path=rel_path, size=size)
-        if cur is None:
-            plan['new'].append(item)
-        elif cur == size:
-            plan['same'].append(item)
-        else:
-            item['now'] = cur
-            plan['clash'].append(item)
-
-    for zname, prefix, target, strip, root_kind in rules:
-        zp = src / zname if src.is_dir() else None
-        if not zp or not zp.is_file():
-            continue
-        base = st if root_kind == 'station' else P.RAW_ROOT
-        try:
-            with zipfile.ZipFile(zp) as z:
-                for info in z.infolist():
-                    if info.is_dir():
-                        continue
-                    nm = place.gbk_name(info)
-                    if prefix and not nm.startswith(prefix):
-                        continue
-                    tail = nm[len(prefix):] if prefix else nm
-                    if strip:
-                        tail = tail.split('/', 1)[1] if '/' in tail else tail
-                    rel = f'{target}/{tail}'
-                    if prefix and nm == prefix.rstrip('/'):
-                        continue
-                    classify(rel, info.file_size, base)
-        except zipfile.BadZipFile:
-            res.append(('!', zp.name, '不是合法 zip (现场包常被截断; 重新拷一份)', RC_STRUCT))
-    for fn, target in loose:
-        fp = src / fn
-        if fp.is_file():
-            classify(f'{target}/{fn}', fp.stat().st_size, st)
-    return res, plan
-
-
-def guidance(plan: dict, stat: dict) -> list[str]:
-    """R8 放置指导: 按新增源类给出该跑的重算命令与预期变化。"""
-    out = []
-    new = [x['path'] for x in plan.get('new', [])]
-    kinds = {p.split('/', 1)[0] for p in new}
-    if not new:
-        out.append('没有新增源件: 不必重算 (若刚换过数据, 用 scripts/rebuild_all.py --dry-run 看计划)')
-        return out
-    out.append(f'新增 {len(new)} 件, 涉及源类: {sorted(kinds)}')
-    if 'scada_10min' in kinds or 'scada_1min' in kinds:
-        out.append('· SCADA 侧变了 → `python scripts/rebuild_all.py` (含 10 个构建器, 约 15 分钟); '
-                   '台账/报警/油样一起换时同一条命令即可')
-    if kinds & {'故障报警', '风机故障记录', '油样报告'}:
-        out.append('· 台账类变了 → `python scripts/rebuild_all.py --skip-scada` (约 2 分钟)')
-    if kinds & {'windcms', 'm5_cms_tcm'}:
-        out.append('· 振动侧变了 → 重算链第 ④b 步会自动摄入 (`scripts/vib_raw_build.py`); '
-                   '注意它只做索引/谱, 不会重生成报告 (本包缺六层链, 见 docs §7)')
-    out.append('· 放完先自检: `python scripts/raw_data_check.py --deep` → 再看 `/detail/` 左栏「系统维护」两屏')
-    out.append('· 重算后核对锚点: 报警 39211 · 工单 5876 · 油样 404 · temp_monthly 19494 · 本体 9702 对象')
-    return out
-
-
-def main() -> int:
-    ap = argparse.ArgumentParser()
-    ap.add_argument('--farm', default=None)
-    ap.add_argument('--src', default=None, help='现场数据包目录 (给了就做增量体检: 新增/相同/冲突三清单)')
-    ap.add_argument('--deep', action='store_true', help='加可读性抽样 (CSV 编码/表头, TCM json 字段)')
-    ap.add_argument('--json', action='store_true')
-    ap.add_argument('--write-doc', action='store_true', help='把结论写进 docs/输入数据放置指导_v0.1.md')
-    a = ap.parse_args()
-
-    st = station_dir(a.farm)
-    res, notes, stat = check_tree(st, deep=a.deep)
-    plan = {}
-    if a.src:
-        src = pathlib.Path(a.src)
-        if not src.is_dir():
-            print(f'[X] --src 不是目录: {src}')
-            return RC_STRUCT
-        r2, plan = increment_src(src, st)
-        res += r2 + increment_findings(plan, src.name)
-    guide = guidance(plan, stat)
-
-    if a.json:
-        print(json.dumps(dict(station=P.rel(st), results=[dict(level=x[0], item=x[1], note=x[2], rc=x[3]) for x in res],
-                              notes=[dict(level=x[0], item=x[1], note=x[2]) for x in notes], stat=stat, plan=plan,
-                              guidance=guide), ensure_ascii=False, indent=1, default=str))
-    else:
-        print(f'== 输入数据放置体检 · {P.rel(st)} ==')
-        if stat.get('exists'):
-            print(f'   {stat["files"]:,} 件 / {stat["bytes"]/2**30:.2f} GB')
-            for k, v in sorted(stat['subdirs'].items()):
-                print(f'     {k:14s} {v["files"]:7,d} 件  {v["bytes"]/2**20:9.1f} MB')
-        for level, item, note, rc in res:
-            print(f'   [{level}] {item}: {note}')
-        seen = set()
-        for level, item, note, rc in notes:
-            if note in seen:
-                continue
-            seen.add(note)
-            print(f'   [{level}] {item}: {note}')
-        if plan:
-            print('\n== 增量放置三清单 (源包 → data/raw) ==')
-            print(f'   新增 {len(plan["new"]):,} 件 (将拷入) · 相同 {len(plan["same"]):,} 件 (同尺寸, 跳过) · '
-                  f'冲突 {len(plan["clash"]):,} 件 (同名不同大小, **需人确认**)')
-            for k, label in (('new', '新增'), ('clash', '冲突')):
-                for x in plan[k][:5]:
-                    extra = f' (现有 {x.get("now"):,} B)' if k == 'clash' else ''
-                    print(f'     [{label}] {x["path"]}  {x["size"]:,} B{extra}')
-                if len(plan[k]) > 5:
-                    print(f'     … 另有 {len(plan[k]) - 5} 件{label}')
-        print('\n== 放置指导 ==')
-        for g in guide:
-            print(f'   {g}')
-    rc_map = {x[3] for x in res if x[0] in ('X', '!') and x[3]}
-    rc = max(rc_map) if rc_map else (RC_NOTE if notes else RC_OK)
-    if not a.json:
-        print(f'\n结论: {"合规" if rc == 0 else "见上"}; 退出码 {rc}')
-    if a.write_doc:
-        write_doc(res, notes, stat, plan, guide)
-    return rc
-
-
-DOC = P.ROOT / 'docs' / '输入数据放置指导_v0.1.md'
-BEGIN, END = '<!-- RAWCHECK:BEGIN -->', '<!-- RAWCHECK:END -->'
-
-
-def write_doc(res, notes, stat, plan, guide) -> None:
-    body = [BEGIN, '### 体检结论(自动生成,勿手改)', '',
-            f'· 源件目录 `{P.rel(station_dir())}`:{stat.get("files", 0):,} 件 / {stat.get("bytes", 0)/2**30:.2f} GB']
-    for k, v in sorted((stat.get('subdirs') or {}).items()):
-        body.append(f'  - `{k}/`:{v["files"]:,} 件 / {v["bytes"]/2**20:.1f} MB')
-    if res:
-        body += ['', '| 级别 | 项 | 说明 |', '|---|---|---|']
-        body += [f'| `{x[0]}` | `{x[1]}` | {x[2]} |' for x in res[:40]]
-    else:
-        body += ['', '· 结构与命名检查:**无不一致**(提示见下)']
-    for x in notes[:12]:
-        body.append(f'· `{x[0]}` {x[1]}:{x[2]}')
-    if plan:
-        body += ['', f'· 增量(源包 → data/raw):新增 {len(plan["new"]):,} · 相同 {len(plan["same"]):,} · '
-                     f'冲突 {len(plan["clash"]):,}']
-    body += ['', '**放置后该做什么**:'] + [f'· {g}' for g in guide] + ['', END]
-    block = '\n'.join(body)
-    if DOC.is_file():
-        t = DOC.read_text(encoding='utf-8')
-        if BEGIN in t and END in t:
-            t = re.sub(re.escape(BEGIN) + r'.*?' + re.escape(END), lambda m: block, t, flags=re.S)
-        else:
-            t = t.rstrip() + '\n\n' + block + '\n'
-    else:
-        t = '# 输入数据放置指导\n\n' + block + '\n'
-    DOC.write_text(t, encoding='utf-8')
-    print(f'   已写入 {P.rel(DOC)}')
-
-
-if __name__ == '__main__':
-    from src import console
-    console.soft()
-    sys.exit(main())
+#!/usr/bin/env python3
+# -*- coding: utf-8 -*-
+r"""输入数据放置体检 + 增量放置指导 (2026-09-17 用户令 3)。
+
+用户要的是两件事:
+  ① **检查**放进去的输入数据(特别是**增量**放的)是否符合规则;
+  ② 给出**放置指导**(该放哪、命名怎么写、放完跑什么)。
+
+## 规则 (每条都能指到具体文件, 依据 = 数据目录约定 + 各摄入脚本的真实读法)
+
+    R1 顶层只许 场站目录 / 共享资料(西门子4.0技术资料) / 说明文件(*.txt|*.md);
+       常见错误: 把 10 个 csv 直接解到 data/raw 根下, 或把整包 zip 丢在根上。
+    R2 场站目录下只许约定子目录 (src/windscada/config.py::STATION_SUBDIRS);
+       常见错误: 解压时多套一层包名目录 (如 `10分钟数据/10分钟数据/WTG01.csv`)。
+    R3 各源类的结构与命名:
+       scada_10min/  scada_1min/  平铺 `<机组>.csv`, 机组集合要与场配置一致
+       故障报警/      .xls/.xlsx/.xml, 不许多套层
+       风机故障记录/  年目录 + .xls/.xlsx/.rar/.jpg
+       油样报告/      两级 `台号/部件/*.pdf` (维护页按此分目录)
+       windcms/       含 CMS 原始导出: `*_decode.json` 或 `<包名>/measurement/**`
+       m5_cms_tcm/    含 handoff 正本 (handoff_vibration_v2.json / component_history.json) + 厂家报告/
+    R4 可读性抽样: CSV 能按 utf-8/gbk 解出非空表头; TCM `*_decode.json` 能解析出场站/机组字段。
+        (`--deep` 才做, 默认只查结构与命名, 免得在 GB 级目录上耗时)
+    R5 **增量语义** (给了 --src 时): 逐件对比"源包会落到哪"与"现在有什么", 出三清单:
+       新增 / 相同(将跳过) / **冲突**(同名不同大小 ⇒ 会覆盖已摄入的数据, 必须人来定夺)。
+    R6 时间空洞: scada_* 的文件名里若带年月, 检查是否缺月/重月 (数据层页面的时间轴会缺一段)。
+    R7 体量提示: 单文件 > 2 GB、目录 > 60 GB、或混进 zip/rar (应先解压再放)。
+    R8 放置指导: 按"新增了哪类源"给出该跑的重算命令 (台账/SCADA/振动/全量)。
+
+## 退出码
+    0 合规(可能带提示) · 5 结构或命名违例 · 6 增量冲突(同名不同大小) · 7 缺源类/时间空洞 · 8 仅有提示
+
+## 用法
+    python scripts/raw_data_check.py                      # 体检现有 data/raw/<场站>
+    python scripts/raw_data_check.py --deep               # 加可读性抽样
+    python scripts/raw_data_check.py --src <现场包目录>     # 增量体检: 出新增/相同/冲突三清单
+    python scripts/raw_data_check.py --src <现场包> --json
+    python scripts/raw_data_check.py --write-doc          # 把规则表与结论写进放置指导文档
+"""
+from __future__ import annotations
+
+import argparse
+import json
+import os
+import pathlib
+import re
+import sys
+import zipfile
+
+ROOT = pathlib.Path(__file__).resolve().parents[1]
+sys.path.insert(0, str(ROOT))
+from src import paths as P                                          # noqa: E402
+
+RC_OK, RC_STRUCT, RC_CLASH, RC_GAP, RC_NOTE = 0, 5, 6, 7, 8
+
+TOP_OK_DIRS = ('西门子4.0技术资料',)          # data/raw 下与"场站目录"并列的共享资料
+TOP_OK_FILES = ('.txt', '.md', '.pdf')        # 说明类文件
+ZIP_MOD = {'.zip', '.rar', '.7z', '.tar', '.gz', '.zst'}
+
+# Access 读写 .mdb 时留下的锁/临时文件 (不是源数据; 2026-09-19 实逮: MDB 重建期间 scada_mdb 下出现 .laccdb)
+ACCESS_TMP_SUFFIX = {".laccdb", ".ldb", ".tmp"}
+BIG_FILE = 2 * 1024 ** 3
+BIG_DIR = 60 * 1024 ** 3
+
+# 各类源的期望 (后缀集合, 说明) —— 与摄入脚本的真实读法一致 (见 place_raw_data.py 的映射表注释)
+EXPECT = {
+    'scada_10min': ({'.csv'}, '平铺 <机组>.csv (每台一个文件, 不许再套层)'),
+    'scada_1min': ({'.csv'}, '平铺 <机组>.csv'),
+    '故障报警': ({'.xls', '.xlsx', '.xml'}, '年度/季度报警导出 (.xls) + XML 报警导出'),
+    '风机故障记录': ({'.xls', '.xlsx', '.rar', '.jpg', '.png'}, '年目录 + 月度汇总表/现场照片'),
+    '油样报告': ({'.pdf'}, '两级: <台号>/<部件>/*.pdf'),
+    'windcms': ({'.json', '.pdf', '.docx'}, 'CMS 原始导出 *_decode.json(可套 <包名>/measurement/ 层) + 厂家报告/'),
+    'm5_cms_tcm': ({'.json', '.docx', '.pdf', '.md'}, 'handoff 正本 + 厂家报告/'),
+    # 2026-09-17 用户令: 现场交来的年度 MDB 归档(25年.zip / 26年.zip)就地放在这里 —— 它就是
+    # 原始收资(Access 按月分类),取代此前的"已导出 CSV"两个 zip(10分钟数据.zip / 1分钟数据.zip)。
+    # 允许 .zip(原样归档)+ .mdb(已解出时);解出的目录层次 `<年>/<月>/` 由 R2 的子目录规则放行。
+    'scada_mdb': ({'.zip', '.mdb'}, '年度归档 <年>年.zip(内层 <年>/<月>/<年-月-类>.zip → .mdb)'),
+}
+DATE_RE = re.compile(r'(20\d{2})[-_年]?(0[1-9]|1[0-2])')
+
+
+def station_dir(farm: str | None = None) -> pathlib.Path:
+    """当前场的原始件目录: data/raw/<场站目录名>。"""
+    from src.windscada import config as C
+    try:
+        d = C.raw_station_dir(farm)
+    except BaseException:
+        d = P.RAW_ROOT / (farm or P.farm())
+    return pathlib.Path(d)
+
+
+def sizes_map(base: pathlib.Path) -> dict:
+    """{相对路径: 字节数} —— 只算文件。"""
+    out = {}
+    if base.is_dir():
+        for f in base.rglob('*'):
+            if f.is_file():
+                out[f.relative_to(base).as_posix()] = f.stat().st_size
+    return out
+
+
+def check_tree(st: pathlib.Path, deep=False) -> tuple[list, list, dict]:
+    """规则体检 → (问题列表[(level, 项, 说明, rc)], 提示列表, 统计)。"""
+    res, notes = [], []
+    stat = dict(exists=st.is_dir(), files=0, bytes=0, subdirs={}, units=[])
+
+    # R1 data/raw 顶层
+    if P.RAW_ROOT.is_dir():
+        for p in sorted(P.RAW_ROOT.iterdir()):
+            if p.is_dir():
+                if p.name != st.name and p.name not in TOP_OK_DIRS:
+                    res.append(('!', f'data/raw/{p.name}', f'顶层目录未在约定内 (场站目录应为 {st.name}/; '
+                                                           f'共享资料放 {TOP_OK_DIRS})', RC_STRUCT))
+            elif p.suffix.lower() not in TOP_OK_FILES:
+                res.append(('!', f'data/raw/{p.name}', 'data/raw 顶层只放场站目录与说明文件; '
+                                                       '源件要放进 <场站>/<源类>/', RC_STRUCT))
+    if not st.is_dir():
+        notes.append(('i', f'{P.rel(st)}', '尚未放置输入数据 —— 页面会显示"无产物/无数据" (属预期); '
+                                           '放置见 docs/输入数据放置指导_v0.1.md', RC_OK))
+        return res, notes, stat
+
+    # R2 场站内子目录 + R3 各源类
+    from src.windscada import config as C
+    allowed = set(C.STATION_SUBDIRS)
+    for d in sorted(st.iterdir()):
+        if d.is_file():
+            if d.suffix.lower() in ZIP_MOD:
+                res.append(('!', P.rel(d), '压缩包直接放在场站目录下 —— 先解压再按源类落位', RC_NOTE))
+            else:
+                notes.append(('i', P.rel(d), '场站目录下的散装文件 (多半是现场随手拷的; 建议归入对应源类子目录)', RC_OK))
+            continue
+        if d.name not in allowed:
+            res.append(('!', f'{P.rel(d)}', f'未约定的子目录 (约定: {sorted(allowed)})'
+                                           ' —— 常见成因: 解压时多套了一层包名', RC_STRUCT))
+            continue
+        files = [f for f in d.rglob('*') if f.is_file()]
+        nbytes = sum(f.stat().st_size for f in files)
+        stat['subdirs'][d.name] = dict(files=len(files), bytes=nbytes)
+        stat['files'] += len(files)
+        stat['bytes'] += nbytes
+        exts, bad = {}, []
+        for f in files:
+            e = f.suffix.lower()
+            exts[e] = exts.get(e, 0) + 1
+            want = EXPECT.get(d.name, (None, ''))[0]
+            # ★2026-09-19: Access 的**锁文件/临时文件**(.laccdb/.ldb/~$*) 不是源数据 ——
+            #   读/写 .mdb 时 Access 会在同目录留下它(实测: MDB 重建期间 scada_mdb 下出现 .laccdb),
+            #   按"后缀不在约定内"报结构问题会把人引到错的地方。忽略并如实注明。
+            if e in ACCESS_TMP_SUFFIX or f.name.startswith('~$'):
+                notes.append(('i', P.rel(f), 'Access 锁/临时文件 (读写 .mdb 时产生), 不算源件', RC_OK))
+                continue
+            if want and e not in want and e not in ZIP_MOD:
+                bad.append(f)
+        if bad:
+            sample = ', '.join(sorted({f.suffix.lower() for f in bad})[:4])
+            res.append(('!', P.rel(d), f'{len(bad)} 个文件后缀不在该源类的约定内 ({sample}); '
+                                       f'约定: {sorted(EXPECT[d.name][0])}', RC_STRUCT))
+        if not files:
+            notes.append(('i', P.rel(d), '目录是空的 (放了数据才有用)', RC_OK))
+        # 各类专项
+        if d.name in ('scada_10min', 'scada_1min'):
+            names = {f.name for f in files}
+            nested = [f for f in files if len(f.relative_to(d).parts) > 1]
+            if nested:
+                res.append(('!', P.rel(d), f'{len(nested)} 个文件在子目录里 —— 摄入按 <源类>/<机组>.csv 平铺读, '
+                                           f'多一层就读不到', RC_STRUCT))
+            try:
+                want = set(C.farm(None)['turbines'])
+                n_want = len(want)
+            except BaseException:
+                want, n_want = set(), 0
+            got = {f.stem for f in files if f.parent == d}
+            # ★ 两类命名都对: scada_10min 用显示机组号 (WTG01.csv); scada_1min 的现场导出用的是
+            #   **内部台号** (01E.csv…37B.csv, 见 place_raw_data.py 的映射说明) —— 后者按"台数一致"判完整,
+            #   不再按名字判 (2026-09-17 实测: 按 WTG 名字判会把 38/38 齐全的 1min 目录误报成"缺 38 台")。
+            if d.name == 'scada_10min':
+                miss = sorted(want - got) if want else []
+                extra = sorted(got - want) if want else []
+                stat['units'].append((d.name, len(got)))
+                if miss:
+                    res.append(('!', P.rel(d), f'缺 {len(miss)} 台机组的数据: {miss[:6]}'
+                                               f'{" …" if len(miss) > 6 else ""} (页面按台取数, 缺台就是空白)', RC_GAP))
+                if extra:
+                    res.append(('!', P.rel(d), f'{len(extra)} 个文件名不是本场机组: {extra[:6]}', RC_STRUCT))
+            else:
+                stat['units'].append((d.name, len(got)))
+                if n_want and len(got) != n_want:
+                    res.append(('!', P.rel(d), f'台数不符: 现有 {len(got)} 个 <台号>.csv, 场配置是 {n_want} 台'
+                                               f' (1min 导出用内部台号 01E/37B 这类, 按台数判完整)', RC_GAP))
+            # R6 时间空洞 (文件名带年月时才判)
+            months = sorted({m.group(0) for m in (DATE_RE.search(f.stem) for f in files) if m})
+            if months:
+                stat['months'] = months
+        if d.name == '油样报告':
+            flat = [f for f in files if len(f.relative_to(d).parts) < 2]
+            if flat:
+                res.append(('!', P.rel(d), f'{len(flat)} 个 PDF 直接躺在源类目录下 —— 约定是 <台号>/<部件>/*.pdf '
+                                           f'(维护页按此分组)', RC_STRUCT))
+        if d.name == 'windcms':
+            js = [f for f in files if f.suffix.lower() == '.json']
+            dec = [f for f in js if f.name.endswith('_decode.json')]           # TCM 导出本体
+            if not dec:
+                res.append(('!', P.rel(d), '没有 *_decode.json (CMS 原始导出) —— 振动侧就没有可重算的源件', RC_GAP))
+            if deep and dec:
+                # 抽样核对"摄入脚本真正要的结构" —— 依据 scripts/rudong_tcm_index.py:
+                #   doc['body']['body'][<ISO时间戳>] = [ {Record: {...}}, … ], 机组取 Record.Location / Record.LocationName
+                # (顶层还有 method/controller/buildId 等信封字段, 别拿它们当判据 —— 2026-09-17 实测踩过这个坑)
+                import random as _rnd
+                _rnd.seed(0)
+                sample = _rnd.sample(dec, min(12, len(dec)))
+                ok_n, bad_n, shapes = 0, 0, {}
+                for f in sample:
+                    try:
+                        o = json.loads(f.read_text(encoding='utf-8', errors='replace'))
+                        body = ((o or {}).get('body') or {}).get('body') if isinstance(o, dict) else None
+                        if not isinstance(body, dict) or not body:
+                            shapes[str(sorted(o)[:4]) if isinstance(o, dict) else type(o).__name__] = \
+                                shapes.get(str(sorted(o)[:4]) if isinstance(o, dict) else type(o).__name__, 0) + 1
+                            bad_n += 1
+                            continue
+                        recs = next((v for v in body.values() if isinstance(v, list) and v), None)
+                        keys = sorted((recs[0] or {}).get('Record', {}) or {}) if recs else []
+                        if recs and (('Location' in keys) or ('LocationName' in keys)):
+                            ok_n += 1
+                        else:
+                            shapes[f'Record 缺 Location/LocationName: {keys[:6]}'] = \
+                                shapes.get(f'Record 缺 Location/LocationName: {keys[:6]}', 0) + 1
+                            bad_n += 1
+                    except Exception as e:
+                        shapes[f'解析失败 {type(e).__name__}'] = shapes.get(f'解析失败 {type(e).__name__}', 0) + 1
+                        bad_n += 1
+                stat['tcm_sample'] = dict(ok=ok_n, bad=bad_n, shapes=shapes)
+                if bad_n:
+                    res.append(('!', P.rel(d), f'TCM 抽样 {len(sample)} 件里 {bad_n} 件结构对不上摄入脚本的取法 '
+                                               f'(body.body[时间戳]=[{{Record:…}}], 机组取 Record.Location/LocationName): '
+                                               f'{list(shapes.items())[:3]}', RC_STRUCT))
+                else:
+                    stat['tcm_ok'] = True
+        if d.name == 'm5_cms_tcm':
+            hand = [f for f in files if f.name in ('handoff_vibration_v2.json', 'component_history.json')]
+            if not hand:
+                notes.append(('i', P.rel(d), '缺 handoff 正本 (handoff_vibration_v2.json / component_history.json) —— '
+                                             '融合面判级与 /cms/ 只能用随包快照 (docs §7 的已知缺口)', RC_OK))
+
+    # R4 可读性抽样 (CSV 编码/表头)
+    if deep:
+        for name in ('scada_10min', 'scada_1min'):
+            d = st / name
+            if d.is_dir():
+                f = next(iter(sorted(x for x in d.glob('*.csv'))), None)
+                if f:
+                    head = None
+                    for enc in ('utf-8-sig', 'gbk'):
+                        try:
+                            with f.open(encoding=enc, errors='strict') as fh:
+                                head = fh.readline().strip()
+                            stat[f'{name}_enc'] = enc
+                            break
+                        except Exception:
+                            continue
+                    if not head:
+                        res.append(('!', P.rel(f), 'CSV 表头读不出来 (utf-8/gbk 都不行)', RC_STRUCT))
+                    elif len(head.split(',')) < 5:
+                        res.append(('!', P.rel(f), f'表头列数异常 ({len(head.split(","))} 列): {head[:60]}', RC_STRUCT))
+
+    # R7 体量
+    if stat['bytes'] > BIG_DIR:
+        notes.append(('?', P.rel(st), f'源件合计 {stat["bytes"]/2**30:.1f} GB —— 摄入与重算会很慢, 建议分批', RC_NOTE))
+    return res, notes, stat
+
+
+def increment_findings(plan: dict, src_name: str = '') -> list:
+    """把增量三清单里的**冲突**变成一条可拦人的问题 (供本模块与 place_raw_data 共用)。"""
+    out = []
+    if plan.get('clash'):
+        c = plan['clash'][0]
+        out.append(('!', f'--src {src_name}'.strip(), f'{len(plan["clash"])} 件**同名不同大小**: 直接放会覆盖已摄入的源件 '
+                                                        f'(示例 {c["path"]}: 现有 {c.get("now", 0):,} B → 源包 '
+                                                        f'{c["size"]:,} B) —— 先确认哪一份是对的, 再决定 --force 覆盖还是改名并存',
+                    RC_CLASH))
+    return out
+
+
+def increment_src(src: pathlib.Path, st: pathlib.Path) -> tuple[list, dict]:
+    """R5 增量三清单: 遍历源包 zip 成员/散件, 算出它们会落到哪, 与现有树对比。"""
+    res, plan = [], dict(new=[], same=[], clash=[], unmapped=[])
+    import importlib.util
+    spec = importlib.util.spec_from_file_location('_place', ROOT / 'scripts' / 'place_raw_data.py')
+    place = importlib.util.module_from_spec(spec)
+    try:
+        spec.loader.exec_module(place)
+    except SystemExit:
+        pass
+    rules = list(getattr(place, 'RULES', [])) + list(getattr(place, 'RULES_MECH', [])) + list(getattr(place, 'RULES_VIB', []))
+    loose = list(getattr(place, 'RULES_VIB_LOOSE', []))
+    have = sizes_map(st)
+    have_root = sizes_map(P.RAW_ROOT)
+
+    def classify(rel_path: str, size: int, base: pathlib.Path):
+        target = base / rel_path
+        cur = have.get(rel_path) if base == st else have_root.get(rel_path)
+        item = dict(path=rel_path, size=size)
+        if cur is None:
+            plan['new'].append(item)
+        elif cur == size:
+            plan['same'].append(item)
+        else:
+            item['now'] = cur
+            plan['clash'].append(item)
+
+    for zname, prefix, target, strip, root_kind in rules:
+        zp = src / zname if src.is_dir() else None
+        if not zp or not zp.is_file():
+            continue
+        base = st if root_kind == 'station' else P.RAW_ROOT
+        try:
+            with zipfile.ZipFile(zp) as z:
+                for info in z.infolist():
+                    if info.is_dir():
+                        continue
+                    nm = place.gbk_name(info)
+                    if prefix and not nm.startswith(prefix):
+                        continue
+                    tail = nm[len(prefix):] if prefix else nm
+                    if strip:
+                        tail = tail.split('/', 1)[1] if '/' in tail else tail
+                    rel = f'{target}/{tail}'
+                    if prefix and nm == prefix.rstrip('/'):
+                        continue
+                    classify(rel, info.file_size, base)
+        except zipfile.BadZipFile:
+            res.append(('!', zp.name, '不是合法 zip (现场包常被截断; 重新拷一份)', RC_STRUCT))
+    for fn, target in loose:
+        fp = src / fn
+        if fp.is_file():
+            classify(f'{target}/{fn}', fp.stat().st_size, st)
+    return res, plan
+
+
+def guidance(plan: dict, stat: dict) -> list[str]:
+    """R8 放置指导: 按新增源类给出该跑的重算命令与预期变化。"""
+    out = []
+    new = [x['path'] for x in plan.get('new', [])]
+    kinds = {p.split('/', 1)[0] for p in new}
+    if not new:
+        out.append('没有新增源件: 不必重算 (若刚换过数据, 用 scripts/rebuild_all.py --dry-run 看计划)')
+        return out
+    out.append(f'新增 {len(new)} 件, 涉及源类: {sorted(kinds)}')
+    if 'scada_10min' in kinds or 'scada_1min' in kinds:
+        out.append('· SCADA 侧变了 → `python scripts/rebuild_all.py` (含 10 个构建器, 约 15 分钟); '
+                   '台账/报警/油样一起换时同一条命令即可')
+    if kinds & {'故障报警', '风机故障记录', '油样报告'}:
+        out.append('· 台账类变了 → `python scripts/rebuild_all.py --skip-scada` (约 2 分钟)')
+    if kinds & {'windcms', 'm5_cms_tcm'}:
+        out.append('· 振动侧变了 → 重算链第 ④b 步会自动摄入 (`scripts/vib_raw_build.py`); '
+                   '注意它只做索引/谱, 不会重生成报告 (本包缺六层链, 见 docs §7)')
+    out.append('· 放完先自检: `python scripts/raw_data_check.py --deep` → 再看 `/detail/` 左栏「系统维护」两屏')
+    out.append('· 重算后核对锚点: 报警 39211 · 工单 5876 · 油样 404 · temp_monthly 19494 · 本体 9702 对象')
+    return out
+
+
+def main() -> int:
+    ap = argparse.ArgumentParser()
+    ap.add_argument('--farm', default=None)
+    ap.add_argument('--src', default=None, help='现场数据包目录 (给了就做增量体检: 新增/相同/冲突三清单)')
+    ap.add_argument('--deep', action='store_true', help='加可读性抽样 (CSV 编码/表头, TCM json 字段)')
+    ap.add_argument('--json', action='store_true')
+    ap.add_argument('--write-doc', action='store_true', help='把结论写进 docs/输入数据放置指导_v0.1.md')
+    a = ap.parse_args()
+
+    st = station_dir(a.farm)
+    res, notes, stat = check_tree(st, deep=a.deep)
+    plan = {}
+    if a.src:
+        src = pathlib.Path(a.src)
+        if not src.is_dir():
+            print(f'[X] --src 不是目录: {src}')
+            return RC_STRUCT
+        r2, plan = increment_src(src, st)
+        res += r2 + increment_findings(plan, src.name)
+    guide = guidance(plan, stat)
+
+    if a.json:
+        print(json.dumps(dict(station=P.rel(st), results=[dict(level=x[0], item=x[1], note=x[2], rc=x[3]) for x in res],
+                              notes=[dict(level=x[0], item=x[1], note=x[2]) for x in notes], stat=stat, plan=plan,
+                              guidance=guide), ensure_ascii=False, indent=1, default=str))
+    else:
+        print(f'== 输入数据放置体检 · {P.rel(st)} ==')
+        if stat.get('exists'):
+            print(f'   {stat["files"]:,} 件 / {stat["bytes"]/2**30:.2f} GB')
+            for k, v in sorted(stat['subdirs'].items()):
+                print(f'     {k:14s} {v["files"]:7,d} 件  {v["bytes"]/2**20:9.1f} MB')
+        for level, item, note, rc in res:
+            print(f'   [{level}] {item}: {note}')
+        seen = set()
+        for level, item, note, rc in notes:
+            if note in seen:
+                continue
+            seen.add(note)
+            print(f'   [{level}] {item}: {note}')
+        if plan:
+            print('\n== 增量放置三清单 (源包 → data/raw) ==')
+            print(f'   新增 {len(plan["new"]):,} 件 (将拷入) · 相同 {len(plan["same"]):,} 件 (同尺寸, 跳过) · '
+                  f'冲突 {len(plan["clash"]):,} 件 (同名不同大小, **需人确认**)')
+            for k, label in (('new', '新增'), ('clash', '冲突')):
+                for x in plan[k][:5]:
+                    extra = f' (现有 {x.get("now"):,} B)' if k == 'clash' else ''
+                    print(f'     [{label}] {x["path"]}  {x["size"]:,} B{extra}')
+                if len(plan[k]) > 5:
+                    print(f'     … 另有 {len(plan[k]) - 5} 件{label}')
+        print('\n== 放置指导 ==')
+        for g in guide:
+            print(f'   {g}')
+    rc_map = {x[3] for x in res if x[0] in ('X', '!') and x[3]}
+    rc = max(rc_map) if rc_map else (RC_NOTE if notes else RC_OK)
+    if not a.json:
+        print(f'\n结论: {"合规" if rc == 0 else "见上"}; 退出码 {rc}')
+    if a.write_doc:
+        write_doc(res, notes, stat, plan, guide)
+    return rc
+
+
+DOC = P.ROOT / 'docs' / '输入数据放置指导_v0.1.md'
+BEGIN, END = '<!-- RAWCHECK:BEGIN -->', '<!-- RAWCHECK:END -->'
+
+
+def write_doc(res, notes, stat, plan, guide) -> None:
+    body = [BEGIN, '### 体检结论(自动生成,勿手改)', '',
+            f'· 源件目录 `{P.rel(station_dir())}`:{stat.get("files", 0):,} 件 / {stat.get("bytes", 0)/2**30:.2f} GB']
+    for k, v in sorted((stat.get('subdirs') or {}).items()):
+        body.append(f'  - `{k}/`:{v["files"]:,} 件 / {v["bytes"]/2**20:.1f} MB')
+    if res:
+        body += ['', '| 级别 | 项 | 说明 |', '|---|---|---|']
+        body += [f'| `{x[0]}` | `{x[1]}` | {x[2]} |' for x in res[:40]]
+    else:
+        body += ['', '· 结构与命名检查:**无不一致**(提示见下)']
+    for x in notes[:12]:
+        body.append(f'· `{x[0]}` {x[1]}:{x[2]}')
+    if plan:
+        body += ['', f'· 增量(源包 → data/raw):新增 {len(plan["new"]):,} · 相同 {len(plan["same"]):,} · '
+                     f'冲突 {len(plan["clash"]):,}']
+    body += ['', '**放置后该做什么**:'] + [f'· {g}' for g in guide] + ['', END]
+    block = '\n'.join(body)
+    if DOC.is_file():
+        t = DOC.read_text(encoding='utf-8')
+        if BEGIN in t and END in t:
+            t = re.sub(re.escape(BEGIN) + r'.*?' + re.escape(END), lambda m: block, t, flags=re.S)
+        else:
+            t = t.rstrip() + '\n\n' + block + '\n'
+    else:
+        t = '# 输入数据放置指导\n\n' + block + '\n'
+    DOC.write_text(t, encoding='utf-8')
+    print(f'   已写入 {P.rel(DOC)}')
+
+
+if __name__ == '__main__':
+    from src import console
+    console.soft()
+    sys.exit(main())

+ 13 - 1
scripts/windscada_serve.py

@@ -585,7 +585,19 @@ def vibcms_results():
         hset = {r.turbine for _, r in ft.iterrows() if _f.verdict_class(str(r.振动结论))[1] in ('bad', 'warn')}
         cset = {r[0] for r in action}
         diff = dict(only_cms=sorted(cset - hset), only_handoff=sorted(hset - cset))
-        return dict(date=date, overall=overall, action=action, grades=grades, diff=diff)
+        # ★2026-09-19: 把**这个窗真实覆盖的日期区间**一并回给页面。原来只回 date(出件日) 与表格,
+        #   页面上看不到"数据从哪天到哪天" ⇒ 用户拿 2026-03~04 的导出却只看到 4 月的谱时,
+        #   无法判断是"数据没呈现"还是"源件本身就那么点"(实测源件最早只到 2026-03-16, 逐台更晚)。
+        _span, _win = '', ''
+        try:
+            from src.windcms import config as _wcfg, data as _wdata
+            _w2 = _wcfg.farm(str(CFG.get('key') or 'rudong'))
+            _span = _wdata.span_text(_w2) or ''
+            _win = ','.join(_wdata.windows(_w2))
+        except Exception:
+            _span = _win = ''
+        return dict(date=date, overall=overall, action=action, grades=grades, diff=diff,
+                    时间范围=_span, 窗=_win)
     except Exception as e:
         return dict(error=f'windcms 报告解析失败: {e}')  # 响亮, 不静默空表
 

+ 49 - 0
src/windcms/data.py

@@ -28,6 +28,55 @@ def windows(cfg):
     return dict(sorted(out.items()))
 
 
+_SPAN_CACHE: dict = {}
+
+
+def window_spans(cfg):
+    """{窗: dict(rows, t_min, t_max, turbines, start_late)} —— 每个窗**真实覆盖的日期区间**。
+
+    为什么要有它(用户令 2026-09-19 实逮): 窗名 `w0316` 只是"窗起始日 03-16"的代号, 页面上到处只有
+    这个名字与"证据窗末 2026-04-21", **看不出这个窗从哪天到哪天** —— 用户拿到的振动导出是
+    `CMS_RuDong_CGN_202603-04`(2026-03 至 2026-04), 于是他看到逐台页只有 4 月的谱, 就以为
+    "3 月的数据没被呈现"。实测源件本身最早只到 **2026-03-16 17:03**(逐台起始不等: 最晚 03-30),
+    也就是说数据没错、是**呈现上没把覆盖区间说出来**。这里把区间算出来, 供报告/页面如实写明。
+
+    只读每个窗索引的 `trigger_time` + `turbine` 两列(百万行级也只几秒), 进程内缓存。
+    """
+    key = tuple(sorted(windows(cfg)))
+    if _SPAN_CACHE.get('key') == key:
+        return _SPAN_CACHE['val']
+    out = {}
+    for w, info in windows(cfg).items():
+        try:
+            d = pd.read_parquet(info['index'], columns=['trigger_time', 'turbine'])
+            t = pd.to_datetime(d['trigger_time'], errors='coerce')
+            per_first = t.groupby(d['turbine']).min()
+            out[w] = dict(rows=int(len(d)), t_min=str(t.min()), t_max=str(t.max()),
+                          turbines=int(d['turbine'].nunique()),
+                          first_min=str(per_first.min()), first_max=str(per_first.max()),
+                          start_late=bool(str(per_first.min())[:10] != str(per_first.max())[:10]))
+        except Exception as e:
+            out[w] = dict(rows=0, t_min='', t_max='', turbines=0, err=f'{type(e).__name__}: {e}')
+    _SPAN_CACHE.update(key=key, val=out)
+    return out
+
+
+def span_text(cfg, win: str | None = None) -> str:
+    """给人看的一句话: `w0316 2026-03-16 → 2026-04-21`(多窗时逗号分隔; 逐台起始不等时如实写出来)。"""
+    sp = window_spans(cfg)
+    ws = [win] if win else list(sp)
+    parts = []
+    for w in ws:
+        s = sp.get(w) or {}
+        if not s.get('rows'):
+            continue
+        rng = f"{str(s['t_min'])[:10]} → {str(s['t_max'])[:10]}"
+        if s.get('start_late'):
+            rng += f"(逐台起始不等: 最早 {str(s['first_min'])[:10]} · 最晚 {str(s['first_max'])[:10]})"
+        parts.append(f'{w} {rng}')
+    return ' · '.join(parts)
+
+
 def _read_index(p):
     cols = ['turbine', 'sensor_name', 'meas_name', 'trigger_time', 'rpm', 'condition_key', 'alarm_type', 'ds_size', 'scalar_value', 'overload']
     d = pd.read_parquet(p)

+ 19 - 3
src/windcms/report.py

@@ -254,7 +254,19 @@ def turbine_page(cfg, tid, ctx):
             f'<tr><td>{SENSOR_CN.get(r.sensor_name, r.sensor_name)}</td><td>{r.meas_name}</td><td class="num">{r.median:.4g}</td><td class="num">{r.x_fleet:.2f}</td><td>{r.mask_status}</td></tr>' for r in top.itertuples()) + '</tbody></table>')
     b.append('</div>')
     b.append(baseline_card_html(cfg, tid))   # 5 三层基线对照 (2026-08-26)
-    return ui.shell(f'{tid} · {cfg["name"]}', '\n'.join(b), farm_name=cfg['name'], machine=cfg['machine'], active='', base='..', subtitle=f'{tid} 逐台页 · {cfg["machine"]}')
+    # ★2026-09-19: 逐台页原来只写 'WTGxx 逐台页 · 机型', 用户看不出**这一台**的数据从哪天到哪天 ⇒
+    #   "3 月的振动数据没呈现"这类疑问只能靠猜。这里把"本台实测区间"与"窗区间"都写到副标题上
+    #   (本台区间从标量表实算, 不写死; 窗区间见 data.span_text)。
+    _own = ''
+    try:
+        _ts = pd.to_datetime(s[s.turbine == tid]['trigger_time'], errors='coerce')
+        if len(_ts.dropna()):
+            _own = f'本台数据 {_ts.min():%Y-%m-%d} → {_ts.max():%Y-%m-%d} · '
+    except Exception:
+        _own = ''
+    _span = data.span_text(cfg)
+    return ui.shell(f'{tid} · {cfg["name"]}', '\n'.join(b), farm_name=cfg['name'], machine=cfg['machine'],
+                    active='', base='..', subtitle=f'{tid} 逐台页 · {_own}{_span or cfg["machine"]}')
 
 
 def build(cfg, turbines=None, verbose=True):
@@ -490,12 +502,16 @@ def build(cfg, turbines=None, verbose=True):
         wb = workbench.workbench_body(cfg, reg, reg_detail, ctx, cell, engineer=eng)
         (out / fname).write_text(ui.shell(f'{cfg["name"]} 振动诊断 · {vn}', wb, farm_name=cfg['name'],
                                           machine=cfg['machine'], active='/',
-                                          subtitle=f'{vn} · 证据窗末 {win_end}'), encoding='utf-8')
+                                          # ★2026-09-19: 首页原来只写"客户版 · 证据窗末 YYYY-MM-DD" ——
+                                          #   没有**起始日**, 用户看不出这个窗覆盖 3 月还是只有 4 月。
+                                          subtitle=f'{vn} · {win_span} · 证据窗末 {win_end}'), encoding='utf-8')
     for t in turbines:
         (out / 'turbines' / f'{t}.html').write_text(turbine_page(cfg, t, ctx), encoding='utf-8')
         if verbose:
             print('.', end='', flush=True)
-    md = [f'# {cfg["name"]} CMS 振动诊断 · 自动报告(结构化层)\n', f'机型:{cfg["machine"]};数据窗:{", ".join(cfg["windows"])}\n', '## 逐台(融合级 ≠ 正常)\n', '| 台 | 融合 | CMS | 模型 | 机制 | 模型依据 |', '|---|---|---|---|---|---|']
+    md = [f'# {cfg["name"]} CMS 振动诊断 · 自动报告(结构化层)\n',
+          f'机型:{cfg["machine"]};数据窗:{data.span_text(cfg) or ", ".join(cfg["windows"])}\n',
+          '## 逐台(融合级 ≠ 正常)\n', '| 台 | 融合 | CMS | 模型 | 机制 | 模型依据 |', '|---|---|---|---|---|---|']
     for _, r in fusion.iterrows():
         if r['融合'] != '正常':
             md.append(f'| {r["台"]} | {r["融合"]} | {r["CMS"]} | {r["模型"]} | {r["机制"]} | {r["模型依据"]} |')

+ 12 - 3
src/windcms/report_std.py

@@ -245,11 +245,15 @@ def build(cfg, llm_polish=False):
     reg, detail, c = registry(cfg)
     date = datetime.date.today().strftime('%Y年%m月%d日')
     ws = list(data.windows(cfg))
+    # ★2026-09-19: 窗名只是"起始日"的代号(w0316), 页面上只有它 + "证据窗末" ⇒ 看不出这个窗**从哪天到哪天**。
+    #   用户拿到的振动导出是 2026-03 至 2026-04, 看到逐台页只有 4 月的谱就以为"3 月数据没呈现";
+    #   实测源件本身最早只到 2026-03-16(逐台起始不等) ⇒ 是**呈现上没说覆盖区间**。这里把实算区间写进报告正文。
+    span = data.span_text(cfg)
     n_ok = int((reg['状态等级'] != '不可判').sum())
     md = []
     md.append(f'# {cfg["name"]}_CMS 振动状态评估报告(自动生成版 {datetime.date.today().isoformat()})\n')
-    md.append(f'> 三轴口径:状态等级(优秀/良好/报警/危险/不可判)× 证据状态(确诊/准定论·预警/候选/INSUFFICIENT)× 行动等级建议(P0/P1/P2/P3,供业主行动排程参考,最终行动等级由综合报告线裁定)。数据窗 {", ".join(ws)};机型 {cfg["machine"]}。本报告由 windcms 结构化层生成,判级全部来自确定性判据;量化门槛{"不公开(内部版可见)" if PUBLIC else "见 3.1"}。\n')
-    md.append(f'## 1 项目概况\n\n| 项目 | 内容 |\n|---|---|\n| 风电场 | {cfg["name"]} |\n| 机组数量 | {cfg["n_turbines"]} 台 |\n| 机型 | {cfg["machine"]} |\n| 监测系统 | Gram & Juhl TCM,8 测点/台({"、".join(SENSOR_CN[s] for s in SENSORS)}) |\n| 数据窗 | {", ".join(ws)}(离线导出包) |\n| 监测数据完整性 | 可评定 {n_ok}/{cfg["n_turbines"]} 台;不可判台见附录 |\n')
+    md.append(f'> 三轴口径:状态等级(优秀/良好/报警/危险/不可判)× 证据状态(确诊/准定论·预警/候选/INSUFFICIENT)× 行动等级建议(P0/P1/P2/P3,供业主行动排程参考,最终行动等级由综合报告线裁定)。数据窗 {span};机型 {cfg["machine"]}。本报告由 windcms 结构化层生成,判级全部来自确定性判据;量化门槛{"不公开(内部版可见)" if PUBLIC else "见 3.1"}。\n')
+    md.append(f'## 1 项目概况\n\n| 项目 | 内容 |\n|---|---|\n| 风电场 | {cfg["name"]} |\n| 机组数量 | {cfg["n_turbines"]} 台 |\n| 机型 | {cfg["machine"]} |\n| 监测系统 | Gram & Juhl TCM,8 测点/台({"、".join(SENSOR_CN[s] for s in SENSORS)}) |\n| 数据窗 | {span}(离线导出包) |\n| 监测数据完整性 | 可评定 {n_ok}/{cfg["n_turbines"]} 台;不可判台见附录 |\n')
     md.append('## 2 评价标准\n\n参照 VDI 3834 与 NB/T 31129—2018《风力发电机组振动状态评价导则》的分级思路,结合本场机群横向比较与各机组自身历史轨迹,将振动状态分为四级,另设“不可判”。\n\n| 状态等级 | 判定与处置原则 |\n|---|---|\n' + '\n'.join(f'| {a} | {b} |' for a, b in STATE_DESC) + '\n')
     thr = [('冲击轴:Peak 逐日中位/基线(末窗 vs 末窗前)', '持续恶化 → 报警', '稳定高位(相对机群) → 报警', '反复发作(高值日≥▪)→ 报警;孤立尖峰 → 良好'),
            ('特征频率族:六道前提闸 + 同域绝对锚 + 证据族', '准定论·预警 → 报警', '候选·新发 → 报警;候选 → 良好', '参考/INSUFFICIENT → 优秀(记基线)'),
@@ -272,7 +276,12 @@ def build(cfg, llm_polish=False):
         md.append(f'### {t}\n\n图 {t} 机组诊断分析(见逐台页 turbines/{t}.html:标量趋势、最新谱与 OEM 特征线)\n\n分析:{analysis}\n\n{concl}{(" " + act) if act else ""}\n\n{accept}\n')
     if llm_polish:
         md.append(f'> 本章分析段由离线模型按书面工程汉语润色 {polished} 段(数字与台号经接地闸核对, 未过闸者保留模版原文)。\n')
-    md.append('## 5 说明事项\n\n1. 评估不以单一峰值确诊部件,不命名滚道/圈侧;状态等级、证据状态与行动等级建议三者不得相互替代;确诊只认实物或换件闭环。\n2. 数据止于数据窗末日;盲区台按 A1 三源合判列出。主轴承前后通道 8–13 kHz 存在同路电气调幅成分,该带不用于机械源定位;rms_200 未越限不能作正常依据。\n3. 本报告为振动单轴;润滑/温度/实物由其他分册承接,P0/P1 由综合报告线裁定。相对判据均配同域绝对量,“独有/特异”均经全场基率核对。\n')
+    md.append(f'## 5 说明事项\n\n1. 评估不以单一峰值确诊部件,不命名滚道/圈侧;状态等级、证据状态与行动等级建议三者不得相互替代;确诊只认实物或换件闭环。\n'
+              f'2. **本报告覆盖的数据区间是 {span}**(窗名 `{", ".join(ws)}` 只是"窗起始日"的代号,不是区间本身)。'
+              f'窗内不含更早的数据 —— 若现场导出的原始件本身从更早开始,需重新导出并重算;'
+              f'逐台起始时刻不等时以上区间已写明最早/最晚,读"逐台"结论前先看该台自己的覆盖。\n'
+              f'3. 数据止于数据窗末日;盲区台按 A1 三源合判列出。主轴承前后通道 8–13 kHz 存在同路电气调幅成分,该带不用于机械源定位;rms_200 未越限不能作正常依据。\n'
+              f'4. 本报告为振动单轴;润滑/温度/实物由其他分册承接,P0/P1 由综合报告线裁定。相对判据均配同域绝对量,"独有/特异"均经全场基率核对。\n')
     md.append('## 附录 A 38 台状态等级表(状态管理,不代表确诊数量)\n\n| 机组号 | 主轴承前 | 主轴承后 | 齿轮箱 | 发电机 | 综合 | 融合级(模型) | CMS 红/黄(末窗) | 行动等级建议 |\n|---|---|---|---|---|---|---|---|---|')
     for _, r in reg.iterrows():
         md.append(f'| {r["机组"]} | {r["主轴承前"]} | {r["主轴承后"]} | {r["齿轮箱"]} | {r["发电机"]} | {r["状态等级"]} | {r["融合级"]} | {r["CMS红"]}/{r["CMS黄"]} | {r["行动等级建议"]} |')

+ 1 - 1
src/windcms/serve.py

@@ -315,7 +315,7 @@ def run(cfg, port=8030, model=None):
                         print(f'[warn] 内嵌图失败 {_tid}: {_e}', flush=True); panels = ''
                     return panels + f'<p class="small"><a href="/turbines/{_tid}.html">→ {_tid} 逐台页 (全部测点趋势/谱/证据)</a></p>'
                 body = dl + '<div class="card report">' + _re.sub(r'<p>图[\s\u3000]+(WTG\d+)[^<]*见逐台页[^<]*</p>', _inline, html_md) + '</div>'
-                return self._send(ui.shell('评估报告', body, farm_name=cfg['name'], machine=cfg['machine'], active='/report.html', subtitle='标准模版报告 · 结构化层生成'))
+                return self._send(ui.shell('评估报告', body, farm_name=cfg['name'], machine=cfg['machine'], active='/report.html', subtitle='标准模版报告 · 结构化层生成 · ' + (__import__('src.windcms.data', fromlist=['x']).span_text(cfg) or '')))
             if self.path.startswith('/files/'):
                 from urllib.parse import unquote
                 p = out / unquote(self.path[7:])

+ 14 - 2
src/windcms/workbench.py

@@ -122,6 +122,12 @@ def _trend_block(c, t, sen, meas, engineer, title=''):
     ★阈值线纪律: masks 里确有 yellow/red, 但 mask_time 多为 **2017-12-28** — 8 年前设定,
     与 2024~2026 的数据不同期。画可以, **必须把设定日期标出来**, 否则读者会当成现行阈值。
     Peak 等测量厂商根本没设阈值 (th 为空) ⇒ **不画假线**, 明写"厂商未设阈值"。
+
+    ★2026-09-19 修 (用户报"重算后振动产物没呈现 2026-03 数据"): 原来按**窗**聚合
+    (`groupby('window')` → 中位), 而本机只有一个窗(w0316 = 2026-03-16 → 2026-04-21) ⇒ 整条趋势
+    只画**一个点**, 时间戳还是窗内中位时刻(≈4 月上旬), 于是"3 月"在图上完全不见。实测点密度:
+    WTG29/主轴承前/Peak 窗内 9 条记录、跨 2026-03-18 → 2026-04-16 共 9 天 —— 数据在, 是**聚合粒度**
+    把 3 月压没了。现在按 (窗, 日) 聚合: 单窗也画逐日序列, 3 月起止如实出现在轴上。
     """
     try:
         q, th = tcm.trend(c['scalars'], c['masks'], 'CGN Rudong', t, sen, meas)
@@ -131,10 +137,16 @@ def _trend_block(c, t, sen, meas, engineer, title=''):
         return '<p class="muted small">该测点在当前数据窗内无趋势点。</p>', {}
     q = q.copy()
     q['ts'] = pd.to_datetime(q['trigger_time'], errors='coerce')
-    per = q.dropna(subset=['ts']).groupby('window').agg(ts=('ts', 'median'), v=('scalar_value', 'median')).sort_values('ts')
+    # 逐 (窗, 日) 中位 —— 不用 groupby('window'): 单窗时那只会得到一个点, 把整段区间压成一个时刻
+    # (2026-09-19 实逮: 3 月的数据在图上"消失"就是这个原因, 不是数据没进来)
+    q = q.dropna(subset=['ts'])
+    q['day'] = q['ts'].dt.floor('D')
+    per = (q.groupby(['window', 'day'], as_index=False)
+             .agg(ts=('ts', 'median'), v=('scalar_value', 'median'))
+             .sort_values('ts'))
     if not len(per):
         return '<p class="muted small">趋势点时间戳缺失。</p>', {}
-    series = [{'name': '本机逐窗中位', 'x': list(per['ts']), 'y': [float(x) for x in per['v']]}]
+    series = [{'name': '本机逐日中位', 'x': list(per['ts']), 'y': [float(x) for x in per['v']]}]
     hl = []
     try:
         ref = tcm.fleet_ref(c['scalars'], sen, meas)

+ 5 - 0
src/windscada/lang.py

@@ -855,6 +855,11 @@ _add({
     'fus.cms_head':    ('振动状态评估 (windcms 报告 {d} 结果转录)', 'Vibration condition assessment ({d} reviewed report)'),
     'fus.cms_note':    ('状态管理判级, 不代表确诊数量; 判据/谱线/计算过程不在本页 — 深潜走底部独立系统。',
                         'Condition-management ratings, not a count of confirmed faults; criteria, detailed signal markers and calculations are not on this page — use the standalone system linked below.'),
+    # ★2026-09-19: 页面上原来只有"出件日 + 证据窗末", 没有**起始日** ⇒ 用户拿 2026-03~04 的导出
+    #   只看到 4 月的谱时会问"3 月的数据呢"。窗名(w0316)只是起始日代号, 这里把实算区间摆出来并说明
+    #   为什么页面上的谱多显示末期日期 (末窗状态=取窗内最新记录, 是口径不是缺数据)。
+    'fus.cms_window':  ('数据窗 {r} —— 本窗**实算**覆盖区间 (窗名只是起始日代号); 页面上的谱与"末窗"状态取自窗内最新记录, 故多显示末期日期。',
+                        'Data window {r} — the measured coverage of this window (the window name is only its start-date code). Spectra and "latest-window" states come from the newest record inside the window, so later dates dominate the display.'),
     'fus.cms_action':  ('需行动机组 (报警 / 不可判)', 'Units needing action (alarm / undeterminable)'),
     'fus.cms_focus': ('状态焦点', 'Focus'), 'fus.cms_focus_note': ('(传动链: 主承前 · 主承后 · 齿轮箱 · 发电机两端)', '(drivetrain: main bearing fore / aft · gearbox · generator both ends)'),
     'fus.cms_evidence': ('证据强度', 'Evidence strength'), 'fus.cms_act': ('行动', 'Action'),

+ 3 - 1
src/windscada/ui/app.js

@@ -916,7 +916,9 @@
     const V = FUS.cms; if (V.error) return `<div class="panel mb"><p class="small" style="color:var(--critical)">${dv(V.error)}</p></div>`;
     const cp = st => stateChip(st);
     const CK = { '主轴承前': 'mbf', '主轴承后': 'mba', '齿轮箱': 'gb', '发电机': 'gen', '发电机DE': 'gde', '发电机NDE': 'gnde', 'Main bearing, front': 'mbf', 'Main bearing, rear': 'mba', 'Gearbox': 'gb', 'Generator': 'gen', 'Generator DE': 'gde', 'Generator NDE': 'gnde' };
-    return `<div class="panel mb"><h3>${t('fus.cms_head', { d: esc(V.date || '') })}</h3><p class="small">${dv(V.overall || '')} ${t('fus.cms_note')}</p>
+    return `<div class="panel mb"><h3>${t('fus.cms_head', { d: esc(V.date || '') })}</h3>
+      ${V['时间范围'] ? `<p class="small" style="margin:2px 0 6px">${t('fus.cms_window', { r: esc(V['时间范围']) })}</p>` : ''}
+      <p class="small">${dv(V.overall || '')} ${t('fus.cms_note')}</p>
       <h3 style="margin-top:10px">${t('fus.cms_action')}</h3><div class="tblwrap"><table><tr><th>${t('ui.unit')}</th><th>${t('fus.loop_state')}</th><th>${t('fus.cms_focus')} <span class="small muted" style="font-weight:400">${t('fus.cms_focus_note')}</span></th><th>${t('fus.cms_evidence')}</th><th>${t('fus.cms_act')}</th></tr>
       ${(V.action || []).map(r => { const cm = {}; String(r[2]).split(/[;;]/).forEach(seg => { const m = seg.split(/[::]/); if (m.length > 1) { const k = CK[m[0].trim()] || null; if (k) cm[k] = gc(m[1].trim()); } });
         return `<tr><td><b>${unitNo(r[0])}</b></td><td>${cp(r[1])}</td><td title="${esc(r[2])}">${chainStates(cm)} <span class="small muted">${dv(r[2])}</span></td><td style="min-width:150px">${evLadder(r[3])}</td><td>${/P0/.test(r[4]) ? `<b style="color:var(--critical)">${esc(r[4])}</b>` : esc(r[4])}</td></tr>`; }).join('')}</table></div>