Install
openclaw skills install @tianzhiceng297-boop/eia-process-intelInvestigation workflow that turns Chinese EIA filings (环评报告/环评公示) into investment due-diligence intelligence: structured extraction (product scheme, per-step process flow, equipment list, material balance), adversarial field-credibility grading, physics-based cross-validation (low-risk fields as trust anchors to back-calculate capacity/yield/emissions), process-to-equipment mapping inference, and an append-only triple ledger with provenance. Use when: (1) analyzing a company EIA filing or 受理公示, (2) verifying capacity claims / 产能真实性核验, (3) inferring equipment selection from a 设备清单, (4) back-calculating yield from 物料平衡 / material balance, (5) auditing EIA numbers for gaming patterns (批小建大), (6) building a process/equipment fact ledger across deals. Triggers: 环评, EIA report, material balance, equipment list, capacity verification, yield back-calculation, regulatory gaming audit.
openclaw skills install @tianzhiceng297-boop/eia-process-intelTurn Chinese EIA filings (环境影响评价报告 / 报告表 / 受理公示) into investment due-diligence intelligence. The EIA is a rare document the company itself wrote under legal liability and disclosed before construction: equipment lists, material balances and product schemes are mandatory disclosures. This skill reads it adversarially — extract, grade every number for regulatory gaming, cross-validate with low-risk trust anchors, infer process routes and equipment tiers, and book everything into a provenance-tracked triple ledger.
Core move: EIAs almost never disclose per-step yields. Back-calculate end-to-end yield from the material balance (key inputs such as substrates vs chip outputs), then localize the bottleneck step. Worked demo: examples/format-demo-ir-detector.md (fictional case, all numbers marked illustrative).
| Situation | Action |
|---|---|
| Analyzing a company EIA filing (环评报告/报告表/公示) | Run the full Phase 0–6 workflow below |
| Verify capacity claims (产能真实性核验) | Capacity back-calc recipes + external permit records |
| Infer process route or equipment tier | Three-face route evidence + references/mappings/<track>.md |
| Back-calculate yield from material balance (物料平衡) | Yield recipe + bottleneck localization |
| Audit EIA numbers for distortion | Field grading table + threshold-hugging signatures |
| Reconcile EIA caliber vs BP / interviews / prospectus | Caliber conflict table (report section 2) |
| Build a reusable process/equipment fact base | Ledger schema, append-only JSONL |
references/ledger-schema.md).| # | Question | Primary evidence |
|---|---|---|
| 1 | Actual capacity vs financing-story capacity, and build-out pace | Product scheme, construction phases, approval info |
| 2 | Which process route is actually in use | Process sequence + materials + pollutants (three faces) |
| 3 | Equipment selection tier (origin / generation / density) | Equipment list, models, counts vs capacity ratio |
| 4 | Yield and unit consumption; which step is the bottleneck | Material balance |
| 5 | Do pollution fingerprints reveal undisclosed processes | Hazardous chemicals / hazardous waste / special gases |
| 6 | Where EIA caliber contradicts company claims | All fields vs BP / interviews / announcements |
| 7 | Credibility grading and back-calculation of each number | Field grading table below |
Principle: the more a number affects approval outcome, environmental spend, or regulatory burden, the more it is shaped by gaming between the company and the regulator. Update grades as cross-validation evidence accumulates; a falsified grade changes the table, with the case noted in references/field-credibility.md.
| Field | game_risk | Typical distortion direction | Cross-validation path |
|---|---|---|---|
| Design capacity 设计产能 | High | Both ways: inflate (grab emission quotas, reserve expansion headroom) or deflate (stay under high-level approval thresholds) | Equipment counts × industry cycle time; material inputs; annual hours; permit execution reports |
| Product mix 产品结构 | High | Split SKUs, disguise mass production as pilot/R&D, hide restricted products | Hazardous-waste types, waste-liquid composition, byproducts reveal real products |
| Pollutant generation/emission 污染物 | High | Understate → lower monitoring frequency and treatment specs | Treatment facility design capacity back-calc; industry per-unit factors; execution-report actuals |
| Raw material usage 原辅材料 | Med-high | Understate toxic/hazardous chemicals → smaller buffer zone, lower regulatory class | Material-balance self-consistency; hazwaste ratio; warehouse size |
| Operating hours 设备利用率 | Med-high | Understate annual run hours → lower emission estimates | Implied cycle time from equipment density; power/water consumption |
| Process route 工艺路线 | Medium | Vague or blended wording to dodge specialized review | Pollutant fingerprints; equipment models |
| Equipment count 设备数量 | Med-low | Understate advanced units, omit cross-period equipment | Ratio vs capacity; permit application cross-check |
| Physical info 厂房/排放口/环保设施 | Low | Essentially trustworthy (site-verifiable, later documents cross-check) | Use as trust anchors to back-calculate everything above |
Discovery output: gaming traces (批小建大、化整为零、hidden process steps、pilot-name mass production) go into a dedicated list AND are booked under the ledger relation 博弈痕迹 — they are also governance-DD evidence.
The grading table is a prior, not a verdict. Every high-risk field evolves within a case:
status tag as cross-validation proceeds (申报值 / 已交叉验证 / 与外部冲突).Worked instance: power consumption, prior 中高 → mean-load check found it inconsistent with the line's heating/RF/UPW configuration → posterior downgraded to "conflicting evidence", two competing explanations recorded (caliber distortion vs under-declared utilization), pending utility bills.
Prior-table feedback rule: a single-case posterior never edits the grading table. The prior moves only when the same distortion direction is confirmed across ≥2 independent cases, or a hard counter-example appears — with the case list cited in references/field-credibility.md.
The phases above are the discipline skeleton, not the route — real investigations loop. Define an anomaly as anything that contradicts the current working hypothesis:
Protocol on hitting an anomaly:
failed-queries.md. Never resolve a loop in memory alone;Worked instance (nanoimprint acceptance report, 2026-09-10): bottle-scale special-gas usage (anomaly 4) triggered a material-balance re-check → solvent-vs-emission self-consistent, but the same loop surfaced the mean-load power inconsistency (anomaly 3) — one loop, two written artifacts.
Search in order (field-tested; append as you learn):
Query patterns: company full name; 项目名 环评报告表; 项目名 环境影响报告书 受理 + city name. The same project usually has three versions (受理公示 / 拟批准公示 / 审批决定) — the acceptance version is the fullest. 报告表 (short) and 报告书 (full) differ substantially; grab both.
Extract section by section into fact rows, each with source (file + page + verbatim quote) and disclosure date. OCR scanned PDFs first; transcribe tables from table text, never paraphrase.
Tag every numeric field with game_risk (高/中高/中/中低/低, defaults from the grading table) and a distortion direction (虚增/虚减/拆分). Run the three interest-review questions on every high-risk field. Sweep for the threshold-hugging signatures.
Use low-risk trust anchors (equipment counts, floor area, outfall locations, treatment facility sizes — site-verifiable, later documents cross-check) to back-calculate high-risk fields. Then narrow with external sources, in order of availability:
Tag every field status: 申报值 (declared only) / 已交叉验证 (≥1 independent path, recorded) / 与外部冲突 (conflicting evidence exists — never overwrite the number, the conflict itself goes into the report).
references/mappings/<track>.md, walk 工序→设备类别→档次信号→国别/代表厂商. A step missing from the mapping gets tagged 【映射缺失】and enters the backlog — never improvise a mapping on the spot.Everything inferred is tagged 【推断】with the chain written out: which facts → which rule → which conclusion.
Append triples per references/ledger-schema.md. Rules: append-only JSONL, one file per track, stored in the DD project workspace {project dir}/eia-ledger/<track>.jsonl — never in this skill directory. Every record carries source / game_risk / status / disclosure date / booking date. Queries the ontology cannot answer go to failed-queries.md in the same directory; the schema only moves when real failed queries justify it.
{"s":"X-Detek (fictional)","p":"拥有","o":"T2SL detector line project (fictional)","q":{"建设性质":"新建"},"source":{"file":"x-detek-eia-20XX-acceptance.pdf (fictional)","page":12,"quote":"本项目新建…"},"game_risk":"低","status":"申报值","disc_date":"20XX-06","added":"2026-09-09"}
Ledger skeleton: 8 entities (公司/项目/产品/工序/设备/设备类别/物料/污染物) × 8 relations (拥有/包含工序/使用设备/归类/投入/产出/排放/博弈痕迹), each fact tagged game_risk + status. Full definitions, attribute scopes, and the schema-growth mechanism: references/ledger-schema.md.
The ledger becomes an intelligence system only when cases call each other. At the end of every engagement, after Phase 6, run three standard aggregations over the track's ledger file(s) — plain grep/join over JSONL, no new infrastructure:
Two-layer write rule — never fuse them:
{project dir}/eia-ledger/cross-case-results.md (dated; query type; cases involved; raw result). This is data, not knowledge.references/mappings/)One file per track; row shape: EIA signal (process / pollutant / hazchem keywords) → equipment class (vocabulary) → tier signal → country / representative vendors → rationale. Append rows after every engagement; no evidence, no row; uncertain rows marked 【待验证】with the verification path. Renaming an equipment class is a vocabulary change — sync references/ledger-schema.md in the same commit.
| File | Contents |
|---|---|
references/field-credibility.md | Grading table rationale, v1 prior, update log |
references/cross-validation.md | Full recipes: capacity (cycle-time / material / energy), yield, emissions, product-mix fingerprints; status rules |
references/ledger-schema.md | Ontology minimal set (8×8), JSONL format, failed-query growth mechanism |
references/mappings/semiconductor-equipment.md | Seed mapping ~22 rows (semiconductor / IR detector / nanoimprint) |
examples/format-demo-ir-detector.md | Fictional worked example: yield 18–26% back-calculated from material balance |