Install
openclaw skills install @seulkilu/results-claim-hedging-checkerAcademic writing revision advisor for the Results section of psychology and STEMM research papers. Checks claim-hedging alignment, causal language appropriateness, statistical reporting completeness, and interpretation boundaries. Triggers when the user asks to review, revise, or check a Results section draft, or mentions hedging, claim strength, overclaiming, or causal language in academic writing.
openclaw skills install @seulkilu/results-claim-hedging-checkerDiagnose and revise the Results section of psychology / STEMM research papers for five recurring problem types: (1) missing or excessive hedging, (2) claim strength mismatched to statistical evidence, (3) causal language inappropriate for the design, (4) theoretical interpretation leaking into Results, and (5) subjective evaluative language. Produce a sentence-level audit with risk ratings and concrete revision suggestions, grounded in 18 curated examples from 8 classic psychology papers.
For every sentence, check against the five diagnostic dimensions. A sentence may trigger more than one dimension.
| Dimension | What to detect | Reference examples |
|---|---|---|
| D1. Hedging | Missing hedging where indirect evidence is used; or excessive hedging that undermines clear findings | F-01, F-05, F-07, F-09, F-16 |
| D2. Claim Strength vs. Evidence | Overclaims beyond what the statistics support; strong claim words ("strong evidence," "robust finding") used without corresponding evidence; causal leaps from correlational data | F-03, F-10, F-15 |
| D3. Causal Language | Strong causal verbs ("caused," "produced," "leads to," "determines") used when design does not support causal inference, or even when it does but Results convention favors neutral phrasing | F-06, F-11, F-13 |
| D4. Interpretation in Results | Theoretical mechanisms, "because" explanations, hypothesis restatements, or discussion-level conclusions appearing in Results | F-04, F-08, F-12, F-18 |
| D5. Subjective / Evaluative Language | Evaluative adjectives ("instructive," "conspicuous," "remarkable"), subjective hedges ("It is helpful to consider," "Interestingly,"), or absolute expressions ("leave no room for doubt," "prove") | F-02, F-14, F-15, F-17 |
Apply the scoring rubric (below) to assign a 1—5 分 per dimension per sentence(5 分为最高,1 分为最低)。如果某个维度没有问题,该维度应得 5 分。不得在输出中使用"2 分 / Pass"这种含糊表述——直接给出 1—5 的数字分数。如果某句在所有五个维度上均为 5 分,则该句无需出现在报告中(除非用户要求全文审计)。
For each flagged sentence:
Assemble all diagnostic results into the output format specified below. 最终输出必须严格使用以下六个标题,不得使用其他自定义标题(Dimension Score / Key Problems / Evidence from Draft / Example-based Comparison / Revision Suggestions / Priority Level),以便汇总 Skill 统一整合各子 Skill 的诊断结果。不得添加额外的自定义标题或省略任何一个标准标题。
Dimension Score 部分必须:
每个维度采用 1—5 分制,5 分为最高(无问题),1 分为最低(严重问题)。每个维度独立评分后,综合各维度给出一个总体分(1—5 分),并附一句话总评。
| 分值 | 含义 | 行动要求 |
|---|---|---|
| 5 | 无问题:声称强度与证据匹配,hedging 恰当,因果语言与设计一致,无理论解释渗入,无主观评价词 | 无需修改 |
| 4 | 轻微:风格偏好层面的微小问题,不影响科学准确性 | 可选修改;低优先级 |
| 3 | 中度:声称略超出证据支持,或 hedging 缺失/过度 | 建议修改;说明原因 |
| 2 | 较严重:因果语言与设计不匹配,或相关数据写成因果断言,或 Results 中出现理论解释 | 推荐修改;解释误导风险 |
| 1 | 严重:绝对化断言(prove, definitely),因果跳跃严重,使用强动词且无对应证据 | 必须修改;解释误导风险 |
注意:如果某个维度没有问题,该维度应得 5 分,不得使用"Pass""2 分 / Pass"等含糊表述。
D1 — Hedging
D2 — Claim Strength vs. Evidence
职责边界:本维度不负责统计报告格式的完整性检查(如效应量、置信区间的具体格式)。单纯缺少效应量或置信区间且未使用强声称词时,不作为扣分项。
D3 — Causal Language
行为用法豁免:当触发词(如 "produced")用于描述被试行为(如 "participants produced responses")而非因果推断时,不视为问题,该维度记 5 分。
D4 — Interpretation in Results
D5 — Subjective / Evaluative Language
最终输出必须严格使用以下六个标题,不得使用其他自定义标题,以便汇总 Skill 统一整合各子 Skill 的诊断结果。
# Results Claim & Hedging Audit
## Dimension Score
评分:1—5 分(**5 分为最高,1 分为最低**)。先给出 D1—D5 每个维度的分数,再给出一个总体分,并附一句话总评。不得使用"2 分 / Pass"等含糊表述。
> 示例(高质量草稿,各维度均无问题):
> - D1 Hedging: 5 分 — hedging 使用恰当,无过度或缺失。
> - D2 Claim Strength vs. Evidence: 5 分 — 声称强度与证据匹配,无过度声称。
> - D3 Causal Language: 5 分 — 因果语言与设计类型匹配,无问题。
> - D4 Interpretation: 5 分 — 纯数据报告,无理论解释渗入。
> - D5 Subjective Language: 5 分 — 用词客观,无评价性语言。
> - **总体分: 5 分** — 草稿质量较高,声称强度与证据匹配,无过度声称或因果跳跃。
> 示例(存在问题的草稿):
> - D1 Hedging: 3 分 — 间接证据处缺少 hedging。
> - D2 Claim Strength vs. Evidence: 2 分 — 使用 "strong evidence" 但未提供对应统计证据。
> - D3 Causal Language: 4 分 — 因果语言基本匹配,个别动词可更中性。
> - D4 Interpretation: 5 分 — 无理论解释渗入。
> - D5 Subjective Language: 3 分 — 出现 "instructive" 等评价性形容词。
> - **总体分: 3 分** — 存在多处声称-证据不匹配和 hedging 缺失,建议逐句修改。
## Key Problems
主要问题:逐条列出。
> 示例:
> 1. S4 使用 "strong evidence" 但未提供效应量或置信区间支撑。
> 2. S6 的 "produced" 虽在结构启动文献中属惯用行为动词,但匹配 D3 触发词清单。
## Evidence from Draft
原文证据:摘录相关句子。
> 示例:
> - **S4**: "It showed a strong main effect of prime type (Estimate = .56, SE = .16, z = 3.60, p < .001)."
> - **S6**: "Pair-wise comparisons indicated that participants produced fewer DO responses following PO-An primes."
## Example-based Comparison
与 examples 的对照:引用具体 Example ID 并说明对应关系。
> 示例:
> - **S4 vs. F-07**(Elkin et al., 1989):F-07 使用 "there was evidence of significant superiority" 进行适度 hedging,而 S4 直接使用 "strong",声称强度超出证据支持。
> - **S6 vs. F-01**(Milgram, 1963):F-01 使用 "showed signs of" 描述被试行为,是规范的行为描述;S6 的 "produced" 同样用于描述被试行为,属行为用法豁免。
## Revision Suggestions
修改建议:逐句给出可操作的替换方案。
> 示例:
> - **S4 修改前**: "It showed a strong main effect of prime type (Estimate = .56, SE = .16, z = 3.60, p < .001)."
> **S4 修改后**: "It showed a main effect of prime type (Estimate = .56, SE = .16, z = 3.60, p < .001)."
> - **S6 修改前**: "participants produced fewer DO responses following PO-An primes"
> **S6 修改后**: "participants gave fewer DO responses following PO-An primes"
## Priority Level
修改优先级:按高、中、低排列。
> 示例:
> - **高**: S4 — 强声称词 "strong" 缺乏证据支撑,可能误导读者高估效应大小。需删除 "strong" 或补充效应量。
> - **中**: S6 — "produced" 触发 D3 但属行为用法豁免,替换为 "gave" 可消除歧义,非必须。
> - **低**: 全文风格统一性检查,建议后续通读确认所有声称词与证据匹配。
references/examples/examples_memberF.md — 18 curated examples (F-01 ~ F-18) from 8 classic psychology papers. Load this file before generating revision suggestions to match user sentences against canonical patterns.references/rubric.md — (Placeholder) Extended scoring rubric with domain-specific rules.references/checklist.md — (Placeholder) Pre-submission checklist for authors.results-statistics-convention-checker 负责。