This project runs a real meta-analysis / systematic-review pipeline with an agent doing the grunt work and a human standing at the decisions that carry authorship. That design only means something if the claims about "human oversight" are true and are stated at exactly the strength they hold — no more. This document is the single source of truth for that stance. It is referenced, verbatim in intent, from the README and from the skills' supervised-mode section.
The core commitment, in one sentence:
AI drafts it · a human signs off at seven irreversible gates · the signing author writes the theory and takes full responsibility for every claim. Fully-unsupervised auto-submission is out of scope — by design, for technical and accountability reasons, not as a missing feature.
Everything below explains what that sentence guarantees, and — just as important — what it does not guarantee.
Human oversight in this project happens in two distinct phases with different jobs. Conflating them is exactly how tools end up claiming more certainty than they have, so we keep them separate.
While the pipeline runs, it stops at each irreversible point and hands you an audit card built for that decision. The card carries the agent's recommendation (marked recommended), the reason for it, the consequence of each option, a drill-down link into the underlying evidence, and the freedom to send it back. The recommended mark is deliberate: it is the agent's informed recommendation based on knowledge you may not hold — that is value, not a bias to be hidden.
Be precise about what this phase is:
- It is a flow-first informed confirmation, not an independent re-verification by you. We do not inject known-wrong items to test whether you are really looking, and we add no "prove-you-read-it" friction. Running the pipeline in one continuous flow is a feature, not a gap to be patched.
- What is machine-enforced here is the agent's truthfulness toward you, not your independence.
The enforcement chain (
signoff_preflight.pyplus the structured decision block) guarantees that what the agent shows you is real: the question you signed is the question that was posed (same-source hash), every figure or table the card points at actually exists, and the narrative does not contradict the numbers. If any of those fail, the downstream stage cannot start. - So Phase 1 answers "is the agent showing me the truth?" — it does not claim "the human independently verified this." That claim belongs to Phase 2.
The independent human review happens after the draft is complete, when the signing author works through a structured audit package from scratch. The package reorganizes material that would otherwise be scattered across 90+ files into a bounded, reachable, educational audit path: a responsibility map (what you must do yourself · what you must spot-check · what you may trust the verification chain for), a small set of audit stations, and a three-tier signature page. The budgeted effort is on the order of 9–13 hours, splittable into 3–4 half-days — a defined checklist, not an unbounded "read everything again."
This is where independent judgment actually occurs. It is also where the disclosure that goes into the manuscript is generated — mechanically, from what was actually signed, so the stated strength of human oversight can never exceed the fact.
The two phases are disclosed together, in the manuscript, in these exact terms:
in-run informed confirmations with drill-down access + post-hoc independent audit via structured audit package
Nothing stronger. The wording is derived from sign-off facts (which gates were signed, whether evidence was opened, which signature tiers were completed), never authored to impress.
The operational contracts for both phases are the authority, not this summary:
- In-run gates →
skills/review-methodology-foundations/references/SUPERVISED-MODE-GATES.md - Post-hoc audit package →
skills/review-methodology-foundations/references/SIGNING-AUTHOR-AUDIT-PACKAGE.md - Supervised-mode overview →
skills/meta-analysis/SKILL.md§9.1
Supervised mode runs the full publication-grade methodology and adds seven structural sign-off gates at the irreversible points. Each gate stops the pipeline, renders an audit card, and waits.
| Gate | Irreversible point it guards | What it blocks until you sign |
|---|---|---|
| SP1 | Freeze the search strategy + eligibility criteria | Running the search and screening |
| SP2 | Screening disagreements + Cohen's κ | Moving to extraction |
| SP3 | Calibration before extraction goes live | Full-volume extraction |
| SP4 | Low-confidence extractions + key-number verification | Effect-size computation |
| SP5 | Model + heterogeneity (I²/τ²/PI) | Moderators, sensitivity, publication bias |
| SP6 | Interpretation, moderators, publication bias, claims-vs-data | Writing the draft |
| SP7 | Citation final check (HTTP-verified) | Delivering the final manuscript |
Three properties make these real gates, not prompts to "please review":
- They actually block. Each gate's verdict lands in a
decisions/file. Before the next stage runs,signoff_preflight.pyreads it — unsigned, rejected, or failing a hard check → non-zero exit code, and the downstream stage cannot start. - Red lights are computed by the engine, not read by the model. Every statistical flag (I² > 90%, one study carrying > 30% of the weight, a non-significant PET intercept, funnel asymmetry, κ below threshold, seed recall under 100%, an unverified citation) is derived by R / the preflight script from the fitted object's fields. The model never eyeballs a printout to decide what is significant.
- Hard checks come before the human sign-off. A red light, a κ below threshold, or an unverified citation is evaluated first and blocks even a signed gate — so a rubber-stamp cannot wave a broken result through. The evaluation order is fixed: missing file → hard check → sign-off → pass.
The division of labor is deliberate and is not going to move:
- The agent does the grunt work — search ingestion, dual-blind screening, dual extraction, running the statistics in R, drafting figures and prose. This is the work that used to consume hundreds of research-assistant hours.
- The human makes the calls that cannot be delegated — the seven sign-off judgments, the post-hoc audit, the construct-validity decisions, the strength of the stated conclusions, and the parts of the paper that are irreducibly authorial: the introduction, the discussion, and the theoretical framing.
- The human puts their name on it and is responsible for every claim. Authorship is not a formatting step at the end; it is the reason the gates exist.
We do not ship — and will not ship — a fully-autonomous "no human" mode that drafts and submits on your behalf. This is out of scope for two reasons, and only these two: a technical one (the mechanism has not been validated to a level where unsupervised output is trustworthy — see §5) and an accountability one (a scientific claim needs a responsible human author; a pipeline cannot be that author). It is not a feature we ran out of time to build.
This project's standing rule is to never claim more certainty than the evidence supports, and to state oversight at exactly its true strength. Applied to the tool itself:
- Disclosure never exceeds fact. The strength of human oversight written into any output is generated from what was actually signed, and is monotone in the facts — it can only under-claim, never over-claim. In-run gates are disclosed as informed confirmation, not as independent verification, because that is what they are.
- The enforcement chain prevents collapse, not malice. It reliably catches the real failure mode — an agent taking shortcuts (fabricating a terminal summary, copying a template, writing the sign-off itself): those trip a hash mismatch, a missing-evidence check, or a narrative-vs-number contradiction, and the pipeline stops. It does not defend against a fully adversarial agent that manufactures a self-consistent decision block, computes the correct hash, and points its evidence at real files — that would pass. The in-run hash locks integrity (nothing was altered after signing), not authenticity (that the block came from a genuine presentation). The post-hoc audit, done from scratch by the signing author, is the backstop for exactly this.
- The audit package guarantees a path, not attention. It guarantees the signing author has a bounded, reachable, educated audit route and a disclosure that does not exceed what was signed. It does not guarantee the author truly looked — a checkbox is not proof of reading. That residual is carried by the attestation wording ("attests to having…", i.e. the recorded signing act, not a verified reading) and optional check-questions, not by pretending a machine can enforce diligence.
- Mechanism validated on synthetic data; the real-data last mile is in progress. The supervised engine — gates, calibration, red-light computation — has been exercised end-to-end on a synthetic pilot. Validating the full chain on real PDFs and real database searches is ongoing. The synthetic pilot's accuracy numbers are an optimistic upper bound, not real-world accuracy; treat them accordingly.
- It will not fabricate to prop up a result. No invented data, no invented citations, no invented statistics to make a finding look stronger or a gate look passed. Suspected fabricated citations are flagged red at SP7; unverifiable numbers block, they do not get waved through.
- It will not soften a red light into a "feature." Red lights are computed by R. A signed rubber-stamp cannot clear one — the hard check runs before the sign-off and blocks a signed gate just the same. "The result is red but the author accepted it anyway" is not a supported path.
- It will not submit or claim authorship on your behalf. The tool drafts and checks; submission, correspondence, and authorship are yours.
Guidance, in the spirit of the RAISE principle this project follows — human oversight · transparent disclosure · final responsibility on the human:
- Run supervised mode for anything you intend to submit. It carries the full strict-mode methodology and adds the seven structural gates on top. The other modes are legitimate for coursework and proposals, but each discloses its own simplifications — read those disclosures.
- Complete the post-hoc audit package before you declare the work done. The audit entry point and the three-tier signature page are a delivery gate, not a suggestion; the independent human review is the point of the whole design.
- Let the disclosure be generated from the sign-off facts. Do not hand-write a human-oversight paragraph that reads stronger than what you actually signed. The pipeline generates it from the facts precisely so it cannot drift above them.
- You are the author. Every claim in the manuscript is yours to stand behind. The gates and the audit package exist to make that responsibility discharge-able in bounded time — they do not transfer it to the tool.
- 核心承诺:AI 初编 · 人在 7 个不可逆点签核 · 署名人写理论并对每一条主张负全责。不支持无人审核的自动投稿——出于技术与问责双重理由,是设计边界,不是缺失的功能。
- 两段式监督(各司其职,不可混淆):
- 运行中 = 心流优先的知情确认(Agent 推荐+理由+后果+可下钻+可打回;保留"(推荐)"预标=Agent 基于你不掌握的知识给的知情建议)。它不是你的独立复核,也不加"考验人"的摩擦。机器强制的是 Agent 侧的真实性(同源呈现 / 证据存在 / 叙事不与数字矛盾),不是逼你独立判断。
- 事后 = 署名人经审核包从 0 独立走查(9–13h,可拆 3–4 个半天)——这才是真正的独立人类审核,也是稿件披露措辞的机读事实源。
- 披露原句(如实、不夸大):
in-run informed confirmations with drill-down access + post-hoc independent audit via structured audit package。 - 7 门是真门:缺签 / 打回 / 硬检查不过 → preflight
exit≠0,下游启动不了;红灯由 R 算(不是 LLM 读 summary);硬检查先于软签,红灯拦得住已签的门(橡皮图章救不回)。 - 诚实边界(绝不重蹈"披露高于事实"的实战病根):披露强度单调不超过事实;强制链防"塌缩"非防"恶意";审核包保证有界可达被教育的动线+披露不超过签了什么,不保证署名人真用心看了(勾选≠真看);机制在合成数据上验证过,真实数据最后一公里进行中,合成 pilot 准确率是乐观上界而非真实准确率。
- 不做什么:不伪造数据/引用/统计撑门面;不把红灯软化成 feature;不替你投稿或署名。
- 运行契约以下列为准(本文件不复制其内容):
skills/review-methodology-foundations/references/SUPERVISED-MODE-GATES.md(运行中 7 门)·skills/review-methodology-foundations/references/SIGNING-AUTHOR-AUDIT-PACKAGE.md(事后审核包)·skills/meta-analysis/SKILL.md§9.1(监督模式总览)。