Repository navigation
Expand file tree
/
Copy pathnssd_helper.py
More file actions
626 lines (532 loc) · 26 KB
/
Copy pathnssd_helper.py
File metadata and controls
626 lines (532 loc) · 26 KB
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
438
439
440
441
442
443
444
445
446
447
448
449
450
451
452
453
454
455
456
457
458
459
460
461
462
463
464
465
466
467
468
469
470
471
472
473
474
475
476
477
478
479
480
481
482
483
484
485
486
487
488
489
490
491
492
493
494
495
496
497
498
499
500
501
502
503
504
505
506
507
508
509
510
511
512
513
514
515
516
517
518
519
520
521
522
523
524
525
526
527
528
529
530
531
532
533
534
535
536
537
538
539
540
541
542
543
544
545
546
547
548
549
550
551
552
553
554
555
556
557
558
559
560
561
562
563
564
565
566
567
568
569
570
571
572
573
574
575
576
577
578
579
580
581
582
583
584
585
586
587
588
589
590
591
592
593
594
595
596
597
598
599
600
601
602
603
604
605
606
607
608
609
610
611
612
613
614
615
616
617
618
619
620
621
622
623
624
625
626
"""NSSD helper — independent primary search for Chinese social-sciences &
humanities (SSH) literature (国家哲学社会科学文献中心 / ncpssd.cn).
Role (v2.3.0 Phase 3): an A-tier, fully self-contained Chinese source that emits
``UnifiedPaperEntity`` in the SAME shape as ``ss_helper --search`` /
``openalex_helper`` so STEP 4/5 federated fusion ingests NSSD JSON exactly like
openalex.json. It is a discipline-scoped supplemental source for the ``zh``
search space (social sciences & humanities), alongside OpenAlex (baseline) and
yiigle (medicine). Pure ``requests`` (already a dependency) — zero new deps.
Endpoint (verified live 2026-07-12, no login / no signature / no WAF cookie):
POST https://www.ncpssd.cn/searchHandler/search
headers: Content-Type: application/x-www-form-urlencoded; charset=UTF-8
X-Requested-With: XMLHttpRequest
(ordinary browser User-Agent)
body: search=<PQ>&pageNum=1&pageSize=N&sort=
PQ (keyword-style; each concept block ORs its synonyms across 3 field codes
— title IKTE / subject IKST / abstract IKSE — and blocks are ANDed):
single concept: (IKTE="<t>" OR IKST="<t>" OR IKSE="<t>")
two concepts: (IKTE="A" OR IKST="A" OR IKSE="A")
AND (IKTE="B" OR IKST="B" OR IKSE="B")
The multi-block form recovers recall a single-string query loses: the
endpoint tokenises a multi-word IKxx="a b" value and forces both tokens
into the SAME field, dropping records that split the concepts across
fields (live-verified 519 vs 222, A3_nssd_verification.md 2026-07-16).
response {"result": bool, "code": 200, "data": {"total": int, "rows": [...]}}
Field mapping (row -> UnifiedPaperEntity), all verified against live rows:
title -> title
creator -> authors (split on ';' / ';'; strip '[1]' / '[1,2]'
affiliation markers; a minority of records use
whitespace-separated CJK names — handled
conservatively so foreign names like "John
Smith" stay one author)
cbw_name -> venue (journal / 集刊 title)
issn -> issn ONLY when it matches the ISSN pattern. For 集刊文章
(serial-book articles) this slot carries an ISBN
instead (e.g. "978-7-5642-1400-5",
"7-81049-896-7" — verified live); those are
rejected so an ISBN never masquerades as an ISSN.
remark -> abstract
years -> year (int)
id -> source_native_id = "nssd:<id>"
subject -> keywords real controlled subject terms, e.g.
"乡村振兴;乡村振兴战略" (';'-delimited) or the
whitespace-delimited variant "情绪调节 体育教学",
split into individual keywords.
range -> keywords inclusion/quality signal (e.g.
"1;CSSCI_C2017_2018;NSSD;RWSKHX") stored as a
single namespaced token "nssd_range:<raw>" so it
is not lost; B2 wires it into 分区 later.
type -> type mapped to the unified vocab — "article" for the
journal / 集刊 articles NSSD indexes. The RAW NSSD
type is ALSO kept as a namespaced
"nssd_type:<raw>" keyword so a downstream zh-space
router can identify / filter foreign-journal rows
("外文期刊文章") without this helper hard-coding that
policy. NB "外文期刊文章" is a JOURNAL classification,
not a promise the record is English — live rows
show Chinese-titled 外文期刊文章 — so we annotate,
never hard-drop.
doi -> doi lowercased + URL-prefix stripped to the bare
"10.x/y" form (matches yiigle / openalex so the
cross-source dedup key lines up). Usually empty;
a free dedup win when present, lifts paper_id to
the DOI.
(constant) -> sources = ["nssd"].
Compliance (hard requirements):
* Metadata + abstract ONLY. NEVER download or cache subscription full text
(NSSD full text needs a free login — deliberately not touched here).
* Nothing is written to disk / git — this module only fetches and returns.
* Source attribution string for downstream display:
"国家哲学社会科学文献中心(NSSD)" (see ``ATTRIBUTION``), printed to stderr on each
CLI search so a consuming pipeline can surface the source credit (parity
with yiigle_helper). Provenance also travels on ``sources=["nssd"]``.
Robustness: every network / HTTP / JSON failure degrades gracefully — the run
returns whatever was collected (empty on total failure) and prints a short note
to stderr; it never raises.
"""
from __future__ import annotations
import re
import sys
from typing import Dict, List, Optional
from urllib.parse import urlencode
from .types import Author, UnifiedPaperEntity
# ---------------------------------------------------------------------------
# Endpoint / request constants
# ---------------------------------------------------------------------------
SEARCH_URL = "https://www.ncpssd.cn/searchHandler/search"
# Human-readable source attribution (compliance). Printed to stderr on each CLI
# search (parity with yiigle_helper); provenance also travels on sources=["nssd"].
ATTRIBUTION = "国家哲学社会科学文献中心(NSSD)"
_SOURCE = "nssd"
_HEADERS = {
"Content-Type": "application/x-www-form-urlencoded; charset=UTF-8",
"X-Requested-With": "XMLHttpRequest",
"User-Agent": (
"Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) "
"AppleWebKit/537.36 (KHTML, like Gecko) Chrome/124.0 Safari/537.36"
),
}
_DEFAULT_TIMEOUT = 30
# The endpoint honours the requested pageSize; cap per request to stay polite.
_MAX_PAGE_SIZE = 100
# Hard cap on paged requests so a huge result set can never run away.
_MAX_PAGES = 5
# Field codes the keyword PQ addresses, in the canonical order title → subject →
# abstract. Each concept block ORs its synonyms across all three; the byte order
# of this tuple is load-bearing for R-19 (a single-token query must reproduce the
# pre-v2.4 "(IKTE=… OR IKST=… OR IKSE=…)" string exactly).
_FIELD_CODES = ("IKTE", "IKST", "IKSE")
# ISSN form "XXXX-XXXX" with an optional trailing check-digit X (e.g.
# "1674-344X"). An ISBN-10/13 (with or without hyphens) never matches this, so
# 集刊 rows whose ``issn`` slot actually holds an ISBN are rejected.
_ISSN_RE = re.compile(r"^\d{4}-\d{3}[\dXx]$")
# Affiliation markers appended to author names: "[1]", "[1,2]" (half-width) and
# the full-width variant "[1]".
_AFFIL_RE = re.compile(r"[\[[][^\]]]*[\]]]")
# A pure-CJK author token (allows the interpunct used in transliterated names).
_CJK_NAME_RE = re.compile(r"^[一-鿿㐀-䶿·・]+$")
# ---------------------------------------------------------------------------
# Field parsing (pure functions — unit-testable without network)
# ---------------------------------------------------------------------------
def _strip_term(term: object) -> str:
"""Trim a term and strip embedded double quotes (half/full-width) so they can
never break the ``IKxx="..."`` PQ delimiters.
This is byte-for-byte the sanitisation the pre-v2.4 single-string
``_build_pq`` applied (``.strip()`` then drop " “ ”), factored out so every
synonym in every block is sanitised identically."""
return (str(term) if term is not None else "").strip().replace('"', "").replace("“", "").replace("”", "")
def _block_clause(synonyms: List[str]) -> str:
"""Assemble ONE concept block into ``(IKTE="s1" OR IKST="s1" OR IKSE="s1"
OR IKTE="s2" OR ...)`` — every synonym ORed across the 3 field codes.
Empty / quote-only synonyms are dropped; an empty block yields ``""`` (the
caller filters those out so no stray ``()`` reaches the PQ)."""
ors = [
f'{fc}="{s}"'
for syn in synonyms
if (s := _strip_term(syn))
for fc in _FIELD_CODES
]
return "(" + " OR ".join(ors) + ")" if ors else ""
def _build_pq(query: str = "", blocks: Optional[List[List[str]]] = None) -> str:
"""Build the NSSD keyword PQ. Two entry modes (D-20: semantics to the LLM,
string assembly to code):
* ``blocks`` — a list of concept blocks, each a list of synonyms. Concept
boundaries and synonym generation are the UPSTREAM LLM's job (the
``query_plan`` concept blocks); this function only MECHANICALLY assembles
"OR the synonyms across 3 field codes inside a block, AND the blocks":
``(block1) AND (block2) AND ...``. Live-verified to recover ~519 hits for
数字经济 × 共同富裕 vs the 222 the single-string form returned
(A3_nssd_verification.md, 2026-07-16). When ``blocks`` is given ``query`` is
ignored (the caller has already done the splitting).
* ``query`` (fallback, ``blocks is None``) — mechanically split on
whitespace, each token becoming a single-synonym block. This uses ONLY the
whitespace boundaries the caller already supplied — it does NOT do semantic
concept recognition (that stays with the LLM). A SINGLE-token query yields
ONE block whose output is BYTE-FOR-BYTE identical to the pre-v2.4
single-string ``_build_pq`` (R-19 zero regression); a multi-token query
gains the cross-field AND recall the old single-string form lost.
Embedded double quotes are stripped so they cannot break the PQ delimiters.
"""
if blocks is None:
# Mechanical fallback: whitespace tokens -> one single-synonym block each.
blocks = [[tok] for tok in (query or "").split()]
clauses = [c for b in blocks if (c := _block_clause(b))]
return " AND ".join(clauses)
def _clean_str(v: object) -> Optional[str]:
"""Trim to a non-empty string, or None. Guards the rare literal 'None'."""
if v is None:
return None
s = str(v).strip()
if not s or s.lower() == "none":
return None
return s
def _clean_issn(raw: object) -> Optional[str]:
"""Return the value only if it is a real ISSN; else None.
This is the ISBN discriminator: 集刊文章 rows put an ISBN (e.g.
"978-7-5642-1400-5") in the ``issn`` slot, which must not surface as an ISSN.
"""
s = _clean_str(raw)
if s is None:
return None
return s if _ISSN_RE.match(s) else None
def _normalize_doi(raw: object) -> Optional[str]:
"""Bare, lowercased DOI ("10.x/y") — strips any URL / "doi:" prefix.
Mirrors ``yiigle_helper._normalize_doi`` and the OpenAlex / SS convention so
the cross-source dedup key (federated_kg_resolver keys on the lowercase DOI)
lines up. Previously NSSD only lowercased, so a "https://doi.org/10.x" row
would not collide with the bare "10.x" form emitted by the other sources.
"""
s = _clean_str(raw)
if s is None:
return None
for prefix in ("https://doi.org/", "http://doi.org/", "doi.org/", "doi:"):
if s.lower().startswith(prefix):
s = s[len(prefix):]
break
s = s.lower().strip()
return s or None
# Subject-term separators: half/full-width ';' ',' and the CJK enumeration '、'.
_SUBJECT_SEP_RE = re.compile(r"[;;,,、]")
def _parse_subject(raw: object) -> List[str]:
"""Split the NSSD ``subject`` slot into individual controlled subject terms.
NSSD delimits subjects inconsistently (verified live): usually ';' (e.g.
"乡村振兴;乡村振兴战略"), occasionally whitespace (e.g.
"情绪调节 体育教学模式 体育教学"). Split on ';' ',' '、' first; if that yields a
single whitespace-containing token, fall back to a whitespace split (same
conservative heuristic as ``_parse_authors``) so multi-word terms are not
over-fragmented in the common delimited case. De-duplicated, order-preserved.
"""
s = _clean_str(raw)
if s is None:
return []
parts = [p.strip() for p in _SUBJECT_SEP_RE.split(s)]
parts = [p for p in parts if p]
if len(parts) == 1 and re.search(r"\s", parts[0]):
parts = parts[0].split()
out: List[str] = []
for p in parts:
if p and p not in out:
out.append(p)
return out
def _parse_year(raw: object) -> Optional[int]:
"""Extract a 4-digit year as int, or None."""
s = _clean_str(raw)
if s is None:
return None
m = re.search(r"\d{4}", s)
return int(m.group()) if m else None
def _parse_authors(creator: object) -> List[Author]:
"""Parse the NSSD ``creator`` string into a list of ``Author``.
- Primary separator is ';' (also the full-width ';').
- Each name may carry a trailing affiliation marker ("[1]", "[1,2]") — stripped.
- A minority of records separate CJK author names by whitespace instead of
';'. We split on whitespace ONLY when the ';' split yielded a single name
AND every whitespace token is pure-CJK — so multi-token foreign names
("John Smith") are preserved as one author.
"""
s = _clean_str(creator)
if s is None:
return []
s = s.replace(";", ";") # full-width semicolon -> half-width
names: List[str] = []
for part in s.split(";"):
name = _AFFIL_RE.sub("", part).strip()
if name:
names.append(name)
if len(names) == 1 and re.search(r"\s", names[0]):
tokens = names[0].split()
if len(tokens) > 1 and all(_CJK_NAME_RE.match(t) for t in tokens):
names = tokens
return [Author(name=n) for n in names]
def _row_to_entity(row: Dict) -> UnifiedPaperEntity:
"""Convert one NSSD ``data.rows[i]`` record into a ``UnifiedPaperEntity``.
Produces the same entity shape openalex_helper / ss_helper emit, so the
federated resolver ingests NSSD JSON identically.
"""
rid = _clean_str(row.get("id")) or _clean_str(row.get("data_id"))
doi = _normalize_doi(row.get("doi"))
keywords: List[str] = []
# Real controlled subject terms first — these are genuine keywords.
keywords.extend(_parse_subject(row.get("subject")))
rng = _clean_str(row.get("range"))
if rng:
# Interim home for the CSSCI/CSCD/NSSD inclusion signal (B2 relocates it
# into 分区). Namespaced so it is unambiguously distinguishable and easy
# to strip: any consumer filters `kw.startswith("nssd_range:")`.
keywords.append(f"nssd_range:{rng}")
raw_type = _clean_str(row.get("type"))
if raw_type:
# Surface the RAW NSSD document/journal class (e.g. "外文期刊文章") as a
# namespaced, filterable signal so a downstream zh-space router can drop
# foreign-journal rows WITHOUT this helper hard-coding that policy. Same
# escape-hatch pattern as nssd_range:; any consumer filters
# `kw.startswith("nssd_type:")`.
keywords.append(f"nssd_type:{raw_type}")
return UnifiedPaperEntity(
doi=doi,
source_native_id=(f"nssd:{rid}" if rid else None),
title=_clean_str(row.get("title")) or "",
abstract=_clean_str(row.get("remark")),
authors=_parse_authors(row.get("creator")),
year=_parse_year(row.get("years")),
venue=_clean_str(row.get("cbw_name")),
issn=_clean_issn(row.get("issn")),
# Every NSSD record ingested here is a journal / 集刊 article, so the
# unified category is "article"; the Chinese-vs-foreign distinction is a
# journal classification (not a document type) and rides on the
# nssd_type:<raw> keyword above, never on this field.
type="article",
keywords=keywords,
doi_url=(f"https://doi.org/{doi}" if doi else None),
sources=[_SOURCE],
)
# ---------------------------------------------------------------------------
# HTTP + search
# ---------------------------------------------------------------------------
def _fetch_page(
query: str,
page_num: int,
page_size: int,
session=None,
blocks: Optional[List[List[str]]] = None,
) -> Optional["tuple[List[Dict], int]"]:
"""Fetch one page. Returns (rows, total), or None on any failure.
``session`` (a requests-like object with ``.post``) is injected by tests so
no network is touched; production passes None and uses ``requests``.
``blocks`` (optional concept blocks) is forwarded to ``_build_pq``; None
keeps the whitespace-fallback / single-string behaviour.
"""
if session is not None:
sess = session
else:
import requests # local import: only when a real network call is needed
sess = requests
body = urlencode(
{
"search": _build_pq(query, blocks=blocks),
"pageNum": page_num,
"pageSize": page_size,
"sort": "",
}
)
try:
res = sess.post(SEARCH_URL, data=body, headers=_HEADERS, timeout=_DEFAULT_TIMEOUT)
except Exception as exc: # network error, DNS, timeout, ...
print(f"[nssd_helper] request failed: {exc}", file=sys.stderr)
return None
if getattr(res, "status_code", None) != 200:
print(
f"[nssd_helper] NSSD returned HTTP {getattr(res, 'status_code', '?')}",
file=sys.stderr,
)
return None
try:
payload = res.json()
except Exception as exc:
print(f"[nssd_helper] could not parse NSSD JSON: {exc}", file=sys.stderr)
return None
# Shape guard: a legal-JSON body of an unexpected shape (a top-level list, or
# a `data` slot carrying a list instead of the {"rows": [...]} object) must
# NOT raise on `.get(...)` — degrade to an empty page (never-raise contract).
if not isinstance(payload, dict):
print(
f"[nssd_helper] unexpected NSSD JSON shape: {type(payload).__name__}",
file=sys.stderr,
)
return None
if not payload.get("result"):
print(
f"[nssd_helper] NSSD result=false (code={payload.get('code')})",
file=sys.stderr,
)
return None
data = payload.get("data")
if not isinstance(data, dict):
data = {}
rows = data.get("rows")
if not isinstance(rows, list):
rows = []
try:
total = int(data.get("total") or 0)
except (TypeError, ValueError):
total = 0
return rows, total
def search(
query: str,
n: int = 25,
year_min: Optional[int] = None,
session=None,
blocks: Optional[List[List[str]]] = None,
) -> List[UnifiedPaperEntity]:
"""Independent NSSD primary search — returns UnifiedPaperEntity[].
``n`` : max number of results to return.
``year_min`` : keep only papers with year >= this (client-side; the NSSD
form exposes no verified year param, so we over-fetch and
filter). Records with an unparseable year are dropped when a
``year_min`` is set.
``session`` : injected requests-like object (tests); None uses ``requests``.
``blocks`` : optional concept blocks (list of synonym lists) built by the
upstream LLM's ``query_plan``. When given, the PQ is the
multi-concept "block-internal OR × cross-block AND" form (P2-5
recall fix) and ``query`` is only the fallback label; when None
the flat ``query`` is whitespace-split — a single-token query
is byte-identical to pre-v2.4 (R-19).
Any failure degrades gracefully to whatever was collected (empty on total
failure); never raises.
"""
# Nothing to search? (empty query AND no usable block terms.) Bail before any
# network call. Using _build_pq here keeps the "is there anything to search"
# test in one place and correctly admits a blocks-only call with empty query.
if not _build_pq(query, blocks=blocks):
return []
target = max(1, int(n))
# Over-fetch when year-filtering client-side so we can still return ~target.
fetch_target = target * 4 if year_min is not None else target
page_size = min(max(fetch_target, 1), _MAX_PAGE_SIZE)
entities: List[UnifiedPaperEntity] = []
seen: set = set()
for page in range(1, _MAX_PAGES + 1):
result = _fetch_page(query, page, page_size, session=session, blocks=blocks)
if result is None:
break
rows, total = result
if not rows:
break
for row in rows:
if not isinstance(row, dict):
continue # shape guard: skip a non-dict row (never-raise)
ent = _row_to_entity(row)
# Filter by year FIRST, THEN mark `seen`: a year-dropped record must
# not poison the dedup set and hide a later same-key record whose year
# IS valid (e.g. two id-less rows that share a title fallback key).
if year_min is not None and (ent.year is None or ent.year < year_min):
continue
key = ent.source_native_id or ent.doi or ent.title
if key in seen:
continue
seen.add(key)
entities.append(ent)
if len(entities) >= target:
return entities[:target]
if len(rows) < page_size:
break
if page * page_size >= total:
break
return entities[:target]
# ---------------------------------------------------------------------------
# Serialization + CLI
# ---------------------------------------------------------------------------
def _to_dict(entity: UnifiedPaperEntity) -> Dict:
"""Serialise a UnifiedPaperEntity to a plain dict, flattening authors —
identical in shape to openalex_helper._to_dict / ss_helper's search
serializer, so downstream ingests NSSD JSON like openalex.json."""
out: Dict = {}
for field_name, value in entity.__dict__.items():
if field_name == "authors":
out[field_name] = [a.__dict__ for a in value]
else:
out[field_name] = value
return out
def _cli_resolve_blocks(
blocks_json: Optional[str], block_flags: Optional[List[str]]
) -> Optional[List[List[str]]]:
"""Resolve the CLI ``--blocks`` / ``--block`` flags into a concept-block
structure, or None (fall back to whitespace-splitting ``--search``).
* ``--blocks`` — a JSON list-of-lists, e.g. ``[["数字经济","数据要素"],["共同富裕"]]``.
* ``--block`` — repeatable; each occurrence is one block, synonyms
``'|'``-separated (e.g. ``--block "数字经济|数据要素" --block 共同富裕``).
``--blocks`` (JSON) wins when both are given. Malformed / wrong-shaped JSON
degrades to None with a stderr note (never raises), matching the module's
graceful-degradation contract."""
if blocks_json:
import json
try:
parsed = json.loads(blocks_json)
except Exception as exc:
print(f"[nssd_helper] ignoring --blocks (bad JSON: {exc})", file=sys.stderr)
return None
if isinstance(parsed, list) and parsed and all(isinstance(b, list) for b in parsed):
return [[str(s) for s in b] for b in parsed]
print(
"[nssd_helper] ignoring --blocks (want a non-empty JSON list-of-lists)",
file=sys.stderr,
)
return None
if block_flags:
return [flag.split("|") for flag in block_flags]
return None
if __name__ == "__main__":
import argparse
import json
parser = argparse.ArgumentParser(
description=(
"NSSD helper — independent primary search for Chinese social-"
"sciences & humanities literature (ncpssd.cn). Outputs "
"UnifiedPaperEntity[] like `ss_helper --search`."
)
)
parser.add_argument(
"--search",
metavar="QUERY",
default=None,
help=(
"keyword query (Chinese) — searched over title / subject / abstract. "
"Whitespace-split into concept blocks unless --block/--blocks is given. "
"Optional only when --block/--blocks supplies the concepts."
),
)
parser.add_argument(
"--n", type=int, default=25, help="max results to return (default 25)."
)
parser.add_argument(
"--year-min",
type=int,
help="minimum publication year, inclusive (client-side filter).",
)
parser.add_argument(
"--block",
action="append",
metavar="SYN1|SYN2|...",
help=(
"One concept block (repeatable). Synonyms within a block are "
"'|'-separated and ORed across title/subject/abstract; separate "
"--block flags are ANDed. Semantic concept-splitting is the caller's "
"job — this only assembles. Overrides whitespace-splitting of --search."
),
)
parser.add_argument(
"--blocks",
metavar="JSON",
help=(
'Concept blocks as a JSON list-of-lists, e.g. '
'\'[["数字经济","数据要素"],["共同富裕"]]\'. Takes precedence over --block.'
),
)
parser.add_argument(
"--output-file", help="write JSON here (defaults to stdout)."
)
args = parser.parse_args()
blocks = _cli_resolve_blocks(args.blocks, args.block)
if not (args.search and args.search.strip()) and blocks is None:
parser.error("provide --search QUERY and/or --block/--blocks")
results = search(args.search or "", n=args.n, year_min=args.year_min, blocks=blocks)
output = [_to_dict(p) for p in results]
payload = json.dumps(output, indent=2, ensure_ascii=False)
if args.output_file:
with open(args.output_file, "w", encoding="utf-8") as f:
f.write(payload)
else:
print(payload)
# Compliance attribution — stderr only (parity with yiigle_helper), so stdout
# stays a pure JSON document.
print(ATTRIBUTION, file=sys.stderr)