diff --git a/CHANGELOG.md b/CHANGELOG.md index 22deec316..bcdb33074 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -9,7 +9,16 @@ `aworld-cli` image behavior. - Added the asynchronous HTTP service, direct URL/object-reference ingestion, bounded queuing, persistent PaddleOCR-VL worker lifecycle, resumable PDF - batches, Document IR v2 formatting repair, and video evidence artifacts. + batches, Document IR v3 grounding/formatting repair, and video evidence + artifacts. +- Preserved Paddle's detector geometry and confidence before VLM block merging, + normalized contextual page labels, and removed HTML/Markdown presentation + tokens from grounding attribution text. On a fixed 20-case Layout sample, + the official score improved from 62.31% to 76.71% with zero regressions. +- Completed the clean 500-case Visual Grounding evaluation at 71.6970% and + updated the five-dimension equal-weight overall to 65.3520%. The report also + distinguishes that leaderboard aggregation from the 65.4367% case-weighted + mean across 2,534 numeric results. - Added public deployment, hardware, configuration, operations, and security documentation plus a pinned ParseBench evaluation report. - Removed environment-specific deployment history, internal hosts, repository diff --git a/aworld-tools/filex/README.md b/aworld-tools/filex/README.md index 448c4ec37..a0a903d1b 100644 --- a/aworld-tools/filex/README.md +++ b/aworld-tools/filex/README.md @@ -13,7 +13,7 @@ The repository ships two deployment shapes: Supported inputs include PDF, Markdown/text, Word, PowerPoint, Excel/CSV, images, audio, video, HTTP(S) files, and YouTube transcript sources. PDF output -uses Document IR v2; video output can include timestamped keyframes, OCR +uses Document IR v3; video output can include timestamped keyframes, OCR evidence, and a storyboard. Video evidence is not yet a full semantic video understanding model. @@ -261,10 +261,13 @@ Useful model variables are: ## Evaluation -The latest validated component baseline on the pinned ParseBench 2,553-case -suite is **52.07% five-dimension equal-weight overall**. Strongest performance -is content faithfulness (82.88%); the main remaining gap is visual grounding -(5.30%). +The pinned ParseBench suite contains 2,553 cases and all five dimensions have +reached a trusted terminal state. The older 5.30% layout number is retired +because it mixed 458 legacy contract zeros with 42 current Document IR results; +it is not a valid score for the current runtime. A fixed 20-case current-layout +A/B improved from **62.31% to 76.71%** after the Document IR v3 repair (11 +improved, 9 unchanged, 0 regressed). The subsequent clean 500-case layout run +scored **71.70%**. | Dimension | Cases | FileX score | Official PaddleOCR-VL-1.6 reference | | --- | ---: | ---: | ---: | @@ -272,13 +275,12 @@ is content faithfulness (82.88%); the main remaining gap is visual grounding | Charts | 568 | 56.14% | 54.24% | | Content faithfulness | 506 | 82.88% | 82.71% | | Semantic formatting | 476 | 48.40% | 54.64% | -| Visual grounding / layout | 500 | 5.30% | 77.80% | -| Equal-weight overall | 2,553 | **52.07%** | **67.43%** | +| Visual grounding / layout | 500 | **71.70%** | 77.80% | +| Equal-weight overall | 2,553 | **65.35%** | **67.43%** | Nineteen formatting cases returned `not_scored` and are excluded rather than -counted as zero. The table combines the newest trusted campaign for each -dimension, so it is a component baseline rather than a claim that every row was -produced by one immutable release. See +counted as zero. The overall row is the equal-weight mean of the five dimension +means; the case-weighted mean over 2,534 numeric results is 65.44%. See [the ParseBench evaluation report](docs/parsebench-evaluation.md) for pinned revisions, methodology, limitations, and the optimization roadmap. diff --git a/aworld-tools/filex/docs/parsebench-evaluation.md b/aworld-tools/filex/docs/parsebench-evaluation.md index c59d9e673..78b447633 100644 --- a/aworld-tools/filex/docs/parsebench-evaluation.md +++ b/aworld-tools/filex/docs/parsebench-evaluation.md @@ -17,9 +17,9 @@ planning, LLM quality, or downstream task completion. mean across the five dimensions System/control-plane errors were tracked separately from parser scores. A -`not_scored` case was not converted to zero. The final baseline contains 2,534 -numeric scores, 19 `not_scored` formatting cases, and no unresolved execution -failures. +`not_scored` case was not converted to zero. All five dimensions reached a +trusted terminal state with no unresolved execution failures. Nineteen +formatting cases remain `not_scored` and are excluded from score means. ## Results @@ -29,13 +29,31 @@ failures. | Charts | 568 | 568 | 0 | 56.1396% | 54.24% | +1.8996 pp | | Content faithfulness | 506 | 506 | 0 | 82.8802% | 82.71% | +0.1702 pp | | Semantic formatting | 476 | 457 | 19 | 48.4013% | 54.64% | -6.2387 pp | -| Visual grounding / layout | 500 | 500 | 0 | 5.2995% | 77.80% | -72.5005 pp | -| Equal-weight overall | 2,553 | 2,534 | 19 | **52.0726%** | **67.43%** | **-15.3574 pp** | - -The official reference values are the ParseBench published -PaddleOCR-VL-1.6 Full Pipeline results. FileX is close to the reference on -tables and content, exceeds it on this chart run, trails on formatting, and -does not yet satisfy the benchmark's grounding contract. +| Visual grounding / layout | 500 | 500 | 0 | 71.6970% | 77.80% | -6.1030 pp | +| Equal-weight overall | 2,553 | 2,534 | 19 | **65.3520%** | **67.43%** | -2.0780 pp | + +The official reference values are the ParseBench published PaddleOCR-VL-1.6 +Full Pipeline results. FileX is close to the reference on tables and content, +exceeds it on this chart run, and trails on formatting and visual grounding. +The overall value is an equal-weight mean of the five dimension means, matching +the published leaderboard aggregation. A case-weighted mean over the 2,534 +numeric results is 65.4367%; it answers a different question and should not be +compared directly with the 67.43% leaderboard overall. + +The previously reported 5.2995% layout value is retired: it combined 458 +legacy hard-coded contract zeros with 42 newly scorable Document IR cases. It +measured a migration state, not the current parser. On the same fixed 20 cases, +Document IR v3 improved the official mean from 62.3079% to 76.7097% (+14.4018 +percentage points): 11 cases improved, 9 were unchanged, and none regressed. +The two worst cases moved from 0% to 100% and from 25% to 75%. + +The clean 500-case layout result also corrects an evaluation-adapter issue in +64 PDFs containing both `layout` annotations and legacy `order` rows. The +upstream JSONL loader otherwise constructs a parse test case instead of a +layout-detection test case. Only `layout` rows are passed into the official +layout evaluator; reading order remains scored by that evaluator from each +annotation's `ro_index`. Those 64 cases all produced numeric scores, averaging +79.6376%, and the complete 500-case layout mean is 71.6970%. ## Interpretation @@ -55,10 +73,10 @@ scores do not measure service reliability or throughput. ### Main gaps -1. **Visual grounding**: FileX does not yet emit reliable element bounding - boxes in the coordinate system required by ParseBench. A correct fix needs - page geometry, element identity, coordinate normalization, and confidence; - guessing boxes would inflate neither correctness nor usefulness. +1. **Visual grounding**: Document IR v3 now preserves Paddle detector boxes, + confidence, page geometry, and parser text/order metadata. Remaining errors + are concentrated in detector misses, class ambiguity, and content + attribution for image-only controls. 2. **Semantic formatting**: headings, lists, emphasis, superscript/subscript, and reading-order boundaries still lose information in difficult pages. 3. **Not-scored formatting cases**: these require separate contract diagnosis; @@ -66,21 +84,23 @@ scores do not measure service reliability or throughput. ## Recommended optimization path -1. Extend Document IR v2 with page-relative and pixel-space bounding boxes, - source image dimensions, rotation, and stable element IDs. -2. Carry Paddle layout detections through normalization instead of rebuilding - geometry from Markdown. +1. Keep the fixed 20-case A/B as the release regression set and rerun the full + 500-case layout suite after detector or mapping changes. +2. Tune detector thresholds and label aliases against a separate development + split, especially small pictures and page furniture, without changing the + pinned official scorer. 3. Add a formatting state machine that reconciles OCR spans with the PDF text layer and preserves nested lists, heading levels, emphasis, and scripts. -4. Create regression sets per formatting rule and grounding object class before - another full-suite run. -5. Re-run only the affected dimensions with pinned data/scorer revisions, then - publish a release-level score from one immutable FileX revision. +4. Create regression sets per formatting rule and grounding object class. +5. Publish a release-level score only from one immutable FileX revision and + complete campaign manifest. ## Reproducibility boundary This is the newest trusted result for each component, not a single campaign -executed from one immutable FileX commit. It should be used as an engineering -baseline and optimization guide. A release claim should pin the FileX image -digest, parser configuration, hardware profile, data revision, scorer revision, -and all five campaign manifests. +executed from one immutable FileX commit. The 500 layout cases use one repaired +FileX revision, while the five-dimension overall combines the newest trusted +campaign per dimension. It should be used as an engineering baseline and +optimization guide. A release claim should pin the FileX image digest, parser +configuration, hardware profile, data revision, scorer revision, and all five +campaign manifests. diff --git a/aworld-tools/filex/src/document_parse_service/pdf/paddle_ocr_pdf_provider.py b/aworld-tools/filex/src/document_parse_service/pdf/paddle_ocr_pdf_provider.py index 2e7b75f5a..04d3b31c7 100644 --- a/aworld-tools/filex/src/document_parse_service/pdf/paddle_ocr_pdf_provider.py +++ b/aworld-tools/filex/src/document_parse_service/pdf/paddle_ocr_pdf_provider.py @@ -11,7 +11,7 @@ import threading import time from dataclasses import dataclass, field -from html import escape +from html import escape, unescape from pathlib import Path from typing import Any @@ -178,46 +178,14 @@ def _build_document_ir( pages: list[dict[str, Any]] = [] for fallback_index, result in enumerate(raw_results): payload = cls._json_payload(result) - elements: list[dict[str, Any]] = [] - blocks = payload.get("parsing_res_list") - if not isinstance(blocks, list): - blocks = [] - for fallback_order, block in enumerate(blocks, start=1): - if not isinstance(block, dict): - continue - bbox = block.get("block_bbox") or block.get("bbox") or [] - if not isinstance(bbox, (list, tuple)) or len(bbox) != 4: - continue - try: - normalized_bbox = [float(value) for value in bbox] - except (TypeError, ValueError): - continue - order = block.get("block_order") - elements.append( - { - "id": str( - block.get("global_block_id") - or block.get("block_id") - or f"p{fallback_index}-b{fallback_order}" - ), - "type": str( - block.get("block_label") - or block.get("label") - or "unknown" - ), - "bbox": normalized_bbox, - "text": str( - block.get("block_content") - or block.get("content") - or "" - ), - "reading_order": ( - int(order) if isinstance(order, (int, float)) else None - ), - "group_id": block.get("global_group_id") - or block.get("group_id"), - } - ) + blocks = cls._parsed_blocks(payload) + page_height = cls._numeric_dimension(payload.get("height")) + elements = cls._layout_elements( + payload, + blocks, + fallback_index, + page_height=page_height, + ) page_index = payload.get("page_index") pages.append( { @@ -233,11 +201,212 @@ def _build_document_ir( } ) return { - "schema_version": "filex-document-ir-v2", + "schema_version": "filex-document-ir-v3", "coordinate_system": "pixel_top_left_xyxy", "pages": pages, } + @classmethod + def _layout_elements( + cls, + payload: dict[str, Any], + parsed_blocks: list[dict[str, Any]], + page_index: int, + *, + page_height: float | None, + ) -> list[dict[str, Any]]: + """Keep detector geometry and enrich it with parser text/order metadata. + + PaddleOCR-VL may merge adjacent detector regions before VLM parsing. The + merged regions are useful for Markdown, but they erase small layout + objects and are therefore unsuitable as the canonical grounding output. + """ + + raw_layout = payload.get("layout_det_res") + raw_boxes = raw_layout.get("boxes") if isinstance(raw_layout, dict) else None + if not isinstance(raw_boxes, list) or not raw_boxes: + return [ + cls._parsed_element(block, page_index, order, page_height=page_height) + for order, block in enumerate(parsed_blocks, start=1) + ] + + elements: list[dict[str, Any]] = [] + matched_parsed: set[int] = set() + for order, raw_box in enumerate(raw_boxes, start=1): + if not isinstance(raw_box, dict): + continue + bbox = cls._bbox(raw_box.get("coordinate") or raw_box.get("bbox")) + if bbox is None: + continue + match_index = cls._best_text_match(bbox, parsed_blocks) + parsed = parsed_blocks[match_index] if match_index is not None else {} + if match_index is not None: + matched_parsed.add(match_index) + parsed_order = parsed.get("block_order") + confidence = raw_box.get("score") + element = { + "id": str( + raw_box.get("id") + or f"p{page_index}-d{order}" + ), + "type": cls._normalize_layout_label( + raw_box.get("label") or parsed.get("block_label"), + bbox=bbox, + page_height=page_height, + ), + "bbox": bbox, + "text": cls._plain_text( + parsed.get("block_content") or parsed.get("content") + ), + "reading_order": ( + int(parsed_order) + if isinstance(parsed_order, (int, float)) + else None + ), + "group_id": parsed.get("global_group_id") or parsed.get("group_id"), + "source": "layout_detection", + } + if isinstance(confidence, (int, float)): + element["confidence"] = float(confidence) + elements.append(element) + + # Preserve parser-only blocks such as generated tables/formulas when no + # detector region represents them. Do not duplicate blocks already used + # to enrich a detector prediction. + for index, block in enumerate(parsed_blocks): + if index not in matched_parsed: + elements.append( + cls._parsed_element( + block, + page_index, + index + 1, + page_height=page_height, + ) + ) + return elements + + @classmethod + def _parsed_blocks(cls, payload: dict[str, Any]) -> list[dict[str, Any]]: + blocks = payload.get("parsing_res_list") + if not isinstance(blocks, list): + return [] + return [block for block in blocks if isinstance(block, dict) and cls._bbox( + block.get("block_bbox") or block.get("bbox") + ) is not None] + + @classmethod + def _parsed_element( + cls, + block: dict[str, Any], + page_index: int, + fallback_order: int, + *, + page_height: float | None, + ) -> dict[str, Any]: + order = block.get("block_order") + bbox = cls._bbox(block.get("block_bbox") or block.get("bbox")) + return { + "id": str( + block.get("global_block_id") + or block.get("block_id") + or f"p{page_index}-b{fallback_order}" + ), + "type": cls._normalize_layout_label( + block.get("block_label") or block.get("label"), + bbox=bbox, + page_height=page_height, + ), + "bbox": bbox, + "text": cls._plain_text(block.get("block_content") or block.get("content")), + "reading_order": int(order) if isinstance(order, (int, float)) else None, + "group_id": block.get("global_group_id") or block.get("group_id"), + "source": "document_parsing", + } + + @staticmethod + def _plain_text(value: Any) -> str: + """Return semantic block text without Markdown/HTML presentation tokens.""" + + text = str(value or "") + text = re.sub(r"<[^>]+>", " ", text) + text = re.sub(r"!\[([^]]*)\]\([^)]+\)", r" \1 ", text) + text = re.sub(r"\[([^]]+)\]\([^)]+\)", r" \1 ", text) + text = re.sub(r"(?m)^\s{0,3}#{1,6}\s+", "", text) + return re.sub(r"\s+", " ", unescape(text)).strip() + + @staticmethod + def _normalize_layout_label( + value: Any, + *, + bbox: list[float] | None, + page_height: float | None, + ) -> str: + label = str(value or "unknown").strip().lower().replace("-", "_").replace(" ", "_") + aliases = { + "figure": "image", + "picture": "image", + "footer_image": "image", + "header_image": "image", + "caption": "figure_title", + "title": "paragraph_title", + "page_footer": "footer", + "footer_text": "footer", + "page_header": "header", + "header_text": "header", + } + label = aliases.get(label, label) + if label in {"number", "page_number"} and bbox and page_height: + if bbox[1] >= page_height * 0.9: + return "footer" + if bbox[3] <= page_height * 0.1: + return "header" + return label + + @staticmethod + def _bbox(value: Any) -> list[float] | None: + if not isinstance(value, (list, tuple)) or len(value) != 4: + return None + try: + bbox = [float(item) for item in value] + except (TypeError, ValueError): + return None + if bbox[2] <= bbox[0] or bbox[3] <= bbox[1]: + return None + return bbox + + @classmethod + def _best_text_match( + cls, + detector_bbox: list[float], + parsed_blocks: list[dict[str, Any]], + ) -> int | None: + best_index: int | None = None + best_iou = 0.0 + for index, block in enumerate(parsed_blocks): + parsed_bbox = cls._bbox(block.get("block_bbox") or block.get("bbox")) + if parsed_bbox is None: + continue + iou = cls._bbox_iou(detector_bbox, parsed_bbox) + if iou > best_iou: + best_iou = iou + best_index = index + # A strict IoU avoids copying one merged parser region's text into each + # of several small detector boxes contained by that region. + return best_index if best_iou >= 0.5 else None + + @staticmethod + def _bbox_iou(left: list[float], right: list[float]) -> float: + x1 = max(left[0], right[0]) + y1 = max(left[1], right[1]) + x2 = min(left[2], right[2]) + y2 = min(left[3], right[3]) + intersection = max(0.0, x2 - x1) * max(0.0, y2 - y1) + if intersection <= 0: + return 0.0 + left_area = (left[2] - left[0]) * (left[3] - left[1]) + right_area = (right[2] - right[0]) * (right[3] - right[1]) + return intersection / (left_area + right_area - intersection) + @staticmethod def _json_payload(result: Any) -> dict[str, Any]: json_value = getattr(result, "json", None) diff --git a/aworld-tools/filex/tests/document_parse_service/test_paddle_ocr_pdf_provider.py b/aworld-tools/filex/tests/document_parse_service/test_paddle_ocr_pdf_provider.py index 939e7e129..28ad87ea9 100644 --- a/aworld-tools/filex/tests/document_parse_service/test_paddle_ocr_pdf_provider.py +++ b/aworld-tools/filex/tests/document_parse_service/test_paddle_ocr_pdf_provider.py @@ -185,7 +185,7 @@ def concatenate_markdown_pages(markdown_list): assert provider._pipeline_kwargs()["use_chart_recognition"] is True assert result.document_ir == { - "schema_version": "filex-document-ir-v2", + "schema_version": "filex-document-ir-v3", "coordinate_system": "pixel_top_left_xyxy", "pages": [ { @@ -195,11 +195,12 @@ def concatenate_markdown_pages(markdown_list): "elements": [ { "id": "7", - "type": "title", + "type": "paragraph_title", "bbox": [10.0, 20.0, 300.0, 80.0], "text": "Heading", "reading_order": 1, "group_id": 7, + "source": "document_parsing", } ], "spans": [], @@ -208,6 +209,85 @@ def concatenate_markdown_pages(markdown_list): } +def test_paddle_ocr_document_ir_prefers_unmerged_layout_detections() -> None: + module = _load_provider_module() + payload = { + "page_index": 0, + "width": 800, + "height": 600, + "layout_det_res": { + "boxes": [ + {"label": "image", "score": 0.91, "coordinate": [700, 20, 730, 50]}, + {"label": "image", "score": 0.89, "coordinate": [700, 60, 730, 90]}, + {"label": "chart", "score": 0.97, "coordinate": [20, 20, 650, 500]}, + ] + }, + "parsing_res_list": [ + { + "block_label": "aside_text", + "block_content": "merged icons", + "block_bbox": [695, 15, 735, 100], + "block_id": 1, + "block_order": 2, + }, + { + "block_label": "chart", + "block_content": "chart text", + "block_bbox": [20, 20, 650, 500], + "block_id": 2, + "block_order": 1, + }, + ], + } + + document_ir = module.PaddleOcrPdfProvider._build_document_ir([payload]) + + elements = document_ir["pages"][0]["elements"] + assert document_ir["schema_version"] == "filex-document-ir-v3" + assert [element["type"] for element in elements[:3]] == ["image", "image", "chart"] + assert elements[0]["text"] == "" + assert elements[1]["text"] == "" + assert elements[2]["text"] == "chart text" + assert elements[2]["reading_order"] == 1 + assert elements[2]["confidence"] == 0.97 + assert elements[3]["type"] == "aside_text" + assert elements[3]["source"] == "document_parsing" + + +def test_document_ir_uses_plain_semantic_text_and_page_context_labels() -> None: + module = _load_provider_module() + payload = { + "page_index": 0, + "width": 1000, + "height": 1000, + "layout_det_res": { + "boxes": [ + {"label": "table", "score": 0.9, "coordinate": [20, 20, 800, 850]}, + {"label": "number", "score": 0.8, "coordinate": [850, 950, 950, 980]}, + {"label": "footer_image", "score": 0.7, "coordinate": [20, 900, 120, 980]}, + ] + }, + "parsing_res_list": [ + { + "block_label": "table", + "block_content": "
| Lincoln | City |