42 KiB
OCR-v2 Project Management Log
Branch:
master| Last Updated: 2026-07-02 Active work: Figure caption prefix recovery + inline HTML fix + table matching verification completed. 585 regression tests pass.0. Executive Summary
Current state: All 10-paper audit issues resolved (3 audit fix commits + rotated prematch refactor + asset-internal figure number recovery). U746UJ7G
figure_unknown_000→figure_002with recovered label "Plot of Criteria Time". 428 targeted tests pass. Next: Monitor production OCR for regressions, archive stale project docs.1. Architecture
1.1 The problem (pre-v2)
raw_label/text → final role → normalize → rescue → demote/promote → attach → renderEarly role guesses caused: title/author/footnote misclassification, body paragraphs typed as references, reference continuations leaking into body, left/right column interleave bugs, duplicate figure captions.
1.2 The solution (v2 target)
raw observations → structural signatures → stable anchors/families → zone inference → role resolution → figure/table validation → render + health1.3 Core principles
- seed_role is a proposal, never a final role —
assign_block_role()proposes;normalize_document_structure()decides- VERIFY_REQUIRED roles (paper_title, authors, section_heading, reference_item, figure_caption, etc.) must have
role_verification_status == "ACCEPT"with non-emptyrole_sourceandrole_evidence- Zone != Role — being in the body reading environment does not imply
body_paragraph- Reference tail first — reference sections protected from tail contamination
- Frontmatter source-backed — OCR localizes but must not invent/discard canonical frontmatter
Key design documents:
- Architecture spec:
docs/superpowers/specs/2026-06-08-ocr-anchor-first-structured-parsing-design.md- Implementation plan:
docs/superpowers/plans/2026-06-08-ocr-anchor-first-structured-parsing-plan.md- Role gate plan:
docs/superpowers/plans/2026-06-11-ocr-verified-structural-role-gate.md- Latest spec realignment:
docs/superpowers/specs/2026-06-13-ocr-real-paper-regression-and-spec-realignment-design.md2. Current Status
2.1 Test Suite
| Suite | Result | |-------|--------| | Figure stack (figures + reader + containment + backmatter boundary) | **286 passed** ✅ | | Document + roles + render + gate + rebuild + state machine + trace | **330 passed, 0 failed** ✅ | | Author bio detection (ocr_bio) | **37 passed** ✅ | | Spec contracts | All passed ✅ | | Real-paper regressions | All passed ✅ | | **Total OCR tests** | **1018 passed, 275 skipped (fixture unavailable), 0 failed** ✅ |
| Structural gate | Installed | | Role assignment (seed only) | ✅ | | Zone inference + fallback | ✅ | | Figure inventory | Global distance clustering — correct | | Backmatter boundary | ✅ Ref-anchored partition committed | | Pre-ref tail zone | ✅ Fixed | | Abstract detection | ✅ Fixed | | Frontmatter heading normalization | ✅ Fixed — no text matching | | non_body_insert clustering | ✅ Fixed — page-1 guard threshold | | Render (backmatter headings) | ✅ Correct heading rendering | | State machine | ✅ Accepts done_degraded as terminal | | Rebuild (span backfill skip) | ✅ Version match check fixed | | Author bio detection (Passes B+C) | ✅ P0: post-ref text-only bios as backmatter_body. P1: figure-residual portrait assets as author_bio_asset + figure_caption support |2.3 Fix Status
# Paper Issue Type Fix Commit 1 — P21 zone fix (4AG67PBH "Conflict of Interest" heading) Pipeline code infer_zones()frontmatter_side_blocks page gate: addedfirst_reference_page is not None and page >= first_reference_page - 135aabae2 4AG67PBH Author bio text in post-ref reference_item (b8/b10) New module post_ref_bio_cleanupreclassifies bios asbackmatter_body;_bio_text_scorecategory-weighted 0-5 scoringe2f0c8a3 — author_bio_asset role contract Pipeline code author_bio_assetadded to render_default=False, index_default=False skip sets7810eb14 — Pass C pipeline wiring Pipeline code Insert post_ref_bio_cleanup+prune_figure_inventory_after_bioafterwrite_back_figure_roles7810eb15 — P1 residual author bio pass (Pass B) New function residual_author_bio_passdetects portrait unmatched_assets/unresolved_clusters with nearby bio text7a1cc5e6 — P1 figure_caption support in Pass C Role expansion post_ref_bio_cleanupextended forfigure_captionrole7a1cc5e7 — tag_figure_contained_text author_bio guard Protection Skip author_bioblocks andauthor_bio_assetrole in figure containment7a1cc5e8 — Bio word limit 80→200 Pipeline code Real bio text 90-100 words exceeded 80-word limit in _bio_text_scoreae081a49 4AG67PBH Barbara bio as structured_insert_candidate missed by Pass C Role expansion Added structured_insert_candidateto Pass C role listae081a410 4AG67PBH Page 25 portrait id=5 missing from unmatched_assets Pipeline bug fix ocr_figures.py4479/4529 used page-agnostic bare block_id filter, hitting collision when same id exists on different pages. Changed to(page, block_id)tuples.ae081a411 4AG67PBH Acknowledgment text absorbed as reference_item (p21) Regex fix _is_reference_item_candidatetightenedfe9cc7013 WV2FF4NV Fig 10/6 locator caption bridge New feature _is_previous_page_legend_locator+ bridge inbuild_figure_inventory: connects locator → previous full legend → current visual group. Recovers misclassified legends from rejected_legends.3f61f4a14 — Issue 3: Backfill word leakage beyond bbox Word-level filter Added _word_belongs_to_block+_word_center_inside_rectfilters after expanded clip796e8bb15 — Issue 1A: Validation-first bare table skips same-page asset Continue guard Validation-first branch only early-exits when no same-page assets exist 796e8bb16 — Issue 1B: Split table caption continuation escapes ownership Continuation materialization _find_table_caption_continuation+_materialize_table_captioninsidebuild_table_inventory796e8bb17 — Issue 4: Short papers (≤2p) incorrectly red + needs_rebuild Health profile Added _health_profile(page_count)→short_formwaives abstract/heading gates796e8bb18 — Issue 2: Demoted-caption figure inner-text leakage Container bbox regions Validated _container_bboxregions intag_figure_contained_textvia 3 helpers + containment-only integration0e4ecbc19 — Issue 5: Cross-column page_assets groups falsely accepted Column-homogeneity gate _column_band_id+ rejection in_is_safe_page_assets_group4ab227e20 — Issue 6: Same asset consumed by figure AND table Post-hoc arbitration resolve_media_asset_conflictsresolves asymmetric cases; weak/weak stays inownership_conflicts4ab227eP2#1a — Previous-page legend locator bridge (✅ Fixed
3f61f4a)WV2FF4NV Fig 10: locator "See legend on previous page" on p16 not bridged to full legend (in rejected_legends, misclassified as body_paragraph)
Remaining P2
- Figure containment gaps —
cluster_bboxcontainment only runs on matched figures; composite figures with demoted caption upstream never enter containment. Inner-text detection missesvision_footnote/paragraph_titleraw_labels.- Short "Table N" caption matching — bare
"Table 1."labels miss table_asset in ownership pipeline. Existing code partially handles it; residual cases remain.- Page-1 body paragraph OCR text extraction failure (was: width check) — W33NLVJ2 right col Introduction block: OCR detected text region (raw_label=text) but extracted empty string. Pipeline previously mis-signaled as
unknown_structural. Fixed: nowocr_text_missing.- Body text backfill overlap —
backfill_missing_text_from_pdfusesget_text("words", clip=expanded)returning words beyond block's bbox; can cause text duplication at render level.- 2HEUD5P9 reference ordering — ✅ Fixed by F12 (raw_label=reference_content signal)
- Mixed-column tail page zone absorption — ✅ Fixed: tighten year-period regex
P3 (Boundary / Edge Cases)
- DW biography page mismatch — pages 32-34 vs expectations 33-34.
- AJR side-caption recovery — deliberately deferred; not part of generic inventory refactor.
- Chinese Windows encoding — non-ASCII PDF filenames garbled via GBK codepage. Glob fallback added; residual
meta.jsoncorruption from earlier runs.backfill_missing_text_from_pdfbbox-exact filtering — need decision on render-level dedup vs backfill-level bbox clamp.- Multi-caption page column collapse —
page_assetsgroup on multi-column pages merges separate figures in different columns.- Figure/table shared-consumption registry — no shared consumed registry for ambiguous image-like blocks.
- Short papers (<3 pages) health falsely red — no headings/abstract in Letter/Editor formats.
4. Active Queue
- ✅ P1 backmatter boundary (ref-anchored partition)
- ✅ Pre-ref tail zone fix (4KCHGV2Z)
- ✅ Gate 5 frontmatter fix series (24YKLTHQ)
- ✅ Stale trace-vs-expectation fixtures cleared (10 assertion updates)
- ✅ All stale test expectations reconciled (non_body_insert, caffard, legend_like, structural gate, render, state machine, rebuild, truth docs)
- ✅ P0 author bio detection (post_ref_bio_cleanup for reference_item)
- ✅ P1 author bio detection (residual_author_bio_pass + figure_caption support + tag_figure_contained_text protection)
- NEXT: Run lint (ruff) then merge ocr-v2 → master
4.1 Immediate Next Steps
- 25 fixed tests (10 distinct issues)
- Full OCR regression sweep: 1018 passed, 0 failed
- P0 bio detection: 30 new tests, 0 regressions
- P1 bio detection: 7 new tests, 0 regressions
- tag_figure_contained_text author_bio protection
- Run lint (ruff)
- Merge
ocr-v2intomaster
5. Key File Map
5.1 Production Code (OCR Pipeline)
File Role paperforge/worker/ocr_blocks.pyRaw + structured block generation; preserves seed_role only paperforge/worker/ocr_roles.pyassign_block_role()— seed/proposal logic only (NOT final)paperforge/worker/ocr_document.pynormalize_document_structure()— anchor/family/zone + gate + final rolespaperforge/worker/ocr_structural_gate.pyVERIFY_REQUIREDrole decision + abstract span + reference zone + healthpaperforge/worker/ocr_orchestrator.pyBody reorder, column validation, layered assembly paperforge/worker/ocr_render.pyrender_fulltext_markdown()— consumes verified artifacts ONLYpaperforge/worker/ocr_health.pyHealth reporting, merged gate summary paperforge/worker/ocr_profiles.pySpan extraction, profile aggregation, cross-validation paperforge/worker/ocr_figures.pyFigure reader pipeline paperforge/worker/ocr_tables.pyTable inventory matching paperforge/worker/ocr_scores.pyScore functions (spatial, structured_insert, etc.) paperforge/worker/ocr_rebuild.pyDerived rebuild entry point paperforge/worker/ocr_pdf_spans.pyPDF span backfill for OCR-missed blocks paperforge/worker/ocr_bio.pyAuthor biography detection utilities and passes 5.2 Test Files
File Purpose tests/test_ocr_figures.pyFigure inventory, matching, ownership tests/test_ocr_document.pyDocument structure, normalize, gate tests/test_ocr_render.pyFulltext render contract tests/test_ocr_figure_reader.pyReader contract tests/test_ocr_trace_vs_expectations.pyReal-paper trace vs expectations gap report (8 gold papers) tests/test_ocr_real_paper_regressions.pyPage-level + document-level regression on real papers tests/test_ocr_real_paper_audit_contracts.pyGold-fixture quality gate tests/test_ocr_spec_contracts.pyArchitecture contract tests tests/test_ocr_structural_gate.pyGate unit tests tests/test_ocr_v2_structural_regressions.pyStructural regression guards tests/test_ocr_layout_first_regressions.pyLayout-first behavior guards 5.3 Test Fixtures
tests/fixtures/ocr_real_papers/{CAQNW9Q2,DWQQK2YB,TSCKAVIS,A8E7SRVS,K7R8PEKW,6FGDBFQN,SAN9AYVR,2GN9LMCW}/5.4 Design Documents
File Purpose docs/superpowers/specs/2026-06-08-ocr-anchor-first-structured-parsing-design.mdArchitecture spec docs/superpowers/specs/2026-06-13-ocr-real-paper-regression-and-spec-realignment-design.mdCurrent phase spec docs/superpowers/specs/2026-06-10-ocr-figure-reader-contract-design.mdFigure reader spec docs/superpowers/specs/2026-06-23-ocr-visual-grammar-hardening-design.mdVisual grammar hardening docs/superpowers/specs/2026-06-27-figure-containment-and-backmatter-boundary-design.mdCurrent P0/P1 spec docs/superpowers/plans/2026-06-18-ocr-v2-readiness-master-plan.mdReadiness gates master plan docs/superpowers/plans/2026-06-17-ocr-v2-closeout-single-plan.mdClose-out single plan docs/superpowers/plans/2026-06-15-group-first-figure-inventory-plan.mdGroup-first figure refactor (deferred) docs/superpowers/specs/README-ocr.mdOCR design index 5.5 Active Truth Files
File Role project/current/ocr-v2-active-queue.mdActive queue — next-work authorit project/current/ocr_rebuild_audit.mdEvidence source for queue project/current/ocr-v2-generalization-boundary.mdArchitecture boundary note project/current/ocr-v2-remaining-issues-2026-06-18.mdHistorical readiness residuals
6. Decision Log
Date Decision Rationale 2026-06-08 Adopt anchor-first OCR architecture Early role guessing caused cascading errors; needed structural discovery first 2026-06-10 Separate figure reader from body prose Figure info must be reader-visible without body pollution 2026-06-11 Install verified structural role gate Spec was being bypassed; gate enforces seed≠final contract 2026-06-13 Dual-gate regression + spec-contract testing Tests protected helper behavior, not real-paper outcomes 2026-06-13 Root-cause approach: no renderer patches Heading merge goes in raw blocks; boundary detection fixed in infer_zones 2026-06-13 Figure sequential matching as cross-page tradeoff Caption-asset pairs on different pages get lower confidence 2026-06-15 Expand deterministic gold set to 8 papers Needed broader regression surface before changing figure inventory 2026-06-15 Do not solve AJR side-caption recovery in group-first refactor Keep scope generic; AJR-specific rescue is later phase 2026-06-15 Group-first matching is next architectural target Existing clusters/visual groups too late in pipeline 2026-06-17 Single-thread close-out note authoritative queue moved to project/current/ocr-v2-closeout-priority.md2026-06-21 Replace greedy region-growth with global distance clustering Human sees assets as perceptual groups, not competing candidates 2026-06-23 Pre-ref=body flow, post-ref=backmatter CRediT/Ethics above References → body_zone 2026-06-23 Figure containment: render-hygiene pass build_figure_inventory Containment shouldn't affect matching, only rendering 2026-06-23 Reference sort: two capture groups Prevents false matches on plain year numbers 2026-06-26 P0 before P1 Fix 2/4/5 are <30 lines fully understood; Fix 1/3 need new specs 2026-06-28 Pre-ref tail zone: strip from region_bus not re-apply _apply_zone_labelsre-applied stale tail zone frominfer_zones()after ref partition. Fix: strip pre_ref block IDs from region_bus before zone re-apply.2026-06-28 Author byline: require lowercase letters _looks_like_initial_lastname_bylinematched all-caps journal taglines. Fix: require any lowercase letter in matched text.2026-06-28 Page-1 body_start: metadata headings should not trigger _is_first_page_body_starttreated ANY section_heading as body start. Fix: only real body section headings (introduction, methods, etc.) trigger body_start on page 1.2026-06-28 Frontmatter heading normalization: no text matching Metadata sidebar labels rejected by structural gate fell to unknown_structural. Fix: normalize held heading blocks in frontmatter_main_zone to frontmatter_noise using only zone + gate decision + seed_role, no text matching. 2026-06-28 Author bio detection: three-pass cascade, P0 first Strong structure first, residual explanation second. Real figures must never be preempted. P0: post-ref text-only. P1: figure residual. P2+: P1 profile card pre-pass. 2026-06-28 Category-weighted bio scoring career=+3, education=+2, research=+2, institution=+1, publication=+1. Returns (score, categories) tuple. Threshold: score ≥ 4 AND categories ≥ 2. 2026-06-28 author_bio_asset role: non-rendered, non-indexed Bio artifacts removed from figure_inventory entirely, never returned to unmatched_assets. Clean prune before reader. 2026-07-01 Asset-internal figure number recovery: metadata-only pass Recovery must NOT split OCR blocks or mutate chart text — only patches figure_number, figure_id, recovered_label_text on existing matched figures. Coordinate normalization is caller responsibility. 2026-07-01 Broaden recovery gate to handle normal prematch unknown figures Synthetic-figure gate ( bbox_only_assetflag) excludedfigure_unknown_NNNfrom normal rotated prematch path. Gate now allows figure_unknown figures without synthetic flags.
7. Agent Instructions
7.1 Project Folder Management
The authoritative prompt for project record management lives in
.omp/AGENTS.md. It is auto-loaded by omp into every session context and defines when and how to updatePROJECT-MANAGEMENT.md,project/current/*, andproject/archive/.TL;DR: PROJECT-MANAGEMENT.md updated every session end. project/current/ updated at milestones only. project/archive/ gets moved-to (not deleted) when stale.
7.2 How to Update PROJECT-MANAGEMENT.md
- Update section 2 (Current Status) — test counts, component status, fix status
- Update section 3 (Remaining Issues) — remove resolved items, add new ones
- Update section 4 (Active Queue) — check/adjust next steps
- Add new decisions to section 6 (Decision Log)
- Add a compressed entry to section 8 (Session Timeline)
- Update "Last Updated" date at top
7.3 Before Starting Any OCR-v2 Work
- Read section 4 (Active Queue) for current priority
- Read the relevant design doc from section 5.4
- Understand what the tests currently expect
- Work one repair at a time; verify before moving on
7.4 Test Commands
# Full OCR test suite python -m pytest tests/test_ocr_*.py -v --tb=short # Figure stack only python -m pytest tests/test_ocr_figures.py tests/test_ocr_figure_reader.py tests/test_ocr_render.py -q # Real-paper regression only python -m pytest tests/test_ocr_real_paper_regressions.py -v --tb=short # Spec-contract tests only python -m pytest tests/test_ocr_spec_contracts.py -v --tb=short # Structural gate only python -m pytest tests/test_ocr_structural_gate.py tests/test_ocr_document.py -v --tb=short # Rebuild a single paper python scripts/dev/ocr_rebuild_paper.py DWQQK2YB # Rebuild + regenerate block_trace python scripts/dev/ocr_rebuild_paper.py --trace DWQQK2YB # Lint python -m ruff check paperforge/worker/ocr_*.py7.5 Design Doc Reading Order
docs/superpowers/specs/README-ocr.md— indexdocs/superpowers/specs/2026-06-08-ocr-anchor-first-structured-parsing-design.md— architecturedocs/superpowers/specs/2026-06-13-ocr-real-paper-regression-and-spec-realignment-design.md— current phase
8. Session Timeline (Compressed)
Date Session Key Results Detailed Archive 2026-06-16 Gold fixture expansion + 10 pipeline fixes 98 bug annotations, 8 papers audited, 8 fixes applied (F1-F10) §9.1 2026-06-17 OCR-v2 boundary close-out pass Zone-boundary fixes, tail/backmatter shrink, correspondence routing. 202P/1F/43S §9.2 2026-06-18 Readiness-gates implementation Gates 1-4 complete + blind audit protocol. 249/249 figure/health/document pass §9.3 2026-06-19 Blind audit (5 unseen papers) All PASS — no new failure families. OCR-v2 declared "state healthy" §9.3 2026-06-21~22 Figure merge: greedy → global distance clustering Union-find clustering, 261 tests. Caption-independent semantic grouping, ownership registry, local pairing, conflict detection §9.5 2026-06-22 Cross-page caption consumption fix Reader/render contract breach fixed — cross-page matches now consume caption on legend page §9.4 2026-06-23 Visual grammar hardening Composite parent detection, dense page arbitration, figure/table separator veto, dedup refinement §9.5 2026-06-26 Rebuild production run (699 papers) --resumecheckpoint, glob fallback, N6XCZD25 body text fix§9.6 2026-06-26 Figure number inference + container admission Leading [1]gap filled, blue sidebar box rendered as[!NOTE]§9.7 2026-06-26 UI polish (plugin) Dashboard CSS, maintenance tab redesign, Vercel-style polish §9.8 2026-06-27 Rebuild audit + index repair 6 hard-block/high bugs fixed (sync, workspace, field registry) §9.9 2026-06-27 Deep investigation — 5-fix spec P0 all committed (ref sort, caption insert, figure containment). P1 backmatter in progress §9.10 2026-06-28 P1 backmatter boundary committed Ref-anchored partition ( 3e33e5b). Pre-ref=body flow confirmed (9b72783). 16/16 tests pass. All 5 audit papers verified.§9.11 2026-06-28 Gate 5 blind audit + pre-ref tail zone fix Gate 5: 24YKLTHQ (13p) + 4KCHGV2Z (9p) rebuilt post-P1. Found pre-ref body pages misclassified as tail_nonref_hold_zone. Root cause: _apply_zone_labels re-applies stale region_bus after ref partition. Fixed by stripping pre_ref block IDs from tail zone. 4KCHGV2Z P7: tail=20 → body=2+disp=5. All 286 figure/backmatter tests green. §9.12 2026-06-28 Gate 5 frontmatter fix series (3 fixes) 24YKLTHQ: author byline lowercase guard, metadata body_start fix, frontmatter heading normalization (no text matching). All 3 fixes verified on real paper, 461 tests pass, 0 new regressions. §9.13 2026-06-28 Test fix session: 25 tests reconciled Fixed 10 stale test issues: non_body_insert guard, caffard abstract, legend_like role, structural gate, backmatter heading render, state machine (done_degraded), body_zone anchor, rebuild backfill skip, truth surface docs, trace-vs-expectations (10 assertions). 616 OCR tests, 0 failed. Expectations updated for post-P1 behavior. §9.14 2026-06-28 Data-driven truth audit (2 papers) 2HEUD5P9 (27p) + 4AG67PBH (25p) — no vision (model limit). Found 3 pipeline defect patterns: zone_leak_frontmatter_to_body (2 papers), reference_boundary_body_mix (2 papers), title_repeat_page2 (1 paper). 12 ghost unknown_structural blocks in 2HEUD5P9. Findings saved to audit/2026-06-28-data-audit-findings.json. §9.15 2026-06-28 P0 author bio detection implementation Created ocr_bio.py with category-weighted bio scoring, Pass C (post_ref_bio_cleanup), figure match guards. Wired author_bio_asset role contract + pipeline. 30 new tests pass. 1041 total OCR tests, 0 regressions. Commits: e2f0c8a,7810eb1.§9.16 2026-06-28 P1 author bio detection implementation Added residual_author_bio_pass (figure-residual portrait assets), extended post_ref_bio_cleanup for figure_caption, tag_figure_contained_text protection. 7 new P1 tests. 1018 total OCR tests, 0 regressions. Commit: 7a1cc5e.§9.17 2026-07-01 Audit fix commits + orientation-aware rotated figure normalization Commit 1 ( 2d40ad9) + Commit 2 (21bdfd0) + Commit 4 (7670227) landed, then refactored rotated-caption handling out of synthetic fallback into normal figure pre-match. Added PyMuPDFdir/wmodecapture, same-page rotated settlement, rotated crop render. U746UJ7G now matches viasame_page_rotated; KUR9PBJC unchanged. 422 regression tests pass.§9.18 2026-07-01 Asset-internal figure number recovery implementation Added extract_pdf_lines_normalizedhelper,_recover_missing_figure_numbers_from_assetspass inbuild_figure_inventory, 5 gate functions, 2 pattern constants. U746UJ7Gfigure_unknown_000→figure_002with recovered label "Plot of Criteria Time". 6 new tests. 428 regression tests pass.§9.19 2026-07-02 Round 2 truth audit + 3 targeted bug fixes for 37LK5T97 Batch-audited 10 fresh papers (5 GREEN / 4 YELLOW / 1 RED). 37LK5T97 found with Figure 1 broken (sidecar caption demoted) + 6 unmatched rotated tables. Fixes: (1) _is_sidecar_candidateguard in candidate_resolution, (2)adjacent_x+y_overlapin score_table_match for rotated captions, (3) rotated table render bbox+270° correction. Also: rotated figure crop quality fix (4x zoom + coordinate normalization in_crop_asset_from_pdf). Commits:59cd01a,bd3f3b6,86e0d14. 428 regression tests pass.§9.20 2026-07-02 Zone/role robustness completion — Figure caption prefix recovery + inline table fix + table matching audit Figure caption prefix recovery from PDF text layer ( _recover_figure_heading_prefix): 5S7UI34M 4→9 matched figures, 33→1 unmatched. Inline<table>HTML role fix: 650 blocks nowtable_htmlafter rebuild (priority bug: raw_label=table fired before<table>check). Table matching audit: 620 remainingmedia_assetpending full rebuild. 585 figure/table/role tests pass.§9.21 9. Historical Detail Archive
Fixed records and verbose session logs preserved below. The archive is read-only reference — active work is tracked in sections 2-4 above.
9.1 Gold Fixture Expansion + Bug Fixes (2026-06-16)
Full-day debugging session across 8 gold papers. 98 bug annotations, 8 pipeline fixes (F1-F10).
Root cause categories identified:
- Frontmatter noise unrecognized (ISSN, journal citation) → body pollution
- Cross-page text fragmentation
- Tail zone: body prose incorrectly converted to backmatter
- Backmatter heading gaps
- Heading merge/split logic
- Heading rescue from unknown_structural
- Heading prefix by role (not font size)
- Permissive figure matching
- All same-page assets included in group match
- Composite region text requirement removed
Fixes applied:
ocr_roles.py,ocr_document.py,ocr_blocks.py,ocr_render.py,ocr_scores.py,ocr_figures.py,ocr.pyLayout-first Phase 1 pass added — table inventory relaxed for
media_assetblocks,_should_keep_formal_caption_seed()added,tests/test_ocr_layout_first_regressions.pycreated.
9.2 Boundary Close-Out Pass (2026-06-17)
Plan:
docs/superpowers/plans/2026-06-17-ocr-v2-closeout-single-plan.mdTasks 1-5 executed:
- Same-page reference/body boundary split (block-level vertical split by heading position)
- False tail backmatter conversion reduction
- Page-1 correspondence line → frontmatter_support
- Preproof page-1 frontmatter: first-surviving-page logic, margin-band watermark detection, figure inner label extension
- Active truth-file cleanup: reconciled P0-P2 close-out
Result: 202P/1F/43S (sole failure = pre-existing DW figure ownership)
Key commits:
6f68bf2,b7d369e,c7a9c93,827a2cc,9329843
9.3 Readiness-Gates (2026-06-18 ~ 06-19)
Plan:
docs/superpowers/plans/2026-06-18-ocr-v2-readiness-master-plan.mdGates implemented:
- Gate 1 (Completeness):
_summarize_page_text_coverage()+_classify_region_text_completeness()+audit_rendered_text_coverage()- Gate 2 (Figure ownership): 8 tasks — previous-page sequential fallback, DW Fig 3 xfail→pass, sidecar partition, fallback tightening. Final: 249/249 pass
- Gate 3 (Ordering):
_ref_number_sort_keyregex extended, reference boundary normalizers- Gate 4 (Layout coverage):
audit/coverage_ledger.jsonwith readiness-class taxonomy, contract tests enforce named representativesBlind audit (5 unseen papers): ALL PASS. Papers: 8VB9ZVQG, U746UJ7G, L6ALWJFP, PZ8B59K4, GU9R8EPE. No new failure families discovered.
Remaining residuals: ~40 stale audit truth blocks, ~50 edge-case misclassifications (low severity).
9.4 Cross-Page Caption Consumption Fix (2026-06-22)
Paper: SAN9AYVR Figure 24 double-emit. Cross-page figure ownership repair — reader/render chain consumed caption blocks on asset page instead legend page.
Fixes applied:
- Reader payload records cross-page caption consumption on legend page
- Render path suppresses original caption block on asset page
_recompute_final_unmatched_assets()added — orphan truth now matches final ownership truth- Cross-page duplicate
![[render/figures/figure_N.md]]eliminatedTest status: 150 passed (figures + reader + render)
9.5 Figure Merge Refactor + Visual Grammar Hardening (2026-06-21 ~ 06-23)
Core change: Replaced greedy region-growth with global distance clustering (union-find):
_cluster_page_assets(): horizontal <12% pw, vertical <8% ph, text-separator-aware- Caption-as-boundary: each legend claims assets by y-band
- Composite parent detection:
_build_composite_parent_figure_groups_visual_only()- Dense page arbitration:
_build_dense_composite_parent_candidates()when ≥4 visual fragments- Ownership registry:
FigureOwnershipRegistrywith conflict detection- Figure/table separation:
asset_family_hint+ table-like veto (confidence ≥0.70)- Dedup refinement:
_normalized_caption_body— internal punctuation preserved, terminal punctuation stripped_ref_number_sort_keywith[N]bracket supportKey fix (merge-gate closeout): Same-number-distinct dedup + grid collapse fix (highest-score selection, not first-match)
Test status: 216→225→261 passed (figure stack)
9.6 Rebuild Production Run (2026-06-26)
Full
ocr rebuild --all(699 papers):
--resumecheckpoint support (interrupted rebuilds resume)- ASCII tqdm progress bar fix (Windows)
- Removed rogue
[DEBUG]print inocr_render.py:1295Fixes found:
- N6XCZD25: Body paragraphs misclassified
structured_insert→body_spine_matchflag inbody_zone- Chinese Windows encoding: glob fallback in
resolve_pdf_path; garbledmeta.jsonauto-corrected on rebuild
9.7 Figure Number Inference + Container Admission (2026-06-26)
Figure number inference (leading
[1]gap): 8-step algorithm in_infer_missing_main_figure_numbers(). N6XCZD25figure_unknown_005→figure_001.Container admission rewrite (evidence-driven):
- 7-method container extraction: page-sized/crop-like excluded, line-like→grouping only, vertical component merge
- Three-phase per-page loop replaces lazy-cache
- Blue sidebar box score 0.45→0.80 → rendered as
[!NOTE]calloutKey commits:
45cf65e,bab0167,5ba1a6d
9.8 UI Polish (2026-06-26)
Plugin dashboard cleanup:
- Vercel-style CSS (cards, collapsible headers, status grid, issue summary funnel)
- Title tooltip with
title_fullfield- Table sorted by time descending
- Click-to-copy on paper paths (pf-copy)
- Global "Start Working" cleanup — Doctor/Repair only in Issues section
- Run OCR button shows pending count
- Redo OCR → Maintenance button (opens settings→maintenance tab)
- DESIGN.md reference file added
9.9 Rebuild Audit + Index Repair (2026-06-27)
6 bugs fixed:
Issue Severity Status sync_service.py:_timeUnboundLocalErrorHARD BLOCK FIXED Workspace fulltext never synced from OCR output HIGH FIXED build_indexearly-return skips_build_entryHIGH FIXED Field registry missing 9 common fields → 6000+ WARN MEDIUM FIXED "700 papers missing fulltext" — false alarm MISLEADING NOTED Base legacy fields — user choice INFO DEFERRED Field registry additions:
aliases,tags,journal,first_author,pmid,impact_factor,abstract,keywords,ocr_time.
9.10 Deep Investigation — 5 Fix Spec (2026-06-27)
5-paper vision audit: 25K5KZAQ, NC66N4Q3, 9TW98JH8, YGH7VEX6, XD2BPCMG.
Fixes identified and resolved:
Fix Issue Root Cause Resolution Fix 1 Figure-internal text containment No spatial containment in build_figure_inventoryRender-hygiene pass: 6 helpers + 19 tests ✅ Fix 2 Reference sorting [N]bracket gapregex only matches N./N), not[N]Two capture groups `r"^\s*(?:\d+ Fix 3 Backmatter boundary (CRediT/Ethics) 7.97pt < 11pt threshold; no container keywords Ref-anchored partition design 🔄 Fix 4 Figure caption → non_body_insertfigure_captionin_INSERT_CANDIDATE_ROLESRemoved list ✅ Fix 5 Demoted body paragraph in figure legends body_paragraphre-enters matching via legend detectionFilter before _is_validation_first_legend_candidate✅Spec:
docs/superpowers/specs/2026-06-27-figure-containment-and-backmatter-boundary-design.mdPlan:docs/superpowers/plans/2026-06-27-figure-containment-implementation-plan.md9.18 10-Paper Truth Audit + GPT Cross-Validation (2026-07-01)
Scope: 10 papers sampled from vault (3 green, 4 yellow, 3 red), batch-audited via
ocr_truth_audit.py+ annotated page vision agents.
1 complete vision audit: U746UJ7G — 0 matched figures, root cause = vector-rendered Figure 1+2 (52 drawing paths, 0 embedded images).
Key bugs found:
_TABLE_PREFIX_PATTERN3 处只认\d+,不认罗马数字 I/II/III(KUR9PBJC pages 4/5/7)raw_label=figure_title的 Table N caption 无 guard → 进figure_captionvision_footnote里 "This figure..." 被吞成 footnote(U746UJ7G p8:4)unmatched_assets重复计数(matched 后又在 unmatched 列表)- HTML table block 断连(2YW2MJBL)
- Vector figure pipeline 盲区(U746UJ7G)
False positives:
reference_span_error~90% noise,same_page_boundary_error100% unreliable.GPT 二审修正 8 处: Roman+S prefix 三处同步、score hard gate、tie-breaking、dedup、
_is_near_figure_mediahelper。
Outputs:
- Spec:
docs/superpowers/specs/2026-07-01-ocr-audit-findings-for-gpt.md- Plan:
docs/superpowers/plans/2026-07-01-ocr-audit-gpt-fix-plan.md- Per-paper vision reports:
local://*-vision-report.mdExecution: 3/4 commits implemented (2d40ad9,21bdfd0,7670227). Commit 3 (unmatched dedup) pre-existing from PR3 (4ab227e). Follow-up refactor moved rotated-caption handling earlier: PyMuPDFdir/wmodenow preserved inspan_metadata; rotatedvision_footnotecandidates enter normal legend matching;same_page_rotatedmatches carryrotation_correction_deginto crop/render. U746UJ7G now produces a normal matched figure (figure_unknown_000,settlement_type=same_page_rotated) instead of synthetic fallback. 422 regression tests pass.- Block reviews:
audit/<KEY>/block_review.jsonl(U746UJ7G, KUR9PBJC, CGGYTEEQ, 7FNV9AW2 etc.)9.19 Asset-Internal Figure Number Recovery (2026-07-01)
Problem: U746UJ7G rotated figure — "Figure 2. Plot of Criteria Time..." was absorbed into chart asset bbox during OCR assembly. The rotated prematch refactor (9.18) produced a matched figure (
figure_unknown_000) but withoutfigure_number. No way to assign Figure 2 to the figure.Solution: Metadata-only recovery pass after synthetic vector fallback, before dedup.
extract_pdf_lines_normalizedinocr_pdf_spans.pyextracts PDF rawdict text lines page-by-page, normalizing coordinates to OCR space_recover_missing_figure_numbers_from_assetsiterates matched_figures without numbers, scans each matched asset's bbox for internal PDF line labels ("Figure N.", "Fig. N:")_needs_asset_internal_figure_number_recovery: gates by figure_id prefix (figure_unknown_*orsynthetic_figure_*) + text description signal_looks_like_internal_figure_label: rejects full-sentence patterns ("Figure X shows...") via regex match_asset_edge_band_score: geometric rejection for lines in asset center (not label position) or covering >15% of asset area- Coordinate normalization is caller responsibility (done in ocr_rebuild.py)
Files changed:
paperforge/worker/ocr_figures.py: 7 new functions + 2 pattern constants + signature change tobuild_figure_inventory+ recovery pass callpaperforge/worker/ocr_pdf_spans.py:extract_pdf_lines_normalizedhelperpaperforge/worker/ocr_rebuild.py: call helper, passpage_pdf_lines_by_pageto inventory buildertests/test_ocr_figures.py: 6 new tests (basic recovery, duplicate rejection, normal-fig unaffected, multi-label conflict, center rejection, overlap gate)Spec:
docs/superpowers/plans/2026-07-01-ocr-asset-internal-figure-number-recovery-plan.mdExecution:
- U746UJ7G verified:
figure_number == 2,figure_id == "figure_002",recovered_label_textcontains "Plot of Criteria Time", flags containfigure_number_recovered_from_asset_text- 428 regression tests pass (422 existing + 6 new)
9.20 Round 2 Truth Audit + 37LK5T97 Bug Fixes (2026-07-02)
Scope: 10 new papers batch-audited via
ocr_truth_audit.py(high-risk mode). 5 GREEN / 4 YELLOW / 1 RED.
RED paper: 37LK5T97 — "Both IM and EC ossification occurs during the bone-healing process"Three bugs found and fixed:
Figure 1 sidecar demotion: Caption (left column, 246px) and image (right column, 693px) had zero x_overlap.
_is_near_figure_media()missed it,_looks_like_figure_narrative_prose()caught the long description, and candidate_resolution demoted tobody_paragraph. Fix: Added_is_sidecar_candidate()guard — checks vertical overlap with media_asset when horizontal overlap is absent. Caption preserved asfigure_caption_candidate.
- File:
paperforge/worker/ocr_document.py- Result: Figure 1 matched as
figure_001(caption block_id=2, asset block_id=9)Rotated table caption matching: Tables 1-3 had rotated captions (span dir=[0,-1], vertical text beside table body).
score_table_matchrequired x_overlap + asset_below_caption, both failed for rotated sidecar layout. Fix: Addedadjacent_x+y_overlap_with_assetscoring branch when caption has rotated text and x_overlap < 0.5.
- File:
paperforge/worker/ocr_scores.py- Result: All 6 tables matched (has_asset=true, score >= 0.65)
Rotated table render orientation: Table body also had dir=[0,-1] (rotated 90° content on portrait page). Rendered JPEG showed vertical text. Fix: Added
_table_has_rotated_content()helper; computes unionrender_bbox(caption+asset) andrender_rotation_deg=270in table entry; render loop passesrotation_degto_crop_asset_from_pdf.
- Files:
paperforge/worker/ocr_tables.py,paperforge/worker/ocr_objects.py- Result: Tables 1-5 rendered at correct orientation (1908×2858 → 2858×1908)
Additional quality fix: Rotated figure crop quality improved in prior session —
Matrix(2,2)→Matrix(4,4), PIL rotation from pix.tobytes("png") (single pass, no double JPEG), and OCR→PDF coordinate conversion fix. U746UJ7G figure_002: 2618×1914 px, 5.0M px (4.5x improvement).Commits:
59cd01a(figure quality+rotation coord fix),bd3f3b6(sidecar+table match),86e0d14(table render rotation)
Tests: 428 regression tests pass (no new tests added, existing cover the affected paths)
Analysis:docs/superpowers/analysis/2026-07-02-37lk5t97-figure1-sidecar-bug-analysis.md
Audit findings:docs/superpowers/specs/2026-07-02-ocr-truth-audit-round2-findings.md9.21 Figure Caption Prefix Recovery + Inline Table Fix (2026-07-02)
Problem 1: PaddleOCR fails to detect standalone "Figure N" / "FIGURE N" headings in bold/small-caps fonts. Caption body is captured as
figure_caption_candidatebut lacks the figure number prefix →_is_formal_legend()fails → caption never enters matching pool.Fix:
_recover_figure_heading_prefix()inocr_figures.pychecks the PDF text layer (via existingpage_pdf_lines_by_pageinfrastructure — no extra PDF open) for "Figure N" lines. If the next PDF line (by y-order) shares ≥15 common-prefix chars with the OCR caption text, the heading is prepended. Runs BEFORE the zone/style filter so recovered captions pass through to legend matching.Result: 5S7UI34M (PVA综述): 4→9 matched figures, 33→1 unmatched (p1 logo). HQAQBSBP: Figure 5 recovered. 372 figure tests pass.
Problem 2: Inline
<table>HTML blocks withraw_label=tablehit theraw_label="table" → media_assetfallback (ocr_roles.py:1355) before reaching the<table>→table_htmlcheck (line 1456, gated byraw_label="text"). Plusocr_document.py:6120-6121convertedtable_htmltotable_html_candidate— a role with no downstream handler → structural gate downgraded tounknown_structural.Fix: (1) Moved inline
<table>check beforeraw_label=tablefallback inassign_block_role(). (2) Addedtable_htmlverifier in structural gate (self-identifying, accepts if text starts with<table>). (3) Removed deadtable_html → table_html_candidateconversion.Result: AH6Q7DLC (worst case, 30 blocks): 29/30 →
table_html(1 = reference_item). 585 figure/table/role tests pass.Files changed:
paperforge/worker/ocr_figures.py—_recover_figure_heading_prefix()+ body_zone filter guardpaperforge/worker/ocr_roles.py— inline<table>before raw_label=tablepaperforge/worker/ocr_document.py— removedtable_html_candidatedead pathpaperforge/worker/ocr_structural_gate.py—table_htmlverifier