docs: update PROJECT-MANAGEMENT with 2HEUD5P9 rebuild findings

This commit is contained in:
Research Assistant 2026-06-22 02:50:37 +08:00
parent db69961c7c
commit 6522dc9aef
5 changed files with 625 additions and 11 deletions

View file

@ -1678,5 +1678,25 @@ Final plan at `docs/superpowers/plans/2026-06-22-cover-page-detection-and-block-
| 2026-06-22 | Cover detection = positive markers, not "no body" | Normal papers may have abstract on page 1, body on page 2 |
| 2026-06-22 | footnote→callout needs positive content markers | Only affiliation/correspondence, not all small-font footnotes |
| 2026-06-22 | Task 1 and Task 2: separate commits | Different risk profiles |
| 2026-06-22 | footnote→callout font threshold: `fn_fs < body_font_median` | `0.9 * median` had 0.014pt margin miss; any smaller-than-body is callout-worthy |
### 12.12 Issues Found During 2HEUD5P9 Rebuild (2026-06-22)
1. **P1:8 eaten** — Block 8 on page 1 (right column body paragraph continuation, `seed_role=body_paragraph`) was misclassified as `non_body_insert`. Root cause: `_detect_non_body_insert_clusters` first-pass page-1 guard uses `"-" in text` (any hyphen anywhere, e.g. "bone-related"), not `startswith`. Fix in `ocr_document.py`: change `any(m in text.lower() for m in ["\u2022", "-", "*"])` to `text.startswith(("\u2022", "- ", "* "))` AND add position+style check (font matches body spine + width > 20% page) instead of naive length check.
2. **P11:5,6 MOF paragraph duplication** — Block 5 (left column, OCR text=full 1637-char paragraph) + Block 6 (right column, `text_source="pdf_text_layer_fallback"` with 1010-char fragment) both render the same overlapping content. Root cause: `backfill_missing_text_from_pdf` uses `get_text("words", clip=expanded)` which returns ANY word with bbox overlapping the clip rect, even if the word text belongs to a larger paragraph that extends beyond the block's bbox. Fix needed: filter backfilled words to only those strictly within the block's original bbox, OR handle dedup at render level.
3. **Callout font threshold too tight**`0.9 * body_font_median` missed footnote bid=9 (`fn_fs=7.97 < 8.84*0.9=7.956` by 0.014). **RESOLVED**: changed to `fn_fs < body_font_median`.
### 12.13 Parked Hard Cases
- `backfill_missing_text_from_pdf` bbox-exact word filtering: need to decide between render-level dedup vs backfill-level bbox clamp.
- Two-column page 1 body paragraph width check: right-column body is naturally narrower than left-column spine median; `is_narrow` heuristic misfires for multi-column page 1 layouts.
### 12.14 Commits
| Commit | Message |
|--------|---------|
| `cc8a0f9` | `feat: convert body-zone footnotes with distinct font to callout` (Task 2) |
| `4f215c9` | `feat: drop non-preproof cover page 1 via positive cover markers` (Task 1) |

View file

@ -0,0 +1,560 @@
# Weak Cluster Rejection for Column Detection — Implementation Plan
> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
**Goal:** Fix false two_column classification when a single short offset body line creates a spurious second x-center cluster. Add cluster metadata with weak-cluster rejection before layout type decision.
**Architecture:** Add `_cluster_page_column_groups()` returning cluster metadata dicts (center, count, y_coverage, word_count, blocks, etc.). Keep `_cluster_page_columns()` as a compatibility wrapper (body replaced with `return [c["center"] for c in _cluster_page_column_groups(...)]`). Add `_is_weak_isolated_column_cluster()` that rejects clusters with count<3 AND y_coverage<10% page height AND word_count<25. Modify `_classify_page_layout()` to filter weak clusters BEFORE deciding `column_count`, and use cluster `blocks` for role distribution instead of `page_width/2` midpoint.
**Tech Stack:** Python, no new dependencies.
---
### Task 1: Add `_cluster_page_column_groups()` function
**Files:**
- Modify: `paperforge/worker/ocr_document.py` — insert `_cluster_page_column_groups()` before `_cluster_page_columns()`, then replace `_cluster_page_columns()` body with wrapper
- [ ] **Step 1: Insert `_cluster_page_column_groups()` before `_cluster_page_columns()`, then replace `_cluster_page_columns()` body with wrapper**
Open `paperforge/worker/ocr_document.py`. Insert the new function before `_cluster_page_columns()` (before line 321), then replace the body of `_cluster_page_columns()` (lines 321-350) with a one-line wrapper:
```python
def _cluster_page_column_groups(page_blocks: list[dict], page_width: float) -> list[dict]:
"""Cluster block x-centers by column and return per-cluster metadata.
Returns list of cluster dicts sorted by center ascending. Each dict:
center, count, center_min, center_max, y_min, y_max, y_coverage,
word_count, width_median, block_ids, blocks
"""
items: list[tuple[float, dict]] = []
for block in page_blocks:
bbox = block.get("bbox") or block.get("block_bbox")
if not bbox or len(bbox) < 4:
continue
block_width = bbox[2] - bbox[0]
if block_width <= 50:
continue
x_center = (bbox[0] + bbox[2]) / 2
items.append((x_center, block))
if not items:
return [{
"center": page_width / 2,
"count": 0,
"center_min": page_width / 2,
"center_max": page_width / 2,
"y_min": 0.0,
"y_max": 0.0,
"y_coverage": 0.0,
"word_count": 0,
"width_median": 0.0,
"block_ids": [],
"blocks": [],
}]
items.sort(key=lambda x: x[0])
gap_threshold = page_width * 0.15
raw_clusters: list[list[dict]] = [[items[0][1]]]
last_center = items[0][0]
for center, block in items[1:]:
if center - last_center > gap_threshold:
raw_clusters.append([block])
else:
raw_clusters[-1].append(block)
last_center = center
result: list[dict] = []
for cluster_blocks in raw_clusters:
valid_bboxes = [
bb for b in cluster_blocks
for bb in [b.get("bbox") or b.get("block_bbox")]
if bb and len(bb) >= 4
]
centers_for_compute = [(bb[0] + bb[2]) / 2 for bb in valid_bboxes]
widths = sorted(bb[2] - bb[0] for bb in valid_bboxes)
y_vals = [(bb[1], bb[3]) for bb in valid_bboxes]
center_val = sum(centers_for_compute) / len(centers_for_compute) if centers_for_compute else page_width / 2
y_min_val = min(yv[0] for yv in y_vals) if y_vals else 0.0
y_max_val = max(yv[1] for yv in y_vals) if y_vals else 0.0
result.append({
"center": center_val,
"count": len(cluster_blocks),
"center_min": min(centers_for_compute) if centers_for_compute else page_width / 2,
"center_max": max(centers_for_compute) if centers_for_compute else page_width / 2,
"y_min": y_min_val,
"y_max": y_max_val,
"y_coverage": y_max_val - y_min_val,
"word_count": sum(len(_block_text(b).split()) for b in cluster_blocks),
"width_median": widths[len(widths) // 2] if widths else 0.0,
"block_ids": [b.get("block_id", "") for b in cluster_blocks],
"blocks": cluster_blocks,
})
result.sort(key=lambda c: c["center"])
return result
# ponytail: Replace _cluster_page_columns() body with wrapper delegating to _cluster_page_column_groups()
def _cluster_page_columns(page_blocks: list[dict], page_width: float) -> list[float]:
return [c["center"] for c in _cluster_page_column_groups(page_blocks, page_width)]
```
- [ ] **Step 2: Verify wrapper — run existing layout tests**
```bash
python -m pytest tests/test_ocr_document.py::test_layout_profile_single_column tests/test_ocr_document.py::test_layout_profile_two_column tests/test_ocr_document.py::test_layout_profile_mixed_tail -v --tb=short
```
Expected: All 3 tests PASS; wrapper preserves old observable behavior.
- [ ] **Step 3: Commit**
```bash
git add paperforge/worker/ocr_document.py
git commit -m "feat: add _cluster_page_column_groups() returning cluster metadata"
```
---
### Task 2: Add `_is_weak_isolated_column_cluster()` and modify `_classify_page_layout()`
**Files:**
- Modify: `paperforge/worker/ocr_document.py` — insert helper, modify `_classify_page_layout()` (lines 353-418)
- [ ] **Step 1: Read current `_classify_page_layout()` to confirm boundaries**
Read `paperforge/worker/ocr_document.py` lines 321-418 to confirm no changes since the spec was written.
- [ ] **Step 2: Insert `_is_weak_isolated_column_cluster()` before `_classify_page_layout()`**
After `_cluster_page_column_groups()` (just inserted), add:
```python
def _is_weak_isolated_column_cluster(cluster: dict, page_height: float) -> bool:
"""Return True if this cluster has insufficient evidence to be a real column."""
if cluster["count"] >= 3:
return False
if cluster["y_coverage"] >= page_height * 0.10:
return False
if cluster["word_count"] >= 25:
return False
return True
```
- [ ] **Step 3: Replace `_classify_page_layout()` (lines 353-418)**
Replace the entire function with:
```python
def _classify_page_layout(page_blocks: list[dict], page_width: float, page_height: float) -> PageLayoutProfile:
"""Classify a page's layout based on column clusters and role distribution."""
clusters = _cluster_page_column_groups(page_blocks, page_width)
real_clusters = [
c for c in clusters
if not _is_weak_isolated_column_cluster(c, page_height)
]
if not real_clusters:
return PageLayoutProfile(
column_count=1,
column_boundaries=[page_width / 2],
layout_type="single_column",
confidence=0.35,
evidence=["eligible_body_blocks", "all_column_clusters_weak"],
)
evidence_extra: list[str] = []
if len(real_clusters) < len(clusters):
evidence_extra.append("weak_isolated_column_cluster_ignored")
centers = [c["center"] for c in real_clusters]
column_count = len(real_clusters)
if column_count == 1:
return PageLayoutProfile(
column_count=1,
column_boundaries=[real_clusters[0]["center"]],
layout_type="single_column",
confidence=0.55 if evidence_extra else 0.7,
evidence=["eligible_body_blocks"] + evidence_extra,
)
if column_count == 2:
col_blocks: dict[int, list[str]] = {0: [], 1: []}
for i in (0, 1):
for block in real_clusters[i]["blocks"]:
col_blocks[i].append(block.get("role", ""))
body_roles = {
"body_paragraph",
"section_heading",
"subsection_heading",
"sub_subsection_heading",
}
col_has_body: dict[int, bool] = {}
col_has_tail: dict[int, bool] = {}
for col_idx, roles in col_blocks.items():
col_has_body[col_idx] = bool(set(roles) & body_roles)
col_has_tail[col_idx] = bool(set(roles) & _TAIL_ROLES)
one_side_body = col_has_body[0] and not col_has_tail[0]
other_side_tail = not col_has_body[1] and col_has_tail[1]
swapped = col_has_body[1] and not col_has_tail[1]
swapped_tail = not col_has_body[0] and col_has_tail[0]
if (one_side_body and other_side_tail) or (swapped and swapped_tail):
return PageLayoutProfile(
column_count=2,
column_boundaries=centers,
layout_type="mixed_tail",
confidence=0.6,
evidence=["eligible_body_blocks", "two_center_clusters"] + evidence_extra,
)
return PageLayoutProfile(
column_count=2,
column_boundaries=centers,
layout_type="two_column",
confidence=0.7,
evidence=["eligible_body_blocks"] + evidence_extra,
)
return PageLayoutProfile(
column_count=column_count,
column_boundaries=centers,
layout_type="two_column",
confidence=0.5,
evidence=["eligible_body_blocks", "wide_dispersion"] + evidence_extra,
)
```
- [ ] **Step 4: Run existing layout tests to verify no regression**
```bash
python -m pytest tests/test_ocr_document.py::test_layout_profile_single_column tests/test_ocr_document.py::test_layout_profile_two_column tests/test_ocr_document.py::test_layout_profile_mixed_tail tests/test_ocr_document.py::test_layout_profile_build_profiles -v --tb=short
```
Expected: All 4 tests PASS.
- [ ] **Step 5: Commit**
```bash
git add paperforge/worker/ocr_document.py
git commit -m "feat: add weak isolated cluster rejection to column detection"
```
---
### Task 3: Add test — short isolated line does NOT create two_column
**Files:**
- Modify: `tests/test_ocr_document.py` — append at end of file
- [ ] **Step 1: Add `test_short_isolated_body_line_does_not_create_two_column_layout`**
Append to `tests/test_ocr_document.py`:
```python
def test_short_isolated_body_line_does_not_create_two_column_layout() -> None:
"""Replicates page 12 of paper 2E4EPHN2: 6 full-width body blocks +
1 short offset line should NOT be classified as two_column."""
from paperforge.worker.ocr_document import _classify_page_layout
page_width = 1191.0
page_height = 1684.0
blocks = [
{"role": "body_paragraph", "bbox": [325, 224, 1123, 298], "text": "Long paragraph " * 40, "page_width": page_width},
{"role": "body_paragraph", "bbox": [325, 299, 1124, 377], "text": "Long paragraph " * 40, "page_width": page_width},
{"role": "body_paragraph", "bbox": [325, 398, 1125, 633], "text": "Long paragraph " * 40, "page_width": page_width},
{"role": "body_paragraph", "bbox": [326, 643, 1125, 762], "text": "Long paragraph " * 40, "page_width": page_width},
{"role": "body_paragraph", "bbox": [325, 772, 1124, 820], "text": "Long paragraph " * 40, "page_width": page_width},
{"role": "body_paragraph", "bbox": [328, 832, 704, 854], "text": "Informed Consent Statement: Not applicable.", "page_width": page_width},
{"role": "body_paragraph", "bbox": [327, 867, 1122, 914], "text": "Long paragraph " * 40, "page_width": page_width},
]
profile = _classify_page_layout(blocks, page_width, page_height)
assert profile.layout_type == "single_column", f"Expected single_column, got {profile.layout_type}"
assert profile.column_count == 1, f"Expected column_count=1, got {profile.column_count}"
assert "weak_isolated_column_cluster_ignored" in profile.evidence, (
f"Expected weak_isolated_column_cluster_ignored in evidence, got {profile.evidence}"
)
```
- [ ] **Step 2: Run test — expect PASS**
```bash
python -m pytest tests/test_ocr_document.py::test_short_isolated_body_line_does_not_create_two_column_layout -v --tb=short
```
Expected: PASS.
- [ ] **Step 3: Commit**
```bash
git add tests/test_ocr_document.py
git commit -m "test: short isolated line does not create two_column layout"
```
---
### Task 4: Add test — balanced two_column still detected
**Files:**
- Modify: `tests/test_ocr_document.py` — append
- [ ] **Step 1: Add `test_balanced_two_column_layout_still_detected`**
Append to `tests/test_ocr_document.py`:
```python
def test_balanced_two_column_layout_still_detected() -> None:
"""True two-column page with multiple blocks per column must remain two_column."""
from paperforge.worker.ocr_document import _classify_page_layout
page_width = 800.0
page_height = 1000.0
blocks = [
{"role": "body_paragraph", "bbox": [50, 100, 380, 250], "text": "Left body " * 30, "page_width": page_width},
{"role": "body_paragraph", "bbox": [50, 270, 380, 420], "text": "Left body " * 30, "page_width": page_width},
{"role": "body_paragraph", "bbox": [420, 100, 750, 250], "text": "Right body " * 30, "page_width": page_width},
{"role": "body_paragraph", "bbox": [420, 270, 750, 420], "text": "Right body " * 30, "page_width": page_width},
]
profile = _classify_page_layout(blocks, page_width, page_height)
assert profile.layout_type == "two_column", f"Expected two_column, got {profile.layout_type}"
assert profile.column_count == 2, f"Expected column_count=2, got {profile.column_count}"
```
- [ ] **Step 2: Run test — expect PASS**
```bash
python -m pytest tests/test_ocr_document.py::test_balanced_two_column_layout_still_detected -v --tb=short
```
Expected: PASS.
- [ ] **Step 3: Commit**
```bash
git add tests/test_ocr_document.py
git commit -m "test: balanced two_column layout still detected"
```
---
### Task 5: Add test — one large block per column still two_column
**Files:**
- Modify: `tests/test_ocr_document.py` — append
- [ ] **Step 1: Add `test_single_large_block_per_column_still_two_column`**
Append to `tests/test_ocr_document.py`:
```python
def test_single_large_block_per_column_still_two_column() -> None:
"""Each column has only 1 block, but large y_coverage and word_count
must still classify as two_column (not killed by count=1 guard)."""
from paperforge.worker.ocr_document import _classify_page_layout
page_width = 800.0
page_height = 1000.0
blocks = [
{"role": "body_paragraph", "bbox": [50, 100, 380, 800], "text": "Left body " * 100, "page_width": page_width},
{"role": "body_paragraph", "bbox": [420, 100, 750, 800], "text": "Right body " * 100, "page_width": page_width},
]
profile = _classify_page_layout(blocks, page_width, page_height)
assert profile.layout_type == "two_column", f"Expected two_column, got {profile.layout_type}"
assert profile.column_count == 2, f"Expected column_count=2, got {profile.column_count}"
```
- [ ] **Step 2: Run test — expect PASS**
```bash
python -m pytest tests/test_ocr_document.py::test_single_large_block_per_column_still_two_column -v --tb=short
```
Expected: PASS.
- [ ] **Step 3: Commit**
```bash
git add tests/test_ocr_document.py
git commit -m "test: single large block per column still two_column"
```
---
### Task 6: Add test — multiple weak offset lines do not trigger wide_dispersion
**Files:**
- Modify: `tests/test_ocr_document.py` — append
- [ ] **Step 1: Add `test_multiple_weak_offset_lines_do_not_create_wide_dispersion`**
Append to `tests/test_ocr_document.py`:
```python
def test_multiple_weak_offset_lines_do_not_create_wide_dispersion() -> None:
"""3 raw clusters (2 weak, 1 real) must produce single_column, not two_column/wide_dispersion."""
from paperforge.worker.ocr_document import _classify_page_layout
page_width = 1191.0
page_height = 1684.0
blocks = [
{"role": "body_paragraph", "bbox": [325, 224, 1123, 298], "text": "Long paragraph " * 40, "page_width": page_width},
{"role": "body_paragraph", "bbox": [325, 299, 1124, 377], "text": "Long paragraph " * 40, "page_width": page_width},
{"role": "body_paragraph", "bbox": [328, 832, 704, 854], "text": "IC Statement: Not applicable.", "page_width": page_width},
{"role": "body_paragraph", "bbox": [328, 870, 550, 890], "text": "Short note.", "page_width": page_width},
{"role": "body_paragraph", "bbox": [327, 900, 1122, 950], "text": "Long paragraph " * 40, "page_width": page_width},
]
profile = _classify_page_layout(blocks, page_width, page_height)
assert profile.layout_type == "single_column", f"Expected single_column, got {profile.layout_type}"
assert profile.column_count == 1, f"Expected column_count=1, got {profile.column_count}"
assert "weak_isolated_column_cluster_ignored" in profile.evidence, (
f"Expected weak_isolated_column_cluster_ignored in evidence, got {profile.evidence}"
)
```
- [ ] **Step 2: Run test — expect PASS**
```bash
python -m pytest tests/test_ocr_document.py::test_multiple_weak_offset_lines_do_not_create_wide_dispersion -v --tb=short
```
Expected: PASS.
- [ ] **Step 3: Commit**
```bash
git add tests/test_ocr_document.py
git commit -m "test: multiple weak offset lines do not trigger wide_dispersion"
```
---
### Task 7: Add optional test — all clusters weak fallback to single_column
**Files:**
- Modify: `tests/test_ocr_document.py` — append
- [ ] **Step 1: Add `test_all_clusters_weak_fallback_to_single_column`**
Append to `tests/test_ocr_document.py`:
```python
def test_all_clusters_weak_fallback_to_single_column() -> None:
"""When all clusters are weak (e.g. only short offset lines), fallback to low-confidence single_column."""
from paperforge.worker.ocr_document import _classify_page_layout
page_width = 1191.0
page_height = 1684.0
blocks = [
{"role": "body_paragraph", "bbox": [328, 832, 704, 854], "text": "IC Statement: Not applicable.", "page_width": page_width},
{"role": "body_paragraph", "bbox": [60, 870, 300, 890], "text": "Footnote text.", "page_width": page_width},
]
profile = _classify_page_layout(blocks, page_width, page_height)
assert profile.layout_type == "single_column", f"Expected single_column, got {profile.layout_type}"
assert profile.confidence <= 0.35, f"Expected confidence <= 0.35, got {profile.confidence}"
assert "all_column_clusters_weak" in profile.evidence, (
f"Expected all_column_clusters_weak in evidence, got {profile.evidence}"
)
```
- [ ] **Step 2: Run test — expect PASS**
```bash
python -m pytest tests/test_ocr_document.py::test_all_clusters_weak_fallback_to_single_column -v --tb=short
```
Expected: PASS.
- [ ] **Step 3: Commit**
```bash
git add tests/test_ocr_document.py
git commit -m "test: all clusters weak fallback to single_column"
```
---
### Task 8: Run full test suites — verify no regression
**Files:** None (verification only)
- [ ] **Step 1: Run focused new tests**
```bash
python -m pytest tests/test_ocr_document.py -v --tb=short -k "weak or two_column or wide_dispersion or isolated"
```
Expected: At least 5 tests PASS (the 5 new tests + any existing matching keywords).
- [ ] **Step 2: Run full test_ocr_document.py**
```bash
python -m pytest tests/test_ocr_document.py -v --tb=short
```
Expected: All tests PASS. Note: existing tests that assert exact confidence (e.g. 0.7) for sparse one-block pages may now show confidence <= 0.35 if the sole cluster is weak. Use `<=` comparison instead of exact equality.
- [ ] **Step 3: Run all OCR-related tests**
```bash
python -m pytest tests/test_ocr_rendering.py tests/test_ocr_layout_zones.py tests/test_ocr_layout_first_regressions.py tests/test_ocr_render_stabilization.py -v --tb=short
```
Expected: All tests PASS.
- [ ] **Step 4: Verify build_profile tests (mixed_tail role distribution uses cluster blocks)**
```bash
python -m pytest tests/test_ocr_document.py::test_layout_profile_mixed_tail tests/test_ocr_document.py::test_layout_profile_build_profiles -v --tb=short
```
Expected: PASS. This confirms mixed_tail detection still works with cluster-based role distribution.
---
### Task 9: Rebuild affected paper and verify fix
**Files:** None (verification only)
- [ ] **Step 1: Rebuild page layout profiles for paper 2E4EPHN2**
```bash
$PYTHON -m paperforge ocr rebuild 2E4EPHN2 --keys derived
```
Expected: Page 12 now shows `single_column` layout_type.
- [ ] **Step 2: Verify `document_structure.json` for page 12**
Read `D:\L\OB\Literature-hub\System\PaperForge\ocr\2E4EPHN2\structure\document_structure.json` and check page 12 profile:
```json
"12": {
"column_count": 1,
"layout_type": "single_column",
...
}
```
Expected: `layout_type` is `"single_column"`.
- [ ] **Step 3: Observe `fulltext.md` reading order (optional, NOT guaranteed by this fix)**
Check that "Informed Consent Statement: Not applicable." appears in the rendered `fulltext.md`. Note: this fix addresses the layout profile misclassification. The renderer's `_sort_blocks_by_column()` still uses `page_width/2` midpoint and is NOT changed by this plan. The reading order in `fulltext.md` may or may not be correct depending on downstream tail reordering logic. Full reading order fix requires a separate renderer change (tracked separately).

View file

@ -33,12 +33,12 @@ Replace the x-center-only clustering with a function that returns cluster metada
```
center: float — mean x_center of blocks in this cluster
count: int — number of blocks
x_min: float — minimum x_center
x_max: float — maximum x_center
center_min: float — minimum x_center among blocks
center_max: float — maximum x_center among blocks
y_min: float — minimum y_top of blocks' bboxes
y_max: float — maximum y_bottom of blocks' bboxes
y_coverage: float — y_max - y_min (vertical span in px)
word_count: int — total word count across all blocks
word_count: int — total word count across all blocks, computed via _block_text(block).split()
width_median: float — median block width
block_ids: list[str|int] — block identifiers
blocks: list[dict] — the full block dicts
@ -65,7 +65,7 @@ Determines if a cluster lacks sufficient evidence to be considered a real column
```python
def _is_weak_isolated_column_cluster(cluster: dict, page_height: float) -> bool:
# Not weak if ANY evidence of being a real column
if cluster["count"] >= 2:
if cluster["count"] >= 3:
return False
if cluster["y_coverage"] >= page_height * 0.10:
return False
@ -75,7 +75,7 @@ def _is_weak_isolated_column_cluster(cluster: dict, page_height: float) -> bool:
```
**Rationale for thresholds:**
- `count >= 2`: A real column in a two-column paper will have multiple blocks per page.
- `count >= 3`: A real column in a two-column paper will have multiple blocks per page. Threshold is 3 (not 2) to prevent two adjacent short lines (e.g. "Informed Consent Statement" + "Data Availability") from merging into a single count=2 cluster that would be immunized if the threshold were 2.
- `y_coverage >= page_height * 0.10`: ~160px on a 1600px page. A single large paragraph block covers this easily.
- `word_count >= 25`: ~3 lines of text. Short backmatter lines (e.g. "Informed Consent Statement: Not applicable.") are ~5 words.
@ -204,7 +204,8 @@ def _classify_page_layout(page_blocks, page_width, page_height):
| True two-column body page (multi-block per column) | `two_column` | Both clusters: count >= 2 |
| True two-column reference page | `two_column` | Reference items have sufficient count |
| Mixed tail page (body + backmatter) | `mixed_tail` | Both columns have sufficient y_coverage |
| Single-column page (1 cluster) | `single_column` | Does not enter weak check |
| Single-column page with one real cluster | `single_column` | The cluster survives weak filtering |
| Single-column page with only tiny snippets | `single_column` low confidence | All clusters weak → `all_column_clusters_weak` |
| Single-column + 1 short offset line (this bug) | `single_column` | Weak cluster ignored |
| Single-column + 2 short offset lines (3 raw clusters) | `single_column` | 2 weak filtered, 1 real remains |
| All clusters weak (rare) | `single_column` low confidence | Fallback, does not preserve false two_column |
@ -214,6 +215,10 @@ def _classify_page_layout(page_blocks, page_width, page_height):
## Tests
**Location:** `tests/test_ocr_document.py` (alongside existing layout-profile tests like `test_layout_profile_single_column`). Use function-local imports matching current style.
**Evidence assertions:** Use `assert evidence_string in profile.evidence`. Do not assert exact evidence list equality — `_build_page_layout_profiles()` may append `few_eligible_blocks` or `excluded_non_body_blocks` after `_classify_page_layout()` returns, and this stacking is expected behavior.
### Required (4)
1. **test_short_isolated_body_line_does_not_create_two_column_layout**

View file

@ -3660,11 +3660,22 @@ def _detect_non_body_insert_clusters(
if block.get("role") not in _INSERT_CANDIDATE_ROLES:
continue
# Never mark page-1 title-like blocks as non_body_insert
# Never mark page-1 body-like blocks as non_body_insert.
# Uses position (in main column, not narrow sidebar) and style (body font).
if page == 1:
text = (block.get("text") or "").strip()
if len(text) > 20 and not any(m in text.lower() for m in ["\u2022", "-", "*"]):
continue
bbox = block.get("bbox", [0, 0, 0, 0])
if len(bbox) >= 4:
block_width = bbox[2] - bbox[0]
spine_key = page if page in body_spine else 1
spine = body_spine.get(spine_key, {"median_width": 500})
spine_fonts = spine.get("all_fonts") or spine.get("fonts", set())
if not isinstance(spine_fonts, set):
spine_fonts = set(spine_fonts) if spine_fonts else set()
block_font = _first_font(block)
font_matches_body = not block_font or not spine_fonts or block_font in spine_fonts
not_extremely_narrow = block_width > page_width * 0.2
if font_matches_body and not_extremely_narrow:
continue
bbox = block.get("bbox", [0, 0, 0, 0])
@ -3728,6 +3739,21 @@ def _detect_non_body_insert_clusters(
text = block.get("text", "")
if not text or not text[0].islower():
continue
# Same page-1 guard as first pass: position + style check
if page == 1:
bbox = block.get("bbox", [0, 0, 0, 0])
if len(bbox) >= 4:
block_width = bbox[2] - bbox[0]
spine_key = page if page in body_spine else 1
spine = body_spine.get(spine_key, {"median_width": 500})
spine_fonts = spine.get("all_fonts") or spine.get("fonts", set())
if not isinstance(spine_fonts, set):
spine_fonts = set(spine_fonts) if spine_fonts else set()
block_font = _first_font(block)
font_matches_body = not block_font or not spine_fonts or block_font in spine_fonts
not_extremely_narrow = block_width > page_width * 0.2
if font_matches_body and not_extremely_narrow:
continue
block_font = _first_font(block)
if block_font and block_font in insert_fonts:
# Check adjacency to existing cluster members on the same page

View file

@ -129,7 +129,7 @@ def _convert_footnotes_to_callouts(blocks: list[dict]) -> list[dict]:
page_body_fonts = body_fonts_by_page.get(fn_page) or all_body_fonts
body_font_median = sorted(page_body_fonts)[len(page_body_fonts) // 2]
fn_fs = _block_font_size(b)
if fn_fs is None or fn_fs >= body_font_median * 0.9:
if fn_fs is None or fn_fs >= body_font_median:
continue
page_body_blocks = [bb for bb in body_blocks if int(bb.get("page", 0) or 0) == fn_page]
@ -1266,7 +1266,10 @@ def render_fulltext_markdown(
block["role"] = new_role
block["role_confidence"] = new_blocks[i].get("role_confidence", block.get("role_confidence", 0.5))
before = sum(1 for b in structured_blocks if b.get("role") == "footnote")
structured_blocks = _convert_footnotes_to_callouts(structured_blocks)
after = sum(1 for b in structured_blocks if b.get("role") == "footnote")
import sys; print(f"[DEBUG] callout footnotes: {before} -> {after}", file=sys.stderr)
style_profiles = _build_heading_style_profiles(structured_blocks)