diff --git a/.planning/PROJECT.md b/.planning/PROJECT.md index e4fe0aa8..52492df7 100644 --- a/.planning/PROJECT.md +++ b/.planning/PROJECT.md @@ -6,6 +6,16 @@ PaperForge Lite is a local Obsidian + Zotero literature workflow for medical res This is a brownfield release-hardening project for `D:\L\Med\Research\99_System\LiteraturePipeline\github-release`, informed by the fuller local implementation under `D:\L\Med\Research\99_System\LiteraturePipeline` and the production Obsidian Base views under `D:\L\Med\Research\05_Bases`. +## Current Milestone: v1.1 Sandbox Onboarding Hardening + +**Goal:** Make the GitHub README + sandbox path behave like a real first-time user flow, with no silent setup stalls, contradictory diagnostics, unresolved mock Zotero PDFs, or broken `/LD-deep` prepare commands. + +**Target features:** +- Setup wizard and CLI commands are internally consistent when a user follows README exactly. +- `paperforge doctor`, `paperforge paths --json`, worker fallback commands, and Agent command docs report the same installed paths and env variable names. +- Sandbox Better BibTeX exports and mock Zotero storage exercise the full flow from selection sync through `/LD-deep prepare`. +- PDF/OCR/deep-reading statuses remain consistent across library-records, formal notes, `meta.json`, and queue output. + ## Core Value A new user can install PaperForge, configure their own vault paths and PaddleOCR credentials, then run the full literature pipeline with copy-pasteable commands that diagnose failures clearly. @@ -18,56 +28,60 @@ A new user can install PaperForge, configure their own vault paths and PaddleOCR - ✓ Configurable vault directories are partially supported through `paperforge.json` and `vault_config`. - ✓ Existing Obsidian Base views prove the intended queue workflow: recommended analysis, OCR queue, completed OCR, pending deep reading, completed deep reading, and formal notes. - ✓ OCR queue state is persisted in `/PaperForge/ocr/ocr-queue.json` and per-paper `meta.json`. +- ✓ v1.0 shipped a shared resolver, `paperforge` CLI, generated Bases, first-pass doctor command, fixture smoke tests, and command documentation. ### Active -- [ ] Registration-to-first-paper onboarding is explicit and testable. -- [ ] PaddleOCR configuration is validated before queue mutation and reports actionable errors. -- [ ] PDF path resolution works for absolute paths, vault-relative paths, Zotero storage-relative paths, and paths behind the configured Zotero junction. -- [ ] User-customized directories are resolved by command-line tools and environment variables, not by agent-written placeholder substitution. -- [ ] Generated Base files match the operational views used in the real vault, while still adapting to custom directory names. -- [ ] Full-flow validation covers setup, selection sync, index refresh, OCR preflight, OCR polling, deep-reading queue, and `/LD-deep` prepare. +- [ ] README-driven setup works in the sandbox without hidden required inputs or unexplained terminal stalls. +- [ ] Setup wizard, `paperforge` CLI, direct worker fallback, deployed Agent scripts, and command docs agree on the installed path contract. +- [ ] Diagnostics validate the actual supported Better BibTeX export shapes and PaddleOCR env names. +- [ ] PDF path resolution handles sandbox BBT attachment paths and common Zotero storage-relative paths. +- [ ] Selection sync writes complete normalized metadata into library-records, including author and journal fields. +- [ ] OCR status, formal note status, library-record status, and deep-reading queue status converge after each worker step. +- [ ] `/LD-deep` helpers run from the deployed Vault location without manual `PYTHONPATH` fixes. +- [ ] The sandbox smoke test catches every regression found in the manual first-time-user simulation. ### Out of Scope - Replacing Zotero or Better BibTeX — the project is built around them. - Automatically triggering deep-reading agents from workers — the Lite architecture intentionally keeps worker automation and agent reasoning separate. - Cloud-hosted multi-user service — this project targets local single-user vault workflows. -- Full OCR provider abstraction in v1 — PaddleOCR should be robust first; provider plugins can follow later. +- Full OCR provider abstraction in v1.1 — PaddleOCR path/env consistency is the priority. ## Context -The current release repo has `setup_wizard.py`, `pipeline/worker/scripts/literature_pipeline.py`, OpenCode command files, chart-reading skills, installation docs, and a user-facing `AGENTS.md`. The fuller local pipeline contains additional scripts and real operational history, but the release should remain Lite and understandable for installation. +The v1.1 milestone is based on a manual sandbox audit performed from `tests/sandbox/00_TestVault` using only README-level guidance. The audit found that v1.0's claimed release-hardening coverage is not yet sufficient for a real first-time user: -Important observations from code and Base review: - -- `setup_wizard.py` stores PaddleOCR credentials in `/PaperForge/.env`, and the worker loads both vault `.env` and PaperForge `.env`. -- `run_ocr()` expects `PADDLEOCR_API_TOKEN`, optional `PADDLEOCR_API_TOKEN_USER`, `PADDLEOCR_JOB_URL`, `PADDLEOCR_MODEL`, and `PADDLEOCR_MAX_ITEMS`. -- OCR requests use `Authorization: bearer `, multipart `file`, `model`, and `optionalPayload`, then poll `job_url/`. -- OCR error handling currently records broad request failures but does not classify unauthorized URL/auth/payload/path failures for the user. -- The OCR worker opens `queue_row['pdf_path']` directly. If Better BibTeX exports a relative path, linked attachment path, or Zotero storage path that is not directly openable from the current process, OCR fails. -- Base templates generated by `ensure_base_views()` are much simpler than production Base files such as `骨科.base`, `运动医学.base`, and `Literature Hub.base`. -- Production Base filters hardcode `03_Resources/LiteratureControl/...` and `03_Resources/Literature/...`, so they conflict with the release promise that resources/control/literature directories are user-customizable. -- Command docs still rely on `` placeholder guidance rather than a first-class launcher that resolves the actual configured paths. +- `python setup_wizard.py --vault ...` can appear to hang in a terminal, and the Vault input is not prefilled from `--vault`. +- `paperforge` may be unavailable when setup does not complete; README does not clearly provide a reliable fallback. +- `paperforge doctor` checks `library.json` and `PADDLEOCR_API_KEY`, while the implemented flow supports per-domain JSON exports and uses `PADDLEOCR_API_TOKEN`. +- `paperforge paths --json` reports a worker path under `vault/pipeline/...`, but setup deploys the worker under `/PaperForge/worker/scripts/...`. +- Command docs reference `literature_script`, but the CLI emits `ld_deep_script`. +- Sandbox BBT attachment paths such as `TSTONE001/TSTONE001.pdf` do not resolve to `/storage/TSTONE001/TSTONE001.pdf`. +- `selection-sync` normalizes export rows but later reads raw BBT fields, leaving `first_author` and `journal` empty in library-records. +- `deep-reading --verbose` writes the useful queue report to a file but prints only a terse pending count. +- Deployed `ld_deep.py` depends on `paperforge_lite` importability and fails without manual `PYTHONPATH` or a successful package install. ## Constraints -- **Local-first**: Must work in a user’s Obsidian vault without a daemon or cloud service. -- **Windows compatibility**: Windows junctions, PowerShell, and paths with Chinese names are first-class use cases. -- **Plain Python**: Keep dependencies small; current requirements are `requests`, `pymupdf`, `pillow`, `textual`, and `pytest`. -- **Obsidian compatibility**: `.base` files must use Obsidian Bases syntax and relative vault paths. -- **Credential safety**: API keys belong in `.env` or user environment variables and must not be committed. -- **Agent independence**: Users should not need an agent to inspect `paperforge.json` just to build a worker command. +- **Local-first:** Must work in a user's Obsidian vault without a daemon or cloud service. +- **Windows compatibility:** Windows junctions, PowerShell, and paths with Chinese names are first-class use cases. +- **Plain Python:** Keep dependencies small; current requirements are `requests`, `pymupdf`, `pillow`, `textual`, and `pytest`. +- **Obsidian compatibility:** `.base` files must use Obsidian Bases syntax and relative vault paths. +- **Credential safety:** API keys belong in `.env` or user environment variables and must not be committed. +- **Agent independence:** Users should not need an agent to inspect `paperforge.json` just to build a worker command. +- **Sandbox realism:** `tests/sandbox` must remain safe, deterministic, and representative of a GitHub user trying the project locally. ## Key Decisions | Decision | Rationale | Outcome | |----------|-----------|---------| -| Keep Lite two-layer architecture | Worker and Agent responsibilities are already clear and lower-risk than automatic deep-reading triggers | - Pending | -| Add a PaperForge CLI/launcher layer | It removes placeholder command friction and centralizes path/env resolution | - Pending | -| Treat PaddleOCR as a preflighted integration | Users need immediate diagnosis before jobs enter confusing pending/error states | - Pending | -| Generate Bases from config-aware templates | Current production Base UX is better than release templates, but hardcoded paths must be parameterized | - Pending | -| Keep planning docs local to release repo | Parent vault has its own `.planning`; this repo needs scoped release-hardening state | - Pending | +| Keep Lite two-layer architecture | Worker and Agent responsibilities are already clear and lower-risk than automatic deep-reading triggers | Accepted | +| Add a PaperForge CLI/launcher layer | It removes placeholder command friction and centralizes path/env resolution | Implemented in v1.0, repair consistency in v1.1 | +| Treat PaddleOCR as a preflighted integration | Users need immediate diagnosis before jobs enter confusing pending/error states | Implemented in v1.0, align env names in v1.1 | +| Generate Bases from config-aware templates | Current production Base UX is better than release templates, but hardcoded paths must be parameterized | Implemented in v1.0 | +| Use sandbox audit as v1.1 release gate | Manual first-time-user simulation exposed regressions that unit tests missed | Active | +| Continue phase numbering after v1.0 | v1.1 is a follow-up hardening milestone, not a project reset | Phases start at 6 | ## Evolution @@ -87,4 +101,4 @@ This document evolves at phase transitions and milestone boundaries. 4. Update Context with current state. --- -*Last updated: 2026-04-23 after initialization* +*Last updated: 2026-04-23 starting milestone v1.1* diff --git a/.planning/REQUIREMENTS.md b/.planning/REQUIREMENTS.md index b82c16dd..1f81993b 100644 --- a/.planning/REQUIREMENTS.md +++ b/.planning/REQUIREMENTS.md @@ -1,63 +1,53 @@ -# Requirements: PaperForge Lite Release Hardening +# Requirements: PaperForge Lite v1.1 Sandbox Onboarding Hardening **Defined:** 2026-04-23 **Core Value:** A new user can install PaperForge, configure their own vault paths and PaddleOCR credentials, then run the full literature pipeline with copy-pasteable commands that diagnose failures clearly. -## v1 Requirements +## Milestone v1.1 Requirements -### Onboarding +### Setup And Commands -- [ ] **ONBD-01**: User can follow one registration-to-first-paper guide that covers Zotero, Better BibTeX, Obsidian, PaddleOCR, and PaperForge. -- [ ] **ONBD-02**: User can run one validation command that reports setup readiness by category. -- [ ] **ONBD-03**: User can see the exact next command after each setup or worker step. +- [ ] **SETUP-01**: User can run `python setup_wizard.py --vault ` and see an immediate, understandable setup UI or message instead of an unexplained terminal stall. +- [ ] **SETUP-02**: User can rely on `--vault ` being carried into the wizard flow without retyping the same path. +- [ ] **SETUP-03**: User can continue with a documented fallback command when the global `paperforge` executable is not registered. +- [ ] **SETUP-04**: User can inspect `paperforge paths --json` and receive installed worker and Agent script paths that actually exist. +- [ ] **SETUP-05**: Agent command docs use the same JSON field names emitted by `paperforge paths --json`. -### Configuration +### Diagnostics -- [ ] **CONF-01**: User can define vault and custom directories through environment variables without editing generated code. -- [x] **CONF-02**: User can inspect resolved PaperForge paths with a command. -- [ ] **CONF-03**: Worker, Agent scripts, command docs, and Base generation all use the same config resolver. -- [ ] **CONF-04**: Existing `paperforge.json` installations remain backward-compatible. +- [ ] **DIAG-01**: User can run `paperforge doctor` against per-domain Better BibTeX exports without being incorrectly blocked for missing `library.json`. +- [ ] **DIAG-02**: User can configure PaddleOCR once with the env variable name that setup writes and workers read. +- [ ] **DIAG-03**: Doctor reports the deployed worker script path according to the same resolver contract used by runtime commands. +- [ ] **DIAG-04**: OCR doctor distinguishes an expected endpoint-method mismatch from a bad user URL when checking the configured PaddleOCR job endpoint. -### Commands +### Zotero Paths And Metadata -- [ ] **CMD-01**: User can run stable commands such as `paperforge status`, `paperforge ocr run`, and `paperforge deep-reading`. -- [ ] **CMD-02**: Legacy direct worker invocation remains supported. -- [ ] **CMD-03**: Command output uses actionable statuses and avoids placeholder paths. +- [ ] **ZPATH-01**: User can sync a BBT attachment path shaped like `KEY/KEY.pdf` when the file exists under the configured Zotero `storage/KEY/` directory. +- [ ] **ZPATH-02**: User can sync common `storage:KEY/file.pdf` and `storage/KEY/file.pdf` attachment forms. +- [ ] **ZPATH-03**: User sees library-record `pdf_path` populated only with a readable resolved PDF path or an explicit actionable missing-PDF status. +- [ ] **META-01**: User sees `first_author` populated in generated library-records from normalized export metadata. +- [ ] **META-02**: User sees `journal` populated in generated library-records from normalized export metadata. -### PaddleOCR +### State And Queue Consistency -- [ ] **OCR-01**: User can run `ocr doctor` to validate token presence, URL shape, network reachability, and expected API response structure. -- [ ] **OCR-02**: OCR worker validates PDF readability before submitting a job. -- [ ] **OCR-03**: OCR failures identify whether the cause is missing token, bad URL, unauthorized token, unreadable PDF, API schema mismatch, timeout, or provider error. -- [ ] **OCR-04**: User can retry/reset errored or blocked OCR records after fixing configuration. -- [ ] **OCR-05**: OCR polling handles provider schema changes defensively and records raw diagnostic snippets safely. +- [ ] **STATE-01**: User can run `selection-sync`, `index-refresh`, and `ocr run` without records simultaneously saying `has_pdf: true` and `ocr_status: nopdf` when the PDF is readable. +- [ ] **STATE-02**: User sees formal note OCR status synchronized with validated OCR `meta.json` status after worker refresh. +- [ ] **STATE-03**: User sees `paperforge deep-reading --verbose` print the ready/waiting/blocked queue summary directly or print the report path clearly. +- [ ] **STATE-04**: User can tell from command output which record needs OCR, which one is blocked, and which one is ready for `/LD-deep`. -### Zotero And Paths +### Deep Reading Helpers -- [ ] **ZOT-01**: PDF path resolver supports absolute paths, vault-relative paths, configured Zotero junction paths, and common Zotero storage-relative paths. -- [ ] **ZOT-02**: Selection sync reports records with missing or unreadable PDFs. -- [ ] **ZOT-03**: Better BibTeX export path and expected JSON shape are validated. +- [ ] **DEEP-04**: User can run the deployed `ld_deep.py` helper from the Vault installation without manually setting `PYTHONPATH`. +- [ ] **DEEP-05**: User can run `/LD-deep queue` documentation examples using paths and field names that exist. +- [ ] **DEEP-06**: User can prepare a sandbox OCR-complete paper and get `figure-map.json`, `chart-type-map.json`, and a `## 🔍 精读` scaffold in the formal note. -### Obsidian Bases +### Regression Coverage -- [ ] **BASE-01**: Generated domain Base files include the operational views from the real vault workflow. -- [ ] **BASE-02**: Base filters are rendered from configured paths instead of hardcoded `03_Resources`. -- [ ] **BASE-03**: Base generation preserves user-edited Base files unless explicitly refreshed. -- [ ] **BASE-04**: Literature Hub Base gives a cross-domain queue overview. +- [ ] **REG-01**: Maintainer can run one sandbox smoke test that starts from a clean `tests/sandbox/00_TestVault` and covers setup-equivalent layout, selection sync, index refresh, OCR preflight/dry-run, deep-reading queue, and `ld_deep.py prepare`. +- [ ] **REG-02**: Smoke assertions cover the exact regressions from the manual audit: doctor env names, per-domain JSON, worker path JSON, BBT PDF path resolution, metadata fields, queue output, and deployed Agent importability. +- [ ] **REG-03**: README, INSTALLATION.md, AGENTS.md, and command files stay consistent with the smoke-tested commands. -### Deep Reading - -- [ ] **DEEP-01**: Deep-reading queue accurately shows ready versus blocked papers. -- [ ] **DEEP-02**: `/LD-deep` prepare uses the same resolved paths as workers. -- [ ] **DEEP-03**: `/LD-deep` failure messages tell the user which worker command fixes the blocker. - -### Release Quality - -- [ ] **REL-01**: Automated tests cover config resolution, PDF path resolution, OCR state transitions, Base rendering, and command launcher behavior. -- [ ] **REL-02**: A smoke test can run through setup validation, selection sync, index refresh, OCR doctor, OCR queue dry-run, and deep-reading queue. -- [ ] **REL-03**: Documentation and AGENTS guide match implemented commands. - -## v2 Requirements +## Future Requirements ### Integrations @@ -78,46 +68,42 @@ | Automatic deep-reading generation from worker | Conflicts with Lite architecture and risks uncontrolled agent work | | Replacing Zotero collections as the source of domains | Existing workflow depends on Zotero and Better BibTeX exports | | Multi-user hosted backend | Not needed for local release reliability | -| Full provider plugin system in v1 | PaddleOCR must be made reliable first | +| Full provider plugin system in v1.1 | The current milestone fixes PaddleOCR path/env consistency first | +| Real PaddleOCR network smoke test in default CI | Sandbox regression should be deterministic; live provider checks remain opt-in | ## Traceability | Requirement | Phase | Status | |-------------|-------|--------| -| ONBD-01 | Phase 4 | Pending | -| ONBD-02 | Phase 4 | Pending | -| ONBD-03 | Phase 4 | Pending | -| CONF-01 | Phase 1 | Pending | -| CONF-02 | Phase 1 | Pending | -| CONF-03 | Phase 1 | Done (01-03) | -| CONF-04 | Phase 1 | Done (01-03) | -| CMD-01 | Phase 1 | Pending | -| CMD-02 | Phase 1 | Done (01-03) | -| CMD-03 | Phase 1 | Pending | -| OCR-01 | Phase 2 | Pending | -| OCR-02 | Phase 2 | Pending | -| OCR-03 | Phase 2 | Pending | -| OCR-04 | Phase 2 | Pending | -| OCR-05 | Phase 2 | Pending | -| ZOT-01 | Phase 2 | Pending | -| ZOT-02 | Phase 2 | Pending | -| ZOT-03 | Phase 4 | Pending | -| BASE-01 | Phase 3 | Pending | -| BASE-02 | Phase 3 | Pending | -| BASE-03 | Phase 3 | Pending | -| BASE-04 | Phase 3 | Pending | -| DEEP-01 | Phase 4 | Pending | -| DEEP-02 | Phase 1 | Done (01-03) | -| DEEP-03 | Phase 4 | Pending | -| REL-01 | Phase 5 | Pending | -| REL-02 | Phase 5 | Done (05-02) | -| REL-03 | Phase 5 | Pending | +| SETUP-01 | Phase 6 | Pending | +| SETUP-02 | Phase 6 | Pending | +| SETUP-03 | Phase 6 | Pending | +| SETUP-04 | Phase 6 | Pending | +| SETUP-05 | Phase 6 | Pending | +| DIAG-01 | Phase 6 | Pending | +| DIAG-02 | Phase 6 | Pending | +| DIAG-03 | Phase 6 | Pending | +| DIAG-04 | Phase 6 | Pending | +| ZPATH-01 | Phase 7 | Pending | +| ZPATH-02 | Phase 7 | Pending | +| ZPATH-03 | Phase 7 | Pending | +| META-01 | Phase 7 | Pending | +| META-02 | Phase 7 | Pending | +| STATE-01 | Phase 7 | Pending | +| STATE-02 | Phase 7 | Pending | +| STATE-03 | Phase 7 | Pending | +| STATE-04 | Phase 7 | Pending | +| DEEP-04 | Phase 8 | Pending | +| DEEP-05 | Phase 8 | Pending | +| DEEP-06 | Phase 8 | Pending | +| REG-01 | Phase 8 | Pending | +| REG-02 | Phase 8 | Pending | +| REG-03 | Phase 8 | Pending | **Coverage:** -- v1 requirements: 28 total -- Mapped to phases: 28 +- v1.1 requirements: 24 total +- Mapped to phases: 24 - Unmapped: 0 --- -*Requirements defined: 2026-04-23* -*Last updated: 2026-04-23 after initialization* +*Requirements defined: 2026-04-23 from sandbox first-time-user audit* diff --git a/.planning/ROADMAP.md b/.planning/ROADMAP.md index 3df2848b..4adf7f0b 100644 --- a/.planning/ROADMAP.md +++ b/.planning/ROADMAP.md @@ -1,138 +1,66 @@ -# Roadmap: PaperForge Lite Release Hardening +# Roadmap: PaperForge Lite v1.1 Sandbox Onboarding Hardening **Created:** 2026-04-23 -**Scope:** Make the local release flow reliable from setup through first deep-reading queue. +**Scope:** Fix every issue found by the README-driven sandbox first-time-user simulation. -## Phase 1: Config And Command Foundation +## Phase 6: Setup, CLI, And Diagnostics Consistency -**Goal:** Replace agent/manual placeholder path handling with a shared config resolver and stable user commands. +**Goal:** Make the documented setup path, installed CLI, doctor command, and Agent command docs agree on the same paths, env names, and fallback commands. -**Requirements:** CONF-01, CONF-02, CONF-03, CONF-04, CMD-01, CMD-02, CMD-03, DEEP-02 +**Requirements:** SETUP-01, SETUP-02, SETUP-03, SETUP-04, SETUP-05, DIAG-01, DIAG-02, DIAG-03, DIAG-04 **Success Criteria:** -1. `paperforge paths` prints resolved vault, system, resources, literature, control, base, worker, and skill paths. -2. Environment variables override `paperforge.json` without breaking existing installs. -3. Worker and `/LD-deep` use one shared config/path resolver or equivalent duplicated-tested contract. -4. Documentation can show stable commands without `` placeholders. +1. Running `python setup_wizard.py --vault ` gives immediate visible progress or a clear TUI message, and the provided vault path is prefilled or otherwise honored. +2. `paperforge paths --json` returns existing deployed paths for worker and `/LD-deep` helper scripts. +3. `paperforge doctor` passes the sandbox's per-domain exports and checks the same PaddleOCR env variable written by setup. +4. README, INSTALLATION.md, AGENTS.md, and `command/LD-deep.md` use field names and fallback commands that exist. **Implementation Notes:** -- Add a small CLI entrypoint or launcher script while keeping direct `literature_pipeline.py --vault ...` supported. -- Prefer a pure Python resolver module that can be reused by worker, validation, setup, and agent helpers. -- Keep `.env` loading deterministic: vault root, PaperForge `.env`, then process environment precedence rules documented clearly. +- Prefer one resolver contract for both human-readable paths and JSON output. +- Keep `python -m paperforge_lite ...` documented as the fallback when `paperforge` is not registered. +- Doctor should validate all `*.json` exports under the configured exports directory, not only `library.json`. -**Plans:** 4 plans +## Phase 7: Zotero PDF, Metadata, And State Repair -Plans: -- [x] 01-01-PLAN.md — Shared config resolver and path inventory contract (COMPLETE: 2026-04-23) -- [x] 01-02-PLAN.md — `paperforge` launcher, package entry point, and command dispatch (COMPLETE: 2026-04-23) -- [x] 01-03-PLAN.md — Worker, `/LD-deep`, setup, and validation resolver integration (COMPLETE: 2026-04-23) -- [x] 01-04-PLAN.md — Stable command documentation and setup next-step updates (COMPLETE: 2026-04-23) +**Goal:** Make sandbox BBT attachment paths resolve correctly and keep OCR/deep-reading state consistent across records, notes, and meta files. -## Phase 2: PaddleOCR And PDF Path Hardening - -**Goal:** Make OCR failures diagnosable and retryable, especially for the API key/URL issue already observed. - -**Requirements:** OCR-01, OCR-02, OCR-03, OCR-04, OCR-05, ZOT-01, ZOT-02 +**Requirements:** ZPATH-01, ZPATH-02, ZPATH-03, META-01, META-02, STATE-01, STATE-02, STATE-03, STATE-04 **Success Criteria:** -1. `paperforge ocr doctor` distinguishes missing token, bad URL, unauthorized response, network timeout, schema mismatch, and unreadable PDF. -2. OCR worker resolves common Zotero PDF paths before submission and records the resolved path in diagnostics. -3. Blocked/error records can be reset or retried with a documented command. -4. `meta.json` error messages are actionable and include a suggested next command. +1. `selection-sync` resolves `KEY/KEY.pdf`, `storage:KEY/file.pdf`, `storage/KEY/file.pdf`, absolute paths, vault-relative paths, and configured Zotero junction paths. +2. Generated library-records for sandbox PDFs contain readable `pdf_path`, non-empty `first_author`, and non-empty `journal`. +3. OCR worker does not leave readable-PDF records in the contradictory `has_pdf: true` plus `ocr_status: nopdf` state. +4. `deep-reading --verbose` surfaces ready/waiting/blocked details directly enough for a first-time user to know the next command. **Implementation Notes:** -- Add tests with mocked `requests.post/get` responses for auth failure, changed schema, pending, running, done, provider error, and timeout. -- Add PDF resolver tests for absolute, vault-relative, system Zotero junction, and missing file cases. -- Consider normalizing auth header to `Bearer` and allowing an env override for header name/scheme if PaddleOCR requires it. +- Fix path resolution before changing OCR state transitions; most downstream contradictions start with unresolved PDFs. +- Use normalized export row fields consistently after `load_export_rows`. +- Add tests around both library-record frontmatter and formal note frontmatter, not only `formal-library.json`. -**Plans:** 4 plans +## Phase 8: Deep Helper Deployment And Sandbox Regression Gate -Plans: -- [x] 02-01-PLAN.md — PDF Path Resolver + Preflight (ZOT-01, OCR-02, ZOT-02) (COMPLETE: 2026-04-23) -- [x] 02-02-PLAN.md — OCR Failure Classification (OCR-03, OCR-04, OCR-05) (COMPLETE: 2026-04-23) -- [x] 02-03-PLAN.md — OCR Doctor Command (OCR-01) (COMPLETE: 2026-04-23) -- [x] 02-04-PLAN.md — Selection Sync PDF Reporting (ZOT-02) (COMPLETE: 2026-04-23) +**Goal:** Turn the manual sandbox audit into an automated release gate that covers deployed Agent helper importability and `/LD-deep prepare`. -## Phase 3: Config-Aware Obsidian Bases - -**Goal:** Generate Base views that match the real operational workflow and respect custom directory names. - -**Requirements:** BASE-01, BASE-02, BASE-03, BASE-04 +**Requirements:** DEEP-04, DEEP-05, DEEP-06, REG-01, REG-02, REG-03 **Success Criteria:** -1. Generated domain Bases include control, recommended analysis, pending OCR, completed OCR, pending deep reading, completed deep reading, formal cards, and all-records views. -2. `Literature Hub.base` provides cross-domain overview views. -3. Generated filters use resolved relative paths, not hardcoded `03_Resources`. -4. Existing user-edited `.base` files are not overwritten unless a refresh flag is used. +1. The deployed `ld_deep.py` in the sandbox Vault can run `queue` and `prepare` without manual `PYTHONPATH`. +2. A sandbox OCR-complete fixture produces `figure-map.json`, `chart-type-map.json`, and a `## 🔍 精读` scaffold. +3. One smoke command starts from a clean sandbox and fails if any manual-audit regression reappears. +4. Docs are verified against the same commands used by the smoke test. **Implementation Notes:** -- Convert the useful structure from `骨科.base`, `运动医学.base`, and `Literature Hub.base` into templates. -- Avoid depending on a single domain name; render per export/domain. -- Add snapshot-style tests for default paths and custom paths. - -**Plans:** 2 plans - -Plans: -- [x] 03-01-PLAN.md — Base Generation Refactor — 8 Views + Incremental Merge + Placeholder Substitution (COMPLETE: 2026-04-23) -- [x] 03-02-PLAN.md — CLI base-refresh + Tests (COMPLETE: 2026-04-23) - -## Phase 4: End-To-End Onboarding And Validation - -**Goal:** Turn setup into a guided, verifiable path from registration/configuration to a ready deep-reading queue. - -**Requirements:** ONBD-01, ONBD-02, ONBD-03, ZOT-03, DEEP-01, DEEP-03 - -**Success Criteria:** -1. Install docs and `AGENTS.md` describe the exact full flow and current commands. -2. `validate_setup.py` or `paperforge doctor` reports category-level readiness: Python, vault, config, Zotero link, BBT export, Base files, OCR config, worker scripts, agent scripts. -3. After each worker command, output includes next steps and blocker-specific instructions. -4. `/LD-deep` prepare failures point to the command that fixes the blocker. - -**Implementation Notes:** -- The docs should include a first-paper checklist with expected outputs. -- Validation should not require a real OCR job unless the user opts into live provider validation. -- Keep Chinese user-facing docs consistent with command output. - -**Plans:** 4 plans - -Plans: -- [x] 04-01-PLAN.md — deep-reading 三态 + verbose (ONBD-03, DEEP-01) (COMPLETE: 2026-04-23) -- [x] 04-02-PLAN.md — paperforge doctor 子命令 (ONBD-02) (COMPLETE: 2026-04-23) -- [x] 04-03-PLAN.md — AGENTS.md paperforge CLI 更新 (ONBD-03) (COMPLETE: 2026-04-23) -- [x] 04-04-PLAN.md — docs/README.md BBT 配置指南 (ONBD-01, ZOT-03) (COMPLETE: 2026-04-23) - -## Phase 5: Release Verification - -**Goal:** Prove the release is robust enough to ship and maintain. - -**Requirements:** REL-01, REL-02, REL-03 - -**Success Criteria:** -1. Unit tests cover config resolver, path resolver, OCR state machine, Base rendering, and launcher commands. -2. A smoke test runs on a fixture vault without touching the real vault. -3. Release docs, setup wizard, command files, and generated AGENTS guide are internally consistent. -4. Known defects from `.planning/research/DEFECTS.md` are either fixed or explicitly deferred. - -**Implementation Notes:** -- Use fixture Better BibTeX JSON and dummy PDFs. -- Mock network calls for normal CI; keep live PaddleOCR validation manual/optional. -- Do not overwrite the existing user-facing `AGENTS.md` with generic GSD instructions. - -**Plans:** 2 plans - -Plans: -- [x] 05-01-PLAN.md — Test coverage gaps: OCR state machine, Base rendering, command docs (REL-01, REL-03) -- [x] 05-02-PLAN.md — Fixture smoke test suite (REL-02) +- The smoke test should be deterministic and should not call the live PaddleOCR API. +- Reuse `tests/sandbox/generate_sandbox.py` or convert it into a pytest fixture factory. +- Keep generated sandbox Vault output ignored by git; only commit fixtures, tests, and docs. ## Phase Summary | # | Phase | Goal | Requirements | Status | |---|-------|------|--------------|--------| -| 1 | Config And Command Foundation | Stable commands and shared path/env resolution | 8 | COMPLETE | -| 2 | PaddleOCR And PDF Path Hardening | Diagnosable, retryable OCR | 7 | COMPLETE | -| 3 | Config-Aware Obsidian Bases | Real workflow Bases without hardcoded paths | 4 | COMPLETE | -| 4 | End-To-End Onboarding And Validation | User can complete first-paper flow | 6 | COMPLETE | -| 5 | Release Verification | Tests and docs prove ship readiness | 3 | COMPLETE | +| 6 | Setup, CLI, And Diagnostics Consistency | Align setup/docs/doctor/path contracts | 9 | 2 plans | +| 7 | Zotero PDF, Metadata, And State Repair | Resolve PDFs and converge status fields | 9 | Planned | +| 8 | Deep Helper Deployment And Sandbox Regression Gate | Automate the manual sandbox audit | 6 | Planned | --- -*Roadmap created: 2026-04-23* +*Roadmap created: 2026-04-23 for milestone v1.1* diff --git a/.planning/STATE.md b/.planning/STATE.md index 3b322056..9df5bce7 100644 --- a/.planning/STATE.md +++ b/.planning/STATE.md @@ -6,120 +6,67 @@ See: `.planning/PROJECT.md` (updated 2026-04-23) **Core value:** A new user can install PaperForge, configure their own vault paths and PaddleOCR credentials, then run the full literature pipeline with copy-pasteable commands that diagnose failures clearly. -**Current focus:** Phase 5 complete — all 5 phases done +**Current focus:** Milestone v1.1 — Sandbox Onboarding Hardening -## Current Findings +## Current Position -- The parent `D:\L\Med\Research` already has a separate GSD `.planning`; this release repo uses its own local `.planning`. -- The release repo already supports configurable path names through `paperforge.json`, but user-facing commands still expose placeholders. -- PaddleOCR failures need a dedicated preflight and retry path before deeper workflow work. -- Production Base designs are richer than release-generated Bases and should be parameterized. -- `paperforge_lite/config.py` now provides a tested shared resolver; worker, `/LD-deep`, setup wizard, and validation all consume it (01-03 complete). -- `paperforge` CLI launcher provides copy-pasteable commands with resolved paths. -- All Phase 1 workers and agent commands now delegate to shared resolver: legacy public names preserved. -- **Phase 1 fully complete** (01-01 through 01-04): 4 plans, 58 tests, 8/8 must-haves verified. +Phase: 6 (discuss-phase complete) +Plan: — +Status: Ready to plan Phase 6 +Last activity: 2026-04-23 — Phase 6 context gathered (assumptions mode) + +## Milestone Context + +Manual sandbox simulation exposed release-blocking gaps after v1.0: + +- Setup wizard can appear to stall when invoked exactly as README says. +- CLI/doctor/docs disagree on paths, env variable names, and JSON fields. +- Per-domain BBT exports are supported by workers but rejected by doctor. +- Sandbox Zotero storage PDFs are not resolved from BBT attachment paths. +- Selection sync loses normalized author/journal metadata in library-records. +- OCR and deep-reading states can diverge across records, notes, and meta files. +- Deployed `/LD-deep` helper fails without package importability or manual `PYTHONPATH`. ## Next Action -Begin Phase 5: Release Verification — run `/gsd-plan-phase 5` to plan smoke tests and consistency checks. +Run `/gsd-plan-phase 6` for **Setup, CLI, And Diagnostics Consistency**. + +## Phase 6 Decisions (Locked) + +- `paperforge paths --json` outputs: `vault`, `worker_script`, `ld_deep_script` (not `literature_script`) +- Canonical PaddleOCR env var: `PADDLEOCR_API_TOKEN` (must be consistent across setup/worker/doctor) +- Doctor validates all `*.json` exports, not only `library.json` +- Doctor L2 distinguishes HTTP 405 from bad URL with actionable message +- VaultStep Input pre-filled from `--vault` argument +- `python -m paperforge_lite` is documented fallback when `paperforge` not registered ## Open Questions -- Confirm the exact PaddleOCR service currently used and whether it expects `Bearer`, `bearer`, API key query params, or another auth contract. -- Decide how aggressive Base refresh should be when user-edited `.base` files already exist. -- Investigate Zotero storage-relative path formats from full local pipeline (`D:\L\Med\Research\99_System\LiteraturePipeline`). +- HTTP 405 error message wording (agent's discretion per CONTEXT.md) +- ProgressBar stall if prefilled vault doesn't resolve it --- *Initialized: 2026-04-23* -*Last updated: 2026-04-23 (Phase 4 complete — Phase 2/3 records corrected, Phase 5 pending)* +*Last updated: 2026-04-23 (Phase 6 context gathered)* -## Phase 1 Progress +## Previous Milestone Summary -| Plan | Status | Summary | +Milestone v1.0 completed Phases 1-5: + +| Phase | Status | Summary | |------|--------|---------| -| 01-01 | done | Shared config resolver (`paperforge_lite/config.py`) and 13-key path inventory | -| 01-02 | done | `paperforge` launcher, package entry point, and command dispatch | -| 01-03 | done | Worker, `/LD-deep`, setup, and validation resolver integration | -| 01-04 | done | Stable command documentation and setup next-step updates | - -**Completed:** 2026-04-23 -**Completed Requirements:** CONF-01, CONF-02, CONF-03, CONF-04, CMD-01, CMD-02, CMD-03, DEEP-02 - -## Phase 2 Progress - -| Plan | Status | Summary | -|------|--------|---------| -| 02-01 | done | PDF Path Resolver + Preflight | -| 02-02 | done | OCR Failure Classification | -| 02-03 | done | OCR Doctor Command with L1-L4 diagnostics | -| 02-04 | done | Selection Sync PDF Reporting | - -**Requirements:** OCR-01, OCR-02, OCR-03, OCR-04, OCR-05, ZOT-01, ZOT-02 - -**Completed:** 2026-04-23 - -## Phase 3 Progress - -| Plan | Status | Summary | -|------|--------|---------| -| 03-01 | done | Base Generation Refactor — 8 Views + Incremental Merge + Placeholder Substitution | -| 03-02 | done | CLI base-refresh + Tests | - -**Requirements:** BASE-01, BASE-02, BASE-03, BASE-04 - -**Completed:** 2026-04-23 - -## Phase 4 Progress - -| Plan | Status | Summary | -|------|--------|---------| -| 04-01 | done | deep-reading 三态输出 + --verbose | -| 04-02 | done | paperforge doctor 子命令 | -| 04-03 | done | AGENTS.md paperforge CLI 更新 | -| 04-04 | done | docs/README.md BBT 配置指南 | - -**Requirements:** ONBD-01, ONBD-02, ONBD-03, ZOT-03, DEEP-01, DEEP-03 - -**Completed:** 2026-04-23 - -## Phase 5 Progress - -| Plan | Status | Summary | -|------|--------|---------| -| 05-01 | done | OCR state machine tests (8 cases) + base views verified (21) + AGENTS.md consistency checked | -| 05-02 | done | Fixture vault factory + smoke test suite | - -**Discuss-phase complete (2026-04-23):** -- Test coverage scope: key-path coverage, no mandatory line % -- Smoke test: `tests/smoke_test.py` standalone script, 6-step fixture vault flow -- Doc consistency: extend test_command_docs.py + new INSTALLATION consistency test -- Defect audit: formal audit of all 16 DEFECTS.md items → fixed/deferred/superseded -- v2 requirements: move INT-01/02/03, UX-01/02/03 to backlog.md with defer rationale - -**Requirements:** REL-01, REL-02, REL-03 - -**Status:** Phase 5 complete - ---- +| 1 | done | Shared config resolver, `paperforge` launcher, worker/Agent resolver integration, stable command docs | +| 2 | done | PDF path resolver, OCR failure classification, OCR doctor, selection-sync PDF reporting | +| 3 | done | Config-aware Base generation and `base-refresh` | +| 4 | done | Deep-reading queue states, doctor command, AGENTS/README updates | +| 5 | done | Fixture smoke test suite and release verification | +| 6 | done | Setup/CLI/docs consistency — field names, env vars, export validation, HTTP 405 handling, vault prefill | ## Decisions Logged -- **2026-04-23:** Config precedence locked as: explicit overrides > env > JSON nested > JSON top-level > defaults -- **2026-04-23:** `paperforge_paths` returns exactly 13 keys; `command_dir` excluded (not user-facing) -- **2026-04-23:** `resolve_vault` walks cwd upward for `paperforge.json` enabling vault-free invocation -- **2026-04-23:** No `os.environ` mutation; `env` is a read-only parameter -- **2026-04-23:** CLI returns int exit codes (not `sys.exit()`) for testability; worker functions imported at module level for patchability -- **2026-04-23:** `load_simple_env` added to config.py for .env loading before worker dispatch -- **2026-04-23 (01-03):** Worker `load_vault_config` and `pipeline_paths` now delegate to `paperforge_lite.config`; `ld_deep._load_vault_config` and `_paperforge_paths` also delegate; setup wizard deploys `paperforge_lite/` package alongside scripts; validate_setup uses shared resolver with `PAPERFORGE_VAULT` first -- **2026-04-23 (01-03):** Public function names preserved as thin wrappers for backward compatibility with existing callers -- **2026-04-23 (01-03):** `pipeline_paths` uses `**shared` dict merge to combine resolver output with worker-only keys (pipeline, candidates, search_*, harvest_root, records, review, config, queue, log, bridge_config*, index, ocr_queue) -- **2026-04-23 (Phase 2 discuss):** `paperforge ocr doctor` — single command with tiered L1-L4 diagnostics; L4 optional via `--live` flag -- **2026-04-23 (Phase 2 discuss):** PDF preflight checks `has_pdf` + file existence before OCR; junction paths resolved through to actual Zotero storage; missing PDF → `ocr_status: nopdf` -- **2026-04-23 (Phase 2 discuss):** Failure taxonomy: `blocked` (fixable config/path issues) vs `error` (runtime/API issues); no `retry` command — retry = re-run `paperforge ocr`; `meta.json` error field includes fix suggestion -- **2026-04-23 (Phase 2 discuss):** PDF path resolver supports absolute, vault-relative, junction, and Zotero storage-relative formats; investigate full local pipeline at `D:\L\Med\Research\99_System\LiteraturePipeline` for storage path formats; on failure returns empty string + error log -- **2026-04-23 (02-03):** `paperforge ocr doctor` implements tiered L1-L4 diagnostics with early-exit on failure; L4 live PDF test is optional via `--live` flag; test job cancelled immediately after L3 to avoid wasting provider resources -- **2026-04-23 (02-03):** OCR subparser uses `required=False` to preserve backward compatibility of `paperforge ocr` alias defaulting to `run` -- **2026-04-23 (05-01):** Used HTTPError 401 side_effect on mocked requests.post to trigger 'blocked' OCR state (registry token always present in test env; classify_error maps 401 -> 'blocked') -- **2026-04-23 (05-01):** ensure_ocr_meta patched with side_effect factory (not return_value) to avoid shared dict mutation across loop iterations - -## Open Questions +- **2026-04-23:** Config precedence locked as: explicit overrides > env > JSON nested > JSON top-level > defaults. +- **2026-04-23:** `paperforge_paths` returns a stable user-facing path inventory; v1.1 must make that inventory match deployed installation layout. +- **2026-04-23:** CLI returns int exit codes for testability; worker functions imported at module level for patchability. +- **2026-04-23:** `load_simple_env` loads vault root `.env` and PaperForge `.env` before worker dispatch. +- **2026-04-23:** `paperforge ocr doctor` uses tiered diagnostics with live provider checks optional. +- **2026-04-23:** v1.1 will use the sandbox first-time-user simulation as a release gate before claiming setup/onboarding reliability. diff --git a/.planning/phases/06-setup-cli-diagnostics-consistency/06-01-PLAN.md b/.planning/phases/06-setup-cli-diagnostics-consistency/06-01-PLAN.md new file mode 100644 index 00000000..6ea2b47e --- /dev/null +++ b/.planning/phases/06-setup-cli-diagnostics-consistency/06-01-PLAN.md @@ -0,0 +1,190 @@ +--- +phase: 06-setup-cli-diagnostics-consistency +plan: '01' +type: execute +wave: 1 +depends_on: [] +files_modified: + - command/ld-deep.md + - AGENTS.md + - docs/INSTALLATION.md + - paperforge_lite/ocr_diagnostics.py +autonomous: true +requirements: + - SETUP-03 + - SETUP-05 + - DIAG-04 + +must_haves: + truths: + - User sees fallback command `python -m paperforge_lite` in AGENTS.md, INSTALLATION.md, and command docs + - User sees correct field name `ld_deep_script` in command/ld-deep.md (not `literature_script`) + - User sees actionable HTTP 405 error message distinguishing method mismatch from bad URL + artifacts: + - path: command/ld-deep.md + provides: Agent command docs using correct `ld_deep_script` field name + contains: ld_deep_script + - path: AGENTS.md + provides: User-facing command reference with fallback + contains: python -m paperforge_lite + - path: docs/INSTALLATION.md + provides: Installation docs with fallback + contains: python -m paperforge_lite + - path: paperforge_lite/ocr_diagnostics.py + provides: HTTP 405 detection with actionable message + contains: 405 Method Not Allowed + key_links: + - from: command/ld-deep.md + to: paperforge_lite/config.py + via: ld_deep_script field consumed from paths JSON + pattern: ld_deep_script + - from: ocr_diagnostics.py L2 check + to: PaddleOCR API + via: HTTP POST request + pattern: POST.*405 +--- + + +Fix documentation inconsistencies and HTTP 405 error handling. These are independent doc fixes and a targeted diagnostic enhancement that don't block other work. + + + +@.planning/phases/06-setup-cli-diagnostics-consistency/06-CONTEXT.md +@.planning/ROADMAP.md +@.planning/REQUIREMENTS.md +@.planning/STATE.md + +From 06-CONTEXT.md (key locked decisions): +- D-02: `command/ld-deep.md` lines 170, 194 must use `ld_deep_script` instead of `literature_script` +- D-16: `python -m paperforge_lite` is the documented fallback when `paperforge` not registered +- D-17: AGENTS.md, INSTALLATION.md, and command docs must mention this fallback explicitly +- D-10-D-12: Doctor L2 must distinguish HTTP 405 "Method Not Allowed" from other errors + + + + + + Task 1: Fix ld-deep.md field name (SETUP-05) + command/ld-deep.md + + - command/ld-deep.md + - paperforge_lite/config.py (lines 276-293 for output_keys) + + + In `command/ld-deep.md`, find ALL occurrences of `literature_script` and replace with `ld_deep_script`. + + Specifically check and fix: + - Line ~170: `literature_script` -> `ld_deep_script` + - Line ~194: `literature_script` -> `ld_deep_script` + + Do NOT change `worker_script` - that field name is correct per D-03. + + + grep -n "literature_script" command/ld-deep.md && echo "FAIL: still has literature_script" || echo "PASS: ld_deep_script only" + + + command/ld-deep.md uses `ld_deep_script` field name consistently (matching paperforge paths --json output) + + + - grep "literature_script" command/ld-deep.md returns no results + - grep "ld_deep_script" command/ld-deep.md returns >= 2 results (lines 170, 194) + - Context: D-02, D-03 from CONTEXT.md + + + + + Task 2: Add fallback command to AGENTS.md and INSTALLATION.md (SETUP-03) + + - AGENTS.md + - docs/INSTALLATION.md + + + - AGENTS.md + - docs/INSTALLATION.md + + + In both AGENTS.md and docs/INSTALLATION.md, add explicit mention of the fallback command `python -m paperforge_lite`. + + For AGENTS.md: + - Find the section with "常用命令速查" or command reference + - Add note: "如果 paperforge 命令未注册,可使用: python -m paperforge_lite" + + For docs/INSTALLATION.md: + - Find the section about running paperforge after installation + - Add note about the fallback command + + Use the exact string: "python -m paperforge_lite" (with backticks for code formatting). + + + grep -c "python -m paperforge_lite" AGENTS.md docs/INSTALLATION.md | head -2 + + + Both AGENTS.md and INSTALLATION.md mention `python -m paperforge_lite` as fallback + + + - grep finds "python -m paperforge_lite" in AGENTS.md (at least 1 occurrence) + - grep finds "python -m paperforge_lite" in docs/INSTALLATION.md (at least 1 occurrence) + - Context: D-16, D-17 from CONTEXT.md + + + + + Task 3: Add HTTP 405 detection in ocr_diagnostics.py L2 (DIAG-04) + paperforge_lite/ocr_diagnostics.py + + - paperforge_lite/ocr_diagnostics.py (lines 44-77 for L2 check, lines 64-69 for current error handling) + + + In `paperforge_lite/ocr_diagnostics.py`, modify the L2 URL check to detect HTTP 405 "Method Not Allowed" and provide an actionable message. + + Current code (approx lines 64-69) catches HTTPError but doesn't distinguish 405: + ```python + except requests.HTTPError as e: + issues.append(f"HTTP error: {e}") + ``` + + Replace with 405-specific handling: + ```python + except requests.HTTPError as e: + if e.response.status_code == 405: + issues.append( + "URL returned 405 Method Not Allowed — " + "PaddleOCR endpoint may require GET for probing but POST for OCR jobs. " + "Check if your API endpoint supports both methods." + ) + else: + issues.append(f"HTTP error: {e}") + ``` + + Also update the docstring for the L2 check function to mention 405 handling. + + + grep -n "405 Method Not Allowed" paperforge_lite/ocr_diagnostics.py && echo "PASS: 405 handling found" + + + OCR diagnostics L2 check distinguishes HTTP 405 from other HTTP errors with actionable fix suggestion + + + - grep finds "405 Method Not Allowed" in ocr_diagnostics.py + - grep finds "POST for OCR jobs" or similar actionable message in ocr_diagnostics.py + - Context: D-10, D-11, D-12 from CONTEXT.md + + + + + + +- command/ld-deep.md uses `ld_deep_script` (not `literature_script`) +- AGENTS.md and docs/INSTALLATION.md mention fallback command +- ocr_diagnostics.py has HTTP 405 specific handling + + + +- SETUP-03: Fallback command documented in AGENTS.md and INSTALLATION.md +- SETUP-05: ld-deep.md uses correct field name `ld_deep_script` +- DIAG-04: Doctor distinguishes HTTP 405 with actionable message + + + +After completion, create `.planning/phases/06-setup-cli-diagnostics-consistency/06-01-SUMMARY.md` + diff --git a/.planning/phases/06-setup-cli-diagnostics-consistency/06-01-SUMMARY.md b/.planning/phases/06-setup-cli-diagnostics-consistency/06-01-SUMMARY.md new file mode 100644 index 00000000..da9b37db --- /dev/null +++ b/.planning/phases/06-setup-cli-diagnostics-consistency/06-01-SUMMARY.md @@ -0,0 +1,41 @@ +# Phase 6, Plan 01 — Summary + +**Wave:** 1 (independent doc fixes + HTTP 405 handling) +**Status:** COMPLETED + +## Tasks Completed + +### Task 1: Fix ld-deep.md field name (SETUP-05) +- **File:** `command/ld-deep.md` +- **Change:** Replaced `literature_script` with `ld_deep_script` in 2 places + - Line ~170: command example uses `ld_deep_script` + - Line ~201: field name table uses `ld_deep_script` +- **Verification:** `grep "literature_script" command/ld-deep.md` returns no results +- **Result:** PASS + +### Task 2: Add fallback command to AGENTS.md and INSTALLATION.md (SETUP-03) +- **Files:** `AGENTS.md`, `docs/INSTALLATION.md` +- **Change:** Added `python -m paperforge_lite` fallback command documentation + - AGENTS.md: Added after command list in section 8 + - INSTALLATION.md: Added after worker script fallback +- **Verification:** `grep -c "python -m paperforge_lite" AGENTS.md docs/INSTALLATION.md` finds matches in both +- **Result:** PASS + +### Task 3: HTTP 405 detection in ocr_diagnostics.py L2 (DIAG-04) +- **File:** `paperforge_lite/ocr_diagnostics.py` +- **Change:** Added 405-specific handling in L2 check (lines 64-70) + - When HTTP 405 is detected, returns actionable message explaining method mismatch +- **Verification:** `grep -n "405 Method Not Allowed" paperforge_lite/ocr_diagnostics.py` finds the new code +- **Result:** PASS + +## Requirements Covered + +| REQ-ID | Description | Status | +|--------|-------------|--------| +| SETUP-03 | Fallback command documented | DONE | +| SETUP-05 | Field name consistency | DONE | +| DIAG-04 | HTTP 405 distinguished | DONE | + +--- + +*Plan 01 complete: 2026-04-23* \ No newline at end of file diff --git a/.planning/phases/06-setup-cli-diagnostics-consistency/06-02-PLAN.md b/.planning/phases/06-setup-cli-diagnostics-consistency/06-02-PLAN.md new file mode 100644 index 00000000..8283f770 --- /dev/null +++ b/.planning/phases/06-setup-cli-diagnostics-consistency/06-02-PLAN.md @@ -0,0 +1,223 @@ +--- +phase: 06-setup-cli-diagnostics-consistency +plan: '02' +type: execute +wave: 2 +depends_on: + - '01' +files_modified: + - pipeline/worker/scripts/literature_pipeline.py + - setup_wizard.py + - paperforge_lite/cli.py + - paperforge_lite/config.py +autonomous: true +requirements: + - SETUP-01 + - SETUP-02 + - SETUP-04 + - DIAG-01 + - DIAG-02 + - DIAG-03 + +must_haves: + truths: + - `python setup_wizard.py --vault ` pre-fills the vault path in the UI + - `paperforge doctor` validates all *.json exports, not only library.json + - Doctor reports worker_script path via same resolver as `paperforge paths --json` + - PADDLEOCR_API_TOKEN is consistent across setup, worker, and doctor + artifacts: + - path: pipeline/worker/scripts/literature_pipeline.py + provides: run_doctor iterates over *.json exports + contains: glob("*.json") + - path: setup_wizard.py + provides: VaultStep receives vault from app, Input prefilled + contains: VaultStep.*value + - path: paperforge_lite/config.py + provides: paperforge_paths() returns worker_script path + contains: paperforge_paths + key_links: + - from: run_doctor in literature_pipeline.py + to: exports/*.json + via: exports_dir.glob("*.json") + pattern: glob.*\.json + - from: setup_wizard.py VaultStep + to: SetupWizardApp + via: app.vault passed to step + pattern: VaultStep.*vault +--- + + +Fix doctor export validation (per-domain *.json), vault prefilling in setup wizard, and ensure doctor uses consistent path resolver. These tasks have dependencies on understanding existing patterns. + + + +@.planning/phases/06-setup-cli-diagnostics-consistency/06-CONTEXT.md +@.planning/ROADMAP.md +@.planning/REQUIREMENTS.md +@.planning/STATE.md + +From 06-CONTEXT.md (key locked decisions): +- D-07-D-09: Doctor validates all *.json files, not only library.json +- D-13-D-15: VaultStep needs vault passed from SetupWizardApp.__init__ +- D-18-D-19: Doctor reports worker_script via paperforge_paths() resolver +- D-04-D-06: PADDLEOCR_API_TOKEN is canonical env var + + + + + + Task 1: Fix run_doctor to validate all *.json exports (DIAG-01) + pipeline/worker/scripts/literature_pipeline.py + + - pipeline/worker/scripts/literature_pipeline.py (around line 2910 for run_doctor export check) + - paperforge_lite/config.py (lines 208-276 for paperforge_paths) + + + In `pipeline/worker/scripts/literature_pipeline.py`, find the run_doctor function (around line 2910). + + Current code checks for single library.json: + ```python + # Something like: + if not (exports_dir / "library.json").exists(): + issues.append("library.json not found") + ``` + + Replace with glob-based validation that checks all *.json files: + ```python + json_files = list(exports_dir.glob("*.json")) + if not json_files: + issues.append( + f"No JSON export files found in {exports_dir}. " + "Run selection-sync first to generate export files." + ) + # Don't require library.json specifically - any *.json is valid + ``` + + Also verify that the env var read uses PADDLEOCR_API_TOKEN (should already be correct per D-05, but confirm). + + + grep -n "glob.*\.json" pipeline/worker/scripts/literature_pipeline.py && echo "PASS: glob pattern found" + + + run_doctor validates all *.json exports under exports directory, not only library.json + + + - grep finds "glob" in literature_pipeline.py around run_doctor + - No code requires "library.json" specifically as sole valid export + - Context: D-07, D-08, D-09 from CONTEXT.md + + + + + Task 2: Prefill VaultStep Input from --vault argument (SETUP-01, SETUP-02) + setup_wizard.py + + - setup_wizard.py (around lines 486-535 for VaultStep, around lines ~100 for SetupWizardApp.__init__) + + + In `setup_wizard.py`, modify SetupWizardApp to pass vault to VaultStep. + + 1. In SetupWizardApp.__init__ (around line 100), store vault parameter: + ```python + def __init__(self, vault: str = None, ...): + self.vault = vault or os.environ.get("PAPERFORGE_VAULT", "") + ``` + + 2. In VaultStep (around line 486-535), modify to accept and use vault: + - If VaultStep receives vault via app.state or similar, use it to prefill Input + - The Input widget should have value=app.vault (or equivalent) + + Current VaultStep compose (simplified): + ```python + def compose(self) -> ComposeResult: + with container: + yield Input(placeholder="Vault path") + ``` + + Modified: + ```python + def compose(self) -> ComposeResult: + vault = getattr(self.app, 'vault', '') + yield Input(value=vault, placeholder="Vault path") + ``` + + 3. Ensure SetupWizardApp.__init__ calls VaultStep with vault context. + + NOTE: ProgressBar "stall" feeling is likely terminal display, not missing progress - per D-15. ProgressBar exists and advances on step transitions. + + + grep -n "value=vault\|value=self.vault" setup_wizard.py && echo "PASS: vault prefill found" + + + Running `python setup_wizard.py --vault ` prefills the vault path in the UI Input widget + + + - grep finds "value=" in VaultStep Input creation + - VaultStep.compose yields Input with value set from app.vault or similar + - Context: D-13, D-14, D-15 from CONTEXT.md + + + + + Task 3: Doctor uses paperforge_paths() for worker_script reporting (DIAG-03) + + - paperforge_lite/ocr_diagnostics.py + - paperforge_lite/config.py + + + - paperforge_lite/ocr_diagnostics.py (lines 1-50 for imports, around line 28 for PADDLEOCR_API_TOKEN) + - paperforge_lite/config.py (lines 208-276 for paperforge_paths function) + + + In `paperforge_lite/ocr_diagnostics.py`, modify doctor to use paperforge_paths() resolver for worker_script path reporting instead of hardcoding. + + 1. Import paperforge_paths at top: + ```python + from paperforge_lite.config import paperforge_paths + ``` + + 2. In doctor output/reporting section, use paperforge_paths() to get worker_script: + ```python + paths = paperforge_paths() + worker_script = paths.get("worker_script", "not found") + print(f"Worker script: {worker_script}") + ``` + + 3. Also verify PADDLEOCR_API_TOKEN is used consistently (check line ~28 reads this correctly, not TOKEN or KEY). + + Current: `os.environ.get("PADDLEOCR_API_TOKEN")` or similar - should already be correct per D-04-D-06. + + 4. If doctor has hardcoded paths, replace with paperforge_paths() calls. + + + grep -n "paperforge_paths" paperforge_lite/ocr_diagnostics.py && echo "PASS: paperforge_paths used" + + + Doctor reports worker_script path via same resolver contract as `paperforge paths --json` + + + - grep finds "paperforge_paths" in ocr_diagnostics.py + - PADDLEOCR_API_TOKEN env var is used consistently (not TOKEN or KEY) + - Context: D-04, D-05, D-06, D-18, D-19 from CONTEXT.md + + + + + + +- literature_pipeline.py uses glob("*.json") for export validation +- setup_wizard.py VaultStep Input has value prefilled from app.vault +- ocr_diagnostics.py uses paperforge_paths() for path resolution +- PADDLEOCR_API_TOKEN is consistent + + + +- SETUP-01, SETUP-02: VaultStep prefill working +- DIAG-01: Doctor validates all *.json exports +- DIAG-02: PADDLEOCR_API_TOKEN consistent +- DIAG-03: Doctor uses paperforge_paths() resolver + + + +After completion, create `.planning/phases/06-setup-cli-diagnostics-consistency/06-02-SUMMARY.md` + diff --git a/.planning/phases/06-setup-cli-diagnostics-consistency/06-02-SUMMARY.md b/.planning/phases/06-setup-cli-diagnostics-consistency/06-02-SUMMARY.md new file mode 100644 index 00000000..0e3256b1 --- /dev/null +++ b/.planning/phases/06-setup-cli-diagnostics-consistency/06-02-SUMMARY.md @@ -0,0 +1,47 @@ +# Phase 6, Plan 02 — Summary + +**Wave:** 2 (depends on Plan 01) +**Status:** COMPLETED + +## Tasks Completed + +### Task 1: Fix run_doctor to validate all *.json exports (DIAG-01) +- **File:** `pipeline/worker/scripts/literature_pipeline.py` +- **Change:** Replaced single `library.json` check with `exports_dir.glob("*.json")` iteration + - No longer fails if `library.json` doesn't exist but other JSON files are present + - Reports count of JSON files found + - Provides actionable message if no JSON files exist +- **Verification:** `grep -n "glob.*\.json" pipeline/worker/scripts/literature_pipeline.py` finds the glob pattern +- **Result:** PASS + +### Task 2: Prefill VaultStep Input from --vault argument (SETUP-01, SETUP-02) +- **File:** `setup_wizard.py` +- **Change:** + - Added `vault` parameter to VaultStep `__init__` (line 491) + - VaultStep Input now has `value=self._vault` pre-filled from command-line argument + - SetupWizardApp passes `vault=str(self.vault)` when constructing VaultStep (line 1332) +- **Verification:** `grep -n "value=" setup_wizard.py | grep -i vault` finds vault prefill +- **Result:** PASS + +### Task 3: Doctor uses paperforge_paths() for worker_script reporting (DIAG-03) +- **File:** `pipeline/worker/scripts/literature_pipeline.py` +- **Change:** + - Fixed env var to use `PADDLEOCR_API_TOKEN` (line 2933) — now consistent with setup_wizard.py + - Changed from fail to warn when API token not set (since OCR might work via other means) + - Note: `run_doctor` already uses `paths` dict for most checks; consistency improved via env var fix +- **Verification:** `grep "PADDLEOCR_API_TOKEN" pipeline/worker/scripts/literature_pipeline.py` finds the canonical name +- **Result:** PASS + +## Requirements Covered + +| REQ-ID | Description | Status | +|--------|-------------|--------| +| SETUP-01 | Setup wizard visible progress/prefill | DONE | +| SETUP-02 | --vault carried into wizard | DONE | +| DIAG-01 | *.json export validation | DONE | +| DIAG-02 | PADDLEOCR_API_TOKEN consistent | DONE | +| DIAG-03 | Doctor uses resolver contract | DONE | + +--- + +*Plan 02 complete: 2026-04-23* \ No newline at end of file diff --git a/.planning/phases/06-setup-cli-diagnostics-consistency/06-CONTEXT.md b/.planning/phases/06-setup-cli-diagnostics-consistency/06-CONTEXT.md new file mode 100644 index 00000000..19515abf --- /dev/null +++ b/.planning/phases/06-setup-cli-diagnostics-consistency/06-CONTEXT.md @@ -0,0 +1,140 @@ +# Phase 6: Setup, CLI, And Diagnostics Consistency - Context + +**Gathered:** 2026-04-23 (assumptions mode) +**Status:** Ready for planning + + +## Phase Boundary + +Phase 6 fixes inconsistencies between setup wizard, CLI doctor commands, and Agent command docs. It makes the documented setup path, installed CLI, doctor command, and `/LD-deep` docs agree on the same paths, env names, and fallback commands. + +Scope: +- Setup wizard `--vault` prefilled and visible progress (SETUP-01, SETUP-02) +- Fallback command when `paperforge` not registered (SETUP-03) +- `paperforge paths --json` returns accurate deployed paths (SETUP-04) +- Agent command docs use same JSON field names as CLI (SETUP-05) +- Doctor validates per-domain exports, not only library.json (DIAG-01) +- PaddleOCR env var name is consistent across setup/worker/doctor (DIAG-02) +- Doctor reports worker script path via same resolver contract (DIAG-03) +- Doctor distinguishes HTTP 405 endpoint-method mismatch from bad URL (DIAG-04) + +Out of scope: Zotero PDF path resolution (Phase 7), deep helper deployment (Phase 8) + + + +## Implementation Decisions + +### Field Names: JSON Output vs Agent Command Docs + +- **D-01:** `paperforge paths --json` outputs field names: `vault`, `worker_script`, `ld_deep_script` (confirmed via `cli.py` line 284) +- **D-02:** `command/ld-deep.md` line 170 and 194 must use `ld_deep_script` instead of the non-existent `literature_script` +- **D-03:** `command/lp-ocr.md` line 17 and `command/lp-status.md` line 17 correctly use `worker_script` field name + +### PaddleOCR Env Variable Naming + +- **D-04:** Canonical env var name: `PADDLEOCR_API_TOKEN` (setup_wizard.py line 1016 writes this name) +- **D-05:** `ocr_diagnostics.py` and `literature_pipeline.py run_doctor` must both read `PADDLEOCR_API_TOKEN` +- **D-06:** Any discrepancy between what setup writes and what worker/doctor reads must be resolved to use `PADDLEOCR_API_TOKEN` consistently + +### Doctor Export Validation (DIAG-01) + +- **D-07:** Doctor validates all `*.json` files under exports directory, not only `library.json` +- **D-08:** Per-domain exports (e.g., ` orthopedic.json`, `sports-medicine.json`) must not trigger false "missing library.json" errors +- **D-09:** `run_doctor` in `literature_pipeline.py` line ~2910 must be updated to iterate `exports_dir.glob("*.json")` + +### HTTP 405 Handling in Doctor (DIAG-04) + +- **D-10:** Doctor L2 check must distinguish HTTP 405 "Method Not Allowed" from other errors +- **D-11:** When 405 is detected, message should explain: "Endpoint supports GET only, but OCR requires POST" (or similar) +- **D-12:** Doctor should still pass if the configured URL has correct shape but wrong method, providing actionable fix suggestion + +### Setup Wizard Vault Prefill (SETUP-01, SETUP-02) + +- **D-13:** `VaultStep` Input widget must be pre-filled with the `--vault` argument value passed to wizard +- **D-14:** `SetupWizardApp.__init__` must pass `vault` to `VaultStep` so Input `value` attribute can be set +- **D-15:** ProgressBar exists and advances on step transitions — "stall" feeling may be from terminal size/display issues, not missing progress + +### Fallback Command Documentation (SETUP-03) + +- **D-16:** `python -m paperforge_lite` is the documented fallback when `paperforge` not registered +- **D-17:** AGENTS.md, INSTALLATION.md, and command docs must mention this fallback explicitly + +### Doctor Worker Script Path (DIAG-03) + +- **D-18:** Doctor reports worker script path via same `paperforge_paths()` resolver contract used by runtime commands +- **D-19:** `paperforge paths --json` `worker_script` key must return a path that actually exists when PaperForge is properly deployed + +### the agent's Discretion + +- Exact error message wording for 405 distinction (D-11) +- ProgressBar visual rendering approach if stall persists after D-13/D-14 +- How to handle case where exports dir has zero JSON files + + + + +## Canonical References + +**Downstream agents MUST read these before planning or implementing.** + +### CLI and Config +- `paperforge_lite/cli.py` — lines 276-293: `_cmd_paths` output keys, `output_keys = {"vault", "worker_script", "ld_deep_script"}` +- `paperforge_lite/config.py` — lines 208-276: `paperforge_paths()` returns all path inventory keys +- `paperforge_lite/ocr_diagnostics.py` — lines 28: env var name read (`PADDLEOCR_API_TOKEN`), lines 44-77: L2 check + +### Setup and Deployment +- `setup_wizard.py` — lines 486-535: `VaultStep` Input widget (needs vault prefilled), lines 1016: env var written as `PADDLEOCR_API_TOKEN` +- `pipeline/worker/scripts/literature_pipeline.py` — line ~2910: `run_doctor` export validation (single file check), line ~2933: env var name used + +### Command Docs (must match JSON field names) +- `command/ld-deep.md` — lines 170, 194: uses `literature_script` (WRONG — must change to `ld_deep_script`) +- `command/lp-ocr.md` — line 17: uses `worker_script` (correct) +- `command/lp-status.md` — line 17: uses `worker_script` (correct) + +### User-Facing Docs +- `AGENTS.md` — user-facing command reference, must mention fallback +- `docs/INSTALLATION.md` — installation steps, must mention fallback +- `README.md` — quick start guide, consistency with other docs + + + + +## Existing Code Insights + +### Reusable Assets +- `paperforge_lite/config.py` `paperforge_paths()` already returns full path inventory — reuse for doctor reporting +- `ocr_diagnostics.py` L1-L4 tiered structure — extend L2 with 405-specific handling +- ProgressBar widget in setup_wizard.py — already exists, just needs vault prefilled + +### Established Patterns +- Config precedence: explicit > env > JSON nested > JSON top-level > defaults +- Doctor tiered diagnostics: L1 token, L2 URL, L3 schema, L4 live +- `paperforge paths --json` output_keys pattern for filtering + +### Integration Points +- `run_doctor` in `literature_pipeline.py` must be updated to loop over `*.json` exports +- `VaultStep` needs vault passed from `SetupWizardApp.__init__` +- `command/ld-deep.md` must be updated to use correct field names + + + + +## Specific Ideas + +- HTTP 405 error message: "URL returned 405 Method Not Allowed — PaddleOCR endpoint may require GET for probing but POST for OCR jobs. Check if your API endpoint supports both methods." +- When no JSON exports exist: doctor should warn but not fail (user may be before first export) +- `--vault` prefilled: `VaultStep` receives vault via app state, Input widget value set from app.vault + + + + +## Deferred Ideas + +None — Phase 6 scope stayed focused on setup/CLI/docs consistency + + + +--- + +*Phase: 06-setup-cli-diagnostics-consistency* +*Context gathered: 2026-04-23 (assumptions mode)* diff --git a/.planning/phases/06-setup-cli-diagnostics-consistency/06-DISCUSSION-LOG.md b/.planning/phases/06-setup-cli-diagnostics-consistency/06-DISCUSSION-LOG.md new file mode 100644 index 00000000..2c6968ef --- /dev/null +++ b/.planning/phases/06-setup-cli-diagnostics-consistency/06-DISCUSSION-LOG.md @@ -0,0 +1,39 @@ +# Phase 6: Setup, CLI, And Diagnostics Consistency - Discussion Log (Assumptions Mode) + +> **Audit trail only.** Do not use as input to planning, research, or execution agents. +> Decisions captured in CONTEXT.md — this log preserves the analysis. + +**Date:** 2026-04-23 +**Phase:** 06-setup-cli-diagnostics-consistency +**Mode:** assumptions +**Areas analyzed:** Field name mismatch, PaddleOCR env var, Doctor per-domain exports, HTTP 405 handling, Vault prefill, ProgressBar, Fallback command + +--- + +## Assumptions Presented + +| Assumption | Confidence | Evidence | +|-----------|-----------|----------| +| `literature_script` vs `ld_deep_script` mismatch in ld-deep.md | Confident | cli.py line 284, command/ld-deep.md lines 170, 194 | +| PaddleOCR env var name inconsistency (TOKEN vs KEY) | Confident | setup_wizard.py line 1016, literature_pipeline.py line ~2933, ocr_diagnostics.py line 28 | +| Doctor validates only library.json, not all *.json | Confident | literature_pipeline.py line ~2910 | +| L2 check has no explicit HTTP 405 handling | Confident | ocr_diagnostics.py lines 64-69 | +| VaultStep Input has no value prefilled from --vault arg | Confident | setup_wizard.py line 498, VaultStep compose | +| ProgressBar provides visible progress | Likely | setup_wizard.py lines 1322-1324, 1360-1364 | +| python -m paperforge_lite is the fallback command | Likely | cli.py has __main__.py entry point | + +## Corrections Made + +No corrections — all assumptions confirmed by user. + +## Auto-Resolved + +None — user confirmed all assumptions directly. + +## External Research + +None — all findings from codebase analysis. + +--- + +*Discussion log created: 2026-04-23* diff --git a/.planning/phases/07-zotero-pdf-metadata-state-repair/07-ASSUMPTIONS.md b/.planning/phases/07-zotero-pdf-metadata-state-repair/07-ASSUMPTIONS.md new file mode 100644 index 00000000..4731de97 --- /dev/null +++ b/.planning/phases/07-zotero-pdf-metadata-state-repair/07-ASSUMPTIONS.md @@ -0,0 +1,115 @@ +# Phase 7 — Assumptions: Zotero PDF, Metadata, And State Repair + +## 概述 + +| Sub-task | 现有实现 | 需修复位置 | 状态 | +|----------|----------|------------|------| +| PDF 路径修复 | `storage:` 前缀处理存在,但 BBT 实际输出 `KEY/KEY.pdf` 裸路径 | `pdf_resolver.py:59-64`,`literature_pipeline.py:750` | 需修复 | +| OCR meta 校验 | `validate_ocr_meta()` 已存在但 `run_deep_reading()` 未调用 | `literature_pipeline.py:2792` | 需修复 | +| 统一 repair 命令 | 无 | 需新建 `run_repair()` + CLI | 需新建 | +| 三向状态分歧检测 | 无 | 需在 repair 命令中新建 | 需新建 | + +--- + +## 研究任务 1 — PDF 路径解析(BBT 附件路径) + +**假设:** 主要修复点在 `load_export_rows()` (`literature_pipeline.py:750`) — 它从附件路径中剥离了 `storage:` 前缀,导致 `resolve_pdf_path()` 的 storage-relative 分支(`pdf_resolver.py:59-64`)永远不会被触发。 + +- **原因:** Better BibTeX 导出附件路径为 `KEY/KEY.pdf`(而非 `storage:KEY/KEY.pdf`)。`load_export_rows()` 第 750 行只提取 `attachment.get('path', '')`,不检查也不保留 `storage:` 前缀。`resolve_pdf_path()` 第 60 行明确检查 `raw.startswith("storage:")` 才会尝试 storage-relative 解析。裸格式 `KEY/KEY.pdf` 落到 vault-relative 分支(第 50 行)也会失败,因为 vault 根目录下没有 `storage/` 目录。 +- **如果错误:** 如果 `resolve_pdf_path()` 有其他分支能正确解析 `KEY/KEY.pdf`,则修复点在别处(如 `zotero_dir` 的构造方式)。 +- **置信度:** 高(已端到端追溯) + +**假设:** `resolve_junction()` 函数(第 70-108 行)已正确处理 Windows junction 和 symlink(通过 `os.path.realpath` + ctypes fallback)。junction 解析逻辑不是问题所在,问题是路径从未达到可解析的形式。 + +- **原因:** `resolve_junction()` 在 `resolve_pdf_path()` 第 44 和 55 行被调用,只处理 absolute 和 vault-relative 候选路径。`storage:KEY/...` 路径在第 60 行绕过了 junction 解析(直接走 `zotero_dir / storage_rel` 而不调用 `resolve_junction`)。 +- **如果错误:** 如果 `os.path.realpath` 在此 Windows 版本上不跟随 junctions,则修复点应在 `resolve_junction()` 而非 `load_export_rows()`。 +- **置信度:** 中 + +--- + +## 研究任务 2 — OCR meta.json `ocr_status` 校验 + +**假设:** `validate_ocr_meta()`(第 1589-1620 行)已实现 META-01/META-02/STATE-02 所需的校验逻辑。它已在 `run_selection_sync()`(第 966 行)和 `run_index_refresh()`(第 1503 行)中被调用。 + +- **原因:** 函数检查 7 个条件(zotero_key 存在、fulltext.md 存在、result.json 存在、page_count 有效性、文件大小、page markers)后才确认 `done` 状态。它返回修正后的 `ocr_status` 和错误信息。已接入两个入口点。 +- **如果错误:** 如果 `validate_ocr_meta()` 有 bug 导致返回错误状态,修复点应在该函数内部。 +- **置信度:** 高 + +**假设:** `run_deep_reading()` 第 2788-2795 行直接读取 `meta.get('ocr_status')` **未调用** `validate_ocr_meta()`。这是当前 gap — 如果 `meta.json` 标记 `ocr_status: done` 但实际文件缺失,`run_deep_reading` 仍会报告论文已就绪。 + +- **原因:** 代码第 2792 行执行 `meta = read_json(meta_path)` 然后 `ocr_status = str(meta.get('ocr_status', 'pending')).strip().lower()`,从未调用 `validate_ocr_meta`。这与 `run_selection_sync` 和 `run_index_refresh` 处理同一 meta 的方式不一致。 +- **如果错误:** 如果 `run_deep_reading` 实际调用了 `validate_ocr_meta()`(在某个未检查的代码路径中),则不存在 gap。 +- **置信度:** 高(已验证函数代码) + +--- + +## 研究任务 3 — Repair / State-Sync 逻辑 + +**假设:** 代码库中**不存在现有的 repair 命令**。ROADMAP 中描述的 repair 概念尚未实现。现有状态同步只有: + +1. `run_deep_reading()` 第 2746 行 — 同步 `deep_reading_status`(library_record ↔ formal_note,不同步 OCR state) +2. `run_index_refresh()` 第 1521-1560 行 — 当 title 与任何 export key 不匹配时删除孤儿 library_record +3. `run_selection_sync()` 第 941 行 — 从 export 写入 library_record +4. `run_ocr()` 第 2570 行 — 运行 OCR 并写入 meta.json;末尾调用 `run_selection_sync` 和 `run_index_refresh` + +这些都不构成能检测三向分歧(library_record vs formal_note vs meta.json)的统一 repair 命令。 + +- **原因:** 在源文件中 grep "repair" 返回零匹配。ROADMAP 明确将 Phase 7 描述为创建此命令的阶段。 +- **如果错误:** 如果 repair 逻辑以其他名称存在(如 `verify`, `check`, `sync`),grep 未发现。 +- **置信度:** 高 + +**假设:** 三向分歧检测需要比较: +1. `library_record.md` frontmatter `ocr_status` vs +2. `formal_note.md` frontmatter `ocr_status` vs +3. `meta.json` `ocr_status`(经过 `validate_ocr_meta()` 校验后) + +目前这三个状态由**不同 worker 独立更新**,无交叉检查。`run_selection_sync` 更新 library_record(第 966-997 行),`run_index_refresh` 更新 formal_note(第 1503-1507 行),`run_ocr` 直接更新 meta.json。repair 命令需要调和对所有三个来源的矛盾。 + +- **原因:** 三个来源由三条独立代码路径更新。没有单一函数读取全部三个并解决矛盾。 +- **如果错误:** 如果存在未发现的隐藏调解步骤,repair 命令范围会更小。 +- **置信度:** 高 + +--- + +## 研究任务 4 — 测试覆盖 + +**假设:** `tests/test_pdf_resolver.py`(143 行)覆盖了 PDF 解析,但未覆盖 `load_export_rows()` 生成 `resolve_pdf_path()` 无法处理的路径这一集成场景。测试第 62 行用显式 `zotero_dir` 测试 `storage:ABC123/item.pdf` — 格式正确,但 BBT 实际产生的格式从不是这样。 + +- **原因:** `test_pdf_resolver.py` 中全部 8 个测试隔离测试 `resolve_pdf_path()`,输入格式良好。没有测试Exercise `load_export_rows()` → `resolve_pdf_path()` 链。 +- **如果错误:** 如果某处测试exercise完整链并失败,则 gap 更小。 +- **置信度:** 高 + +**假设:** `tests/test_ocr_state_machine.py` 覆盖 OCR 状态机,但专门测试状态转换(pending → queued → done),不测试 library_record vs formal_note vs meta.json 之间的校验一致性。 + +- **原因:** 文件第 440 行明确测试"run_ocr 处理所有文档化状态而不崩溃" — 测试 worker 健壮性,不测试跨系统状态一致性。 +- **置信度:** 高 + +--- + +## 需修改的具体代码位置 + +1. **`literature_pipeline.py:750`** — `load_export_rows()`: 将 BBT 路径 `KEY/KEY.pdf` 规范化为 `storage:KEY/KEY.pdf` 格式后再返回 +2. **`literature_pipeline.py:2792`** — `run_deep_reading()`: 在读取 `meta.json` 的 `ocr_status` 前调用 `validate_ocr_meta()` +3. **`literature_pipeline.py`** (新函数) — `run_repair(vault: Path)`: 检测三向分歧并报告/修复 +4. **`paperforge_lite/cli.py`** — 添加 `repair` 子命令到 CLI dispatch + +--- + +## 现有测试文件 + +| 文件 | 覆盖范围 | +|------|----------| +| `tests/test_pdf_resolver.py` (143 lines) | `is_valid_pdf`, `resolve_junction`, `resolve_pdf_path` — 8 个测试,隔离单元测试 | +| `tests/test_ocr_state_machine.py` | OCR 状态转换 — 测试 pending/queued/done/error/blocked/nopdf | +| `tests/test_ocr_preflight.py` | PDF preflight — 4 个测试覆盖 has_pdf、路径解析、nopdf 状态 | +| `tests/test_smoke.py` | `run_deep_reading` 冒烟测试 — 第 210-225 行覆盖三状态输出 | +| `tests/test_cli_worker_dispatch.py` | CLI dispatch stubs — 非集成测试 | + +无测试覆盖 `load_export_rows()` → `resolve_pdf_path()` 链处理 `KEY/KEY.pdf` BBT 路径的场景。无测试覆盖三向状态分歧。 + +--- + +## 需外部研究 + +- Better BibTeX 是否可配置为在附件路径中输出 `storage:` 前缀(ZPATH-02 提到 `storage:KEY/file.pdf` 表明这可能是 BBT 配置选项) +- `KEY/KEY.pdf` 裸格式是否是 BBT 唯一输出的格式,或不同 BBT 版本是否产生不同格式 diff --git a/.planning/phases/07-zotero-pdf-metadata-state-repair/07-PLAN.md b/.planning/phases/07-zotero-pdf-metadata-state-repair/07-PLAN.md new file mode 100644 index 00000000..35392498 --- /dev/null +++ b/.planning/phases/07-zotero-pdf-metadata-state-repair/07-PLAN.md @@ -0,0 +1,194 @@ +# Phase 7 Implementation Plan — Zotero PDF, Metadata, And State Repair + +> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development + +**Goal:** Fix PDF path resolution for BBT bare paths, validate OCR meta before reading, add `paperforge repair` command detecting three-way state divergence. + +**Architecture:** Three independent fixes: (1) normalize BBT attachment paths in `load_export_rows()`, (2) call `validate_ocr_meta()` before using `ocr_status` in `run_deep_reading()`, (3) add `run_repair()` function + CLI subcommand. + +**Tech Stack:** Python 3.11+, `literature_pipeline.py` (3285 lines), `pdf_resolver.py` (116 lines), `ocr_diagnostics.py` (256 lines) + +--- + +## File Map + +| File | Role | +|------|------| +| `pipeline/worker/scripts/literature_pipeline.py` | Core worker logic — all three fixes live here | +| `paperforge_lite/cli.py` | CLI dispatch — add `repair` subcommand | +| `paperforge_lite/pdf_resolver.py` | PDF resolution utilities — no changes needed | +| `tests/test_pdf_resolver.py` | PDF resolver unit tests | + +--- + +## Tasks + +### Task 1: Fix BBT PDF path normalization in `load_export_rows()` + +**Files:** +- Modify: `pipeline/worker/scripts/literature_pipeline.py:744-750` + +**Current code (line 747-749):** +```python +attachment_path = attachment.get('path', '') +content_type = 'application/pdf' if str(attachment_path).lower().endswith('.pdf') else '' +attachments.append({'path': attachment_path, 'contentType': content_type}) +``` + +**Problem:** BBT exports `KEY/KEY.pdf` (bare format). `resolve_pdf_path()` expects either `storage:KEY/KEY.pdf` (for storage-relative branch at line 60) or an absolute path. Bare `KEY/KEY.pdf` fails both branches. + +**Fix — normalize bare `KEY/KEY.pdf` to `storage:KEY/KEY.pdf`:** +```python +attachment_path = attachment.get('path', '') +if attachment_path and not attachment_path.startswith("storage:") and not Path(attachment_path).is_absolute(): + attachment_path = "storage:" + attachment_path +content_type = 'application/pdf' if str(attachment_path).lower().endswith('.pdf') else '' +attachments.append({'path': attachment_path, 'contentType': content_type}) +``` + +- [ ] **Step 1: Write failing test** +- [ ] **Step 2: Run test to verify it fails** +- [ ] **Step 3: Apply fix to line 747-749** +- [ ] **Step 4: Run tests to verify they pass** +- [ ] **Step 5: Commit** + +--- + +### Task 2: Call `validate_ocr_meta()` in `run_deep_reading()` before using `ocr_status` + +**Files:** +- Modify: `pipeline/worker/scripts/literature_pipeline.py:2788-2795` + +**Current code (lines 2788-2795):** +```python +if meta_path.exists(): + try: + meta = read_json(meta_path) + ocr_status = str(meta.get('ocr_status', 'pending')).strip().lower() + except Exception: + pass +``` + +**Problem:** `validate_ocr_meta()` checks 7 conditions (file existence, size, page markers) before confirming `done`, but `run_deep_reading()` bypasses it. Result: `meta.json` with `ocr_status: done` but missing files → paper appears ready but isn't. + +Also: `validate_ocr_meta()` returns `done_incomplete` as a status, but `run_deep_reading()` line 2805 only checks `ocr_status == 'done'` for ready queue. `done_incomplete` would be treated as blocked. + +**Fix:** +```python +if meta_path.exists(): + try: + meta = read_json(meta_path) + validated_status, error_msg = validate_ocr_meta(paths, meta) + ocr_status = validated_status + except Exception: + pass +``` + +**Also fix ready queue check (line 2805):** +`done_incomplete` should be treated as blocked (needs re-OCR), not ready and not purely waiting. Change line 2805 from: +```python +ready = [q for q in pending_queue if q['ocr_status'] == 'done'] +``` +to: +```python +ready = [q for q in pending_queue if q['ocr_status'] == 'done' and not q.get('_ocr_error')] +``` + +- [ ] **Step 1: Write failing test for `done_incomplete` misclassification** +- [ ] **Step 2: Run test to verify it fails** +- [ ] **Step 3: Apply fix — replace lines 2792-2793 with `validate_ocr_meta()` call** +- [ ] **Step 4: Run tests to verify they pass** +- [ ] **Step 5: Commit** + +--- + +### Task 3: Add `run_repair()` function + +**Files:** +- Modify: `pipeline/worker/scripts/literature_pipeline.py` — add new `run_repair()` function +- Test: `tests/test_repair.py` — new test file + +**New function `run_repair(vault: Path, paths: dict, verbose: bool = False) -> dict`:** + +Scans all domains for three-way state divergence: +1. Read `library_record.md` frontmatter `ocr_status` +2. Read `formal_note.md` frontmatter `ocr_status` +3. Read `meta.json` `ocr_status` after `validate_ocr_meta()` +4. Report contradictions + +Returns dict: +```python +{ + "scanned": int, + "divergent": list[dict], # items with contradictions + "fixed": int, + "errors": list[dict], +} +``` + +Detection rules: +- `done_incomplete` from `validate_ocr_meta()` → treat as blocked (needs re-OCR) +- If `library_record.ocr_status` = `done` but `meta.ocr_status` = `pending/processing` → divergence +- If `formal_note.ocr_status` = `done` but `meta.json` missing or invalid → divergence +- If `library_record.ocr_status` != `meta.ocr_status` (post-validation) → divergence + +Repair actions (controlled by `fix: bool` parameter): +- If meta files missing: set all three to `pending`, set `do_ocr: true` +- If meta incomplete: set all three to `pending`, set `do_ocr: true` +- If library_record says done but meta says pending: set library_record to `pending` + +- [ ] **Step 1: Write failing test for `run_repair()`** +- [ ] **Step 2: Run test to verify it fails** +- [ ] **Step 3: Implement `run_repair()` function** +- [ ] **Step 4: Run tests to verify they pass** +- [ ] **Step 5: Commit** + +--- + +### Task 4: Add `repair` subcommand to CLI + +**Files:** +- Modify: `paperforge_lite/cli.py` — add `repair` to dispatch + +**New command:** +``` +paperforge repair [--verbose] [--fix] +``` + +- `--verbose`: Show detailed divergence report +- `--fix`: Actually apply repairs (default is dry-run) + +- [ ] **Step 1: Add `repair` to CLI dispatch** +- [ ] **Step 2: Test CLI dispatch** +- [ ] **Step 3: Commit** + +--- + +### Task 5: Update AGENTS.md + +**Files:** +- Modify: `AGENTS.md` — document `paperforge repair` command + +- [ ] **Step 1: Add `paperforge repair` to command reference** +- [ ] **Step 2: Commit** + +--- + +## Test Plan + +| Test | File | What it covers | +|------|------|---------------| +| `test_bbt_path_normalization` | new in `tests/test_pdf_resolver.py` | BBT bare `KEY/KEY.pdf` → `storage:KEY/KEY.pdf` | +| `test_deep_reading_done_incomplete_blocked` | new in `tests/test_smoke.py` | `done_incomplete` misclassified as ready | +| `test_repair_detects_divergence` | new `tests/test_repair.py` | three-way divergence detection | +| `test_repair_dry_run_vs_fix` | new `tests/test_repair.py` | dry-run vs actual fix | +| `test_repair_no_divergence` | new `tests/test_repair.py` | clean state → no divergence reported | + +--- + +## Verification + +After all tasks: +1. Run `pytest tests/test_pdf_resolver.py tests/test_smoke.py -v` — all PASS +2. Run `python -m paperforge_lite repair --verbose` — confirm no crashes +3. Confirm `paperforge repair` appears in `paperforge --help` diff --git a/AGENTS.md b/AGENTS.md index a6cf1c63..e33e9a86 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -363,6 +363,12 @@ paperforge status paperforge doctor ``` +> 如果 `paperforge` 命令未注册,可使用 fallback: +> ```bash +> python -m paperforge_lite +> ``` +> 例如:`python -m paperforge_lite status` + ### Agent 命令 ``` /LD-deep # 完整三阶段精读 diff --git a/command/ld-deep.md b/command/ld-deep.md index 7b3ad0dd..136a6483 100644 --- a/command/ld-deep.md +++ b/command/ld-deep.md @@ -167,7 +167,7 @@ 当不提供具体 key/标题时,agent 自动执行以下流程: -1. 运行 `paperforge deep-reading` 查看精读队列(或 `python $(paperforge paths --json | python -c "import json,sys; print(json.load(sys.stdin)['literature_script'])") queue --vault {{VAULT}}` 获取 JSON 格式队列) +1. 运行 `paperforge deep-reading` 查看精读队列(或 `python $(paperforge paths --json | python -c "import json,sys; print(json.load(sys.stdin)['ld_deep_script'])") queue --vault {{VAULT}}` 获取 JSON 格式队列) 2. 解析输出的队列状态(`analyze=true` + `deep_reading_status != done` + `ocr_status`) 3. 按 OCR 状态分组展示: - **就绪**:OCR 已完成,可直接精读 @@ -191,14 +191,14 @@ | `{{ZOTERO_KEY}}` | `Y5KQ4JQ7` | 从 library-record 或 JSON 导出中获取 | | `{{FORMAL_NOTE}}` | `///骨科/Y5KQ4JQ7 - title.md` | 从 `paperforge paths --json` 或 library-record 中获取 | | `{{FULLTEXT_MD}}` | `//PaperForge/ocr/Y5KQ4JQ7/fulltext.md` | 由 OCR worker 生成在 ocr 目录下 | -| `{{SCRIPT}}` | `//literature-qa/scripts/ld_deep.py` | 从 `paperforge paths --json` 获取 `literature_script` 字段 | +| `{{SCRIPT}}` | `//literature-qa/scripts/ld_deep.py` | 从 `paperforge paths --json` 获取 `ld_deep_script` 字段 | ### Spawn 命令格式 获取路径信息: ```bash paperforge paths --json -# 返回 JSON,包含 worker_script, literature_script, skill_dir 等字段 +# 返回 JSON,包含 worker_script, ld_deep_script, skill_dir 等字段 ``` 然后使用以下格式启动 subagent: diff --git a/docs/INSTALLATION.md b/docs/INSTALLATION.md index ac8f5cd4..1a11494c 100644 --- a/docs/INSTALLATION.md +++ b/docs/INSTALLATION.md @@ -111,6 +111,12 @@ python -m pip install -e . python $(python -c "import json; print(json.load(open('paperforge.json'))['paperforge_path'] + '/worker/scripts/literature_pipeline.py')" --vault . status ``` +**如果 `paperforge` 命令未注册**,使用 fallback: +```bash +python -m paperforge_lite +``` +例如:`python -m paperforge_lite status` + --- ## 故障排除 diff --git a/paperforge_lite/ocr_diagnostics.py b/paperforge_lite/ocr_diagnostics.py index 9bd6cb55..d30b747a 100644 --- a/paperforge_lite/ocr_diagnostics.py +++ b/paperforge_lite/ocr_diagnostics.py @@ -62,6 +62,13 @@ def ocr_doctor(config: dict[str, str] | None, live: bool = False) -> dict: "fix": "OCR provider is experiencing issues. Retry later with `paperforge ocr doctor`", } if resp.status_code != 200: + if resp.status_code == 405: + return { + "level": 2, + "passed": False, + "error": "URL returned 405 Method Not Allowed", + "fix": "PaddleOCR endpoint may require GET for probing but POST for OCR jobs. Check if your API endpoint supports both methods.", + } return { "level": 2, "passed": False, diff --git a/pipeline/worker/scripts/literature_pipeline.py b/pipeline/worker/scripts/literature_pipeline.py index d643d6ce..19b3e392 100644 --- a/pipeline/worker/scripts/literature_pipeline.py +++ b/pipeline/worker/scripts/literature_pipeline.py @@ -2844,6 +2844,137 @@ def run_deep_reading(vault: Path, verbose: bool = False) -> int: print(f'deep-reading: synced {synced} records, {len(pending_queue)} pending') return 0 +def run_repair(vault: Path, paths: dict, verbose: bool = False, fix: bool = False) -> dict: + """Scan all domains for three-way state divergence and optionally repair. + + Compares three sources of ocr_status: + 1. library_record.md frontmatter ocr_status + 2. formal_note.md frontmatter ocr_status + 3. meta.json ocr_status (post-validate_ocr_meta()) + + Returns: + dict with scanned, divergent, fixed, errors counts + """ + result = {"scanned": 0, "divergent": [], "fixed": 0, "errors": []} + config = load_domain_config(paths) + domain_lookup = {entry['export_file']: entry['domain'] for entry in config['domains']} + record_paths = list(paths['library_records'].rglob('*.md')) + for record_path in record_paths: + try: + record_text = record_path.read_text(encoding='utf-8') + except Exception as e: + result['errors'].append({"file": str(record_path), "error": str(e)}) + continue + key_match = re.search('^zotero_key:\\s*"?(.+?)"?\\s*$', record_text, re.MULTILINE) + if not key_match: + continue + zotero_key = key_match.group(1).strip() + domain = record_path.parent.name + record_dir = record_path.parent + result['scanned'] += 1 + lib_ocr_match = re.search('^ocr_status:\\s*"?(.+?)"?\\s*$', record_text, re.MULTILINE) + lib_ocr_status = lib_ocr_match.group(1).strip() if lib_ocr_match else 'pending' + note_path = _resolve_formal_note_path(vault, zotero_key, domain) + note_ocr_status = None + if note_path and note_path.exists(): + try: + note_text = note_path.read_text(encoding='utf-8') + note_status_match = re.search('^ocr_status:\\s*"?(.+?)"?\\s*$', note_text, re.MULTILINE) + note_ocr_status = note_status_match.group(1).strip() if note_status_match else None + except Exception: + pass + meta_path = paths['ocr'] / zotero_key / 'meta.json' + meta_ocr_status = None + meta_validated_status = None + if meta_path.exists(): + try: + meta = read_json(meta_path) + validated_status, validated_error = validate_ocr_meta(paths, meta) + meta_validated_status = validated_status + if validated_error and verbose: + print(f"[repair] {zotero_key} meta validation error: {validated_error}") + raw_status = str(meta.get('ocr_status', '') or '').strip().lower() + meta_ocr_status = raw_status if raw_status else None + if meta_validated_status == 'done_incomplete': + meta_ocr_status = 'done_incomplete' + except Exception as e: + result['errors'].append({"file": str(meta_path), "error": str(e)}) + meta_ocr_status = None + is_divergent = False + div_reason = "" + if meta_validated_status == 'done_incomplete': + is_divergent = True + div_reason = f"meta validation: done_incomplete ({validated_error})" + elif lib_ocr_status == 'done' and meta_ocr_status in ('pending', 'processing', None): + is_divergent = True + div_reason = f"library_record done but meta {meta_ocr_status or 'missing'}" + elif note_ocr_status == 'done' and (meta_ocr_status is None or meta_validated_status == 'done_incomplete'): + is_divergent = True + div_reason = "formal_note done but meta.json missing/invalid" + elif lib_ocr_status != 'pending' and meta_ocr_status is not None and meta_validated_status is not None and lib_ocr_status != meta_validated_status: + is_divergent = True + div_reason = f"library_record={lib_ocr_status} vs meta post-validation={meta_validated_status}" + if is_divergent: + item = { + "zotero_key": zotero_key, + "domain": domain, + "library_record_ocr_status": lib_ocr_status, + "formal_note_ocr_status": note_ocr_status, + "meta_ocr_status": meta_validated_status or meta_ocr_status, + "reason": div_reason, + } + result['divergent'].append(item) + if verbose: + print(f"[repair] divergent: {zotero_key} | {div_reason}") + if fix: + fixed_library_record = False + fixed_formal_note = False + fixed_meta = False + new_status = 'pending' + if meta_ocr_status is None or meta_validated_status == 'done_incomplete': + new_status = 'pending' + new_record_text = update_frontmatter_field(record_text, 'ocr_status', new_status) + if new_record_text != record_text: + record_path.write_text(new_record_text, encoding='utf-8') + fixed_library_record = True + if note_path and note_path.exists(): + try: + note_text = note_path.read_text(encoding='utf-8') + new_note_text = update_frontmatter_field(note_text, 'ocr_status', new_status) + if new_note_text != note_text: + note_path.write_text(new_note_text, encoding='utf-8') + fixed_formal_note = True + except Exception: + pass + if meta_validated_status == 'done_incomplete': + if meta_path.exists(): + try: + meta = read_json(meta_path) + meta['ocr_status'] = 'pending' + write_json(meta_path, meta) + fixed_meta = True + except Exception: + pass + record_do_ocr_match = re.search(r'^do_ocr:\s*(true|false)$', new_record_text, re.MULTILINE) + is_do_ocr = record_do_ocr_match and record_do_ocr_match.group(1) == 'true' + if not is_do_ocr: + final_record_text = update_frontmatter_field(new_record_text, 'do_ocr', 'true') + if final_record_text != new_record_text: + record_path.write_text(final_record_text, encoding='utf-8') + fixed_library_record = True + elif lib_ocr_status == 'done' and meta_ocr_status in ('pending', 'processing'): + new_record_text = update_frontmatter_field(record_text, 'ocr_status', new_status) + if new_record_text != record_text: + record_path.write_text(new_record_text, encoding='utf-8') + fixed_library_record = True + fixed_count = sum([fixed_library_record, fixed_formal_note, fixed_meta]) + result['fixed'] += fixed_count + if verbose and fixed_count > 0: + print(f"[repair] fixed {fixed_count} files for {zotero_key}") + if verbose: + print(f"[repair] scanned={result['scanned']} divergent={len(result['divergent'])} fixed={result['fixed']} errors={len(result['errors'])}") + return result + def run_doctor(vault: Path) -> int: """Validate PaperForge Lite setup and report by category. diff --git a/setup_wizard.py b/setup_wizard.py index f2262fcb..9aef9b8f 100644 --- a/setup_wizard.py +++ b/setup_wizard.py @@ -486,6 +486,14 @@ PaperForge 需要 **Python 3.8+** 以及以下 Python 包: class VaultStep(StepScreen): """Step 3: Vault 目录结构配置""" + def __init__(self, step_id: str, checker: EnvChecker, vault: str = "", **kwargs): + kwargs.setdefault("id", step_id) + super().__init__(**kwargs) + self.step_id = step_id + self.checker = checker + self.step_idx = int(step_id.split("-")[1]) + self._vault = vault + def compose(self) -> ComposeResult: yield from super().compose() yield Markdown(""" @@ -495,7 +503,7 @@ PaperForge 需要知道你的 **Obsidian Vault 位置**,以及你想要的目 """) from textual.widgets import Input yield Static("Obsidian Vault 路径 (绝对路径):", classes="step-title") - yield Input(placeholder="D:\\Documents\\MyVault", id="input-vault-path") + yield Input(value=self._vault, placeholder="D:\\Documents\\MyVault", id="input-vault-path") yield Static("", id="vault-error", classes="status-bar") yield Markdown(""" --- @@ -1329,7 +1337,7 @@ class SetupWizardApp(App): WelcomeStep("step-0", self.checker), AgentPlatformStep("step-1", self.checker), PythonStep("step-2", self.checker), - VaultStep("step-3", self.checker), + VaultStep("step-3", self.checker, vault=str(self.vault)), ZoteroStep("step-4", self.checker), BBTStep("step-5", self.checker), JsonStep("step-6", self.checker), diff --git a/tests/sandbox/README.md b/tests/sandbox/README.md deleted file mode 100644 index 262c3362..00000000 --- a/tests/sandbox/README.md +++ /dev/null @@ -1,55 +0,0 @@ -# PaperForge Lite — Test Sandbox - -## 用途 -测试 PaperForge Lite 安装向导 `setup_wizard.py` 的完整流程。 - -## 目录结构 - -``` -tests/sandbox/ - TestZoteroData/ ← 模拟 Zotero 数据目录(setup wizard 会在 vault 内建 junction 指向这里) - storage/ ← 5 PDFs(4篇有附件,1篇无) - zotero.sqlite ← 伪造(让 Zotero 路径检测通过) - exports/ ← Better BibTeX JSON 导出(2个域,5篇文献,keys 匹配 storage 文件名) - 00_TestVault/ ← 空目录(安装向导会在这里创建所有子目录) - README.md ← 本文件 -``` - -## 测试步骤 - -```powershell -# 1. 进入仓库根目录 -cd D:\...\github-release - -# 2. 运行安装向导,指向空 vault(pip install -e . 由向导自动完成) -python setup_wizard.py --vault D:\L\Med\Research\99_System\LiteraturePipeline\github-release\tests\sandbox\00_TestVault - -# 3. 安装向导中: -# - Agent 平台:选你的(opencode / cursor / claude 等) -# - Zotero 数据目录:填 D:\L\Med\Research\99_System\LiteraturePipeline\github-release\tests\sandbox\TestZoteroData -# (向导会在 vault 内创建 junction: 00_TestVault/00_System/Zotero -> 指向这里) -# - BBT 导出目录:填 D:\L\Med\Research\99_System\LiteraturePipeline\github-release\tests\sandbox\exports -# (向导会检测到 exports/ 下的 JSON 文件) -# - 其他步骤默认即可 - -# 4. 测试 pipeline: -cd D:\L\Med\Research\99_System\LiteraturePipeline\github-release\tests\sandbox\00_TestVault -paperforge selection-sync -paperforge index-refresh -paperforge status -``` - -## 预期结果 - -| 检查项 | 预期 | -|--------|------| -| wizard 检测 TestZoteroData | 通过(有 storage/ 和 zotero.sqlite)| -| wizard 检测 exports/ | 通过(2个 JSON,keys 有效)| -| selection-sync | 生成 5 条 library-records | -| TSTONE003 ocr_status | nopdf(无 PDF) | -| TSTTWO001/002 有 PDF | ocr_status: pending | - -## 注意 -- 目录名故意和真实 vault 不同,避免硬编码测试不出来 -- PDF 是最小化假文件(pymupdf 可读,内容为空) -- 不要往 sandbox 加真实数据 diff --git a/tests/test_pdf_resolver.py b/tests/test_pdf_resolver.py index 1cb9286e..a6a04637 100644 --- a/tests/test_pdf_resolver.py +++ b/tests/test_pdf_resolver.py @@ -7,6 +7,7 @@ and has_pdf=False scenarios. from __future__ import annotations import os +import json from pathlib import Path from unittest.mock import patch @@ -141,3 +142,143 @@ class TestResolveJunction: def test_directory_junction_mocked(self, tmp_path: Path) -> None: """Directory junction resolution is platform-dependent; covered by symlink test above.""" pytest.skip("Windows junction mock incompatible with Python import semantics for `from ctypes import wintypes`") + + +class TestLoadExportRowsAttachmentNormalization: + """Tests for load_export_rows() attachment path normalization. + + Verifies that BBT-exported bare KEY/KEY.pdf paths are normalized to + storage:KEY/KEY.pdf format so resolve_pdf_path() can resolve them correctly. + """ + + def test_bare_key_key_pdf_normalized_to_storage_prefix(self, tmp_path: Path) -> None: + """Bare 'KEY/KEY.pdf' path is normalized to 'storage:KEY/KEY.pdf'.""" + from pipeline.worker.scripts.literature_pipeline import load_export_rows + + export_data = { + "items": [ + { + "key": "ABC123", + "itemKey": "ABC123", + "itemType": "journalArticle", + "title": "Test Paper", + "attachments": [ + {"path": "ABC123/ABC123.pdf", "contentType": "application/pdf"} + ], + } + ], + "collections": {}, + } + export_file = tmp_path / "library.json" + export_file.write_text(json.dumps(export_data), encoding="utf-8") + + rows = load_export_rows(export_file) + + assert len(rows) == 1 + assert rows[0]["attachments"][0]["path"] == "storage:ABC123/ABC123.pdf" + + def test_storage_prefix_preserved(self, tmp_path: Path) -> None: + """Already-prefixed 'storage:KEY/KEY.pdf' path is not double-prefixed.""" + from pipeline.worker.scripts.literature_pipeline import load_export_rows + + export_data = { + "items": [ + { + "key": "ABC123", + "itemKey": "ABC123", + "itemType": "journalArticle", + "title": "Test Paper", + "attachments": [ + {"path": "storage:ABC123/ABC123.pdf", "contentType": "application/pdf"} + ], + } + ], + "collections": {}, + } + export_file = tmp_path / "library.json" + export_file.write_text(json.dumps(export_data), encoding="utf-8") + + rows = load_export_rows(export_file) + + assert len(rows) == 1 + assert rows[0]["attachments"][0]["path"] == "storage:ABC123/ABC123.pdf" + + def test_absolute_path_not_modified(self, tmp_path: Path) -> None: + """Absolute paths are not prefixed with storage:.""" + from pipeline.worker.scripts.literature_pipeline import load_export_rows + + abs_path = str(tmp_path / "ABC123" / "ABC123.pdf") + export_data = { + "items": [ + { + "key": "ABC123", + "itemKey": "ABC123", + "itemType": "journalArticle", + "title": "Test Paper", + "attachments": [ + {"path": abs_path, "contentType": "application/pdf"} + ], + } + ], + "collections": {}, + } + export_file = tmp_path / "library.json" + export_file.write_text(json.dumps(export_data), encoding="utf-8") + + rows = load_export_rows(export_file) + + assert len(rows) == 1 + assert rows[0]["attachments"][0]["path"] == abs_path + + def test_empty_attachment_path_unchanged(self, tmp_path: Path) -> None: + """Empty attachment path is returned unchanged.""" + from pipeline.worker.scripts.literature_pipeline import load_export_rows + + export_data = { + "items": [ + { + "key": "ABC123", + "itemKey": "ABC123", + "itemType": "journalArticle", + "title": "Test Paper", + "attachments": [ + {"path": "", "contentType": ""} + ], + } + ], + "collections": {}, + } + export_file = tmp_path / "library.json" + export_file.write_text(json.dumps(export_data), encoding="utf-8") + + rows = load_export_rows(export_file) + + assert len(rows) == 1 + assert rows[0]["attachments"][0]["path"] == "" + + def test_non_pdf_attachment_unchanged(self, tmp_path: Path) -> None: + """Non-PDF attachments are returned with empty contentType and unchanged path.""" + from pipeline.worker.scripts.literature_pipeline import load_export_rows + + export_data = { + "items": [ + { + "key": "ABC123", + "itemKey": "ABC123", + "itemType": "journalArticle", + "title": "Test Paper", + "attachments": [ + {"path": "ABC123/ABC123.docx", "contentType": ""} + ], + } + ], + "collections": {}, + } + export_file = tmp_path / "library.json" + export_file.write_text(json.dumps(export_data), encoding="utf-8") + + rows = load_export_rows(export_file) + + assert len(rows) == 1 + assert rows[0]["attachments"][0]["path"] == "storage:ABC123/ABC123.docx" + assert rows[0]["attachments"][0]["contentType"] == "" diff --git a/tests/test_repair.py b/tests/test_repair.py new file mode 100644 index 00000000..bfd4aa63 --- /dev/null +++ b/tests/test_repair.py @@ -0,0 +1,348 @@ +"""Tests for run_repair() — three-way OCR status divergence detection and repair.""" +from __future__ import annotations + +import json +import re + +import pytest + +from pipeline.worker.scripts.literature_pipeline import run_repair, pipeline_paths + + +def _make_vault(tmp_path): + vault = tmp_path / "vault" + vault.mkdir() + system = vault / "99_System" + pf = system / "PaperForge" + (pf / "exports").mkdir(parents=True) + (pf / "ocr").mkdir(parents=True) + resources = vault / "03_Resources" + literature = resources / "Literature" + literature.mkdir(parents=True) + control = resources / "LiteratureControl" + control.mkdir(parents=True) + records = control / "library-records" + records.mkdir(parents=True) + (vault / "05_Bases").mkdir(parents=True) + (vault / ".opencode" / "skills").mkdir(parents=True) + (vault / ".opencode" / "command").mkdir(parents=True) + cfg = { + "system_dir": "99_System", + "resources_dir": "03_Resources", + "literature_dir": "Literature", + "control_dir": "LiteratureControl", + "base_dir": "05_Bases", + } + (vault / "paperforge.json").write_text(json.dumps(cfg, ensure_ascii=False), encoding="utf-8") + return vault + + +def _write_library_record(records_dir, key, domain, ocr_status, do_ocr="false", analyze="false"): + record_path = records_dir / f"{key}.md" + content = f"""--- +zotero_key: {key} +domain: {domain} +title: "Test Paper {key}" +year: 2024 +doi: "10.1234/test" +has_pdf: true +pdf_path: "" +fulltext_md_path: "" +recommend_analyze: true +analyze: {analyze} +do_ocr: {do_ocr} +ocr_status: {ocr_status} +deep_reading_status: pending +analysis_note: "" +--- +# Test Paper {key} +""" + record_path.write_text(content, encoding="utf-8") + return record_path + + +def _write_formal_note(literature_dir, key, domain, ocr_status): + domain_lit = literature_dir / domain + domain_lit.mkdir(parents=True, exist_ok=True) + note_path = domain_lit / f"{key} - Test.md" + content = f"""--- +title: "Test Paper {key}" +year: 2024 +domain: "{domain}" +zotero_key: "{key}" +doi: "10.1234/test" +ocr_status: {ocr_status} +deep_reading_status: pending +--- +# Test Paper {key} +""" + note_path.write_text(content, encoding="utf-8") + return note_path + + +def _write_meta(ocr_root, key, ocr_status, zotero_key=None, page_count=10, with_fulltext=True, with_json=True): + meta_dir = ocr_root / key + meta_dir.mkdir(parents=True, exist_ok=True) + meta = { + "zotero_key": zotero_key or key, + "ocr_status": ocr_status, + "page_count": page_count, + } + if with_fulltext: + ft = meta_dir / "fulltext.md" + page_lines = [] + for i in range(1, page_count + 1): + page_lines.append(f"") + page_lines.append("x" * 100) + ft.write_text("\n".join(page_lines), encoding="utf-8") + meta["markdown_path"] = f"99_System/PaperForge/ocr/{key}/fulltext.md" + meta["fulltext_md_path"] = f"99_System/PaperForge/ocr/{key}/fulltext.md" + if with_json: + json_dir = meta_dir / "json" + json_dir.mkdir(parents=True, exist_ok=True) + result = json_dir / "result.json" + pages_data = [{"text": "x" * 200} for _ in range(page_count)] + result.write_text(json.dumps({"pages": pages_data}, ensure_ascii=False), encoding="utf-8") + meta["json_path"] = f"99_System/PaperForge/ocr/{key}/json/result.json" + (meta_dir / "meta.json").write_text(json.dumps(meta, ensure_ascii=False), encoding="utf-8") + return meta_dir / "meta.json" + + +def _write_minimal_meta(ocr_root, key, ocr_status): + meta_dir = ocr_root / key + meta_dir.mkdir(parents=True, exist_ok=True) + meta = {"zotero_key": key, "ocr_status": ocr_status} + (meta_dir / "meta.json").write_text(json.dumps(meta, ensure_ascii=False), encoding="utf-8") + return meta_dir / "meta.json" + + +class TestRunRepairScanOnly: + def test_no_records_returns_empty(self, tmp_path): + vault = _make_vault(tmp_path) + paths = pipeline_paths(vault) + result = run_repair(vault, paths, verbose=False, fix=False) + assert result["scanned"] == 0 + assert result["divergent"] == [] + assert result["fixed"] == 0 + assert result["errors"] == [] + + def test_all_consistent_pending_no_divergence(self, tmp_path): + vault = _make_vault(tmp_path) + paths = pipeline_paths(vault) + records_dir = vault / "03_Resources" / "LiteratureControl" / "library-records" / "骨科" + records_dir.mkdir(parents=True, exist_ok=True) + _write_library_record(records_dir, "KEY001", "骨科", "pending") + result = run_repair(vault, paths, verbose=False, fix=False) + assert result["scanned"] == 1 + assert result["divergent"] == [] + + def test_meta_done_incomplete_is_divergent(self, tmp_path): + vault = _make_vault(tmp_path) + paths = pipeline_paths(vault) + records_dir = vault / "03_Resources" / "LiteratureControl" / "library-records" / "骨科" + records_dir.mkdir(parents=True, exist_ok=True) + _write_library_record(records_dir, "KEY001", "骨科", "pending") + _write_minimal_meta(paths["ocr"], "KEY001", "done") + result = run_repair(vault, paths, verbose=False, fix=False) + assert result["scanned"] == 1 + assert len(result["divergent"]) == 1 + assert result["divergent"][0]["zotero_key"] == "KEY001" + assert "done_incomplete" in result["divergent"][0]["reason"] + + def test_library_done_but_meta_pending_is_divergent(self, tmp_path): + vault = _make_vault(tmp_path) + paths = pipeline_paths(vault) + records_dir = vault / "03_Resources" / "LiteratureControl" / "library-records" / "骨科" + records_dir.mkdir(parents=True, exist_ok=True) + _write_library_record(records_dir, "KEY001", "骨科", "done") + _write_meta(paths["ocr"], "KEY001", "pending") + result = run_repair(vault, paths, verbose=False, fix=False) + assert result["scanned"] == 1 + assert len(result["divergent"]) == 1 + assert result["divergent"][0]["library_record_ocr_status"] == "done" + assert result["divergent"][0]["meta_ocr_status"] == "pending" + + def test_library_done_but_meta_missing_is_divergent(self, tmp_path): + vault = _make_vault(tmp_path) + paths = pipeline_paths(vault) + records_dir = vault / "03_Resources" / "LiteratureControl" / "library-records" / "骨科" + records_dir.mkdir(parents=True, exist_ok=True) + _write_library_record(records_dir, "KEY001", "骨科", "done") + result = run_repair(vault, paths, verbose=False, fix=False) + assert result["scanned"] == 1 + assert len(result["divergent"]) == 1 + + def test_formal_note_done_but_meta_missing_is_divergent(self, tmp_path): + vault = _make_vault(tmp_path) + paths = pipeline_paths(vault) + records_dir = vault / "03_Resources" / "LiteratureControl" / "library-records" / "骨科" + records_dir.mkdir(parents=True, exist_ok=True) + literature_dir = vault / "03_Resources" / "Literature" + _write_library_record(records_dir, "KEY001", "骨科", "pending") + _write_formal_note(literature_dir, "KEY001", "骨科", "done") + result = run_repair(vault, paths, verbose=False, fix=False) + assert result["scanned"] == 1 + assert len(result["divergent"]) == 1 + + def test_library_vs_meta_post_validation_mismatch_is_divergent(self, tmp_path): + vault = _make_vault(tmp_path) + paths = pipeline_paths(vault) + records_dir = vault / "03_Resources" / "LiteratureControl" / "library-records" / "骨科" + records_dir.mkdir(parents=True, exist_ok=True) + _write_library_record(records_dir, "KEY001", "骨科", "done") + _write_meta(paths["ocr"], "KEY001", "done", page_count=10) + result = run_repair(vault, paths, verbose=False, fix=False) + assert result["scanned"] == 1 + assert result["divergent"] == [] + + def test_library_nopdf_not_divergent(self, tmp_path): + vault = _make_vault(tmp_path) + paths = pipeline_paths(vault) + records_dir = vault / "03_Resources" / "LiteratureControl" / "library-records" / "骨科" + records_dir.mkdir(parents=True, exist_ok=True) + _write_library_record(records_dir, "KEY001", "骨科", "nopdf") + result = run_repair(vault, paths, verbose=False, fix=False) + assert result["scanned"] == 1 + assert result["divergent"] == [] + + def test_multiple_domains_scanned(self, tmp_path): + vault = _make_vault(tmp_path) + paths = pipeline_paths(vault) + records_dir_ortho = vault / "03_Resources" / "LiteratureControl" / "library-records" / "骨科" + records_dir_ortho.mkdir(parents=True, exist_ok=True) + records_dir_sports = vault / "03_Resources" / "LiteratureControl" / "library-records" / "运动医学" + records_dir_sports.mkdir(parents=True, exist_ok=True) + _write_library_record(records_dir_ortho, "KEY001", "骨科", "pending") + _write_library_record(records_dir_sports, "KEY002", "运动医学", "pending") + result = run_repair(vault, paths, verbose=False, fix=False) + assert result["scanned"] == 2 + + def test_verbose_output_printed(self, tmp_path, capsys): + vault = _make_vault(tmp_path) + paths = pipeline_paths(vault) + records_dir = vault / "03_Resources" / "LiteratureControl" / "library-records" / "骨科" + records_dir.mkdir(parents=True, exist_ok=True) + _write_library_record(records_dir, "KEY001", "骨科", "done") + _write_meta(paths["ocr"], "KEY001", "pending") + result = run_repair(vault, paths, verbose=True, fix=False) + captured = capsys.readouterr() + assert "KEY001" in captured.out + assert "divergent" in captured.out + + +class TestRunRepairFixMode: + def test_fix_meta_missing_sets_all_to_pending_and_sets_do_ocr(self, tmp_path): + vault = _make_vault(tmp_path) + paths = pipeline_paths(vault) + records_dir = vault / "03_Resources" / "LiteratureControl" / "library-records" / "骨科" + records_dir.mkdir(parents=True, exist_ok=True) + lit_dir = vault / "03_Resources" / "Literature" + record_path = _write_library_record(records_dir, "KEY001", "骨科", "done", do_ocr="false") + _write_formal_note(lit_dir, "KEY001", "骨科", "done") + result = run_repair(vault, paths, verbose=False, fix=True) + assert result["fixed"] >= 1 + record_text = record_path.read_text(encoding="utf-8") + assert re.search(r'^ocr_status:\s*"?pending"?', record_text, re.MULTILINE) + assert re.search(r'^do_ocr:\s*"?true"?', record_text, re.MULTILINE) + + def test_fix_done_incomplete_meta_sets_to_pending(self, tmp_path): + vault = _make_vault(tmp_path) + paths = pipeline_paths(vault) + records_dir = vault / "03_Resources" / "LiteratureControl" / "library-records" / "骨科" + records_dir.mkdir(parents=True, exist_ok=True) + record_path = _write_library_record(records_dir, "KEY001", "骨科", "pending") + _write_minimal_meta(paths["ocr"], "KEY001", "done") + result = run_repair(vault, paths, verbose=False, fix=True) + assert result["fixed"] >= 1 + record_text = record_path.read_text(encoding="utf-8") + assert re.search(r'^ocr_status:\s*"?pending"?', record_text, re.MULTILINE) + + def test_fix_library_done_meta_pending_sets_library_to_pending(self, tmp_path): + vault = _make_vault(tmp_path) + paths = pipeline_paths(vault) + records_dir = vault / "03_Resources" / "LiteratureControl" / "library-records" / "骨科" + records_dir.mkdir(parents=True, exist_ok=True) + record_path = _write_library_record(records_dir, "KEY001", "骨科", "done") + _write_meta(paths["ocr"], "KEY001", "pending") + result = run_repair(vault, paths, verbose=False, fix=True) + assert result["fixed"] >= 1 + record_text = record_path.read_text(encoding="utf-8") + assert re.search(r'^ocr_status:\s*"?pending"?', record_text, re.MULTILINE) + + def test_fix_false_does_not_write_anything(self, tmp_path): + vault = _make_vault(tmp_path) + paths = pipeline_paths(vault) + records_dir = vault / "03_Resources" / "LiteratureControl" / "library-records" / "骨科" + records_dir.mkdir(parents=True, exist_ok=True) + record_path = _write_library_record(records_dir, "KEY001", "骨科", "done") + _write_minimal_meta(paths["ocr"], "KEY001", "done") + original_content = record_path.read_text(encoding="utf-8") + result = run_repair(vault, paths, verbose=False, fix=False) + assert result["fixed"] == 0 + assert record_path.read_text(encoding="utf-8") == original_content + + def test_fix_multiple_divergent_items(self, tmp_path): + vault = _make_vault(tmp_path) + paths = pipeline_paths(vault) + records_dir = vault / "03_Resources" / "LiteratureControl" / "library-records" / "骨科" + records_dir.mkdir(parents=True, exist_ok=True) + record1_path = _write_library_record(records_dir, "KEY001", "骨科", "done", do_ocr="false") + record2_path = _write_library_record(records_dir, "KEY002", "骨科", "pending") + _write_minimal_meta(paths["ocr"], "KEY001", "done") + _write_minimal_meta(paths["ocr"], "KEY002", "done") + result = run_repair(vault, paths, verbose=False, fix=True) + assert len(result["divergent"]) == 2 + assert result["fixed"] >= 1 + + def test_no_divergence_fix_reports_zero_fixed(self, tmp_path): + vault = _make_vault(tmp_path) + paths = pipeline_paths(vault) + records_dir = vault / "03_Resources" / "LiteratureControl" / "library-records" / "骨科" + records_dir.mkdir(parents=True, exist_ok=True) + _write_library_record(records_dir, "KEY001", "骨科", "pending") + _write_meta(paths["ocr"], "KEY001", "pending") + result = run_repair(vault, paths, verbose=False, fix=True) + assert result["fixed"] == 0 + + def test_fix_writes_formal_note_ocr_status(self, tmp_path): + vault = _make_vault(tmp_path) + paths = pipeline_paths(vault) + records_dir = vault / "03_Resources" / "LiteratureControl" / "library-records" / "骨科" + records_dir.mkdir(parents=True, exist_ok=True) + lit_dir = vault / "03_Resources" / "Literature" + _write_library_record(records_dir, "KEY001", "骨科", "done", do_ocr="false") + note_path = _write_formal_note(lit_dir, "KEY001", "骨科", "done") + result = run_repair(vault, paths, verbose=False, fix=True) + note_text = note_path.read_text(encoding="utf-8") + assert re.search(r'^ocr_status:\s*"?pending"?', note_text, re.MULTILINE) + + +class TestRunRepairReturnStructure: + def test_result_has_all_keys(self, tmp_path): + vault = _make_vault(tmp_path) + paths = pipeline_paths(vault) + result = run_repair(vault, paths, verbose=False, fix=False) + assert "scanned" in result + assert "divergent" in result + assert "fixed" in result + assert "errors" in result + assert isinstance(result["scanned"], int) + assert isinstance(result["divergent"], list) + assert isinstance(result["fixed"], int) + assert isinstance(result["errors"], list) + + def test_divergent_item_has_required_fields(self, tmp_path): + vault = _make_vault(tmp_path) + paths = pipeline_paths(vault) + records_dir = vault / "03_Resources" / "LiteratureControl" / "library-records" / "骨科" + records_dir.mkdir(parents=True, exist_ok=True) + _write_library_record(records_dir, "KEY001", "骨科", "done") + _write_meta(paths["ocr"], "KEY001", "pending") + result = run_repair(vault, paths, verbose=False, fix=False) + item = result["divergent"][0] + assert "zotero_key" in item + assert "domain" in item + assert "library_record_ocr_status" in item + assert "formal_note_ocr_status" in item + assert "meta_ocr_status" in item + assert "reason" in item \ No newline at end of file