- Test pending->queued transition on job submission
- Test processing->done transition when polling returns success
- Test processing->error transition on API error response
- Test processing->blocked transition via HTTPError 401 (classify_error maps to blocked)
- Test sync_ocr_queue skips done/blocked items from existing queue
- Test cleanup_blocked_ocr_dirs removes empty dirs, preserves dirs with payload
- Test all 7 OCR states (pending, queued, running, done, error, blocked, nopdf) don't crash
Key: patch requests.post with side_effect=HTTPError 401 to trigger
classify_error path, since registry token is always present in test env.
- fixture_vault: creates complete vault structure with paperforge.json
- fixture_library_records: 3 minimal library-record .md files
- fixture_bbt_json: Better BibTeX JSON export in exports/
- fixture_with_pdf: vault with real temp PDF for PDF-resolver tests
- All use pathlib.Path, Windows-compatible
- Add 02-03-SUMMARY.md with execution summary and deviations
- Update STATE.md: mark 02-03 done, add decisions, update next action
- Update ROADMAP.md: mark 02-03 and 01-04 complete
- Convert ocr parser to sub-subcommands (run, doctor)
- Add _cmd_ocr_doctor() with formatted tiered report output
- Update dispatch logic to route ocr doctor subcommand
- Add CLI dispatch test for doctor command
- Update command/lp-ocr.md with doctor documentation
- Add paperforge_lite/ocr_diagnostics.py with tiered L1-L4 checks
- L1: token presence, L2: URL reachability, L3: API schema validation
- L4: optional live PDF round-trip test with polling
- Add 7 mocked unit tests for all levels
- Include blank.pdf test fixture
- load_config now tries paperforge_lite.config.load_vault_config first,
falls back to legacy JSON parsing for pre-01-03 installs
- resolve_vault_for_validate checks PAPERFORGE_VAULT before VAULT_PATH
before cwd (matching the shared resolver's precedence)
- Preserves all existing validation checks
- setup_wizard.py now copies the paperforge_lite/ package alongside
the worker and ld_deep scripts during deployment
- Two copy targets: <pf_path>/worker/paperforge_lite/ for
literature_pipeline.py and <skill_dir>/literature-qa/paperforge_lite/
for ld_deep.py
- Existing copy targets for literature_pipeline.py and ld_deep.py
remain unchanged
- Warning emitted if paperforge_lite source not found
- literature_pipeline: load_vault_config and pipeline_paths now
delegate to paperforge_lite.config, preserving public names
- ld_deep: _load_vault_config and _paperforge_paths now delegate
to shared resolver, preserving public names
- Both scripts require paperforge_lite package to be available
(via pip install or copied install in deployed vaults)
- Existing load_simple_env preserved for .env loading before dispatch
- Replace literature_pipeline.load_vault_config body with
paperforge_lite.config.load_vault_config wrapper
- Rebuild pipeline_paths() from shared paperforge_paths() plus
10 worker-only keys (pipeline, candidates, search_*, harvest_root,
records, review, config, queue, log, bridge_config*, index, ocr_queue)
- Replace ld_deep._load_vault_config and _paperforge_paths with
shared resolver wrappers returning ocr, records, literature
- Fix subprocess test to set PYTHONPATH so paperforge_lite is
importable when worker runs as a standalone script
- 34 tests pass (22 config + 8 worker compat + 4 ld_deep compat)
- Test literature_pipeline.load_vault_config matches paperforge_lite.config
- Test PAPERFORGE_SYSTEM_DIR env var override in worker
- Test pipeline_paths returns all 17 expected keys (shared + worker-only)
- Test ld_deep._load_vault_config and _paperforge_paths match shared resolver
- Test direct worker status subprocess smoke test (CMD-02)
- Tests fail against current code: worker has 5 config keys vs shared's 7,
env overrides not implemented, and pipeline_paths lacks shared resolver keys
- Tests main() with --vault vault paths --json emits valid JSON with vault, worker_script, ld_deep_script keys
- Tests main() with --vault vault paths emits text paths without unresolved <system_dir> or <resources_dir> tokens
- Tests ocr and ocr run both dispatch to run_ocr
- Tests status/selection-sync/index-refresh/deep-reading each dispatch to correct worker
- Uses monkeypatch to isolate from real workers and filesystem