diff --git a/PROJECT-MANAGEMENT.md b/PROJECT-MANAGEMENT.md index 40de078f..54181f2d 100644 --- a/PROJECT-MANAGEMENT.md +++ b/PROJECT-MANAGEMENT.md @@ -1,11 +1,11 @@ -> **Branch:** `master` | **Last Updated:** 2026-07-09 -> **Active work:** sqlite-vec migration complete. ChromaDB replaced with sqlite-vec for vector storage. Embedding provider switched to openai SDK. Hash-based OCR change detection deployed. +> **Branch:** `master` | **Last Updated:** 2026-07-10 +> **Active work:** Retrieval Experience recovery Wayfinder active. The live-architecture inventory is complete; non-destructive failure capture and source-to-vault deployment parity are the current frontier. > > --- > -> **Current state:** All 10 open issues implemented and merged to master. ChromaDB fully replaced with sqlite-vec (vec0 k-NN search, companion meta tables, build_state in SQLite). Embedding provider uses openai SDK (requests fallback available). Hash-based OCR staleness detection (two-tier mtime+xxhash). E2E fixture PDFs + test scaffolding ready. Schema at v6. Test suite: 76 passed across embed/OCR/schema/version/provider suites. +> **Current state:** OCR-v2 status is unchanged. Retrieval is **not production-ready** despite the sqlite-vec/search implementation issues being closed: the architecture audit confirmed broken M-search SQL, an unreachable `@` hybrid path, an incompatible retrieve result envelope, write-then-delete vector replacement, legacy Chroma resume/force gates, and SQLite/JSON build-state split-brain. No production fix or test-suite claim was made in this investigation session. > -> Next: push origin, CI validation, downstream tickets (#21 display, #22 cleanup, remaining E2E assertions). +> Next: resolve [Capture a non-destructive retrieval failure matrix](https://github.com/LLLin000/PaperForge/issues/49) and [Audit source-to-vault deployment parity](https://github.com/LLLin000/PaperForge/issues/47), then unblock [Choose the canonical retrieval architecture and ownership boundary](https://github.com/LLLin000/PaperForge/issues/46). ## 1. Architecture ### 1.1 The problem (pre-v2) @@ -69,6 +69,10 @@ raw observations → structural signatures → stable anchors/families → zone | **Quality indicators** (Layer 2) | ✅ `paperforge/worker/ocr_quality.py` — `build_quality_indicators()` with 5 normalizers | | **Readiness policy** (Layer 2) | ✅ `evaluate_readiness()` + YAML policy evaluator with deep-merge, hard-red, use-case gates | | **Feedback sidecar** (Layer 2) | ✅ `paperforge/worker/ocr_quality_feedback.py` — per-mark hash, stale detection, no UI | +| **Retrieval architecture** | 🔴 Wayfinder recovery active; [live architecture audit](https://gist.github.com/LLLin000/aaf5505a991e85ad9bb4cafa922f48bf) completed | +| **M metadata search** | 🔴 sql.js selects nonexistent `paper_fts.year`; CLI fallback is the only viable arm | +| **@ deep search** | 🔴 plugin omits `--deep` and discards CLI `data.chunks` | +| **Vector build control** | 🔴 new vec0 rows are deleted after write; resume/force still gate on legacy Chroma path; UI reads stale JSON build state | ### 2.3 Fix Status @@ -125,6 +129,10 @@ raw observations → structural signatures → stable anchors/families → zone - `paperforge/worker/ocr_quality_feedback.py` — human feedback sidecar (per-mark hash, stale detection) 22 new tests. Contract polish: `status/gates/reasons` output shape, hash validation, user override bypass. +### Retrieval Experience Recovery + +[Wayfinder: Restore PaperForge retrieval end to end](https://github.com/LLLin000/PaperForge/issues/45) is the active recovery map. [Inventory the live retrieval architecture and contract drift](https://github.com/LLLin000/PaperForge/issues/53) established that current failures are cross-contract defects, not one vector-backend bug. The next frontier is runtime failure evidence plus deployed-artifact parity; implementation remains intentionally paused until the canonical backend, build lifecycle, query/result contract, repair policy, UI states, and acceptance gate are decided. + ### Layer 3: Plugin UI `PaperForgeStatusView` dashboard is 2300+ lines in one file. Status display, OCR maintenance UI, and user-facing health presentation not yet polished. @@ -134,14 +142,20 @@ raw observations → structural signatures → stable anchors/families → zone Remaining legacy OCR issues (carried forward): ## 4. Active Queue -1. ✅ **Layout-category truth audit (Workstream X)** — 6 bug patterns identified. -2. ✅ **Layer 2: OCR Quality Report + Readiness Policy** — `96fd9771`. -3. 🟡 **Plugin UI polish** — dashboard cleanup, status clarity -4. 🟡 **Downstream tooling** — section-aware chunking, figure/table separate handling -5. ⏳ **Architecture cleanup** — deferred post-release -6. ⏳ **Compatibility naming cleanup** — deferred post-release +1. 🔴 **Retrieval recovery Wayfinder** — architecture inventory resolved; implementation paused until decisions are complete. +2. 🟡 **Runtime evidence frontier** — non-destructive Literature-hub/test-vault failure matrix. +3. 🟡 **Deployment parity frontier** — Python/plugin/source/deployed artifact audit. +4. ⏳ **Canonical retrieval architecture** — blocked by the three evidence tickets. +5. 🟡 **Plugin UI polish** — retrieval UI behavior will follow the lifecycle/query contracts. +6. 🟡 **Downstream tooling** — section-aware chunking, figure/table separate handling. +7. ⏳ **Compatibility naming cleanup** — deferred post-release. ### 4.1 Immediate Next Steps +- [x] [Wayfinder: Restore PaperForge retrieval end to end](https://github.com/LLLin000/PaperForge/issues/45) +- [x] [Inventory the live retrieval architecture and contract drift](https://github.com/LLLin000/PaperForge/issues/53) +- [ ] [Capture a non-destructive retrieval failure matrix](https://github.com/LLLin000/PaperForge/issues/49) +- [ ] [Audit source-to-vault deployment parity](https://github.com/LLLin000/PaperForge/issues/47) +- [ ] [Choose the canonical retrieval architecture and ownership boundary](https://github.com/LLLin000/PaperForge/issues/46) - [x] #20 Hash-based OCR change detection - [x] #31 OpenAI SDK embedding provider - [x] #33 E2E test fixtures (3 synthetic PDFs) @@ -290,6 +304,8 @@ Remaining legacy OCR issues (carried forward): | 2026-07-08 | sqlite-vec replaces ChromaDB for vector storage | ChromaDB HNSW index corruption, heavy deps (~100MB), and separate storage. sqlite-vec is ~1MB, stores vectors in same paperforge.db. Brute force at ~50k vectors <100ms. | | 2026-07-08 | E2E test fixtures as synthetic PDFs (PyMuPDF) instead of real arXiv PDFs | Avoids license/distribution issues; deterministic and fast to generate. | | 2026-07-09 | Schema v6: hash/policy columns in vec companion meta tables | Resume hash checks need body_units_hash, object_units_hash, retrieval_policy_version persisted alongside vectors. | +| 2026-07-10 | Retrieval scope is the entire user experience, not only vec0 | M metadata search, `@` retrieval, build lifecycle, status, deployment, and persistence share contracts; auditing only the vector table would miss the observed failures. | +| 2026-07-10 | Source Corpus is preserved; Retrieval Artifacts are disposable | Paper/OCR/structured/user-authored content is authoritative, while FTS indexes, embeddings, vector tables, and companion metadata may be cleared and rebuilt instead of migrated. | --- @@ -384,6 +400,7 @@ python -m ruff check paperforge/worker/ocr_*.py | 2026-07-02 | Zone/role robustness completion — Figure caption prefix recovery + inline table fix + table matching audit | Figure caption prefix recovery from PDF text layer (`_recover_figure_heading_prefix`): 5S7UI34M 4→9 matched figures, 33→1 unmatched. Inline `` HTML role fix: 650 blocks now `table_html` after rebuild (priority bug: raw_label=table fired before `
` check). Table matching audit: 620 remaining `media_asset` pending full rebuild. 585 figure/table/role tests pass. | §9.21 | | 2026-07-03 | Figure pipeline vnext cutover completed | Implemented all remaining vnext passes (composite parent, group/classic sequential, unresolved consolidation, accounting), expanded compare harness, curated 5-paper cutover corpus covering all 9 spec categories, generated diff review (improvement=2 / equivalent=2 / parity=1 / regression=0), updated legacy figure tests for vnext contract, and switched `build_figure_inventory(...)` wrapper to vnext on branch `feat/figure-pipeline-vnext`. Verification: 346 tests passed. | §9.22 | | 2026-07-03 | OCR pairing framework merge-unblock pass | Cleared the remaining merge blockers on `feat/ocr-pairing-framework`: moved figure-only rotation enrichment out of generic state, upgraded table cutover validation to semantic parity across 6 runnable real-paper fixtures including `37LK5T97`, and cleaned touched-file lint/format issues. Verification: 357 targeted tests passed; merge-ready. | §9.23 | +| 2026-07-10 | Retrieval architecture Wayfinder inventory | Charted the end-to-end recovery map and resolved the live-architecture ticket. Confirmed four P0 contract breaks: invalid sql.js metadata SQL, missing `--deep`, ignored `data.chunks`, and build write-then-delete. Published the evidence audit; no production fix applied. | [Architecture audit](https://gist.github.com/LLLin000/aaf5505a991e85ad9bb4cafa922f48bf) | ## 9. Historical Detail Archive diff --git a/paperforge/plugin/src/i18n.ts b/paperforge/plugin/src/i18n.ts index f78bbff1..2266d447 100644 --- a/paperforge/plugin/src/i18n.ts +++ b/paperforge/plugin/src/i18n.ts @@ -284,47 +284,45 @@ const LANG: Record> = { version_compare_paragraphs: "{n} paragraphs changed", version_error_read: "Cannot read version data", retrieval_search_placeholder: "Search papers... (@ for deep search)", - retrieval_search_placeholder_deep: "Search papers via deep retrieval...", - retrieval_search_idle_hint: - "Type to search metadata, or start with @ for deep semantic search", + retrieval_search_placeholder_deep: "Search paper content...", + retrieval_search_idle_hint: "Type a keyword or @ to search paper content", retrieval_searching_metadata: "Searching metadata...", - retrieval_searching_deep: "Running deep search...", + retrieval_searching_deep: "Deep searching...", retrieval_search_cancel: "Cancel", - retrieval_results_count: "{n} result{s}", + retrieval_results_count: "{n} result(s)", retrieval_empty: "No matching papers found.", - retrieval_empty_tips: "Try broader terms or @ deep search.", - retrieval_vectors_not_built: "Vectors Not Built", + retrieval_empty_tips: "Try broader terms or use @ deep search.", + retrieval_vectors_not_built: "Vector index not built", retrieval_vectors_not_built_desc: - "Vector index has not been built. Search requires a built vector index.", + "Build vectors to enable @ deep search with semantic matching.", retrieval_open_vector_settings: "Open Vector Settings", - retrieval_backend_unavailable: "Backend Unavailable", + retrieval_backend_unavailable: "Search backend unavailable", retrieval_backend_unavailable_desc: - "The Python CLI search backend is not responding.", + "The Python CLI search backend is not responding correctly.", retrieval_run_doctor: "Run Doctor", retrieval_retry: "Retry", - retrieval_timeout_title: "Search Timed Out", + retrieval_timeout_title: "Search timed out", retrieval_timeout_desc: - "The search took too long. You can retry or try a simpler query.", - retrieval_model_changed: "Model Changed", + "The search took too long. Try a more specific query.", + retrieval_model_changed: "Model changed", retrieval_model_changed_desc: - "The embedding model has changed since vectors were built. Rebuilding is required.", + "The embedding model has changed since vectors were built. Rebuild to use deep search.", retrieval_rebuild_vectors: "Rebuild Vectors", - retrieval_internal_error: - "An internal error occurred. Check the console for details.", - retrieval_build_idle: "Vectors not yet built", - retrieval_build_ready: "Vectors ready", + retrieval_build_idle: "Vector database ready", + retrieval_build_ready: "{n} vector(s) built", retrieval_build_stopping: "Stopping...", - retrieval_build_stopped: "Build stopped. Resume to continue.", + retrieval_build_stopped: "Build stopped ({n}/{t} papers)", retrieval_build_failed: "Build failed", - retrieval_build_stale: - "Previous build was interrupted. Resume to continue.", - retrieval_build_deps_missing: "Dependencies not installed", - retrieval_build_runtime_mismatch: "Plugin and CLI versions differ", + retrieval_build_stale: "Vectors are stale — rebuild recommended", + retrieval_build_deps_missing: + "Dependencies missing. Install chromadb and openai.", + retrieval_build_runtime_mismatch: "Python runtime version mismatch.", retrieval_stop: "Stop", + retrieval_no_python: "Python not found", + retrieval_internal_error: "An internal error occurred", retrieval_force_rebuild: "Force Rebuild", retrieval_rebuild_warning: "Rebuild will replace {n} existing chunk(s). Continue?", - retrieval_no_python: "No Python found. Check Installation tab.", }, zh: { @@ -585,42 +583,41 @@ const LANG: Record> = { version_compare_title: "{vA} vs {vB}", version_compare_paragraphs: "{n} 段有变化", version_error_read: "无法读取版本数据", - retrieval_search_placeholder: "搜索论文... (@ 进行深度搜索)", - retrieval_search_placeholder_deep: "通过深度检索搜索论文...", - retrieval_search_idle_hint: - "输入关键字搜索元数据,或以 @ 开头进行深度语义搜索", - retrieval_searching_metadata: "正在搜索元数据...", - retrieval_searching_deep: "正在运行深度搜索...", + retrieval_search_placeholder: "搜索论文...(@ 启动深度搜索)", + retrieval_search_placeholder_deep: "搜索论文内容...", + retrieval_search_idle_hint: "输入关键词搜索论文,或以 @ 开头搜索论文内容", + retrieval_searching_metadata: "搜索元数据中...", + retrieval_searching_deep: "深度搜索中...", retrieval_search_cancel: "取消", - retrieval_results_count: "找到 {n} 条结果", + retrieval_results_count: "{n} 个结果", retrieval_empty: "未找到匹配的论文。", - retrieval_empty_tips: "尝试更宽泛的关键词,或使用 @ 进行深度搜索。", + retrieval_empty_tips: "尝试更宽泛的关键词,或使用 @ 深度搜索论文内容。", retrieval_vectors_not_built: "向量索引未构建", - retrieval_vectors_not_built_desc: - "向量索引尚未构建,搜索需要先构建向量索引。", + retrieval_vectors_not_built_desc: "构建向量索引以启用 @ 深度语义搜索。", retrieval_open_vector_settings: "打开向量设置", - retrieval_backend_unavailable: "后端不可用", - retrieval_backend_unavailable_desc: "Python CLI 搜索后端无响应。", + retrieval_backend_unavailable: "搜索后端不可用", + retrieval_backend_unavailable_desc: "Python CLI 搜索后端未正常响应。", retrieval_run_doctor: "运行诊断", retrieval_retry: "重试", retrieval_timeout_title: "搜索超时", - retrieval_timeout_desc: "搜索耗时过长。你可以重试,或尝试简化查询。", + retrieval_timeout_desc: "搜索耗时过长,请尝试更精确的查询。", retrieval_model_changed: "模型已更换", - retrieval_model_changed_desc: "嵌入模型已更换,需要重建向量。", + retrieval_model_changed_desc: + "嵌入模型已更换,需重建向量后才能使用深度搜索。", retrieval_rebuild_vectors: "重建向量", - retrieval_internal_error: "发生内部错误。请查看控制台了解详情。", - retrieval_build_idle: "向量索引尚未构建", - retrieval_build_ready: "向量索引就绪", - retrieval_build_stopping: "正在停止…", - retrieval_build_stopped: "构建已停止。点击继续以恢复。", + retrieval_build_idle: "向量数据库就绪", + retrieval_build_ready: "已构建 {n} 个向量", + retrieval_build_stopping: "正在停止...", + retrieval_build_stopped: "构建已停止({n}/{t} 篇)", retrieval_build_failed: "构建失败", - retrieval_build_stale: "上次构建被中断。点击继续以恢复。", - retrieval_build_deps_missing: "依赖未安装", - retrieval_build_runtime_mismatch: "插件与 CLI 版本不匹配", + retrieval_build_stale: "向量已过期 — 建议重建", + retrieval_build_deps_missing: "依赖缺失。请安装 chromadb 和 openai。", + retrieval_build_runtime_mismatch: "Python 运行时版本不匹配。", retrieval_stop: "停止", + retrieval_no_python: "未找到 Python", + retrieval_internal_error: "发生内部错误", retrieval_force_rebuild: "强制重建", - retrieval_rebuild_warning: "重建将替换 {n} 个已有的 chunk。是否继续?", - retrieval_no_python: "未找到 Python。请查看安装标签页。", + retrieval_rebuild_warning: "重建将替换 {n} 个现有文本块,是否继续?", }, }; diff --git a/project/current/ocr-v2-active-queue.md b/project/current/ocr-v2-active-queue.md index 30d0f81c..0b160ac3 100644 --- a/project/current/ocr-v2-active-queue.md +++ b/project/current/ocr-v2-active-queue.md @@ -1,53 +1,43 @@ # OCR-v2 Active Queue -> Status: Embedding pipeline overhaul complete. PR9A-C all merged to `master`. -> Last updated: 2026-07-08 +> Status: OCR-v2 remains stable; Retrieval Experience recovery is now the active blocking workstream. +> Last updated: 2026-07-10 -## Completed (this session) +## Current checkpoint -### PR9A: Resume & Rebuild Correctness -- OCR rebuild `--all`: version/artifact-based selection via `_needs_derived_rebuild()` -- `--status` and explicit keys: manual override, no version filter -- `.done.{key}` checkpoint markers removed from selection -- `_apply_post_rebuild_version_flags` now writes `derived_version` -- Embed resume entry: three-gate protection (stale state / missing DB / corrupt DB) -- `_pid_alive()` cross-platform PID health check -- 14 regression tests +- [Wayfinder: Restore PaperForge retrieval end to end](https://github.com/LLLin000/PaperForge/issues/45) is the canonical recovery map. +- [Inventory the live retrieval architecture and contract drift](https://github.com/LLLin000/PaperForge/issues/53) is resolved with a reviewed [architecture audit](https://gist.github.com/LLLin000/aaf5505a991e85ad9bb4cafa922f48bf). +- Source Corpus data remains authoritative and must be preserved. FTS indexes, embeddings, vec0 tables, and companion metadata are disposable Retrieval Artifacts. +- No retrieval implementation fix or production build-success claim was made at this checkpoint. -### PR9B: Embed Parallel Encode -- Four dataclasses: `EmbeddingPayload`, `EncodedPayload`, `PaperEmbeddingJob`, `PaperEncodedBundle` -- `prepare_legacy/body/object_payload` — prepare phase extraction -- `encode_payload` / `encode_paper_job` — worker-thread-safe encode -- `write_encoded_payload` — ChromaDB serial write -- Existing `embed_body_units`/`embed_paper` refactored as wrappers -- 23 regression tests +## Confirmed blocking findings -### Provider Fix -- Switched `OpenAICompatibleProvider` from `openai` client to `requests` -- Fixed SiliconFlow NAT connection hang (openai-python#3269) -- 0.3s vs 2.3s init, no more hanging +1. sql.js metadata search cannot prepare its query because `paper_fts.year` does not exist. +2. The plugin `@` path invokes `retrieve` without `--deep`. +3. Retrieve emits `data.chunks`; the plugin accepts only `data.matches` or `data.results`. +4. Embed build writes new vec0 rows and then deletes every row for that paper. +5. Resume and force still use the legacy Chroma vectors directory as their active-store gate/cleanup target. +6. SQLite `build_state` is live truth, but the plugin renders a nested copy from `vector-runtime-state.json`. +7. The repository and Literature-hub plugin bundles have different checksums; the exact deployed delta is not yet resolved. -### PR9C: Streaming Embed Pipeline -- Sliding-window pipeline: prepare + submit bounded in-flight papers -- `wait(FIRST_COMPLETED)` in main thread for encode results -- `processed_count` = skip + embedded, monotonic EMBED_PROGRESS -- Resume skip and no-payload paths also advance processed_count -- Encode failure fails closed (return 1, no silent skip) -- `write_vector_build_state` fallback when file locked (Windows) -- 4 integration tests +## Frontier -### Plugin Fixes -- Status text shows total chunks across all three collections -- "chunks embedded" text below progress bar also uses total +- [ ] [Capture a non-destructive retrieval failure matrix](https://github.com/LLLin000/PaperForge/issues/49) +- [ ] [Audit source-to-vault deployment parity](https://github.com/LLLin000/PaperForge/issues/47) -## Test Status -| Suite | Result | -|-------|--------| -| Python unit tests (PR9A) | 14/14 pass | -| Python unit tests (PR9B) | 23/23 pass | -| Python unit tests (PR9C) | 4/4 pass | -| Plugin tests | 58/58 pass | -| Full vault embed build | 729/729 papers, 20,655 chunks | +These tickets are independent and may run in parallel. They must not mutate Source Corpus data or apply production fixes. -## Immediate Next -- [ ] Full vault embed build (`--force`) on fresh ChromaDB -- [ ] Verify chunk counts in Obsidian plugin display +## Blocked after the frontier + +- [ ] [Choose the canonical retrieval architecture and ownership boundary](https://github.com/LLLin000/PaperForge/issues/46) +- [ ] [Specify the retrieval build lifecycle and crash recovery semantics](https://github.com/LLLin000/PaperForge/issues/52) +- [ ] [Specify metadata and deep-search contracts](https://github.com/LLLin000/PaperForge/issues/54) +- [ ] [Define safe repair, rebuild, and model-change policy](https://github.com/LLLin000/PaperForge/issues/51) +- [ ] [Prototype retrieval panel states and recovery flows](https://github.com/LLLin000/PaperForge/issues/48) +- [ ] [Define the retrieval acceptance matrix and release gate](https://github.com/LLLin000/PaperForge/issues/50) + +## Verification status + +- Read-only Literature-hub DB probe confirmed that `paper_fts` has no `year` column and the sql.js query fails with `OperationalError: no such column: year`. +- Refreshed code-graph traces confirmed that `get_vector_backend()` is tests-only while production retrieval accesses vec0 directly. +- Evidence review confirmed the execution maps and P0 contract failures; build-control-only and explicit-CLI-deep findings were downgraded below P0. +- No test suite was run because this checkpoint is an architecture investigation, not an implementation completion claim.