lllin000_PaperForge/docs/plans
LLLin000 0dfbf2a68d feat: PR 8 — Object Vector Retrieval (paperforge_objects collection)
New paperforge_objects collection in ChromaDB for figure/table caption
vector search, completing the memory layer retrieval triad:

  paperforge_fulltext → legacy_chunk
  paperforge_body     → body_unit
  paperforge_objects  → object_unit (NEW)

Changes:
- manifest.py: compute_object_units_hash(), manifest records both hashes
- _chroma.py: _COLLECTION_NAMES includes paperforge_objects
- builder.py: get_object_units_for_embedding() + embed_object_units()
- status.py: object_chunk_count, total_chunks
- state_snapshot.py: write_vector_runtime() extended with new fields
- search.py: 3rd collection, dict-based source mapping, object_kind/label
- __init__.py: exports for new functions
- commands/embed.py: _has_object_units_in_db, body+object structured
  path routing with dual-hash resume, extended counters
- Test fix: test_build_paper_manifest data complete for hash functions

Verification: 20 object vectors embedded, all per-paper counts match,
status total consistent (58266 = 57435 + 811 + 20),
object caption queries return object_unit results via merge_retrieve.
Unit tests: 153 pass, 1 skip.
2026-07-06 19:58:46 +08:00
..
memory-layer-functional-test-plan.md fix: PR 7 post-test hardening — object units, node_id, expanded tests 2026-07-06 19:34:05 +08:00
memory-layer-implementation-plan.md feat: PR 1 — nested structure tree + body_units correctness 2026-07-06 16:42:08 +08:00
memory-layer-test-report.md fix: PR 7 post-test hardening — object units, node_id, expanded tests 2026-07-06 19:34:05 +08:00
pr-2-and-pr-3.md feat: improve vnext figure matching and split OCR status 2026-07-06 00:55:40 +08:00
pr8-object-vector-retrieval-plan.md feat: PR 8 — Object Vector Retrieval (paperforge_objects collection) 2026-07-06 19:58:46 +08:00