Commit Graph
12 Commits
Author SHA1 Message Date
Richie 40f707798e fix(protected-phrases): isolate phrase generation per book
treefmt / nix fmt (pull_request) Failing after 5s
pytest / pytest (pull_request) Successful in 28s
test ebook search / test-ebook-search (pull_request) Failing after 35s
build_systems / build-bob (pull_request) Successful in 49s
build_systems / build-brain (pull_request) Successful in 49s
build_systems / build-rhapsody-in-green (pull_request) Successful in 1m3s
build_systems / build-jeeves (pull_request) Successful in 2m24s
Run full-book candidate generation inside worker-owned sessions so each book commits independently during backfills. Abort recalculation when a book has no indexed chapters to preserve existing phrase data, and update admin/UI tests for the new generation flow.
2026-07-12 17:51:38 -04:00
Richie c4993f5a53 refactor(ebook): remove spaCy-ner attributes from PhraseCandidate and related functions 2026-07-12 17:51:38 -04:00
Richie a876d71339 feat(ebook): add additional tokens to junk tokens configuration 2026-07-12 17:51:06 -04:00
Richie bfb3463fd0 feat(ebook): migrate to async DB/HTTP and parallelize phrase pipeline
Convert the ebook-search web app to async end to end and add concurrency
to the protected-phrase extraction and judging pipeline so large books no
longer block the event loop or the UI.

ORM / infra:
- Add get_async_postgres_engine and factor shared URL/connect_args building
  into build_postgres_url (reused by the sync and async engine builders)
- Add async FastAPI session helpers (get_async_db, AsyncDbSession) with
  expire_on_commit=False to avoid implicit IO under asyncio

App:
- Use AsyncEngine/AsyncSession throughout routes, search, ingest, embeddings,
  answer, rerank and LLM calls; convert handlers to async
- Share a single httpx.AsyncClient in app state for LLM requests; size the
  connection pool for concurrent phrase-judging workers
- Add judge_tasks: run per-book judging as tracked background tasks so a
  book already being judged isn't double-queued

Protected phrases:
- Add a process pool (pool.py) and worker-count config
  (extraction/judge book/phrase workers) to parallelize candidate generation
  and judging
- Split admin actions into all/missing variants for generation and judging

Config:
- Add protected_phrase_extraction_workers, phrase_judge_book_workers,
  phrase_judge_phrase_workers
2026-07-12 17:51:06 -04:00
Richie 63e0b0dd3b feat(extraction): add cached YAKE extractor for improved performance 2026-07-12 17:49:52 -04:00
Richie 167aedcefb feat(ebook): add junk tokens for improved phrase matching 2026-07-12 17:49:52 -04:00
Richie bcb8b7b169 Add models and database persistence for protected phrase extraction
- Introduced dataclasses for phrase candidates, judgments, and matches in `models.py`.
- Implemented database operations for candidate and protected phrases in `store.py`, including loading, saving, and deleting phrases.
- Enhanced text normalization functions in `text_normalization.py` with detailed docstrings.
- Refactored search functionality to utilize new models and methods for detecting protected phrases.
2026-07-12 17:49:52 -04:00
Richie 8d8c8cd3dd refactor(protected-phrases): extract config and text normalization helpers 2026-07-12 17:49:52 -04:00
Richie 77692b1414 feat(ebook): update protected phrases with additional tokens and phrases 2026-07-12 17:49:52 -04:00
Richie f4a498e45e feat(ebook): enhance phrase judgment logging with failure tracking 2026-07-12 17:49:52 -04:00
Richie d03daeb66d feat(ebook): implement phrase matching functionality and UI enhancements 2026-07-12 17:49:52 -04:00
Richie f8b5ba82a6 feat(ebook): add protected phrase extraction library with config-driven tuning
Refactor protected phrase handling from a single module into a
python/ebook_search/protected_phrases package covering extraction,
storage, and runtime matching. Phrase filtering is now data-driven via
bundled TOML files: ignored_phrases, bad_starts, bad_ends, and
most_common_words.

Add phrase-tuning settings to EbookSearchConfig so candidate generation,
scoring, LLM judging, and matching are configurable rather than hardcoded:
token bounds, entity token limit, raw n-gram min count, frequency and
chapter-spread score thresholds, candidate/LLM/target caps, confidence
threshold, nesting defaults, and the phrase hit boost.
2026-07-12 17:49:52 -04:00