Vector Search
backend/vector-search/ (FastAPI + ChromaDB, port 3001) powers semantic video search —
the system behind “where did they discuss X?” timeline markers.
Endpoints
| Method | Path | Purpose |
|---|---|---|
POST | /ingest | Ingest transcript chunks (with timestamps) into a collection |
GET | /search | Query by text; returns matching chunks ranked by similarity |
GET | /health | Liveness probe |
GET | /stats | Collection counts and index stats |
Pipeline
- The transcript is chunked client-side by the adaptive semantic chunker — topic-boundary detection via cosine-similarity drops, 30–300 words per chunk
- Chunks (with start/end timestamps) are
POST /ingest-ed - The Conductor’s
search_videofunction issuesGET /search - Hits return with timestamps → the timeline drops markers at the exact moments
Fallback
When the service is unreachable, src/lib/vector-search.ts performs browser-local similarity
search over hash-based 128-dim embeddings — lower quality, zero infrastructure.
ChromaDB data persists in the vector-data Docker volume.