Semantic Search & Multimodal RAG
Ask the Conductor “Where did they discuss the black hole?” and markers drop on the timeline at the exact timestamps. Three systems make that work.
Adaptive semantic chunking
src/lib/semantic-chunker.ts replaces fixed-size word-count chunking with topic-boundary
detection:
- Split the transcript into timestamped sentences
- Embed each sentence
- Compute cosine similarity between consecutive pairs
- Where similarity drops below a percentile-based dynamic threshold, cut a chunk boundary
| Parameter | Value | Purpose |
|---|---|---|
MIN_CHUNK_WORDS | 30 | Prevents micro-fragments |
MAX_CHUNK_WORDS | 300 | Prevents mega-chunks |
BREAK_PERCENTILE | 25 | Lower = fewer, larger chunks |
EMBED_DIM | 128 | Hash-based local vectorizer dimension |
Chunks carry start/end timestamps, which is what lets search results become timeline markers — and the building blocks the Create Course lens and timeline edits assemble from. Search isn’t a destination here; it’s how Aura finds the moments worth creating from.
Vector search
Chunks are ingested into the Vector Search service — FastAPI +
ChromaDB — and queried by the Conductor’s search_video function. When the backend is down,
the client (src/lib/vector-search.ts) falls back to browser-local similarity search.
Multimodal RAG (CLIP)
src/lib/clip-embeddings.ts + the CLIP Embeddings service add
true visual search: video frames are embedded with CLIP ViT-B/32, so a query can match
what is seen, not just what is said. Results are combined with text retrieval via fusion
scoring across modalities.
Worker pool
Frame extraction and heavy processing run in a WebWorker pool (src/lib/worker-pool.ts):
- Pool size
n = navigator.hardwareConcurrency(capped at 8) - Per-worker FIFO queues with work-stealing from the tail — idle workers steal from busy ones
- Pool stats (queue depths, steal counts) feed the telemetry layer