Skip to Content
Core FeaturesSemantic Search & RAG

Semantic Search & Multimodal RAG

Ask the Conductor “Where did they discuss the black hole?” and markers drop on the timeline at the exact timestamps. Three systems make that work.

Adaptive semantic chunking

src/lib/semantic-chunker.ts replaces fixed-size word-count chunking with topic-boundary detection:

  1. Split the transcript into timestamped sentences
  2. Embed each sentence
  3. Compute cosine similarity between consecutive pairs
  4. Where similarity drops below a percentile-based dynamic threshold, cut a chunk boundary
ParameterValuePurpose
MIN_CHUNK_WORDS30Prevents micro-fragments
MAX_CHUNK_WORDS300Prevents mega-chunks
BREAK_PERCENTILE25Lower = fewer, larger chunks
EMBED_DIM128Hash-based local vectorizer dimension

Chunks carry start/end timestamps, which is what lets search results become timeline markers — and the building blocks the Create Course lens and timeline edits assemble from. Search isn’t a destination here; it’s how Aura finds the moments worth creating from.

Chunks are ingested into the Vector Search service — FastAPI + ChromaDB — and queried by the Conductor’s search_video function. When the backend is down, the client (src/lib/vector-search.ts) falls back to browser-local similarity search.

Multimodal RAG (CLIP)

src/lib/clip-embeddings.ts + the CLIP Embeddings service add true visual search: video frames are embedded with CLIP ViT-B/32, so a query can match what is seen, not just what is said. Results are combined with text retrieval via fusion scoring across modalities.

Worker pool

Frame extraction and heavy processing run in a WebWorker pool (src/lib/worker-pool.ts):

  • Pool size n = navigator.hardwareConcurrency (capped at 8)
  • Per-worker FIFO queues with work-stealing from the tail — idle workers steal from busy ones
  • Pool stats (queue depths, steal counts) feed the telemetry layer