CLIP Embeddings
backend/clip-embeddings/ (FastAPI + CLIP ViT-B/32, port 3006) embeds frames and text
into a shared vector space, enabling true visual retrieval.
Endpoints
| Method | Path | Purpose |
|---|---|---|
POST | /embed/images | Batch-embed frames |
POST | /embed/image | Embed a single image |
POST | /embed/text | Embed a text query into the same space |
POST | /search/visual | Visual similarity search over embedded frames |
GET | /health · /stats | Probes and counters |
Multimodal RAG
Because CLIP places images and text in one space, a query like “a whiteboard with equations”
matches frames that look like that — even if nobody said those words. The client
(src/lib/clip-embeddings.ts) fuses visual scores with
text-based vector search via fusion scoring across
modalities, producing a single ranked result list for the timeline.
The model cache persists in the clip-data Docker volume, so the first start downloads weights
once.