Skip to Content
Backend ServicesCLIP Embeddings

CLIP Embeddings

backend/clip-embeddings/ (FastAPI + CLIP ViT-B/32, port 3006) embeds frames and text into a shared vector space, enabling true visual retrieval.

Endpoints

MethodPathPurpose
POST/embed/imagesBatch-embed frames
POST/embed/imageEmbed a single image
POST/embed/textEmbed a text query into the same space
POST/search/visualVisual similarity search over embedded frames
GET/health · /statsProbes and counters

Multimodal RAG

Because CLIP places images and text in one space, a query like “a whiteboard with equations” matches frames that look like that — even if nobody said those words. The client (src/lib/clip-embeddings.ts) fuses visual scores with text-based vector search via fusion scoring across modalities, producing a single ranked result list for the timeline.

The model cache persists in the clip-data Docker volume, so the first start downloads weights once.