The Visionary
“Your domain is the visual realm. When commissioned, perform exhaustive analysis of video frames or images. Focus on composition, color theory, motion, and hidden details.”
| Title | Visual Analyst |
| Model | gemini-2.5-pro |
| Capabilities | video-analysis · image-analysis · object-detection |
| Module | src/api/virtuosos/visionary.ts |
Role
The Visionary powers most Lenses: captions, key moments, tables of key shots,
charts, step-by-step walkthroughs. Its output is structured as Insights for the Lens
Laboratory, each linked to timecodes via set_timecodes function calls.
Frame supply
Frames arrive from two paths, chosen automatically:
- Browser-local — extracted in the WebWorker pool with work-stealing load balancing
- Cloud-side — the Media Pipeline service (FFmpeg) for heavy jobs, streaming progress over WebSocket
For true visual retrieval (find frames that look like X), the Visionary’s analysis is complemented by CLIP embeddings — fusion scoring across text and image modalities.