The Conductor
gemini-2.5-proOrchestra Lead
Interprets your intent and delegates to the specialists. Complex queries route through a ReAct plan–act–observe loop; everything else is a single-pass function call.
Aura is a zero-surface, multi-agent platform for deep video analysis, generative media, and adaptive learning. Speak your intent — the Conductor commissions the rest.
Each agent is an expert in one domain, with its own model, constitution, and toolset. They converse over the Symphony Bus — commissions chain from one specialist to the next, and the Critic gates every result before it reaches you.
Orchestra Lead
Interprets your intent and delegates to the specialists. Complex queries route through a ReAct plan–act–observe loop; everything else is a single-pass function call.
Visual Analyst
Deep multimodal analysis of video frames and images — composition, color theory, motion, hidden details — structured as Insights for the Lens Laboratory.
Research Specialist
Grounds the analysis in the real world via Google Search — fact checks, recent events, external context — and always cites its sources.
Media Creator
Crafts new media with generative video and image models. Text-to-video, image generation, and instruction-driven image editing.
Logic Specialist
Complex reasoning, structured data extraction, and synthesis. Builds courses from raw analysis and tracks learning progress.
Documentation Lead
Summarizes sessions, organizes insights into reports, generates TTS, and prepares exports — Markdown, PDF, and NLE timelines.
Quality Evaluator
The adversarial quality gate. Scores every output on relevance, factual consistency, and quality — and triggers retry loops when standards are not met.
Under the podium: streaming inference, vector search, CRDTs, WASM sandboxes, and a worker pool — all degrading gracefully to browser-local alternatives when the backends are off.
Ask "where did they discuss the black hole?" and markers drop on the timeline at the exact timestamps. Adaptive semantic chunking + ChromaDB vector search.
Every Virtuoso streams token-by-token via generateContentStream — 60–80% lower perceived latency.
Multi-step queries route through a Reason+Act loop: the Conductor plans, executes, observes, and adapts.
Wake word or mic click — speak your commands for a true zero-surface experience.
Agents generate automation scripts for external creative tools, executed in a Pyodide (Python-in-WASM) sandbox behind 3 security layers.
Any video becomes a structured course. Bayesian Knowledge Tracing with temporal decay powers a calibrated Digital Learner Profile.
Real-time multi-user editing via Yjs — conflict-free shared state, WebSocket transport, peer cursors.
CLIP ViT-B/32 frame embeddings enable true visual search next to text retrieval, fused across modalities.
Frame extraction distributed across N workers with work-stealing load balancing.
Send timelines, annotations, and generated assets to Premiere, Final Cut, or Resolve via FCPXML, EDL, or CSV.
Install third-party Virtuosos with sandboxed execution and SHA-256 integrity verification — or build your own in Agent Studio.
Service Worker caching + an IndexedDB mutation queue give full offline operation with background sync.
Lenses are focused analysis programs in the Lens Laboratory — point one at your media and the orchestra plays it back as structure: captions, charts, courses, diagrams, new media.
Six services, one docker-compose. Every backend is optional — the frontend falls back to browser-local processing, so the symphony never stops.
| Service | Stack | Port | Role |
|---|---|---|---|
| Frontend | React 19 + Vite 6 + TypeScript | 3000 | SPA, agent orchestration, frame extraction |
| API Proxy | Express | 3005 | Gemini key isolation, rate limiting, usage metering |
| Vector Search | FastAPI + ChromaDB | 3001 | Semantic search with adaptive chunking |
| Graph Knowledge | Express + SQLite | 3004 | Concept graph traversal, learning paths |
| Media Pipeline | Express + FFmpeg + WebSocket | 3002/3003 | Cloud-side frame extraction, transcription |
| CLIP Embeddings | FastAPI + CLIP ViT-B/32 | 3006 | Multimodal visual search |
Defense-in-depth AI safety: Zod schema validation on all LLM function calls → Critic adversarial quality gate → Valhalla 3-layer sandbox → plugin sandboxed execution.