Project Valhalla
In one line: Aura lets an LLM agent write code and run it — and Project Valhalla is the security layer that makes that safe. Scripts are statically analyzed, then executed in an in-browser WebAssembly sandbox with a runtime import firewall, capability isolation, and CPU/memory budgets. No server, no container, no host access.
Letting a language model generate code is easy. Executing that code without handing an attacker — or a hallucination — the keys to the machine is the hard part, and it’s the problem most “AI writes code” demos quietly skip. Valhalla doesn’t skip it.
Why this is hard
An autonomous agent that emits executable scripts is a remote-code-execution engine pointed at your own runtime. To make it safe you have to answer, concretely:
- How do you stop a generated script from reading the filesystem, opening a socket, or shelling out?
- How do you catch the obfuscated escape —
getattr(x, '__class__'),globals()['os'],__builtins__— that a naïve keyword filter misses? - How do you stop an infinite loop or a memory bomb from freezing the tab?
- And how do you do all of this client-side, with no backend to attack?
Valhalla’s answer is defense in depth: five independent layers, each of which would have to fail for an escape to succeed.
How it works
- Intent → script. You, or the Conductor via
launch_valhalla, issue a command for a target tool (Blender, Ableton, Figma, …). The Analyst Virtuoso writes the automation script. - Static analysis.
analyzeScriptDeep()runs regex and a lightweight AST pass, producing aSafetyReport. Any critical finding blocks the script before a single line runs. - Sandboxed execution. Safe scripts run in Pyodide (CPython compiled to WebAssembly,
v0.27.5), loaded lazily and pre-warmed to hide cold start. - Capture & report.
stdout/stderrare captured, matplotlib figures are returned as base64 PNGs, and every run is logged to telemetry.
export interface SandboxExecutionResult {
success: boolean
stdout: string
stderr: string
executionTimeMs: number
images: string[] // base64 data URIs (e.g. matplotlib output)
returnValue: any
safetyReport: SafetyReport
blocked: boolean // true if the safety analyzer rejected the script
error?: string
}The security model — five layers
| # | Layer | Mechanism | Stops |
|---|---|---|---|
| 1 | Pre-flight static analysis | regex + structural AST, merged & de-duplicated | known-dangerous code before execution |
| 2 | Runtime import firewall | builtins.__import__ monkey-patched against a 22-module blocklist | subprocess, os, sys, socket, ctypes, threading, urllib, … even if static analysis is bypassed |
| 3 | Capability isolation | Pyodide runs purely in-browser WASM | no network, no real filesystem (virtual /home/pyodide only), no subprocess — it can’t reach the host |
| 4 | CPU budget | Promise.race against a 30 s timeout | infinite loops / runaway compute |
| 5 | Memory budget | WASM linear-memory ceiling | memory-exhaustion bombs |
The principle is redundancy, not cleverness. Layer 1 is best-effort detection; layers 2–5 are structural — they hold even if the analyzer is fooled. That’s the difference between “we filter bad keywords” and “the runtime is incapable of the dangerous operation.”
What the analyzer catches
Findings fall into six categories, each with a severity, rolled into a score:
- Categories:
dangerous-import·destructive-call·infinite-loop·network-operation·system-escape·file-access - Score:
100 − (criticals × 30) − (warnings × 10); blocked if anycriticalfinding.
The AST layer (valhalla-ast-analyzer.ts) exists to catch what regex cannot — the obfuscated escapes used to break naïve Python sandboxes:
| Pattern | Example | Verdict |
|---|---|---|
| Dynamic dunder access | getattr(obj, '__class__') | critical · system-escape |
| Namespace dictionary access | globals()['os'] | critical · system-escape |
| Builtins tampering | __builtins__ | critical · system-escape |
| Eval/exec injection | x = eval(user_input) | critical · system-escape |
| Dynamic file path | open(some_var) | warning · file-access |
The bridge to external software
Valhalla exists for round-tripping: a project you build in Aura can be handed to professional tools to be taken further. The implemented flow is script-as-deliverable — for a named target, the Analyst Virtuoso emits the exact, safety-vetted script to apply, alongside an in-browser preview.
This follows the Principle of Most Direct Execution (PMDE) — prefer the target’s API / scripting interface first, CLI second, GUI automation last. A generated script is auditable, deterministic, and reviewable in a way synthetic clicks never are. And because that script touches your tools, the Gateway keeps a human in the loop at every step: it surfaces the script, its safety report, live stdout, and the preview, and you approve before anything is applied.
The sandbox being network-isolated is not a barrier to improving Aura projects with external software — that isolation only governs the in-browser vetting and preview core. The bridge to the outside world is the validated script itself; today it’s delivered for you to apply with full visibility, rather than injected blindly into a live session. A direct connector is a natural next layer.
Tool-agnostic by design
The target is a parameter, not a branch in the code:
executeValhallaCommand(toolName, command)
// ^^^^^^^^ any tool — the prompt is templated on it:
// "You are an expert automation agent controlling ${toolName}…"Adding a destination means naming it, not building an integration. Blender, Ableton, and Figma are examples, not a menu. That turns Valhalla from “an AI plus one tool” into a substrate — a general primitive (generate a safety-vetted automation script for tool X) the user composes.
It composes across tools. Because every hop is the same primitive, they chain:
Build a course or storyboard in Aura → vetted Blender script for 3D elements → vetted Ableton Live script for audio → export.
The only boundary is a scriptable surface — which is why PMDE prioritizes scripting. That covers essentially every serious tool:
| Scripting surface | Tools (examples) | Vetting today |
|---|---|---|
| Python | Blender, Houdini, Maya, Nuke, DaVinci Resolve, FreeCAD, Unreal | ✅ Full |
| Python via bridge | Ableton Live (Max for Live / pylive) | ✅ Full |
| JavaScript / TS | Figma plugins, web tooling | ⏳ Add a JS safety profile |
| ExtendScript / UXP | After Effects, Photoshop, Premiere | ⏳ Per-language profile |
Unbounded across script-driveable software, Python-first today. Reaching a new language adds a safety profile — the analyzer is the only language-specific piece — not a rewrite. There is no hard-coded ceiling on which software a creative user can reach.
Where this goes — a primitive, not a feature
Because the core is “safely run model-authored code,” Valhalla generalizes past creative media into any domain where an expert would otherwise hand-write a script:
- Education — a lesson becomes a Blender animation or a generated dataset + plot (Aura’s adaptive courses, made generative).
- Film / VFX & 3D — procedural scenes, batch renders, rigging helpers.
- Music & audio — generative MIDI, device racks, arrangement scaffolding.
- Design — repetitive Figma layout/spec generation from a brief.
- Science & data — simulations and plots on
numpy/scipy/pandas/sympy, in-browser, zero install. - Anything scriptable — the user names the tool.
Most “AI + tool” projects are one hard-wired integration. Valhalla is the safe-execution substrate they could all have been built on — and it grows with what its users can imagine.
Available packages
Pre-bundled in Pyodide: numpy, scipy, matplotlib, pandas, sympy, scikit-learn, Pillow, networkx, and more. Matplotlib figures (rendered via the Agg backend) are captured automatically and returned as images.
Status & honest roadmap
A proof of concept, deliberately transparent about what exists versus what’s next — the value is the architecture, and it’s meant to grow with its users.
Built today: per-tool script generation (api/valhalla.ts); regex + AST static analysis with scoring (valhalla-analyzer.ts, valhalla-ast-analyzer.ts); the Pyodide sandbox — import firewall, capability isolation, timeout, capture (valhalla-sandbox.ts); the Gateway UI with human override (ValhallaGateway.tsx); per-run telemetry; 34 analyzer tests.
Next layers (extensions, not redesigns): an opt-in live connector behind the same human gate; safety profiles for JS / ExtendScript / UXP; a per-run capability-grant model.
The hard, safety-critical core — making model-authored code something you can run without flinching — already exists.
The Aura Symphony repository is public and intentionally transparent about both what Valhalla does and what it can become — the goal is to share this safe-execution pattern with the community.