Skip to Content
Project Valhalla

Project Valhalla

In one line: Aura lets an LLM agent write code and run it — and Project Valhalla is the security layer that makes that safe. Scripts are statically analyzed, then executed in an in-browser WebAssembly sandbox with a runtime import firewall, capability isolation, and CPU/memory budgets. No server, no container, no host access.

Letting a language model generate code is easy. Executing that code without handing an attacker — or a hallucination — the keys to the machine is the hard part, and it’s the problem most “AI writes code” demos quietly skip. Valhalla doesn’t skip it.

Why this is hard

An autonomous agent that emits executable scripts is a remote-code-execution engine pointed at your own runtime. To make it safe you have to answer, concretely:

  • How do you stop a generated script from reading the filesystem, opening a socket, or shelling out?
  • How do you catch the obfuscated escape — getattr(x, '__class__'), globals()['os'], __builtins__ — that a naïve keyword filter misses?
  • How do you stop an infinite loop or a memory bomb from freezing the tab?
  • And how do you do all of this client-side, with no backend to attack?

Valhalla’s answer is defense in depth: five independent layers, each of which would have to fail for an escape to succeed.

How it works

  1. Intent → script. You, or the Conductor via launch_valhalla, issue a command for a target tool (Blender, Ableton, Figma, …). The Analyst Virtuoso writes the automation script.
  2. Static analysis. analyzeScriptDeep() runs regex and a lightweight AST pass, producing a SafetyReport. Any critical finding blocks the script before a single line runs.
  3. Sandboxed execution. Safe scripts run in Pyodide (CPython compiled to WebAssembly, v0.27.5), loaded lazily and pre-warmed to hide cold start.
  4. Capture & report. stdout/stderr are captured, matplotlib figures are returned as base64 PNGs, and every run is logged to telemetry.
src/lib/valhalla-sandbox.ts
export interface SandboxExecutionResult { success: boolean stdout: string stderr: string executionTimeMs: number images: string[] // base64 data URIs (e.g. matplotlib output) returnValue: any safetyReport: SafetyReport blocked: boolean // true if the safety analyzer rejected the script error?: string }

The security model — five layers

#LayerMechanismStops
1Pre-flight static analysisregex + structural AST, merged & de-duplicatedknown-dangerous code before execution
2Runtime import firewallbuiltins.__import__ monkey-patched against a 22-module blocklistsubprocess, os, sys, socket, ctypes, threading, urllib, … even if static analysis is bypassed
3Capability isolationPyodide runs purely in-browser WASMno network, no real filesystem (virtual /home/pyodide only), no subprocess — it can’t reach the host
4CPU budgetPromise.race against a 30 s timeoutinfinite loops / runaway compute
5Memory budgetWASM linear-memory ceilingmemory-exhaustion bombs

The principle is redundancy, not cleverness. Layer 1 is best-effort detection; layers 2–5 are structural — they hold even if the analyzer is fooled. That’s the difference between “we filter bad keywords” and “the runtime is incapable of the dangerous operation.”

What the analyzer catches

Findings fall into six categories, each with a severity, rolled into a score:

  • Categories: dangerous-import · destructive-call · infinite-loop · network-operation · system-escape · file-access
  • Score: 100 − (criticals × 30) − (warnings × 10); blocked if any critical finding.

The AST layer (valhalla-ast-analyzer.ts) exists to catch what regex cannot — the obfuscated escapes used to break naïve Python sandboxes:

PatternExampleVerdict
Dynamic dunder accessgetattr(obj, '__class__')critical · system-escape
Namespace dictionary accessglobals()['os']critical · system-escape
Builtins tampering__builtins__critical · system-escape
Eval/exec injectionx = eval(user_input)critical · system-escape
Dynamic file pathopen(some_var)warning · file-access

The bridge to external software

Valhalla exists for round-tripping: a project you build in Aura can be handed to professional tools to be taken further. The implemented flow is script-as-deliverable — for a named target, the Analyst Virtuoso emits the exact, safety-vetted script to apply, alongside an in-browser preview.

This follows the Principle of Most Direct Execution (PMDE) — prefer the target’s API / scripting interface first, CLI second, GUI automation last. A generated script is auditable, deterministic, and reviewable in a way synthetic clicks never are. And because that script touches your tools, the Gateway keeps a human in the loop at every step: it surfaces the script, its safety report, live stdout, and the preview, and you approve before anything is applied.

The sandbox being network-isolated is not a barrier to improving Aura projects with external software — that isolation only governs the in-browser vetting and preview core. The bridge to the outside world is the validated script itself; today it’s delivered for you to apply with full visibility, rather than injected blindly into a live session. A direct connector is a natural next layer.

Tool-agnostic by design

The target is a parameter, not a branch in the code:

src/api/valhalla.ts
executeValhallaCommand(toolName, command) // ^^^^^^^^ any tool — the prompt is templated on it: // "You are an expert automation agent controlling ${toolName}…"

Adding a destination means naming it, not building an integration. Blender, Ableton, and Figma are examples, not a menu. That turns Valhalla from “an AI plus one tool” into a substrate — a general primitive (generate a safety-vetted automation script for tool X) the user composes.

It composes across tools. Because every hop is the same primitive, they chain:

Build a course or storyboard in Aura → vetted Blender script for 3D elements → vetted Ableton Live script for audio → export.

The only boundary is a scriptable surface — which is why PMDE prioritizes scripting. That covers essentially every serious tool:

Scripting surfaceTools (examples)Vetting today
PythonBlender, Houdini, Maya, Nuke, DaVinci Resolve, FreeCAD, Unreal✅ Full
Python via bridgeAbleton Live (Max for Live / pylive)✅ Full
JavaScript / TSFigma plugins, web tooling⏳ Add a JS safety profile
ExtendScript / UXPAfter Effects, Photoshop, Premiere⏳ Per-language profile

Unbounded across script-driveable software, Python-first today. Reaching a new language adds a safety profile — the analyzer is the only language-specific piece — not a rewrite. There is no hard-coded ceiling on which software a creative user can reach.

Where this goes — a primitive, not a feature

Because the core is “safely run model-authored code,” Valhalla generalizes past creative media into any domain where an expert would otherwise hand-write a script:

  • Education — a lesson becomes a Blender animation or a generated dataset + plot (Aura’s adaptive courses, made generative).
  • Film / VFX & 3D — procedural scenes, batch renders, rigging helpers.
  • Music & audio — generative MIDI, device racks, arrangement scaffolding.
  • Design — repetitive Figma layout/spec generation from a brief.
  • Science & data — simulations and plots on numpy / scipy / pandas / sympy, in-browser, zero install.
  • Anything scriptable — the user names the tool.

Most “AI + tool” projects are one hard-wired integration. Valhalla is the safe-execution substrate they could all have been built on — and it grows with what its users can imagine.

Available packages

Pre-bundled in Pyodide: numpy, scipy, matplotlib, pandas, sympy, scikit-learn, Pillow, networkx, and more. Matplotlib figures (rendered via the Agg backend) are captured automatically and returned as images.

Status & honest roadmap

A proof of concept, deliberately transparent about what exists versus what’s next — the value is the architecture, and it’s meant to grow with its users.

Built today: per-tool script generation (api/valhalla.ts); regex + AST static analysis with scoring (valhalla-analyzer.ts, valhalla-ast-analyzer.ts); the Pyodide sandbox — import firewall, capability isolation, timeout, capture (valhalla-sandbox.ts); the Gateway UI with human override (ValhallaGateway.tsx); per-run telemetry; 34 analyzer tests.

Next layers (extensions, not redesigns): an opt-in live connector behind the same human gate; safety profiles for JS / ExtendScript / UXP; a per-run capability-grant model.

The hard, safety-critical core — making model-authored code something you can run without flinching — already exists.

The Aura Symphony repository is public and intentionally transparent about both what Valhalla does and what it can become — the goal is to share this safe-execution pattern with the community.