# Run notes — September 29, 2026

## Source and method

Original eight-second, 1280×720, 30 fps Python/Pillow illustration. The visible contacts are defined by source animation: cup at 2.00 s, pour at 3.00–5.00 s, spoon at 6.20 s. It is deliberately stylized, not camera footage. The source was refined after visual review (antialiasing, cup/saucer contact, pitcher shape and liquid fill). The first MMAudio take used the earlier source; both are retained under `first-pass/`. The public comparison uses only the final source.

Automatic workflow: one complete MMAudio V2 run (`fal-ai/mmaudio-v2`) on the final video, duration 8, seed 420, cfg 4.5, 25 steps. The full scene prompt and negative prompt are in prompts.json.

Layered workflow: one ElevenLabs SFX V2 (`fal-ai/elevenlabs/sound-effects/v2`) run per cue, four requests, durations 2, 3, 2 and 8 seconds, prompt_influence 0.45, mp3_44100_192, looping enabled only for room. Original MP3s and 48 kHz WAV conversions are included. Converting to WAV does not restore detail lost in MP3 encoding. Pour leading trim 0.22 s and duration 2 s, spoon leading trim 0.37 s and duration 1.2 s, each with a 0.15 s ending fade.

## Controlled revision

mix-v1.json deliberately schedules the cup 0.40 s late (12 frames). This is a teaching exercise, not a discovered model failure. In the actual Voyager desktop, pin placed at 2.00 s asks the agent to move the cup cue to that contact and change gain 0.65 → 0.50. The attached review was sent. Agent wrote mix-v2.json and layered-v2.mp4. Exact source/stem/initialrecipe/initialexport hashes remained unchanged; the other three audio entries stayed identical. No generation was needed for the revision.

Exact pinned input:

> Align the cup impact with this contact frame and lower its volume from 0.65 to 0.50. Save mix-v2.json and export layered-v2.mp4. Preserve the pouring, spoon and room entries, all generated source files, mix-v1.json and layered-v1.mp4. Do not generate new audio.

Tested in Voyager development source fffef4cec, Codex 0.158.0, GPT-6-Astra, isolated Linux/Electron profile. Setup preceding the final run used an older local build; final demo and revision use the coherent current build. The desktop demo removes waiting, includes saved-state revisits and still holds, and inserts exported output for clear playback.

## Levels, cost and limitations

The public comparison is matched with a linear FFmpeg loudness pass targeting −24 LUFS, preserving dynamics. Measured delivered outputs: automatic −23.97 LUFS / −3.19 dBTP; layered −23.99 LUFS / −1.64 dBTP. Raw model outputs and working mixes remain unnormalized in this bundle.

Six paid requests total: two eight-second MMAudio runs (first fixture plus refined fixture) and four SFX runs totaling 15 seconds. Public fal documentation retrieved September 29 quotes $0.001/s for MMAudio and $0.002/s for SFX: arithmetic estimate $0.046 total. This is a public-rate estimate, not a receipt for the Moda account charge. No exact workspace charge was returned. Coding-agent access is separate.

This is a workflow demonstration on one illustrated scene. Inputs and amount of direction differ; it is not a controlled model benchmark. Technical review decoded media, measured timing/levels and inspected visual frame sequences. No available reviewer could listen perceptually; do not treat numeric analysis as a verdict on naturalness or sound quality. Listen to the supplied examples when judging that.

Sources: https://hkchengrex.com/MMAudio/ ; https://fal.ai/models/fal-ai/mmaudio-v2 ; https://fal.ai/models/fal-ai/elevenlabs/sound-effects/v2 ; https://elevenlabs.io/docs/overview/capabilities/sound-effects
