Object Worlds
“One idea. One command. One finished video. No human in the loop.”
A fully automated AI video pipeline. You write an idea in Russian into idea.txt. One Python command launches the whole thing:
Sonnet scores the idea's conflict potential 1–10 in a second. Claude Opus expands it into three FLUX scene prompts (hook → reveal → resolution), three narration lines (five words max), a hook line (provocative, eight words max), and a YouTube title with an emoji. Sonnet drafts the description — episode number, two lines, CTA, five hashtags.
FLUX 1.1 Pro renders three 9:16 images. Sonnet QCs each one on a 1–10 scale; passes if it lands at 7 or above. Wan 2.7 image-to-video generates three 7-second 1080p clips. FFmpeg concatenates them in the order Claude specified. A black hook screen with the hook text (Orbitron) and a random SFX drops in front. FFmpeg overlays subtitles, watermark, CTA, and background music. ElevenLabs TTS renders three narration MP3s (female, multilingual v2), which FFmpeg mixes over, ducking the music to 50% during speech.
The output: a 22.5-second short, ready to upload. If any step crashes, run_status.json remembers where it stopped and the restart resumes from there.