← Creative Lab
Personal Venture — 2026

LANTERN

An animal and a premise go in — a finished, narrated, animated children's episode comes out, one stanza per scene

LANTERN turns one line — an animal and a premise — into a finished, narrated, animated children's episode. One stanza becomes one scene. Nothing is stitched until a person has signed off twice.

2
Episodes delivered
≈ $16
Compute per episode
22
Modules, rebuilt from a brief
The hard part

Both human approvals — an animatic pass for timing, a QA gate for motion — sit upstream of the only expensive rendering step, so a mistake is caught while it is still a $0.04 still, never a rendered shot billed by the second.

Client
Personal Venture
Role
Creator / Writer / Pipeline Architect Style System Design / Spec for Agent Rebuild
Timeline
2026
The Problem

A calm children’s story channel needs episodes at a cadence one person cannot hand-animate, at a quality bar where generic AI video is instantly disqualifying — morphing faces, characters that redesign themselves between shots, camera moves with no reason to exist. Models will generate video all day. The actual problem is producing consistent, calm, intentional work repeatably: an editorial and QA problem.

The Build

A ten-stage pipeline with two human gates, built three times. Almost every hard-won piece of it is a constraint rather than a capability — a rule about what may never appear in a prompt, a gate that refuses to run, a decision to cut instead of stretch. The look lives in one file. The providers sit behind one seam. And the whole thing was written down precisely enough that a fresh agent session rebuilt it from a brief, then shipped Episode 2.

Pythonffmpegfal.aiComfyUIKling · Seedream14 CLI commands
The problemit is not "generate video"

A calm children's story channel needs episodes at a cadence one person cannot hand-animate, at a quality bar where generic AI video is instantly disqualifying — morphing faces, characters that redesign themselves between shots, camera moves with no reason to exist, motion that is busy because the model likes motion.

Models will generate video all day. The actual problem is producing consistent, calm, intentional work repeatably — an editorial and QA problem wearing a technical costume. Almost every hard-won piece of this system is a constraint rather than a capability: a rule about what may never appear in a prompt, a gate that refuses to run, a decision to cut instead of stretch.

Stage zero — the input is one linetheme → poem, before any pipeline runs

The earliest running version of this system makes the true entry point plain. Its story worker takes nothing but a theme — an animal and a premise — and writes the entire poem around it: style, stanza count and shape are decided by the prompt, not the author. The poem the rest of the pipeline consumes is itself generated work. Everything below is verbatim from story_worker.py and the artifact it left on disk.

the input, complete

“a firefly who discovers her light can show hidden paths through the enchanted forest”

the style is determined, not asked for

Write a children's story about {theme}
in this EXACT poetic style:
- Light rhyming scheme (natural flow)
- Nature as a living character
- 8–12 stanzas, each 4 lines long
- Animals watching from a distance
  with golden eyes
- Ends with wisdom or discovery

the poem, as generated

“The Whispering Light”

Deep in the jungle, where sunlight is shy,
And vines weave a curtain from the blue of the sky,
A firefly named Luna lived, with a light so bright,
A spark within her, that guided through the night.

10 stanzas · saved with a per-stanza image-prompt file beside it

In the current build the generated poem lands in story.md and gets read by a person before anything else runs — the first of the two approvals, in effect, moved to the very top. From there on it is stanzas in, scenes out.

Episode 2 — “The Moth at Midnight”3:33 · 1080p24 · the locked look
10 scenes 26 shots 26 QA frame-strips 11 narration segments silent here — the narration carries the piece, but this page plays muted

The cuts land on narration pauses because the voice was recorded first and everything downstream is timed to a measurement rather than an estimate. Scene tails are slow push-ins, never freezes. The motion runs on twos inside a 24fps container with a fine luma grain over the top — the reason a painted look survives being animated at all.

The pipelineten stages, two gates — press play, or step it

The two gates are the argument. Both human approvals sit upstream of the expensive step, so the final cut never surprises anyone — and a mistake is caught while it is still a $0.04 still rather than a rendered shot billed by the second.

story one stanza → one scene story.md → scenes.json narration generated FIRST NN.mp3 shot plan from measured VO scenes.json · shots[] scene stills the world, as images NN.png ANIMATIC approval #1 timing angle stills same scene, new camera NN_b.png shot clips native speed, 4–15s NN_a.mp4 QA GATE approval #2 motion episode cuts · cards · bed · 1080p <slug>_final.mp4 provider rail — narration · stills · angles · clips are each routable: --images-provider, --video-provider
step 1/9 A stanza becomes a scene.
counted live

Fully cloud-rendered: ≈ $16 per episode — ≈220s of video at $0.07/s plus stills at $0.04 each.

What was thrown awaythe decision trail, which is the actual work

Anyone can show the good take. The parts of this project worth reading are the ones that were built and then binned — two entire motion systems, four rounds of style search, and a rendering artifact that turned into a rule the tool now enforces by refusing to run.

Rejected — motion, twice

Three motion systems compared: ping-pong looping, time-stretch with optical-flow interpolation, and native-speed shot coverage.
Three motion systems, two discarded. First, ping-pong looping — short clips played forward then reversed to fill a scene; the reversal is instantly legible as fake. Second, time-stretch with optical-flow interpolation — 15s clips slowed roughly 2× to span a 30s scene; interpolation artifacts, and the slowdown itself read as wrong. What shipped was neither: cover the scene with two or three real shots at native speed and cut between them, the last holding its final frame through the breath. The fix was editorial, not technical.

The bug that became a lint

QA gate contact sheet showing real caught failures including a frame with lettering painted into the artwork.
A model painted the character's name into the picture. An illustration prompted with the character's name came back with the name itself lettered into the frame beside her. It propagated — every reference-conditioned still inherited the lettering, and every video model faithfully animated it. The residue is still on disk as 03_take1_stray_M.mp4, a take named for the stray letter that killed it.

Tracing it back produced a permanent house rule — character names live in narration, never in a visual prompt — and in the current system that rule is enforced in three places rather than written on a wall. story.py harvests names out of the poem into a cast array; scenes.py regexes every image, video, motion and angle prompt against it; and generation refuses outright:

cast detected: —

try deleting the name and writing “a small firefly” instead


    

That is the method in miniature. A failure is not fixed once; it is made structurally impossible to repeat.

Search — four rounds to a locked look

Four rounds of style search narrowing from fifteen candidate images to a single locked look.
34 images across four rounds — 15, then 9, then 6, then 4 — narrowing to one look for about $2. The result is not a moodboard: it is a ~10 KB JSON file — prompt prefixes, camera language, retiming, grain, card typography, the music bed. Edit that one file and every episode re-skins.

Direction — notes, not prompts

Art direction notes translated into before and after renders, with a character reference sheet.
Plain-language notes became precise edits to a shared style block, verified on re-render. “Don't make her so cutesy.” “It can just be a nice glowy flower.” That loop — taste, applied through a system, checked against the output — is direction, not prompting until something looks nice.

The gate, in practice

Every shot produces a five-frame strip reviewed against a fixed rubric before anything is stitched, and the verdict is written per shot into scenes.json. The gate is not advisory — a final-stage assembly raises and stops if any shot with footage lacks a pass:

QA gate: these shots have clips but no `pass` verdict: 04b, 07a
Run `storytime review` + `storytime qa`, or --allow-unreviewed.

Episode 1's record: 20 shots across 10 scenes — 3 retaken, 4 trimmed, 14 carrying written notes. The notes read like an edit suite, because that is what they are: “owl blink lands” · “trim 9s to avoid sunset drift” · “retake 2, clean” · “bright wash at end = Luna glow passing”. The gate produces director's notes, not scores. It costs frame extraction and nothing else, and it is the reason the expensive part is not wasted.

Rejected takes are never deleted — 13 sit beside Episode 1, 84 beside Episode 2.

Receipts

A dated census of the archive: modules, assets, takes, tests and episodes across three generations of the system.
The archive census, counted live from disk — every figure on this page is a find, an ls or an ffprobe away from being checked.
Two episodes, two lookssaid plainly, because implied continuity would be a lie
Episode 1 — “The Whispering Light”, 3:59 · 1080p30. The earlier look.
Grain test — the reason a painted surface survives being animated.

The house look was searched, argued over and locked between the two episodes. Episode 1 shipped in the earlier one; Episode 2 is the first in the locked one. Nine generation models were benchmarked on identical inputs before any went into production, and the cheapest was rejected on evidence — it redesigned the character and over-zoomed.

The handoffthe strongest claim here

What was handed over

A written brief. The locked style assets. A lessons file with nine numbered entries, each traceable to a specific failure that had already cost real money.

That was it — handed to a fresh agent session with no memory of the previous two builds.

What came back

A 22-module, 2,005-line pipeline with local and cloud providers behind one seam, a shot planner, a QA reviewer, a doctor, a status view and 14 CLI commands — which then shipped Episode 2 in the locked look.

The provider seam is the only piece of architecture that survived all three rewrites unchanged.

What made that possible was not the agent. It was that the previous two builds had been written down precisely enough to be rebuilt from. The transferable skill is specifying a system well enough that someone who was not in the room can build it correctly — the same skill that makes a good creative director's brief. The rebuild is the proof that the specification was real.

generationmodules · linesepisodeswhat it proved
prototype (2025)20 app iterations0Produced loose stills and clips; never assembled an episode.
v212 · 8701The spine worked end to end. Rougher, and in the earlier look.
v322 · 2,0051Rebuilt from a brief by a fresh session. Local and cloud providers, 14 commands.
Episode 2, The Moth at Midnight, playing on the production machine.
Delivered. Episode 2, “The Moth at Midnight” — 3:33 at 1080p, playing on the production machine. The story was one of the five written for the 2025 prototype; it waited two generations of the system for something that could produce it.

Honest limits

Local generation is supported, not benchmarked. The ComfyUI provider is implemented and routes through the same seam, but no local speed or cost has been measured — so none is claimed.

The narrator is a synthetic preset. It is not presented as a feature here.

Two episodes is two episodes. The cadence is a design target, not observed throughput.

The $16 is modelled, not invoiced — rate-card arithmetic for a clean episode. Actual spend including retakes ran higher.

What is genuinely built

The provider seam. Providers only fill an assets/ directory; nothing downstream knows which one ran. Three implemented — fal, comfyui, assets — routable per stage.

The two gates. Both enforced in code, both upstream of the expensive step.

The look as data. One ~10 KB file, deep-merged per story, controls the whole channel.

The no-names lint. Harvested from the poem, checked on every prompt, blocking.

PythonffmpegfalComfyUIKling Seedreamprovider abstractionQA gatingshot planning editorial systemsagent handoff
Two more artifacts

The bake-off and the command line

Two artifacts from the build that the walkthrough above doesn't show.

  1. Nine models, benchmarked on identical inputs

    The same shot, run across nine video models before any went into production.

    model bake-off · one shot, nine renders
    Model bake-off: the same shot rendered across nine video generation models
  2. The working system

    The command line the rebuild shipped with: the doctor check, the per-shot status view, review and QA, and the episode command.

    storytime doctor · status · review · qa · episode
    The working system: the CLI, the doctor check and the per-shot status view
Limits

Local generation is supported, not benchmarked — the ComfyUI provider routes through the same seam, but no local speed or cost has been measured, so none is claimed. The narrator is a synthetic preset. Two episodes is two episodes; the cadence is a design target, not observed throughput. The $16 is modelled, not invoiced — actual spend including retakes ran higher.

What is built

The provider seam: three providers, routable per stage, invisible downstream. The two gates, both enforced in code, both upstream of the expensive step. The look as data: one ~10 KB file, deep-merged per story, controls the whole channel. The no-names lint: harvested from the poem, checked on every prompt, blocking.

Scope

Project Elements

Pipeline ArchitectureGenerative WritingStyle System DesignAI Illustration & MotionNarrationHuman-in-the-Loop QAMulti-Provider RoutingCLI ToolingAgent Handoff
Next Project
VIOLET Studio →

Have a Project in Mind?

Animation direction, VFX and creative technology. Message me on LinkedIn or get in touch.