LANTERN
An animal and a premise go in — a finished, narrated, animated children's episode comes out, one stanza per scene
LANTERN turns one line — an animal and a premise — into a finished, narrated, animated children's episode. One stanza becomes one scene. Nothing is stitched until a person has signed off twice.
Both human approvals — an animatic pass for timing, a QA gate for motion — sit upstream of the only expensive rendering step, so a mistake is caught while it is still a $0.04 still, never a rendered shot billed by the second.
A calm children’s story channel needs episodes at a cadence one person cannot hand-animate, at a quality bar where generic AI video is instantly disqualifying — morphing faces, characters that redesign themselves between shots, camera moves with no reason to exist. Models will generate video all day. The actual problem is producing consistent, calm, intentional work repeatably: an editorial and QA problem.
A ten-stage pipeline with two human gates, built three times. Almost every hard-won piece of it is a constraint rather than a capability — a rule about what may never appear in a prompt, a gate that refuses to run, a decision to cut instead of stretch. The look lives in one file. The providers sit behind one seam. And the whole thing was written down precisely enough that a fresh agent session rebuilt it from a brief, then shipped Episode 2.
A calm children's story channel needs episodes at a cadence one person cannot hand-animate, at a quality bar where generic AI video is instantly disqualifying — morphing faces, characters that redesign themselves between shots, camera moves with no reason to exist, motion that is busy because the model likes motion.
Models will generate video all day. The actual problem is producing consistent, calm, intentional work repeatably — an editorial and QA problem wearing a technical costume. Almost every hard-won piece of this system is a constraint rather than a capability: a rule about what may never appear in a prompt, a gate that refuses to run, a decision to cut instead of stretch.
The earliest running version of this system makes the true entry point plain. Its story
worker takes nothing but a theme — an animal and a premise — and writes the entire poem
around it: style, stanza count and shape are decided by the prompt, not the author. The poem the rest
of the pipeline consumes is itself generated work. Everything below is verbatim from
story_worker.py and the artifact it left on disk.
the input, complete
“a firefly who discovers her light can show hidden paths through the enchanted forest”
the style is determined, not asked for
Write a children's story about {theme}
in this EXACT poetic style:
- Light rhyming scheme (natural flow)
- Nature as a living character
- 8–12 stanzas, each 4 lines long
- Animals watching from a distance
with golden eyes
- Ends with wisdom or discovery
the poem, as generated
“The Whispering Light”
Deep in the jungle, where sunlight is shy,
And vines weave a curtain from the blue of the sky,
A firefly named Luna lived, with a light so bright,
A spark within her, that guided through the night.
In the current build the generated poem lands in
story.md and gets read by a person before anything else runs — the first of the two
approvals, in effect, moved to the very top. From there on it is stanzas in, scenes out.
The cuts land on narration pauses because the voice was recorded first and everything downstream is timed to a measurement rather than an estimate. Scene tails are slow push-ins, never freezes. The motion runs on twos inside a 24fps container with a fine luma grain over the top — the reason a painted look survives being animated at all.
The two gates are the argument. Both human approvals sit upstream of the expensive step, so the final cut never surprises anyone — and a mistake is caught while it is still a $0.04 still rather than a rendered shot billed by the second.
Fully cloud-rendered: ≈ $16 per episode — ≈220s of video at $0.07/s plus stills at $0.04 each.
Anyone can show the good take. The parts of this project worth reading are the ones that were built and then binned — two entire motion systems, four rounds of style search, and a rendering artifact that turned into a rule the tool now enforces by refusing to run.
Rejected — motion, twice
The bug that became a lint
03_take1_stray_M.mp4, a take
named for the stray letter that killed it.Tracing it back produced a permanent house rule — character
names live in narration, never in a visual prompt — and in the current system that rule is enforced in
three places rather than written on a wall. story.py harvests names out of the poem into a
cast array; scenes.py regexes every image, video, motion and angle prompt
against it; and generation refuses outright:
cast detected: —
try deleting the name and writing “a small firefly” instead
That is the method in miniature. A failure is not fixed once; it is made structurally impossible to repeat.
Search — four rounds to a locked look
Direction — notes, not prompts
The gate, in practice
Every shot produces a five-frame strip reviewed against a fixed rubric before anything is
stitched, and the verdict is written per shot into scenes.json. The gate is not advisory —
a final-stage assembly raises and stops if any shot with footage lacks a pass:
QA gate: these shots have clips but no `pass` verdict: 04b, 07a
Run `storytime review` + `storytime qa`, or --allow-unreviewed.
Episode 1's record: 20 shots across 10 scenes — 3 retaken, 4 trimmed, 14 carrying written notes. The notes read like an edit suite, because that is what they are: “owl blink lands” · “trim 9s to avoid sunset drift” · “retake 2, clean” · “bright wash at end = Luna glow passing”. The gate produces director's notes, not scores. It costs frame extraction and nothing else, and it is the reason the expensive part is not wasted.
Rejected takes are never deleted — 13 sit beside Episode 1, 84 beside Episode 2.
Receipts
find, an ls or an ffprobe away from being checked.The house look was searched, argued over and locked between the two episodes. Episode 1 shipped in the earlier one; Episode 2 is the first in the locked one. Nine generation models were benchmarked on identical inputs before any went into production, and the cheapest was rejected on evidence — it redesigned the character and over-zoomed.
What was handed over
A written brief. The locked style assets. A lessons file with nine numbered entries, each traceable to a specific failure that had already cost real money.
That was it — handed to a fresh agent session with no memory of the previous two builds.
What came back
A 22-module, 2,005-line pipeline with local and cloud providers behind one seam, a shot planner, a QA reviewer, a doctor, a status view and 14 CLI commands — which then shipped Episode 2 in the locked look.
The provider seam is the only piece of architecture that survived all three rewrites unchanged.
What made that possible was not the agent. It was that the previous two builds had been written down precisely enough to be rebuilt from. The transferable skill is specifying a system well enough that someone who was not in the room can build it correctly — the same skill that makes a good creative director's brief. The rebuild is the proof that the specification was real.
| generation | modules · lines | episodes | what it proved |
|---|---|---|---|
| prototype (2025) | 20 app iterations | 0 | Produced loose stills and clips; never assembled an episode. |
| v2 | 12 · 870 | 1 | The spine worked end to end. Rougher, and in the earlier look. |
| v3 | 22 · 2,005 | 1 | Rebuilt from a brief by a fresh session. Local and cloud providers, 14 commands. |
Honest limits
Local generation is supported, not benchmarked. The ComfyUI provider is implemented and routes through the same seam, but no local speed or cost has been measured — so none is claimed.
The narrator is a synthetic preset. It is not presented as a feature here.
Two episodes is two episodes. The cadence is a design target, not observed throughput.
The $16 is modelled, not invoiced — rate-card arithmetic for a clean episode. Actual spend including retakes ran higher.
What is genuinely built
The provider seam. Providers only fill an assets/ directory; nothing downstream
knows which one ran. Three implemented — fal, comfyui, assets
— routable per stage.
The two gates. Both enforced in code, both upstream of the expensive step.
The look as data. One ~10 KB file, deep-merged per story, controls the whole channel.
The no-names lint. Harvested from the poem, checked on every prompt, blocking.
The bake-off and the command line
Two artifacts from the build that the walkthrough above doesn't show.
-
Nine models, benchmarked on identical inputs
The same shot, run across nine video models before any went into production.
-
The working system
The command line the rebuild shipped with: the doctor check, the per-shot status view, review and QA, and the episode command.
Local generation is supported, not benchmarked — the ComfyUI provider routes through the same seam, but no local speed or cost has been measured, so none is claimed. The narrator is a synthetic preset. Two episodes is two episodes; the cadence is a design target, not observed throughput. The $16 is modelled, not invoiced — actual spend including retakes ran higher.
The provider seam: three providers, routable per stage, invisible downstream. The two gates, both enforced in code, both upstream of the expensive step. The look as data: one ~10 KB file, deep-merged per story, controls the whole channel. The no-names lint: harvested from the poem, checked on every prompt, blocking.
Project Elements
Have a Project in Mind?
Animation direction, VFX and creative technology. Message me on LinkedIn or get in touch.