The AI Production Pipeline.
The core teaching. A creative isn't one tool, one prompt, one output. It's five layers of work, each best handled by a different tool, all orchestrated through an LLM. Once you understand the layers, the tools become interchangeable. The system stays the same.
You have your brief from Module 02. Now you produce. Most brands fail at this stage by trying to do everything with one tool. They open their favourite LLM, type a prompt, and expect a finished ad. Or they open Capcut and assume a 20 second video falls out fully edited. Neither happens.
The Pipeline is how serious operators produce ad creatives in 2026. It breaks production into five distinct layers. For each layer, you pick the best tool currently available, generate the asset, then pass it to the next layer. An LLM (Claude or your favourite) sits at the centre, orchestrating prompts and decisions across the whole flow.
Five layers, one orchestrator.
Brief in. Finished creative out. Right tool for every layer.
Script
LLM
Image
Fal.ai
Motion
Fal.ai
Audio
ElevenLabs
Edit
Capcut
Claude (or your favourite LLM) writes prompts, picks tools, drafts scripts, generates voiceover transcripts, manages decisions across all five layers.
The mental shift is simple but powerful: stop thinking about tools, start thinking about layers. Tools change every month. Layers don't. Once you know what you're producing at each layer, swapping a tool for a better one becomes trivial.
Layer 01 — Script
Every creative starts with a script. Even a static image ad needs a hook line and body copy. For video, you need scene-by-scene breakdowns: hook, body, payoff, CTA. Script writing is the cheapest layer to iterate on, so do it properly before producing visuals.
Hook variants, scene-by-scene script for video, on-screen text overlays, captions, and the final CTA. One brief should produce 3 to 5 script variants, not just one.
- ClaudeLong-form scripts, persona-aware copy, brand voice
- ChatGPTAlternative LLM with similar capabilities
- GeminiStrong on research-heavy briefs
Whichever LLM you use, the input is your brief from Module 02. The output should be multiple variants, not a single locked draft. You'll cut down to your favourites once you see what works visually in the next layer.
Layer 02 — Image
For static ads, this is the final visual. For video, these are the keyframes that get animated in Layer 03. Image generation is where the bulk of model innovation has happened in the last 18 months. Quality is now strong enough that a well-prompted image can sit alongside professional photography.
Hero shots of your product. Lifestyle scenes. Before and after frames. Storyboard frames for video. Plus alternative compositions, lighting, and angles for variation.
- Fal.aiMarketplace of latest models (recommended)
- MidjourneyStrong for stylised, cinematic visuals
- Nano BananaPhotorealistic product shots
- FluxOpen model, available on Fal.ai
One platform. The best models. Pay only for what you use.
Fal.ai is a marketplace of the most capable AI models, all accessible from a single interface and a single account. Instead of subscribing to ten different tools, you can use Flux, Kling, Veo, Seedance, Higgsfield, Lora, and dozens of others on demand. Pay per generation, not monthly subscriptions. New models appear every week.
For brands building a Pipeline, Fal.ai is the most efficient way to swap tools as the landscape changes without rebuilding your workflow. fal.ai is where most of your image and motion work will live.
Layer 03 — Motion
If you're producing video, this is where your image generations get animated. Image-to-video models are catching up fast. The best ones can produce 5 to 10 second clips with realistic motion, camera moves, and product handling. Skip this layer entirely if your creative is static.
Short video clips (5 to 10 seconds each) animating your hero images. Camera moves, subject motion, product reveals, environmental shifts. Multiple clips will be edited together in Layer 05.
- KlingTop-tier image-to-video, on Fal.ai
- Veo 3Google's flagship, on Fal.ai
- SeedanceFast, cost-effective option on Fal.ai
- RunwayDirect subscription, strong editing tools
Text-to-video (where you skip the image layer entirely) is also possible. The output is usually less controllable than image-to-video. For ad work where consistency matters, image-to-video almost always produces better results.
How much you control the motion depends on your inputs:
One image (lowest control)
Feed a single keyframe to the model and let it animate. Good for quick lifestyle shots or atmospheric clips. Useful when the idea is loose and any reasonable motion will work.
Start frame + end frame (medium control)
Feed two images: where the clip starts and where it ends. The model interpolates the motion between them. Useful when you know exactly how a shot needs to begin and finish but not every frame in between.
Full storyboard (highest control)
Generate every keyframe in advance, then animate short clips between them. Each clip is short and predictable. Useful when your concept is precise or when you need consistency across multiple shots in the same creative.
Match the precision of your input to the precision of your idea. Loose ideas don't need full storyboards. Tight concepts can't be left to a single image.
Layer 04 — Audio
Audio used to require a recording booth and a voice actor. Now AI voice models produce voiceovers indistinguishable from human in 2 minutes for under a dollar. You can also clone your own voice (or anyone you have permission to record) and have your AI avatar read scripts in that exact voice. Music generation tools handle backing tracks too.
Voiceover narration matching your script, AI avatar footage of a presenter, transcripts for captions, and royalty-free background music tracks.
- ElevenLabsBest AI voiceover, voice cloning
- HeyGenAvatar-led video presentations
- SynthesiaAlternative avatar tool
- SunoAI music generation
Layer 05 — Edit
Even with all the AI tools above, you almost always need a video editor for the final step. This is where you cut together clips, layer voiceover, add captions, sync to a beat, drop in text overlays, and export at the right format and aspect ratio for Meta. Editing is the layer where AI hasn't fully replaced humans yet, and probably won't for a while.
The final ad asset. Vertical 9:16 is now the recommended format for all video ads, including feed placements. Captions burned in. Audio mixed. Length tuned. Multiple variants for testing.
- CapcutBest free editor, AI features built in
- Edit (IG)Instagram's native editor, mobile-first
- CanvaEasy templates, captions, batch exports
Capcut quietly added some of the most useful AI features for ecommerce ads: auto captions (matches the Instagram aesthetic instantly), batch edit (apply changes across many videos at once), AI try-on for clothing brands (swap garments on existing models), background removal, and auto-resize across aspect ratios. Most are free or low-cost. Worth the time to learn the basics.
Claude (or your favourite LLM) as the orchestrator
Five layers, five different tools. Without a central operator, switching between them is exhausting. The brief in Layer 01 needs to map to image prompts in Layer 02. Image specs need to flow into motion prompts in Layer 03. Voiceover scripts in Layer 04 need to match the visuals. The on-screen text and end card copy in Layer 05 need to land the message.
An LLM at the centre solves this. Claude (or your favourite LLM) becomes the orchestrator. You feed it your brief once. From there, it generates everything you need at every layer:
Drafts your script variants in Layer 01
Persona-aware, multiple hooks, scene-by-scene structure for video.
Writes image prompts for Layer 02
Specific to the model you'll use (Flux on Fal.ai, Midjourney, Nano Banana). Includes lighting, composition, mood, and aspect ratio.
Builds motion prompts for Layer 03
Camera moves, subject motion, duration, formatted for Kling or Veo or Seedance.
Generates voiceover transcripts in Layer 04
Matched to your script, with delivery notes (tone, pacing, accent) for ElevenLabs.
Drafts on-screen text and edit notes for Layer 05
Hook overlays, end card copy, suggested cuts and timing.
In practice, this means most of your Pipeline work happens in a single LLM conversation. You're not switching between tools to write prompts. You're writing one brief, then directing the LLM to produce everything you need to feed the production tools.
A purpose-built skill that orchestrates the entire Pipeline. Drop it into a Claude project, paste your brief, and Claude walks through every layer with you. Available with installation instructions in the resources tab.
Run this monthly. Don't run it once.
The Pipeline isn't a one-off project. It's a monthly production rhythm. Your Creative Flywheel needs fresh fuel every month. Once your Pipeline is set up, running a batch should take days, not weeks.
From brief to launched ads in 4 weeks.
Strategy
Refresh personas. Update swipe file. Run the HWC Ad Generator. Write briefs for the month.
Production
Run the Pipeline. Layers 01 to 04 done in Claude + Fal.ai + ElevenLabs. Output: raw assets ready to edit.
Edit
Layer 05. Capcut, Edit, or Canva. Final variants in correct aspect ratios. Captions burned in.
Launch & review
Load creatives into a new ad set inside your Prospecting CBO. Track performance. Plan next month.
The cadence creates compounding effects. Each month builds your visual library, your hook bank, and your understanding of what works in your category. By month 3, your second batch produces in half the time as your first.
The traps to avoid
Trying to do all five layers with one tool. No tool wins at every layer. Use the right one each time.
Skipping the brief. Producing without a clear brief produces output that doesn't connect to a real angle or persona.
Producing one variant per concept. The whole point of AI tools is to generate variants cheaply. One creative per concept misses the value.
Polishing forever. Done is better than perfect. Get it into the ad set. The algorithm will tell you what's working.
Treating AI output as final. Always edit. Always layer in your brand voice and visual language. Raw AI output looks like raw AI output.
The principles that matter
Five layers, not five tools. Layers stay constant. Tools change.
Right tool for the right layer. Don't force one tool to do everything.
Fal.ai for image and motion. One platform, all the latest models.
An LLM is the orchestrator. Claude (or your favourite) sits at the centre.
Capcut, Edit, or Canva for the final assembly. Don't skip the editor.
Run the Pipeline monthly. Fresh fuel for the Creative Flywheel.
Variants over polish. Volume + variety beats perfection every time.