CASE STUDY

1111

pipeline types

Agentic video production across 11 pipeline types

11

pipeline types

0-code

provider swapping

An agentic video pipeline turning scripts into finished video at volume — programmatic assembly with generated components composited against real brand assets, and a human review gate before anything publishes.

Producing video variants at scale — per market, per language, per product — breaks any workflow built on human editors alone.

OpenMontage is an instruction-driven production system: an LLM agent reads declarative YAML manifests and orchestrates 11 specialized pipeline types, from explainers to avatar spokespersons to localization dubs.

A dynamic tool registry enables zero-code provider swapping and cost governance across video, image, TTS and music generation, with modern composition runtimes rendering the result.

What made this hard

Repeatability over craft

The value is producing the same kind of output reliably at volume. A pipeline optimised for a single impressive result solves the wrong problem.

Generated video fails visibly

The failures are obvious to a person and invisible to any automated check, which makes human review a design requirement rather than a precaution.

Render cost scales with output

At volume it becomes a real line item, so cost per rendered output had to be tracked from the first batch.

Brand assets are not replaceable

Generated components had to composite against real footage, logos and typography rather than substitute for them.

How we built it

  1. Template definition

    The structure every output shares — segments, timing, brand treatment — so the pipeline produces variation within it rather than starting fresh.

  2. Content assembly

    Scripts and assets pulled from source systems so output reflects current information rather than a snapshot.

  3. Programmatic composition

    Deterministic editing, timing and captioning, queued and parallelised for volume.

  4. Review gate

    Human approval before publication, on every output.

The stack

FFmpegProgrammatic video assemblyText-to-speechGenerative video componentsQueued render orchestrationPython

What should you take from this?

Frequently asked questions

Will the output look generated?

Some of it will, which is why every output goes through human review before publication. Where quality matters more than volume, this is the wrong tool and we will say so.

What does it cost to run at volume?

Render cost scales with output and becomes a real line item, which is why it belongs in the business case rather than a footnote. We track cost per output from the first pilot batch.

Can it use our brand assets?

That is the normal case. Generated components composite against your real footage, logos and typography rather than replacing them, which is what keeps output recognisably yours.

Related service: AI Video Production Systems

Build something at this level

Tell us what you are building. We reply within one business day.