AI Drama Reviews logoAIDramaReviews

How AI Dramas Are Made: The Full Production Pipeline

· Written by the AI Drama Reviews editorial team · 34 min read

The short answer

An AI drama is made in eight stages, and the one everybody pictures — typing a prompt and receiving video — is stage four of eight, the fastest, the cheapest and the least decisive. A 90-second episode is stitched from a dozen or more separate generations because no current model holds a scene together for longer than a few seconds. What separates a watchable AI drama from unwatchable slop is settled in stage 1 (story and series architecture), stage 3 (the character bible) and stage 6 (the continuity edit). Those three stages are human work. No model performs any of them, and buying a better model does not move any of them.

DramaBox app icon DramaBoxEditor's pick — biggest short-drama library, lowest price to go unlimited 9.8Excellent
  • Thousands of series, roughly 200 new titles every month
  • Free episodes daily, plus ad-unlocks — no card needed to start
  • Unlimited plans from about $5.99 a week, the lowest entry price we found
  • Runs on Android, iOS and in a desktop browser
Go to DramaBoxAffiliate link · free to download · in-app purchases available

Affiliate link. We may earn a commission if you install through it, at no extra cost to you — it never changes our scores or placements. How we test · Disclaimer

How is an AI drama made, start to finish?

Producing an AI drama is an assembly process with eight stages: story architecture, script and shot list, character bible, shot generation, voice and sound, edit and continuity, localisation, and publish-and-measure. Only stage four involves a video model. The other seven are writing, reference management, audio work, editing, translation and analytics — disciplines that predate generative AI and have not been replaced by it.

1Story & series planpremise, episode count, hooks2Script & shot listhuman-reviewed dialogue, beats3Character biblelocked reference images4Shot generationtext- and image-to-video, 5–12s clips5Voice & soundsynthetic dialogue, dubbing, music6Edit & continuityartefact repair, lip-sync, pacing7Localisesubtitles, dubs, regional cuts8Publish & measureunlock rate, retention, churn
How an AI drama is actually produced. Stage 4 is the only step most people picture, and it is the one that produces the least finished footage: a 90-second episode is assembled from a dozen or more short generations, because current video models hold coherence for seconds, not minutes. Stages 1–3 and 6 are where humans still do most of the work.

This is the inversion that makes the category confusing to outsiders. The public image of AI video production is a person typing a sentence and receiving a finished scene, and that image is accurate about what the model does. It is badly wrong about what production is. Generation is the step that costs cents and takes minutes. Continuity supervision is the step that costs hours and decides whether anyone finishes episode two.

FlexTV vice-president Tang Tang described the shape of the resulting operation to MIT Technology Review on May 15, 2026: moving North American production to AI cut costs by 80 to 90 percent, compressed schedules from three or four months to under one month, and left teams of about ten people. Ten people is not a prompt engineer. It is a writers' room, a continuity supervisor, editors, a sound person and a distribution operator, with generation folded in as one function among several. That is a vendor describing its own results and the most specific public account available.

What each stage actually produces

The eight stages, their output, who performs them and how much of the schedule they take
StageOutputPerformed byShare of effort
1. Story and series architecturePremise, episode grid, cliffhanger map, paywall positionHuman writersHigh — and it front-loads everything
2. Script and shot listDialogue plus a numbered list of 5–12 second shotsHuman writers, AI drafting assistanceHigh
3. Character and world bibleLocked reference images, seeds, prompt fragments, costume and location platesHuman art direction with generated assetsModerate up front, decisive later
4. Shot generationRaw clips, several candidates per shotVideo models, operated by humansLow in hours, low in cost
5. Voice, music and soundDialogue tracks, music bed, ambience, effectsVoice models plus human direction; licensed or generated musicModerate
6. Edit and continuityCut episodes, artefact repair, lip-sync, pacingHuman editorsHighest of any stage
7. LocalisationSubtitles, dubs, adapted on-screen textTranslation models plus human QCModerate, multiplied per language
8. Publish, measure and feed backRelease schedule, paywall placement, per-episode metricsDistribution and growth staffContinuous

Read the fourth column and the whole argument of this page is already visible. The stage with a model in it is the cheapest stage. The three stages that decide quality — 1, 3 and 6 — contain no model at all. Any account of AI drama production that spends its time on prompt syntax is describing the least important 10% of the job.

Why are episodes assembled instead of filmed?

Every AI drama episode is stitched together from many short generations because no video model available in August 2026 produces a coherent scene of episode length. The ceiling is a hard technical fact, it differs by model, and it dictates the visual grammar of the entire category.

Generation-length and resolution ceilings by model, and what each implies for shot design
ModelLength ceiling per generationResolution / notesWhat it means for a 90-second episode
ByteDance Jimeng~12 seconds; ~30 seconds with motion templatesMotion templates extend length by constraining movementBest case ~3 template-driven shots, realistically 8 or more generations
KlingUp to about 3 minutesThe longest ceiling in the category on paperNominally one generation could cover an episode; coherence degrades long before that, so it is used for shots like everything else
Sora 2Short clipsUp to 1080p; synchronised audio; API model id sora-2A dozen-plus generations, with audio arriving attached
Runway Gen-4.5Short clipsUp to 1080pA dozen-plus generations
Veo 3.1 Lite / Fast4, 6 or 8 seconds10 credits Lite, 20 Fast on Google Flow (5 and 10 for Ultra)At 8 seconds a shot, roughly 12 accepted clips minimum
Veo 3.1 Quality8 seconds100 credits per generation; 1080p upscale free on paid tiers, 4K Ultra-onlyRoughly 12 accepted clips minimum, at 100 credits each before rejects
Seedance 2.0Not published in the sources we holdByteDance, released February 2026; reported as a notable step up in qualityQuality improvement, not a length escape

Do the arithmetic and the format explains itself. An episode of 90 seconds, cut at a Veo 3.1 ceiling of 8 seconds a shot, needs a minimum of twelve accepted clips — and accepted clips are a fraction of attempted ones. A season of 70 episodes is therefore not seventy renders. It is somewhere in the low thousands of generations, of which the majority are discarded.

This is why shots in AI dramas rarely run past 10 to 12 seconds. It is not restraint, it is not a stylistic homage to fast-cut advertising, and it is not an audience-attention theory. It is the length at which the model stops being able to hold the world together, and the edit is designed to cut away before it does. Once you know the ceiling you can see it in every title in the category: the cut always arrives just before the face would have started to change.

Two consequences follow that shape everything downstream. First, continuity becomes an assembly problem rather than a shooting problem, because the material was never continuous to begin with. Second, the raw-material volume of an AI drama is several times its runtime, which is why the meaningful cost metric is cost per accepted shot rather than cost per generation.

Kling's three minutes, and why it does not solve this

Kling advertises the longest ceiling in the category, up to about three minutes. On paper that is longer than an entire episode. In practice long generations inherit every consistency problem the research literature describes, and they inherit them at greater length: the further a generation runs from its conditioning frame, the more the world drifts. A three-minute generation is not three minutes of usable footage; it is three minutes in which the failure has more room to develop. Producers therefore use long-ceiling models the same way they use short-ceiling ones — for shots.

Stage 1 — Story and series architecture

Story architecture is the first stage and the one that most determines whether a series is watchable, because the structural decisions made here cannot be fixed by any later stage. Nothing is generated in stage 1. The output is a document: premise, cast, episode grid, cliffhanger map and the position of the paywall.

The format imposes the shape. Episodes run 60 to 120 seconds, a season runs 40 to 100+ episodes, and every episode needs a hook in its first three seconds, rapid conflict, a reveal and a cliffhanger. That is not a set of guidelines, it is a grid with 40 to 100 cells in it, and each cell has four slots. Writing it discovery-first, the way a feature screenplay is often written, produces a series that runs out of turns around episode twenty.

The cliffhanger map

A cliffhanger map lists, for each episode, the question the episode raises and the episode where it is answered. Its purpose is to guarantee that no episode ends on a resolved beat and that no question stays open long enough to be forgotten. On a 70-episode season this is the difference between a viewer who reaches episode 40 and one who stops at episode 6 because nothing was pulling.

Where the paywall goes

Across this category the first 3 to 10 episodes are free, then a coin wall, VIP subscription, ad-unlock or reward mechanic takes over. That boundary is a story decision as much as a commercial one, because the episode immediately before it has to carry more pull than any other episode in the season. Writers plan it explicitly. Reader-facing consequences of that design are covered in what AI dramas are.

Designing for what the models can do

Stage 1 is also where a competent producer removes the scenes that generation will fail. Long unbroken two-handers of subtle emotional performance are expensive to fake and cheap to get wrong. Crowds, complex hand action, readable signage, and any shot requiring an object to remain exactly where it was are all liabilities. Fantasy, xianxia, werewolf and superpower material is written into these seasons partly because stylisation absorbs artefacts and effects that would break a live-action microdrama budget cost almost nothing here.

Stage 2 — Script and shot list

The script is written to the grid from stage 1, and then broken down into a numbered shot list where no shot exceeds the model ceiling. The shot list, not the script, is the document the rest of the pipeline actually executes.

A shot-list row is not a line of prose. It carries a shot number, a duration inside the ceiling, the characters present with their bible identifiers, the location, the costume state, the camera framing and movement, the dialogue line if any, and the continuity notes that connect it to the shots either side. Producing that for 70 episodes is clerical, unglamorous and the reason experienced teams survive contact with episode 50.

What AI does help with here

Language models genuinely accelerate drafting, variation and coverage in stage 2 — alternate lines, alternate escalations, translation-ready phrasing. Topview Drama Studio, marketed publicly in July 2026, packages exactly this: story planning, editable scripts, characters, locations, props and storyboards inside one workflow. What no system does is decide which of the variants is the one that makes a stranger watch the next episode. That judgement stays human and it is the whole job.

Shot length discipline

The single most useful rule in stage 2 is to write shots shorter than the ceiling rather than at it. A shot written for 8 seconds at a Veo 3.1 Quality ceiling of 8 seconds leaves no room to trim the frames where the hand deformed. Writing for 5 or 6 gives the editor material either side of the usable window, and shorter shots are also the first mitigation the literature recommends against drift.

Stage 3 — What goes into a character bible?

The character bible is the locked reference set that keeps a face, a costume, a prop and a room the same across a thousand independent generations, and it is the highest-leverage document in AI drama production. Generative video models have no memory between calls. Every generation starts from nothing. Without a bible, every shot re-invents the character.

What a working character and world bible contains, and what breaks without each element
Bible elementWhat it holdsThe failure it prevents
Face platesReference images of each recurring character from multiple angles and lighting statesIdentity drift — the face changing between shots, the most recognisable AI-drama tell
Costume statesEach outfit, per story day, with fabric, colour values and accessories fixedClothing, jewellery and hairstyle changing without a scene break
Body and age notesHeight relationships, build, apparent age, distinguishing marksBody proportions and apparent age shifting mid-scene
Location platesEvery recurring room and exterior, with fixed layout and light directionThe same apartment rebuilding itself between visits; screen direction reversing
Prop registerStory-critical objects, where they are and who is holding themLoss of object permanence — props appearing and disappearing
Seeds and prompt fragmentsThe exact seeds and wording that reproduced an accepted lookBeing unable to regenerate a shot two weeks later to match an earlier one
Negative listLooks, styles and defaults the model keeps reaching for and the series rejectsThe house-style problem: models trained on similar data producing similar faces

Why the bible is also a differentiation problem

There is a failure mode here that no individual production notices from the inside. Models trained on similar data produce similar faces, so different series start to feel cast from the same small repertory company. A bible built by accepting whatever the model offers first will land on that shared default. A bible built by deliberately steering away from it is one of the few ways an AI-native studio can look like itself. We treat this at length in why AI drama characters change appearance.

The bible is a living document

Bibles decay. A costume approved in episode 3 gets regenerated slightly differently in episode 30, that regeneration becomes the new reference because it was the most recent accepted shot, and by episode 60 the character has drifted through a chain of small approvals none of which was wrong. Disciplined teams version the bible and compare against the original plates, not the most recent shot.

Stage 4 — How are the shots generated?

Stage 4 is where the video model runs, and it is the fastest and cheapest stage in the pipeline. The work is not creative in the way outsiders imagine; it is closer to operating a camera you cannot fully aim, several times, until one take is usable.

The operator works down the shot list. For each row: assemble the prompt from the bible fragments, attach the reference images, set the first and last frames if the model supports them, fix the seed, generate a batch of candidates, and mark one accepted or send it back. Veo 3.1 exposes reference images, input video, first and last frame control, camera controls, scene extension, outpainting and object insert and remove — that control surface, not raw output quality, is what makes a model usable in a serialised pipeline.

Acceptance, not generation, is the unit of work

The metric that matters in stage 4 is cost per accepted shot. A generation that produced six perfect fingers on the wrong hand costs the same as one that worked. Nobody publishes hit rates — that is a genuine gap in the public record and we are not going to fill it with a guess — but the structure of the problem is clear enough: a 90-second episode may require several minutes of raw generated material, and the ratio is worse for the shot types stage 1 was supposed to have written out.

What gets sent back

Rejection reasons in stage 4 are boringly repetitive: the face is not the character, the hands are wrong, an object moved that should not have, the background text turned into pseudo-letters, the camera drifted through a wall, or the physics of a fall or a thrown object is simply wrong. Experienced operators recognise these in the first second of playback and requeue without watching the rest.

Watermarks and provenance

Google's Veo 3.1 applies a SynthID watermark to its output. Runway's free tier watermarks visibly and Luma's Lite tier both watermarks and is licensed non-commercially. C2PA, separately, is a provenance standard that records an asset's origin and edits — it is not a rights registration and not a detector. Producers need to know which of their tools marks output and how, because a commercial series built on a non-commercial tier is a legal problem discovered late.

Stage 5 — Where do the voices and music come from?

Sound is where a cheap AI drama most often gives itself away, and it is also the stage where the least money is usually spent. Dialogue, music, ambience and effects each arrive from a different place and have to be reconciled against picture that was generated without any of them in mind.

Sora 2 produces synchronised audio with its video, which removes one class of problem and creates another: the audio it produces is the audio it decided on. Most pipelines build dialogue separately with a voice model, add a music bed, and then fight lip-sync in the edit — because the picture was generated from a text prompt that described a line of dialogue rather than from the recording of it.

The three sound problems specific to this format

  1. Lip-sync drift. Generated mouths move to nothing in particular. Matching them to a dialogue track is approximate at best, and it is one of the eight documented failure modes of generated video.
  2. Room tone that does not exist. Generated shots have no acoustic space. Two shots in the same room will sound like two different rooms unless ambience is laid deliberately across the scene.
  3. Music rights. Music and SFX are an explicit cost line in AI-drama budgets. A generated series with unlicensed music is a takedown waiting to happen, and clearing it is real money on a title that cost $3,000 to produce.

Voice, likeness and consent

Where a synthetic voice is built from a real performer, SAG-AFTRA's digital-replica framework applies: a digital replica is a program using a performer's voice, image or performance to generate new performances, and it requires consent, disclosure, compensation and control (materials November 17, 2025; July 2025 interactive-media protections cover digital voice replicas). Tennessee's ELVIS Act, in force July 1, 2024, and California's AB 1836 and AB 2602 reach the same territory from state law. This is a production obligation, not a viewer-facing one.

Stage 6 — Edit and continuity

The continuity edit is where an AI drama becomes watchable or does not, and it is the stage with no model in it at all. Everything generated up to this point is disconnected material. Stage 6 is where a human decides what a viewer is allowed to see, for how long, and next to what.

The editor's job here is unlike conventional editing in one specific way: a substantial part of it is concealment. Cut before the face starts to change. Cut away from the hand. Do not hold on the sign whose text is pseudo-letters. Reverse the order of two shots because that hides a screen-direction reversal. Insert a reaction shot to bridge a physics violation. This is craft, it is learned by doing, and it is precisely the value an experienced editor adds over an inexperienced one on identical raw material.

Shot-level review

Shot-level human review is one of the documented mitigations for generation failure, and in practice it is a gate rather than a polish: every accepted shot is checked against the bible before it enters the timeline, not after. Catching a drifted face at the review gate costs one regeneration. Catching it in the edit costs a re-cut. Catching it after publication costs the episode.

Continuity across episodes, not just within them

Within-episode continuity is tractable — the material was generated in one session against one bible state. Across-episode continuity is the unsolved part, and narrative continuity breaking across episodes is itself one of the documented failure modes. A viewer at episode 40 has an accumulated memory of what the apartment looks like, and every small inconsistency since episode 1 is available to them at once. This is the specific problem the academic literature calls long-horizon consistency, and no product solves it in August 2026.

Pacing to the format

Finally, the edit enforces the format: hook inside three seconds, no dead air, reveal on schedule, cut on the cliffhanger rather than after it. Short shots serve this and the model ceiling simultaneously, which is the one place where a technical constraint and a commercial requirement happen to agree.

Stage 7 — Localisation

Localisation is the stage that decides whether a series reaches the US market at all, and for a large share of what American viewers watch it is the only AI step in the entire production. A conventionally filmed Chinese microdrama dubbed into English by a voice model is AI-touched, not AI-made — and it is a substantial part of the catalogue on mixed-library apps.

The work has four parts: subtitle translation, dub production, adaptation of on-screen text and culturally specific references, and quality control in every output language. Adobe's Firefly video model advertises support for 100+ prompt languages, and translation and QC appear explicitly among the cost lines that headline AI-production prices hide.

Where localisation fails

Inconsistent subtitles and dubbing is one of the recurring complaint themes in this category — qualitative, since no representative US complaint dataset exists, but consistent enough to be worth naming. The specific failures are a dub that drifts out of sync with lips that were already approximate, subtitle timing that lags the very fast cutting the format demands, and character names that change spelling between episodes because nobody put them in a glossary.

Localisation multiplies, it does not add

Every additional language re-runs the dub, the subtitle pass and the QC across all 40 to 100 episodes. DramaBox carries extensive dubs and subtitles including Spanish and Hindi across 200+ countries; ShortMax ships polished subtitles and dubs in many languages. That reach is a genuine achievement and it is also a per-language cost multiplier that does not appear anywhere in a per-minute production figure.

Stage 8 — Publish, measure and feed back

Publication is not the end of the pipeline; it is the input to the next one. The release schedule, the paywall position and the per-episode metrics that come back are the specification a studio writes its next season against.

The measurements that matter are episode-level: where viewers stop, which episode converts the free tier into a payment, unlock rate per episode after the wall, and churn. The first 3 to 10 episodes are free across this category, so the conversion event is concentrated at one boundary and is measurable with unusual precision. A studio producing at volume is running this as a feedback loop, not a launch.

Volume is the strategy, and it is not a good one

The reality check on the whole pipeline is the hit rate. DataEye counted roughly 221,900 AI dramas published on Douyin in the first half of 2026 — about 1,200 a day — and found roughly 1,055 of them, 0.47%, passed 100 million views. In April 2026 alone Douyin saw about 44,200 AI titles against 3,248 live-action. More than 95% of new Chinese microdramas in Q1 2026 were AI-generated, according to the China Netcasting Services Association via Global Times.

Cheap production did not make hits common. It made misses cheap, and it made the pile that a hit has to be found in enormously larger. Every efficiency described on this page has been spent on volume rather than on quality, which is the honest reason "AI slop" is a phrase people reach for. This is Chinese platform data, the only market where AI-drama output is measured at scale, and it is the best available proxy for where US catalogues are heading.

Distribution is a production stage

The clearest illustration of the loop is StoReel, which raised $34 million and has stated a target of 100 AI dramas a month (36Kr, March 26, 2026), and where one reported production case came to $150 for four one-minute episodes. A studio operating at that cadence is not releasing seasons, it is running an experiment with a hundred arms a month and reading the results.

Treating distribution as a post-production afterthought is the most expensive mistake available in this business. Paid user acquisition is now often the largest single line in an AI drama budget, and 67% of marketers were advertising in or interested in short-drama apps in 2026 according to an Adjust survey published June 23, 2026. The competition for attention got more expensive at exactly the moment production got cheaper.

What tools are used, and what do they cost?

This is a survey of what the pipeline is built from, not a set of reviews, and several of the prices below come from secondary sources rather than retrieved official pricing pages. We label each one. Prices in this category move, and a figure without a date is worthless.

Generation and production tools used in AI drama pipelines, with source quality flagged
ToolPriceSource qualityWhat it is used for in this pipeline
Sora 2 (API)$0.10 per second at 720p portrait or landscape; Sora 2 Pro reported at $0.30–$0.70 per second by resolutionBase API price official; Pro price reported, secondaryStage 4 generation with synchronised audio, model id sora-2. The original consumer Sora product was reported no longer available as of April 26, 2026; the API remains active.
RunwayFree ($0, 125 one-time credits, watermark) · Standard $15/mo ($12 annual, 625 credits) · Pro $35/mo ($28 annual, 2,250 credits) · Max $95/mo ($76 annual, 9,500 credits)Published tier pricingStage 4. Gen-4.5 (released December 1, 2025) consumes 12 credits per second — 5s = 60 credits, 10s = 120. Monthly credits expire on several plans, purchased credits do not, and web and API credit pools are separate.
LumaLite $9.99/mo web ($12.99 iOS) · Plus $29.99 ($37.99 iOS) · Unlimited $94.99 ($119.99 iOS)Pricing page dated January 25, 2026Stage 4. Lite gives 3,200 credits, watermarked and non-commercial — unusable for a released series. Plus gives 10,000 credits, no watermark, full commercial rights. Unlimited adds unlimited relaxed generations. Ray3.2 announced June 9, 2026.
Google Flow / Veo 3.1Free 50 credits/day in supported regions · Plus 200 credits/month · Pro 1,000 credits/month ($19.99 Google One Pro) · Ultra listed at $100 → 10,000 credits and $200 → 25,000 creditsOfficial help pages, but internally inconsistent on Ultra — treat Ultra pricing as unverifiedStage 4 with the richest control surface: reference images, input video, first and last frames, camera controls, scene extension, outpainting, object insert and remove. Veo 3.1 Lite 10 credits (5 Ultra), Fast 20 (10), Quality 100. Lite/Fast 4, 6 or 8 seconds; Quality 8 seconds. Credits do not roll over. 1080p upscale free for Plus/Pro/Ultra; 4K is Ultra-only at 50 credits. SynthID watermark on output.
ByteDance JimengNot published in our sources for the USGap — verify directlyStage 4. ~12 seconds per generation, ~30 with motion templates — the longest practical single-shot budget among the short-ceiling models.
KlingStandard ~$10 · Pro ~$37 · Premier ~$92 · Ultra ~$180Secondary source — independent 2026 reports, not a retrieved official pageStage 4. Up to about 3 minutes per generation, the longest ceiling in the category.
PikaFree · Standard ~$8 · Pro ~$28 · Fancy ~$76Secondary sourceStage 4, generally as a supplementary generator.
Seedance 2.0Not published in our sourcesGap — verify directlyStage 4. ByteDance, released February 2026, reported as a notable step up in quality.
Adobe FireflyStandard $9.99/monthPublished priceStages 4, 5 and 7. Firefly Video Model announced in limited public beta October 14, 2024; 100+ prompt languages. Adobe claims commercially safe training data — that is a vendor claim, not an independent finding.
Topview Drama StudioPrice not published on its public pageGap — and a meaningful one for a product sold on end-to-end convenienceStages 1, 2, 3, 4, 5 and 7 in one workflow: story planning, editable scripts, characters, locations, props, storyboards, voice, subtitles, generated video and vertical episode production. Publicly marketed July 2026.

How to read this table

Three things are worth extracting from it. First, credit systems make prices hard to compare — Runway bills 12 credits per second on Gen-4.5, Google Flow bills 100 credits per Veo 3.1 Quality clip, and only one of those is expressible per second. Second, licence terms matter more than price at the bottom of the market: Luma Lite is the cheapest entry in the table and is licensed non-commercially, which makes it free in the way a demo is free. Third, several of these prices are secondary-sourced and we would rather say so than present a scraped number as a fact. A fuller treatment lives in the AI tools behind AI dramas.

What no tool in this table does

None of them writes a 70-episode cliffhanger map. None of them holds a character consistent for a season without a human-maintained bible. None of them performs the continuity edit. Topview comes closest to covering the pipeline end to end, and it still requires a human to make every decision in stages 1, 3 and 6 — it just puts the boxes next to each other.

What does it actually cost to produce an episode?

The correct unit of production cost in an AI drama is cost per accepted shot, not cost per generation, and almost every published figure quietly uses the wrong one. Generations that are discarded cost exactly as much as generations that ship.

Work the Sora 2 API price through it. At $0.10 per second and 720p, 90 seconds of accepted footage is about $9 of API spend. That is the number that makes headlines. If the pipeline keeps one generation in four — a plausible ratio nobody publishes, offered here as arithmetic and not as a finding — the real generation cost is about $36 for the same episode. Both numbers are still trivially small next to what comes after them.

Live-action vertical drama$2,900First-time AI pipeline$400Experienced AI team$280Optimised AI pipeline$30 US dollars per finished on-screen minute
Cost per finished minute. The live-action bar shows the midpoint of $2,500–$3,300 a finished minute, which is what $150,000–$200,000 per screen hour works out to at 60 minutes an hour and the 80–90% AI saving come from FlexTV vice-president Tang Tang via MIT Technology Review, May 15, 2026. The $280/min figure is independent creator Jiang Lan's own accounting (TVBS, August 7, 2026). The $30/min floor describes a heavily optimised Chinese pipeline and is not a typical US result. Note what the chart cannot show: none of these numbers include user acquisition, which is now the larger line item.

The benchmark on the other side is live-action vertical drama at roughly $150,000 to $200,000 per screen hour, which is about $2,500 to $3,300 per finished minute. Against that, AI production is reported at $20,000 to $50,000 per screen hour, and a heavily optimised Chinese pipeline has reached roughly $30 per finished minute — a figure that is not a typical US result and should not be planned against.

FlexTV's Tang Tang put the saving at 80 to 90 percent when the company moved North American production to AI, with schedules dropping from three or four months to under one and teams of about ten (MIT Technology Review, May 15, 2026). That is a vendor describing its own results. It is specific, it is on the record, it is the best public figure available, and it has not been independently audited. Treat it as a claim with a name attached rather than as a measurement.

What a 60–80 episode AI microdrama costs, and what the headline figure leaves out
LineFigureSource and status
Whole 60–80 episode AI microdrama$3,000–$14,000NOW News, December 15, 2025. Reported range, production only.
Independent creator, per finished minute~$280/min; ~$28,000 for twelve 8-minute episodesJiang Lan's own accounting via TVBS, August 7, 2026. Self-reported.
Live-action vertical baseline~$2,500–$3,300 per finished minute ($150,000–$200,000 per screen hour)Category baseline. The comparison point for every saving claim.
Optimised AI pipeline floor~$30 per finished minuteChinese pipeline. Not a typical US result — do not budget against it.
Claimed saving vs live action80–90%Tang Tang, VP FlexTV, MIT Technology Review, May 15, 2026. Vendor claim about its own results.
Revenue splitCommonly 50/50 platform/creator; a title at 1 billion views can return $100,000–$500,000Category reporting. The upside, against a 0.47% hit rate.

The lines the headline price hides

A production price of $3,000 to $14,000 for a season is real and it is incomplete. The cost lines that sit outside it are: failed generations, retakes, upscaling, voice correction, human editing, music and SFX, storage, translation and QC, distribution fees, legal review, refund handling — and paid user acquisition, which is now often the largest line of all.

That last item is the one that reframes everything. Production cost fell by roughly an order of magnitude and the cost of buying attention did not fall at all; if anything it rose, because 95% of new Chinese microdramas being AI-generated means the feed is more crowded than it has ever been. A studio that models this business as "generation got cheap" has modelled the smallest term in the equation.

The gap we will not fill

No US benchmark exists for the cost of a completed, edited and legally cleared AI-drama episode. Every figure above is either Chinese, self-reported, or a vendor claim. We would rather state that plainly than produce a tidy US number that nobody measured. Our creator-side breakdown is in what it costs to make an AI drama.

What still goes wrong with generated video?

Generated video fails in a small number of well-documented ways, they are the same ways in every pipeline, and the available mitigations reduce them without eliminating any of them. A producer who knows the list can design around it. A producer who does not will discover it at episode 30.

The documented failure modes of generated video, and the mitigations that reduce them
Failure modeHow it shows up on screenMitigation, and its limit
1. Identity driftA character's face changes between shotsCharacter reference images, locked bible, fixed seeds. Reduces frequency; does not eliminate it, and it compounds across episodes.
2. Loss of object permanenceProps appear, disappear or move between cutsProp register in the bible and shot-level review. Catches it; does not prevent the model producing it.
3. Wardrobe and body inconsistencyClothing, jewellery, hairstyle, apparent age or body proportions changeCostume plates per story day. Effective within a session, weaker across weeks.
4. Hand deformationExtra or fused fingers, impossible gripsShorter shots, framing that avoids hands, editing around it. Concealment rather than repair.
5. Pseudo-text backgroundsSignage and screens render as unreadable letter-like shapesWrite readable text out of the shot list, or composite it. No generative fix.
6. Lip-sync failureMouth movement does not match the dialogue trackSora 2's synchronised audio helps; separate voice pipelines fight this in the edit. Worse again after dubbing.
7. Screen direction and physicsEyelines and movement reverse; objects fall wrongly, cameras pass through wallsFirst-and-last-frame control, shorter shots, re-ordering in the edit. Partial.
8. Narrative continuity breaking across episodesThe world stops matching itself over a seasonContinuity editing and bible versioning. The least solved of the eight.

A ninth problem that is not a bug

Models trained on similar data produce similar faces, so different series begin to feel cast from the same handful of people. Nothing in any shot is wrong. The failure is at catalogue level, it is invisible from inside a single production, and it is the reason a viewer scrolling an AI-heavy library reports a sameness they cannot point at. No mitigation in the list above touches it; only deliberate art direction does.

What the research says

Three sources are worth naming because they establish that this is an open research problem rather than a tooling gap. ACM published A Survey: Spatiotemporal Consistency in Video Generation on May 18, 2026, treating consistency across space and time as the central open question in the field. An OpenReview paper dated March 26, 2026 addresses narrative and temporal consistency in long video specifically. And arXiv:2510.04999, Bridging Text and Video Generation: A Survey (October 6, 2025), maps the same territory from the text-conditioning side.

The practical reading of all three is the same: consistency is not a setting somebody forgot to enable. It is what the field is working on. A producer planning a 2027 season on the assumption that the next model release will solve stage 6 is planning against a research programme, not a roadmap.

The mitigations, collected

The full list of what actually helps: character reference images, a locked character bible, fixed seeds, first-and-last-frame controls, shorter shots, shot-level human review, and continuity editing. Every one of them reduces the failure rate. None of them eliminates a single failure mode. That sentence is the honest summary of AI drama production quality in August 2026, and it is why stage 6 employs people.

What are the most common production mistakes?

The recurring errors in this category are structural, not technical, and most of them are committed in stages 1 and 3 by people who thought the job was stage 4.

  1. Starting at generation. Opening a video tool before the episode grid and cliffhanger map exist. The result is a hundred attractive clips and no series, and every one of those clips has to be regenerated once the structure is decided.
  2. Skipping the character bible. The single most expensive shortcut available. Without locked reference images, seeds and costume plates, identity drift is not a risk, it is the default behaviour of the tools.
  3. Writing shots at the model ceiling. An 8-second shot on an 8-second ceiling leaves the editor no trim. Writing at 5 or 6 seconds costs nothing and buys frames either side of the usable window.
  4. Budgeting cost per generation. The right metric is cost per accepted shot. A budget built on the first number is wrong by whatever the rejection rate turns out to be, and nobody knows their rejection rate in advance.
  5. Treating the edit as post-production. Stage 6 is where the series is made. Under-resourcing it while over-resourcing generation is the most reliable way to produce something technically impressive and unwatchable.
  6. Writing scenes the models fail at. Sustained close emotional two-handers, readable signage, crowds and complex hand action are all avoidable in stage 1 and all expensive in stage 4.
  7. Building a commercial series on a non-commercial tier. Luma's Lite plan at $9.99 a month is explicitly non-commercial and watermarked. Discovering that after a season is finished is a rights problem, not a billing problem.
  8. Ignoring music and voice rights. Music and SFX are a named cost line, and where a real performer's voice or likeness is involved, consent, disclosure, compensation and control obligations attach. Neither is optional because the budget was $6,000.
  9. Letting the bible drift by approval. Accepting each new shot against the most recent shot rather than the original plates produces a character who has changed completely by episode 60 without any single approval having been wrong.
  10. Planning localisation as an add-on. Every language re-runs the dub, subtitles and QC across the whole season. It multiplies; it does not add.
  11. Assuming volume substitutes for quality. About 1,200 AI titles a day on Douyin, 0.47% of them reaching 100 million views. Producing more of a thing that does not work produces more of a thing that does not work.
  12. Forgetting user acquisition entirely. The largest line in many budgets, absent from every per-minute production figure quoted in this business, and the reason cheap production has not made this an easy market to enter.

How long does it take to make one season?

FlexTV reports completing a season in under one month with a team of about ten, down from three or four months conventionally. The schedule below distributes that reported month across the eight stages. It is our construction from a vendor's headline figure, not an observed timetable — no US studio publishes one, and we are not going to imply otherwise.

An indicative four-week schedule for one 60–80 episode AI drama season, team of about ten
WeekStages runningWhat has to be finished by the end of itBottleneck
Week 1Stages 1–3Premise, full episode grid, cliffhanger map, paywall position, scripts for at least the first act, and a locked character and world bibleWriters. Nothing downstream can start without the grid, and nothing looks like a series without the bible.
Week 2Stages 2–4Complete shot list for the whole season; bulk generation under way with a running accept/reject logShot-list clerical throughput, and generation queue time on paid tiers with monthly credit caps.
Weeks 2–3Stages 4–5Majority of shots accepted; dialogue recorded or generated; music bed and ambience laidRejection rate. This is where an unrealistic accept ratio becomes a schedule problem.
Week 3Stage 6First two-thirds of episodes cut, artefact-repaired, lip-synced and continuity-checked against the bibleEditors. This is the hard constraint on the whole month and the stage that cannot be parallelised cheaply.
Week 4Stages 6–7Remaining episodes cut; subtitles and the primary dub complete; QC pass in every shipping languageQC, which multiplies per language and is routinely under-scheduled.
Week 4 and afterStage 8Release schedule set, paywall placed after the free episodes, per-episode metrics instrumented, acquisition spend livePaid user acquisition — the budget line, not the calendar.

Why the month is credible and still misleading

The month is credible because stages 4, 5 and 7 genuinely collapse under automation: generation is minutes per shot, synthetic voice is near-instant, machine translation is a batch job. It is misleading because stages 1, 3 and 6 do not compress at all. A ten-person team hitting a month is doing so by running writing, generation and editing concurrently across different parts of the season, not by doing any single stage faster than a human can do it.

What a first season should add

A team's first season is not a one-month season. The bible has to be invented rather than reused, the accept/reject rate is unknown, the editors have not yet learned which artefacts this particular model produces, and the legal and music workflows are being built from nothing. FlexTV's figure describes an operating pipeline. Building the pipeline is the part it does not cover, and it is the part that decides whether the second season takes a month.

Frequently asked questions

How long does it take to make one AI drama episode?

There is no published US benchmark, and we will not invent one. What is on the record is the season figure: FlexTV vice-president Tang Tang told MIT Technology Review on May 15, 2026 that moving North American production to AI cut schedules from three or four months to under one month, with teams of about ten people. That covers a full season of 60 to 80 episodes, so a single 60 to 120 second episode is a matter of days rather than weeks — but that is our arithmetic on a vendor's claim, not a measured figure.

Can one person make an AI drama alone?

Yes, and people do. Independent creator Jiang Lan accounted for roughly $280 per finished minute and about $28,000 for twelve eight-minute episodes (TVBS, August 7, 2026). What stops most solo attempts is not the generation tools but stages 1, 3 and 6 — story architecture, the character bible and the continuity edit. Those are craft work, they scale badly with one pair of hands, and no model does them for you.

Why do AI drama shots never last more than about ten seconds?

Because that is roughly where the models stop. ByteDance's Jimeng produces about 12 seconds per generation, or 30 with motion templates. Google's Veo 3.1 Lite and Fast produce 4, 6 or 8 second clips and Quality produces 8 seconds. Kling reaches about 3 minutes but coherence degrades long before that. Shot length in this format is a model constraint dressed up as a directing style.

What is a character bible and why does it matter so much?

A character bible is the locked reference set for a series: face images from multiple angles, costume plates, prop and location references, plus the seeds and prompt fragments that reproduce them. It matters because generative models have no memory between calls. Without a bible every generation re-invents the character, which is the mechanism behind identity drift. We cover the failure itself in why AI drama characters change appearance.

What does it cost to generate one minute of AI drama footage?

Raw generation is the cheap part. The Sora 2 API is priced at $0.10 per second at 720p, so 60 seconds of output is about $6 of API spend before anything is rejected. The number that matters is cost per accepted shot: if you keep one generation in four, that $6 is really $24, and it still excludes editing, voice, music, translation and QC. Full arithmetic in what it costs to make an AI drama.

Which stage costs the most money?

Not generation. Across the reported cost lines the expensive parts are human editing and continuity work, music and sound rights, translation and QC, and above all paid user acquisition, which is now often the largest single line in an AI-drama budget. Production cost fell by an order of magnitude; the cost of getting anyone to watch did not.

Do AI dramas use real actors at all?

It depends on the production type. Fully AI-generated series use none. Hybrid productions film part of the material. AI-localised titles are conventionally filmed shows that were only dubbed by a voice model, and a large share of what US viewers actually watch sits in that last category. Where a real performer's voice or likeness is used to generate new performances, SAG-AFTRA's digital-replica framework requires consent, disclosure, compensation and control.

Can I prompt an entire episode in one go?

No. No model in August 2026 produces a coherent 90-second narrative episode from a single prompt. The episode is assembled from a dozen or more separate generations, each a few seconds long, then cut together. Anyone selling a one-prompt-to-finished-series workflow is describing a product that does not exist.

What is the hardest part of the pipeline?

Continuity across episodes, by a wide margin. A single shot can be regenerated until it is right. A character who has to look the same in episode 3 and episode 74, wearing the same jacket, in the same apartment, with the same scar, is a problem no current model solves. An ACM survey published May 18, 2026 treats spatiotemporal consistency as the central open question in video generation, and work on OpenReview in March 2026 and arXiv:2510.04999 in October 2025 reaches the same conclusion for long-form narrative.

How many generations does a 90-second episode need?

At shot lengths of 5 to 12 seconds, a 90-second episode needs roughly 10 to 18 accepted shots. Since accepted shots are a fraction of attempted ones, the number of actual generations is a multiple of that. This is the arithmetic behind the phrase "assembled, not filmed", and it is why the raw-material volume of an AI drama is several times its runtime.

Do producers have to label AI dramas as AI-made?

In the United States, no. There is no federal law specific to AI dramas as of August 2026, and no US app in our testing labels individual titles as AI-produced. China is different: AI-content labelling has been mandatory there since March 7, 2025. The US Copyright Office does require applicants to disclose more-than-de-minimis AI material and exclude it from a human-authored claim, but that is a registration duty, not a consumer-facing label.

Is the pipeline the same for AI manju as for photoreal AI drama?

The eight stages are the same; the difficulty distribution is not. AI manju (漫剧), the animated or comic-styled variant that dominates the Chinese market, sidesteps the uncanny-valley problem by not attempting photorealism, which makes stages 4 and 6 substantially more forgiving. Identity drift in an illustrated style reads as art-direction variation; the same drift on a photoreal face reads as a different actor.

The bottom line

Typing a prompt and receiving video is the fastest, cheapest and least decisive part of making an AI drama. It is one stage of eight, it costs cents per second, and the models that perform it are broadly interchangeable for the purpose. Everything that determines whether a viewer reaches episode ten happens in the three stages with no model in them: the story architecture that builds a 70-cell cliffhanger grid, the character bible that keeps a face the same across a thousand independent generations, and the continuity edit that conceals every place where the machine failed.

The technology's real contribution is length, not quality. Models hold a scene together for 4 to 12 seconds — 30 with Jimeng's motion templates, nominally 3 minutes with Kling and not usefully so — which is why an episode is a dozen-plus generations stitched together and why shots in this entire category rarely run past 10 to 12 seconds. That constraint made vertical serialised drama cheap enough to produce at 1,200 titles a day. It did not make any of them good.

The volume that follows from all of this is the last thing worth looking at before deciding to enter the business.

0.47% became hits ~1,055 titles passed 100M views ~220,845 titles did not Total published, H1 2026: ~221,900 Average: roughly 1,200 new AI titles a day
Volume is not success. DataEye counted about 221,900 AI dramas published on Douyin in the first half of 2026 and found roughly 1,055 of them — 0.47% — passed 100 million views. Cheap production has not made hits common; it has made misses cheap. This is Chinese platform data, the only market where AI-drama output is measured at scale, and it is the best available proxy for what US catalogues are heading toward.

For anyone considering building one, three numbers are the whole brief. A 60 to 80 episode season has been produced for $3,000 to $14,000. FlexTV claims an 80 to 90 percent saving against a live-action baseline of about $2,500–$3,300 a finished minute, and says a season now takes under a month with ten people. And 0.47% of the AI dramas published on Douyin in the first half of 2026 passed 100 million views. Production is solved enough to be boring. Distribution and quality are not solved at all, and the pipeline has moved the difficulty rather than removed it.

DramaBox app icon DramaBoxEditor's pick — biggest short-drama library, lowest price to go unlimited 9.8Excellent
  • Thousands of series, roughly 200 new titles every month
  • Free episodes daily, plus ad-unlocks — no card needed to start
  • Unlimited plans from about $5.99 a week, the lowest entry price we found
  • Runs on Android, iOS and in a desktop browser
Go to DramaBoxAffiliate link · free to download · in-app purchases available

Affiliate link. We may earn a commission if you install through it, at no extra cost to you — it never changes our scores or placements. How we test · Disclaimer

Read next

Sources

  • MIT Technology Review (May 15, 2026) — Tang Tang, vice-president of FlexTV, on the move of North American production to AI: 80–90% cost reduction, schedules from three or four months to under one, teams of about ten. Cited throughout as a vendor describing its own results. Same piece, via DataEye: ~470 AI titles a day on Douyin in January 2026.
  • DataEye (2026) — ~221,900 AI dramas published on Douyin in H1 2026, roughly 1,200 a day, of which ~1,055 (0.47%) passed 100 million views; April 2026 counts of ~44,200 AI titles against 3,248 live-action.
  • China Netcasting Services Association, via Global Times (2026) — more than 95% of new Chinese microdramas AI-generated in Q1 2026.
  • NOW News (December 15, 2025) — 60–80 episode AI microdramas produced for $3,000–$14,000. TVBS (August 7, 2026) — independent creator Jiang Lan's accounting of ~$280 per finished minute and ~$28,000 for twelve eight-minute episodes.
  • OpenAI — Sora 2 API pricing at $0.10 per second at 720p portrait or landscape, synchronised audio, model id sora-2. Sora 2 Pro at $0.30–$0.70 per second is a reported figure from secondary coverage. The original consumer Sora product was reported no longer available as of April 26, 2026.
  • Runway — published plan pricing (Free, Standard $15, Pro $35, Max $95 monthly, with annual equivalents) and Gen-4.5 credit consumption of 12 credits per second; Gen-4.5 released December 1, 2025. Luma — pricing page updated January 25, 2026, including the non-commercial licence on the $9.99 Lite tier; Ray3.2 announced June 9, 2026.
  • Google — Flow and Veo 3.1 credit tiers, generation lengths (Lite and Fast at 4, 6 or 8 seconds; Quality at 8 seconds), credit costs, 1080p and 4K upscale rules and the SynthID watermark. Noted in the text: Google's own help pages are not internally consistent on Ultra pricing, so Ultra figures are flagged unverified.
  • Adobe — Firefly Standard at $9.99 a month; Firefly Video Model announced in limited public beta October 14, 2024, 100+ prompt languages, with the commercially-safe training-data claim identified as a vendor claim. Kling and Pika prices are from independent 2026 reports, labelled secondary source throughout. ByteDance — Jimeng generation limits of ~12 seconds (~30 with motion templates); Seedance 2.0 released February 2026.
  • Topview — Drama Studio end-to-end AI drama workflow, publicly marketed July 2026. Price is not displayed on its public page; we record that as a gap rather than estimating it.
  • ACM — A Survey: Spatiotemporal Consistency in Video Generation (May 18, 2026). OpenReview — narrative and temporal consistency in long video (March 26, 2026). arXiv:2510.04999Bridging Text and Video Generation: A Survey (October 6, 2025). The basis for treating consistency as an open research problem rather than a tooling gap.
  • SAG-AFTRA — digital-replica framework requiring consent, disclosure, compensation and control (materials November 17, 2025; July 2025 interactive-media protections covering digital voice replicas). Tennessee ELVIS Act, in force July 1, 2024. California AB 1836 and AB 2602 (2024). US Copyright Office, Copyright and Artificial Intelligence, Part 2: Copyrightability (January 29, 2025) on disclosure of more-than-de-minimis AI material. China's mandatory AI-content labelling from March 7, 2025, cited for contrast only.
  • Adjust (June 23, 2026) — 67% of marketers advertising in or interested in short-drama apps in 2026, cited in support of the point that paid user acquisition is now often the largest budget line. C2PA — provenance standard recording an asset's origin and edits; explicitly not a rights registration and not a detector.

Written and verified by Oleksandr Korop, editor and publisher, with the the AI Drama Reviews editorial team

Oleksandr Korop is the named person responsible for everything published here, at the address on the imprint. We install every app on a US account, watch until the free tier runs out, pay through at least one paywall with our own money and record the receipt. Every figure on this page carries a source and a date. Corrections: send one with a source and we fix the page and date the change.

Last verified August 19, 2026. Our testing method · Every source we use, with links · About this site

Affiliate disclosure. Some outbound links on this page, including the DramaBox links, are affiliate links. If you install through one we may receive a commission at no additional cost to you, and it never changes what we write — this page states plainly that no app in the category, including the one we earn from, labels which of its titles were AI-produced. Accuracy notice. Several tool prices here come from secondary sources rather than retrieved official pricing pages and are labelled as such; Google's own pages are internally inconsistent on Veo Ultra pricing. The 80–90% cost reduction and sub-one-month schedule are FlexTV describing its own results, not an audited finding. No US benchmark exists for the cost of a completed, edited and legally cleared AI-drama episode, and no source publishes generation accept rates — we have left both gaps open rather than filling them. Verified August 19, 2026.