How we test
The method in short
Every app is installed on a US account, watched until the free tier runs out, paid through at least one paywall with our own money, then cancelled in the store and checked. Six weighted axes produce the score: cost to finish a season 25%, catalogue and refresh 20%, production quality 20%, free tier 15%, device coverage 10%, billing transparency 10%. We also classify how much of each catalogue we judge to be AI-produced, and we label that as judgement rather than measurement, because no platform publishes the figure.
The testing protocol
The same five steps are applied to all thirteen apps, in the same order, which is the only thing that makes the scores comparable to each other.
- Install on a US account. App Store and Google Play, plus the desktop browser where one exists. Regional pricing differs, so a US account is not optional.
- Exhaust the free tier. Watch the free openings, take every daily allowance, use the ad-unlocks, and record exactly how far they get you before a paywall appears and where in the episode it lands.
- Pay through at least one paywall. With our own money, at the price the app quotes, and we keep the receipt. The advertised price and the charged price are not always the same number.
- Cancel in the store and verify. Not in the app. Then check that the store shows an expiry date rather than a renewal date, because in this category it frequently does not.
- Assess the catalogue. Sample across genres, apply the six AI-detection signals below, and record what we saw and when.
The six axes
The weights reflect what actually changes a viewer's outcome, not what is easiest to measure. Cost carries the largest share because the same season can cost $5.99 or $50 depending on a single decision at a single paywall.
| Axis | Weight | What we measure | What scores badly |
|---|---|---|---|
| Cost to finish a season | 25% | What it actually costs to watch one 70-80 episode series end to end, by the cheapest available route | Coin-first design, no unlimited option, weekly plans above the category median |
| Catalogue and refresh rate | 20% | Library depth, genre coverage, how much is added and how often | Thin shelves, stale catalogues, one genre carrying everything |
| Production quality | 20% | Picture, performance, dub timing, subtitle accuracy, and the visible AI artefact rate | Dub drift, flattened line readings, high identity-drift rate, unreadable subtitles |
| Free tier | 15% | How much you can watch, and learn, before paying anything | A fixed one-off allowance with no renewal, or a paywall before episode three |
| Device coverage | 10% | Phone, tablet, desktop browser, television, offline | Phone-only, no browser route, no working television option |
| Billing transparency | 10% | Whether you can find out what a season costs before you commit | No official US price page, no published coin rate, prices that change between screens |
How the headline score is built
The headline score is our overall verdict, weighted toward the axes above but not a mechanical average of them. We would rather say that than publish an arithmetic claim we cannot stand behind.
The adjustment we make is for avoidability. A weakness you can neutralise costs less than one you cannot: opaque web pricing stops mattering the moment the app quotes you a figure at checkout, whereas a coin model you cannot escape charges you at every unlock for the whole season. Two apps with the same axis scores can therefore land differently, and where that happens the review says why.
We publish both the axes and the verdict so that a reader who disagrees with our judgement can see exactly where it enters, and substitute their own.
Both numbers, for all thirteen apps
Our published verdict sits above the weighted axis mean for every app we rank, and the two measures produce exactly the same order. That is the claim worth checking, so here is the table that lets you check it.
| Rank | App | Weighted mean of the six axes | Published verdict | Difference |
|---|---|---|---|---|
| 1 | DramaBox | 9.26 | 9.8 | +0.54 |
| 2 | ReelShort | 8.78 | 9.5 | +0.72 |
| 3 | FlexTV | 8.72 | 9.2 | +0.48 |
| 4 | StoReel | 8.71 | 9.1 | +0.39 |
| 5 | ShortMax | 8.65 | 9.0 | +0.35 |
| 6 | Character.AI Series | 8.57 | 8.9 | +0.33 |
| 7 | DramaWave | 7.98 | 8.7 | +0.72 |
| 8 | GoodShort | 7.76 | 8.6 | +0.84 |
| 9 | VibeShort | 7.68 | 8.5 | +0.82 |
| 10 | FlickReels | 7.52 | 8.4 | +0.88 |
| 11 | MoboDrama | 7.37 | 8.2 | +0.83 |
| 12 | MyMuse | 7.28 | 8.1 | +0.82 |
| 13 | NetShort | 7.25 | 8.0 | +0.75 |
Read the difference column as a scale shift, not a thumb on the scale. It is positive for every app, it never reorders anything, and its cause is a single axis. Billing transparency is worth 10% of the mechanical mean and almost the whole category scores badly on it — nine of the thirteen apps give us no usable US price at all. That failure is real, which is why we score it, but it is also the cheapest failure in the category for a reader to neutralise: open the app, read the price, decide. A coin model you cannot escape is not neutralised by anything.
If you would rather use the mechanical number, use it. Both are published for every app, on every review, and we would rather hand you the arithmetic than ask you to trust a single figure.
One thing the table also shows plainly: the app we earn a commission from sits at the top under both measures. We cannot make that conflict disappear by explaining it, so we publish the numbers that would expose it if the ranking had been engineered, and let the arithmetic answer instead of us.
How we judge whether a catalogue is AI-produced
No app in this category publishes what share of its library is AI-produced, and no independent dataset measures it. That leaves observation, so we made the observation systematic and we label the result as judgement rather than fact.
Six signals, applied while watching across at least three genres per app. These are the classification signals we use internally; the reader-facing version of the same method, written as a two-minute test you can run on a single episode, expands them into eight on-screen tells and is linked at the end of this section.
| Signal | What we look for | Reliability |
|---|---|---|
| Shot length | Whether any shot exceeds roughly 10-12 seconds — the practical generation ceiling on current models | High |
| Identity drift | A face changing shape, age or proportion between cuts in the same scene | High |
| Lip sync | Mouth movement drifting against the dialogue track beyond normal dub tolerance | Medium-high |
| Hands and background text | Extra or fused fingers, signage resolving into pseudo-letters | Medium-high |
| Recurring faces | The same handful of faces appearing across unrelated titles and studios | Medium |
| Genre distribution | A catalogue weighted toward fantasy, xianxia, werewolf and superpower stories, where generation performs best | Low on its own |
One signal proves nothing. Cheap live-action microdrama also has poor continuity, fast cutting is a legitimate style, and compression artefacts are not generation artefacts. We treat a classification as supportable when several high-reliability signals appear together across multiple titles, and we say so on the page with a date attached.
We also keep a distinction that most coverage collapses: AI-touched is not AI-made. A series filmed with human actors and dubbed into English by a voice model is AI-touched. A series whose shots were generated is AI-made. Most of what US viewers see on mixed-catalogue apps is the first.
The reader-facing version of this method, written as a two-minute test, is in how to tell if a drama is AI-generated.
How we handle figures
- Every number carries a source and a date in the sentence that uses it, and again in a sources list at the foot of the page.
- Vendor claims are labelled as vendor claims. A company describing its own cost savings is evidence of what the company says, not an audited result.
- Forecasts are labelled as forecasts. A projection is not an outcome, however respectable the firm.
- Derived figures are labelled as calculated. Where we do arithmetic on someone else's numbers, we say the arithmetic is ours.
- Incompatible scopes are kept apart. Omdia measures all microdrama revenue, Deloitte measures in-app purchases, Sensor Tower estimates store IAP only and excludes advertising, third-party Android stores and direct web payments. These are not additive and we never add them.
- Missing prices are stated as missing. We do not estimate a price to fill a table cell.
The limits of this method
Publishing what our method cannot do is part of the method.
- Our AI classifications are judgement, not measurement, and could be wrong for any individual title.
- Catalogue assessments are samples. Nobody, including the platforms, publishes an audited catalogue count.
- Prices we record are the prices we were quoted, on our accounts, on the day. Yours may differ by region, device, store or promotion.
- Four apps — NetShort, MoboDrama, MyMuse and FlickReels — are thinly documented, and our confidence in those assessments is lower. Their pages say so.
- We cannot measure whether audiences mind AI production, because no representative dataset exists and no platform labels its titles.
What would change a score
- A published official US price page would raise a billing transparency score immediately and substantially.
- A published AI catalogue share would replace our classification with the platform's figure, and we would say whose number it is.
- A pricing change — in either direction — moves the cost axis at the next verification pass.
- A device or feature launch moves device coverage; a television route that actually works would move it a lot.
- A reader correction with a source is the fastest route of all. Send one to our contact page.
Independence and funding
This site earns a commission from DramaBox installs and from nothing else. DramaBox also ranks first, and we think the only useful answer to the obvious suspicion is the record rather than a promise. Our DramaBox review names the app's refusal to publish a US price as its worst failing, states that it does not make the best television in this category, and directs readers to ReelShort — from which we earn nothing — if performance quality is what decides it for them.
No app can buy a score, a placement or a change to a review. We accept no free accounts, press access or complimentary coins, because a promotional account does not show you what a real one costs. Full disclosure sits on the about page and on every page that carries a link.
Frequently asked questions
How do you test the apps?
Every app is installed on a US account, watched until the free tier is exhausted, and paid through at least one paywall so the actual charge is recorded rather than the advertised one. We then cancel in the store and verify the cancellation took effect. The same protocol is applied to all thirteen apps, which is what makes the scores comparable.
What are the six scoring axes?
Cost to finish one season (25%), catalogue size and refresh rate (20%), production quality (20%), free tier generosity (15%), device coverage (10%) and billing transparency (10%). The weights reflect what changes a viewer's outcome most, which is why cost carries the largest share.
Why does billing transparency get its own axis?
Because in this category it is the difference between budgeting a season and discovering the price at the till. Nine of the thirteen apps we rank give us no usable US price at all. An app that will not tell you what a season costs is charging you for the uncertainty, and we score that.
Is the headline score an average of the axis scores?
No, and we say so on every review rather than implying an objectivity the arithmetic does not have. The headline is our overall verdict, weighted toward the axes in the published proportions but adjusted for how avoidable each weakness is. Opaque web pricing costs you nothing once the app has quoted you a figure; a coin model you cannot escape costs you at every unlock.
How do you decide whether a title is AI-produced?
Six signals, applied while watching: shot length against the roughly 10 to 12 second generation ceiling, identity drift between cuts, lip-sync accuracy, hand and background-text deformation, recurring faces across unrelated titles, and genre distribution. This produces evidence, not proof, and we date every classification.
Why not just ask the apps what share of their catalogue is AI?
We would publish the answer if any of them gave one. None does, and no independent dataset measures it either. Until that changes, a labelled judgement with a date attached is more useful to a reader than silence.
Do you pay for the subscriptions you test?
Yes. Every paywall in this ranking was paid through with our own money, and every subscription was cancelled afterwards. We do not accept free accounts, press access or complimentary coins, because a promotional account does not show you what a real one costs.
How often do you retest?
Prices, free-episode allowances and catalogues are re-checked on a rolling basis and each page carries its verification date. Scores change when the evidence does, and we date the change rather than editing silently.
Can an app pay to change its score?
No. No payment, affiliate relationship or commercial arrangement can alter a score, a placement or the contents of a review. The one app we earn from is also the one whose failings we name most often, which is the only proof of this we can offer that does not require taking our word for it.
What would make you change a score?
A published official US price would raise a billing transparency score immediately. A published AI catalogue share would replace our classification with their figure. A sustained change in catalogue quality, a pricing change, or a device or feature launch would move the relevant axis. Reader corrections with sources move things faster than anything else.
Why are the scores clustered between 8.0 and 9.8?
Because every app in this ranking is a functioning commercial product that does the basic job. The differences that matter are in price, catalogue and transparency, not in whether the video plays. We do not pad the bottom of the scale to manufacture drama.
Do you review apps you cannot fully verify?
Yes, and we say so on the page. NetShort, MoboDrama, MyMuse and FlickReels are thinly documented, and our confidence in those assessments is correspondingly lower than for DramaBox or ReelShort. Stating the confidence level is more honest than omitting the app.