User Guide

AI Models

Complete reference of all AI models available in OpenStory

OpenStory integrates with a wide range of AI models across four categories: script analysis, image generation, motion/video generation, and music/audio generation. All media models are accessed via Fal.ai, while script analysis uses OpenRouter.

Script Analysis Models

These LLM models analyze your script, extract scenes, characters, and locations, and generate prompts. You can select multiple models to generate parallel sequences for comparison. Next to Generate, Turbo is the default (Luna, Nano Banana 2 Lite, MiniMax H3 Max, ElevenLabs). Quality selects the quality-ranked defaults. Both modes show the full catalog, grouped Fast / Quality.

ModelVendorContext WindowLicense
GPT-5.6 LunaOpenAI1M tokensProprietary (default)
Claude Fable 5Anthropic1M tokensProprietary
Claude Opus 5Anthropic1M tokensProprietary
Claude Opus 5 FastAnthropic1M tokensProprietary (scene-split)
Gemini 3.7 FlashGoogle1M tokensProprietary
Gemini 3.1 ProGoogle1M tokensProprietary
GPT-5.6 SolOpenAI1M tokensProprietary
GLM-5.3 FlashZ.ai1M tokensOpen Weight (MIT)
GPT-5.6 TerraOpenAI1M tokensProprietary
DeepSeek V4 ProDeepSeek1M tokensOpen Weight (MIT)
Claude Sonnet 5Anthropic1M tokensProprietary
Grok 4.6SpaceXAI500K tokensProprietary
Mistral Small 4Mistral262K tokensOpen Weight (Apache 2.0)
Seed 2.0 MiniByteDance262K tokensProprietary

Image Generation Models

These models create the visual images for each scene. You can select multiple models to generate variant images for comparison.

ModelVendorLicenseNotes
Nano Banana 2 LiteGoogleProprietaryTurbo default — fastest Google tier, references, fixed 1K
GPT Image 2OpenAIProprietaryQuality default — text rendering, UI fidelity, up to 4K
Nano Banana 2GoogleProprietaryFast generation and editing
Nano Banana ProGoogleProprietaryEnhanced realism and typography
Grok Imagine Image 2.0SpaceXAIProprietaryNewest Imagine image model, 1K/2K, edit up to 3 refs
Grok Imagine Image QualitySpaceXAIProprietaryQuality Mode — higher fidelity, stronger text
FLUX.2 MaxBlack Forest LabsProprietaryExceptional realism
PhotaPhotaProprietaryCharacter consistency via profiles
Hunyuan Image v3TencentOpen WeightStrong composition
FLUX.2 DevBlack Forest LabsOpen Weight32B open weights with native editing
Qwen Image 2 ProAlibabaOpen Weight (Apache 2.0)Native 2K, text rendering
HiDream I1HiDreamOpen Weight (MIT)17B parameters
Seedream 5.0 ProByteDanceProprietaryFlagship generation and editing
FLUX.2 FlashBlack Forest LabsOpen WeightCheapest distilled FLUX.2 — sub-second, edit up to 4 refs
FLUX.2 TurboBlack Forest LabsOpen WeightDistilled FLUX.2 — ~2s, edit up to 4 refs

Edit Endpoints

Most image models support reference image editing via dedicated edit endpoints. This allows the AI to use character and location reference images when generating scenes, improving visual consistency.

Motion/Video Models

These models animate still images into video clips.

ModelVendorEst. TimeLicenseNotes
MiniMax H3 MaxMiniMax~10sProprietaryTurbo default; native audio
Seedance 2.0ByteDance~3.5 minProprietaryQuality default; native audio
Grok Imagine Video 1.5SpaceXAI~30sProprietaryHighest quality ranking
LTX 2.3 ProLightricks~2 minOpen Weight
Veo 3.1Google~2.5 minProprietary20K max prompt length
MiniMax Hailuo 2.3MiniMax~3 minProprietaryIn the Turbo picker; Seedance-class latency
Kling v3 ProKling~5 minProprietary

Aspect Ratio Compatibility

Not all motion models support all aspect ratios. OpenStory automatically filters to show only compatible models and will switch to a compatible default if your current model doesn't support the selected ratio.

Audio Support

Some motion models can generate audio alongside video. OpenStory checks each model's capabilities to determine audio support.

Music & Audio Models

ModelVendorMax DurationTypeLicense
ElevenLabs MusicElevenLabs600s (10 min)MusicProprietary (default)
ACE-Step 1.5ACE Studio600s (10 min)MusicOpen Weight
ACE-StepACE Studio240s (4 min)MusicOpen Weight

Music vs. Sound Effects

Music models generate background music tracks from text prompts and optional tags.

Capabilities

FeatureElevenLabs MusicACE-Step 1.5ACE-Step
Prompt-basedYesYesYes
Lyrics supportNoYesYes
InstrumentalYesYesYes
Long-formYes (10 min)Yes (10 min)Yes (4 min)