Skip to content

Features

Fifteen subsystems. One window.

VidForge AI is not a wrapper around a single model. Each stage of the pipeline is a purpose-built engine, and every one of them is described here — including what it does when things go wrong.

Write it

4 features

A topic goes in. Research, a script, a shot list and a scene plan come out — all before a single frame is rendered.

AI Script Generator

Turns a topic and a niche into a structured script with a hook, body and payoff, sized to the duration you asked for.

You pick the topic, niche, platform and target length. The script engine writes to a retention structure rather than a word count: an opening hook, escalating body beats and a closing payoff. Output is typed, so every downstream stage — storyboard, voice, subtitles — reads the same object instead of re-parsing prose. Claude and GPT are both supported and switchable per job.

How it runs

  1. You supply the topic, niche, target platform and duration.
  2. The engine researches the subject and drafts to a retention structure.
  3. Output is parsed into a typed script object with titled scenes.
  4. The script title becomes the export filename for every downstream artefact.

What you get

  • Hook-first structure written for retention, not word count
  • Length targeted to the platform you selected, not a generic default
  • Choose Claude or GPT per job, or switch providers app-wide
  • Script title drives export filenames automatically

Use it when

  • Turning a keyword list into a week of scripted videos
  • Producing platform-specific cuts of the same subject at different lengths
  • Drafting a series where every episode follows one structure

AI Research

Grounds each script in real detail about the subject before writing a word of narration.

Before drafting, the engine gathers the facts, figures and framing that make a faceless video worth watching. This is what separates a script that says something from a script that fills sixty seconds. Research is folded straight into the draft so the writing stage has substance to work with.

How it runs

  1. The topic is expanded into the specific questions a viewer would have.
  2. Facts, figures and framing are gathered before any narration is written.
  3. Research is folded into the drafting prompt rather than appended after it.

What you get

  • Scripts carry specifics instead of generic filler
  • Consistent factual grounding across a whole batch
  • Runs inside the same job — no separate research step to manage

Use it when

  • Explainer and educational channels where accuracy is the product
  • Covering unfamiliar niches without becoming an expert first
  • Keeping factual density consistent across a large batch

AI Storyboard

Converts the script into a cinematic shot list with shot type, angle, mood, energy and colour palette per beat.

The storyboard engine replaces flat keyword lists with a typed shot specification. Each shot carries a visual description, search keywords, duration, suggested transition, camera angle, importance score, mood, energy level and colour palette. If the AI call fails, it synthesises a storyboard from the script's own visual keywords so a job never dies at this stage.

How it runs

  1. Each script scene is expanded into one or more shots.
  2. Every shot receives a type, angle, duration, mood, energy and palette.
  3. Search keywords are generated per shot for the footage stage.
  4. If the model call fails, a storyboard is synthesised from the script's own keywords.

What you get

  • Every scene gets a deliberate shot type and angle
  • Colour and mood metadata feed the grade and transition engines
  • Graceful fallback means the pipeline never stalls here
  • Cached per job — re-running a stage is instant

Use it when

  • Giving repetitive listicle formats real visual variety
  • Keeping a consistent visual grammar across an entire series
  • Driving colour grade and transitions from shot metadata rather than presets

AI Scene Planner

Allocates screen time beat by beat so pacing matches the narration instead of splitting the runtime evenly.

The scene planner assigns real durations from the storyboard's importance and energy scores. High-importance beats hold longer; filler gets cut short. Timing is computed against the actual generated narration length, so the video and the voiceover finish together without a trailing silent tail or a clipped final line.

How it runs

  1. Importance and energy scores are read from the storyboard.
  2. Screen time is allocated per beat instead of divided evenly.
  3. Timing is reconciled against the generated narration length.
  4. The final scene is trimmed so audio and video end together.

What you get

  • Pacing follows narration, not a fixed clip length
  • Important beats get the screen time they deserve
  • No dead air at the end of a render

Use it when

  • Short-form where the first three seconds decide retention
  • Long-form where pacing has to vary to hold attention
  • Hitting a hard duration limit without clipping the last line

Build it

6 features

Voice, footage, captions, music and grade are assembled and encoded on your own hardware.

AI Voice Generation

Cloud or fully local narration, with a studio pipeline that keeps long scripts from drifting.

Choose ElevenLabs or OpenAI in the cloud, or run VoxCPM2 or Kokoro entirely on your own machine — including voice cloning from a reference sample. Long scripts are the hard case, so the studio pipeline chunks them at sentence boundaries, quality-checks each chunk, retries failures up to three times, crossfade-merges the result and normalises to −16 LUFS with a −1 dBFS peak ceiling. Generated audio is cached by a hash of provider, voice, speed, text and preset, so nothing is ever paid for twice.

How it runs

  1. The script is chunked at sentence boundaries, around 100 words per chunk.
  2. Each chunk is generated, quality-checked, and retried up to three times.
  3. Chunks are crossfade-merged, then normalised to −16 LUFS with a −1 dBFS ceiling.
  4. The result is cached against a hash of provider, voice, speed, text and preset.

What you get

  • Local models mean narration text never leaves your machine
  • No voice drift or degradation on ten-minute scripts
  • Broadcast-consistent loudness across an entire batch
  • Cache hits make re-renders free
  • Voice library with reusable cloned profiles

Use it when

  • Ten-minute narrations that previously drifted halfway through
  • Keeping one channel voice consistent across hundreds of videos
  • Producing narration with no per-character cost using a local model
  • Working on confidential material where script text must not leave the machine

Subtitle Generator

Transcribes the finished narration on-device with Whisper for word-accurate caption timing.

Captions are generated from the rendered audio, not from the script — which is why the timing actually lines up with the delivery. Whisper runs locally, so audio never leaves your computer for transcription. Styling, positioning and burn-in are configurable per platform aspect ratio.

How it runs

  1. The rendered narration audio is transcribed locally by Whisper.
  2. Word-level timings are derived from the actual delivery, not the script.
  3. Captions are styled and positioned for the target aspect ratio.
  4. Sidecar subtitle files are exported alongside the video.

What you get

  • Word-level timing that matches the spoken delivery
  • Transcription happens on your machine, not in the cloud
  • Caption styling tuned per aspect ratio
  • Sidecar subtitle files exported alongside the video

Use it when

  • Short-form where burned-in captions are effectively mandatory
  • Accessibility compliance requiring an accurate transcript
  • Feeding a translation workflow from a reliable source transcript

Background Music

Picks a bed that matches the storyboard's mood, then ducks it under narration automatically.

The music selector reads the storyboard's overall mood and energy and chooses a matching track from your library. Sidechain ducking lowers the bed whenever narration is present, at a strength you control, so dialogue stays intelligible without manual keyframing. Volume and ducking are per-job settings.

How it runs

  1. The storyboard's overall mood and energy are read.
  2. A matching track is selected from your library.
  3. Sidechain ducking lowers the bed wherever narration is present.
  4. Volume and ducking strength are applied per job.

What you get

  • Track choice follows the video's actual mood
  • Automatic ducking keeps narration clear
  • Per-job volume and ducking strength

Use it when

  • Keeping narration intelligible without manual keyframing
  • Matching music to tone automatically across a varied batch
  • Meeting platform loudness expectations consistently

AI Video Editor

Chooses cuts, transitions, framing, motion and colour the way an editor would — then applies them.

The editor and director engines make the decisions a human editor makes. Beat and rhythm engines set cut timing. The attention engine tracks where the eye should land. The continuity engine avoids jarring adjacent shots. Smart trimming pulls the most useful window out of each clip. Ken Burns motion, transition selection and cinematic colour styling are all driven by the shot's own metadata rather than applied uniformly.

How it runs

  1. Beat and rhythm engines set cut timing against the narration.
  2. The attention engine decides where the eye should land per shot.
  3. The continuity engine rejects jarring adjacent shot pairings.
  4. Smart trimming selects the most useful window inside each clip.
  5. Ken Burns motion, transitions and colour grade are applied from shot metadata.

What you get

  • Cuts land on the narration's rhythm
  • Transitions chosen per shot pair instead of one effect everywhere
  • Colour grade follows the storyboard's palette
  • Continuity checks prevent visual whiplash between shots

Use it when

  • Getting an edited feel without opening an editor
  • Avoiding the single-crossfade-everywhere look of template tools
  • Maintaining a visual identity across a channel automatically

Multi-thread Rendering

GPU encoding on NVENC, QuickSync or AMF, with real test-encode detection and resumable jobs.

VidForge AI does not trust FFmpeg's encoder list. At startup it runs a real one-frame test encode against each candidate and selects the first that genuinely works on your hardware, then prints exactly why anything was rejected. Four render profiles ship, from fast to lossless. Checkpointing means a render interrupted at stage six resumes at stage six rather than starting over.

How it runs

  1. At startup, each candidate encoder is validated by a real one-frame test encode.
  2. The first genuinely working encoder is selected; rejections are logged with reasons.
  3. The chosen render profile sets bitrate, preset and quality targets.
  4. Checkpoints are written per stage so an interrupted render resumes in place.

What you get

  • Hardware encoding verified by test encode, never assumed
  • A clear reason when hardware encoding isn't available
  • Fast, balanced, high-quality and lossless profiles
  • Interrupted renders resume where they stopped

Use it when

  • Rendering overnight batches at GPU speed rather than CPU speed
  • Diagnosing why hardware encoding is unavailable on a specific machine
  • Recovering a long render after a crash or a power loss

Ship it

5 features

Metadata, queueing, scheduling and uploads to every account you connected — in batches, unattended.

Publishing Manager

A separate Publishing Center with its own database, upload queue, scheduler, retries and analytics.

Publishing is an independent module with its own SQLite database, so a publishing failure can never corrupt a render. Jobs watch an export folder and pick up new videos as they appear. The queue throttles uploads, retries transient failures and records every attempt. Nothing uploads on its own: a job stays stopped until you press Start, and an upload interrupted by a crash is put back to pending rather than silently retried.

How it runs

  1. A job watches an export folder and detects newly rendered videos.
  2. Detected files enter a throttled queue with your schedule applied.
  3. Uploads run against each connected account, retrying transient failures.
  4. Every attempt is recorded with its outcome and the platform's own error text.

What you get

  • Isolated database — publishing can't break rendering
  • Folder watchers pick up exports automatically
  • Throttled uploads with retry and full attempt history
  • Nothing uploads until you explicitly start a job
  • Per-account health checks and diagnostics

Use it when

  • Publishing a week of content on a fixed schedule unattended
  • Distributing one render across several accounts and platforms
  • Showing a client a complete record of what was posted and when

Multi-account OAuth

Connect as many YouTube, TikTok and Facebook accounts as you run. Tokens are encrypted and stay on your machine.

Authorisation uses each platform's official installed-application OAuth flow. VidForge AI opens the platform's own consent screen in your browser and receives the redirect on a loopback server on your own computer — nothing is proxied through a VidForge server, because there isn't one. Tokens are encrypted at rest with Fernet using a key held only on your machine, and refreshed in the background. Revoke access at the platform and the connection dies immediately.

How it runs

  1. You register your own OAuth application in the platform's developer console.
  2. The app opens the platform's consent screen in your browser.
  3. The redirect is received by a loopback server on your own machine.
  4. The token is encrypted with a local-only key and refreshed before expiry.

What you get

  • Official OAuth consent screens — passwords never touch the app
  • Tokens encrypted at rest with a local-only key
  • Unlimited connected accounts per platform
  • Chrome profile isolation for separate identities
  • Background token refresh with health classification

Use it when

  • Agencies running many client accounts from one workstation
  • Separating personal and business identities via isolated browser profiles
  • Operating several channels in one niche without credential collisions

Auto SEO Metadata

Generates titles, descriptions, tags and hashtags tuned to each platform's own limits and conventions.

Every platform has different title lengths, tag rules and hashtag etiquette. The SEO engine writes to per-platform profiles, scores the result, validates it against the platform's constraints and retries when validation fails. It runs on its own credentials, isolated from your script-generation keys, and it is off unless you enable it per job.

How it runs

  1. Platform profiles define title length, tag rules and hashtag conventions.
  2. Titles, descriptions, tags and hashtags are generated per destination.
  3. Output is scored and validated against the platform's real constraints.
  4. Validation failures are retried before the upload is attempted.

What you get

  • Metadata written per platform, not copy-pasted across all of them
  • Validated against real character and tag limits before upload
  • Separate credentials from script generation
  • Off by default and enabled per job

Use it when

  • Avoiding truncated titles and rejected tag sets on upload
  • Writing metadata per platform instead of copy-pasting one version
  • Keeping SEO credentials separate from script-generation keys

Batch Generation & Publishing

Queue dozens of topics, render them unattended, and publish the results across every connected account.

Batch mode takes a list of topics and runs the full pipeline across all of them, sharing the footage cache and voice cache to cut both time and API spend. Parallel voice batching keeps the GPU busy. Publishing jobs then distribute finished exports across accounts and platforms on the schedule you set. Each video is still a distinct job with its own history, so a single failure doesn't take down the run.

How it runs

  1. A list of topics is submitted as individual jobs.
  2. Jobs run the full pipeline sequentially, sharing footage and voice caches.
  3. Voice generation batches in parallel to keep the GPU saturated.
  4. Finished exports are handed to publishing jobs on your schedule.

What you get

  • One list of topics becomes a week of content
  • Shared caches cut API cost across the batch
  • Parallel voice generation keeps the GPU saturated
  • One failure doesn't stop the batch

Use it when

  • Turning one planning session into a month of scheduled content
  • Testing many hooks on one subject to find what performs
  • Cutting provider spend across a run through shared caches

Publishing Analytics

One dashboard for upload outcomes, account health and per-platform history across every connected account.

The analytics tab aggregates what actually happened: which uploads succeeded, which failed and why, which tokens are close to expiry, and how each account is performing. Because the history store is local, you keep the full record indefinitely without depending on a platform's retention window.

How it runs

  1. Every upload attempt is written to the local history store.
  2. Outcomes are aggregated per account and per platform.
  3. Token expiry and account health are classified ahead of failure.
  4. History is retained locally for as long as you keep it.

What you get

  • Every upload attempt recorded with its outcome
  • Token expiry and account health surfaced before an upload fails
  • History kept locally and indefinitely

Use it when

  • Spotting an account whose token is about to require reauthorisation
  • Reporting delivery to a client without platform dashboard access
  • Keeping records beyond a platform's own retention window

A note on what this is not

VidForge AI does not host models, and it does not include AI credits. It orchestrates providers you choose and pay directly — which is why there is no per-render fee and no cap on how much you generate. It also means output quality tracks the providers and hardware you configure. The pricing page sets out what a typical video costs at provider rates.

Install it and run one video through.

The fastest way to judge a pipeline is to watch it work. One topic, five minutes, a finished video.