Skip to content

About

We built the tool we needed, then kept the receipts local.

VidForge AI began as an internal pipeline for producing faceless videos at volume without a monthly bill that scaled with output. It grew into a desktop application because that turned out to be the honest architecture for the problem.

The story

Every step already existed. Nothing connected them.

Making a faceless video is not one task; it is eleven. Research the topic. Write a script that holds attention. Decide what appears on screen and for how long. Find footage that actually matches. Generate narration that does not sound synthetic for nine minutes. Time captions to the delivery rather than the page. Choose a music bed and duck it under the voice. Cut, grade, encode. Write metadata for each platform. Upload to each account. Record what happened.

Excellent tools existed for most of those steps. What did not exist was anything that joined them, so the work became a relay of exports between browser tabs, with the editing itself the smallest part of the afternoon.

We started stitching the steps together, and each addition made an architectural fact clearer: the files were already on the machine. The GPU was already there. The AI keys were already ours. A hosted service would have meant uploading everything to a server we would then have to pay for, to do work the local hardware could already do — and billing for the privilege.

So VidForge AI stayed on the desktop. Not as a nostalgic preference, but because it is where the data, the compute and the credentials already live. That decision is why there is no subscription, no render queue, no upload limit, and no copy of your work anywhere but your own disk.

Mission

Remove the production work between having an idea and having a published video — without taking custody of anyone's content to do it.

The measure is simple. Time from topic to uploaded video should be minutes of waiting rather than hours of clicking, and nothing about that convenience should require handing your work, your keys or your audience to an intermediary.

Vision

Capable AI software you own outright, running on hardware you already have, calling models you chose.

Consumer AI is consolidating into monthly rentals of someone else's infrastructure. We think there is a durable alternative: local applications that orchestrate models on your behalf, where switching provider is a dropdown rather than a migration, and where the software keeps working the day a vendor changes its pricing.

What we believe

Four decisions that shaped the build.

These are not marketing positions. Each one is visible in how the software behaves when something goes wrong.

  • 01

    Your machine is the product

    Video work is heavy, and modern desktop hardware is very capable. Sending a gigabyte of footage to a datacentre so it can be composited and sent back is a business model, not an engineering requirement. We render where the files already are.

  • 02

    You should hold your own keys

    Reselling model access means marking it up, metering it, and deciding on your behalf which model you get. Bringing your own key removes all three. You choose the provider, you see the real price, and you keep the relationship when we are not in it.

  • 03

    Automation must stay explicit

    Software that publishes on your behalf has to be predictable. Nothing uploads until you press Start. An upload interrupted by a crash goes back to pending rather than retrying silently. You should never learn what the app did by discovering it on your channel.

  • 04

    Fail loudly, in the open

    When hardware encoding is unavailable, the log says which encoder was rejected and why, not just "using CPU". When a token exchange fails, you get the platform's own error text. Diagnosing a problem should not require reading our source code.

Technology

What it's actually made of.

Published because you are running this on your own machine with your own credentials, and you are entitled to know what that involves.

Desktop

Python and PySide6

A native Qt application, not a packaged web view. It uses the memory a desktop app should.

Storage

Two isolated SQLite databases

The render pipeline and the Publishing Center never share a database, so a publishing failure cannot corrupt a render.

Media

FFmpeg with hardware encoding

NVENC, QuickSync and AMF, each verified by a real one-frame test encode rather than trusting the encoder list.

Text

Anthropic Claude and OpenAI GPT

Switchable per job, with typed outputs so downstream stages never re-parse prose.

Voice

ElevenLabs, OpenAI, VoxCPM2, Kokoro

Cloud or fully local, with a chunk-and-retry pipeline that keeps long scripts from drifting.

Transcription

OpenAI Whisper, on-device

Captions timed against the rendered audio, so they match the delivery rather than the script.

Publishing

Official platform APIs

YouTube Data API v3, TikTok Content Posting API and the Facebook Graph API, each with its own OAuth handler.

Credentials

Fernet encryption at rest

AES-128-CBC with HMAC-SHA256, keyed locally. We hold no copy and cannot recover it.

The interface ships in English, Español, 日本語, ភាសាខ្មែរ. Exact permissions requested from each publishing platform are listed in the Privacy Policy, and every integration we support is documented on the integrations page.

See whether the architecture holds up.

Install it, disconnect your network after the script stage, and watch the voice and subtitle stages finish anyway.