ClipSonic
Cloud or 100% Local AI Runs Fully Offline

AI Video Clipper — Turn Long Videos Into
Shorts. Automatically.

ClipSonic turns YouTube videos, podcasts, streams and long recordings into ready-to-post 9:16 Shorts — AI highlight detection finds the best moments, face tracking keeps the speaker in frame, and animated captions burn in automatically. Everything renders on your computer: no uploads, no monthly fees.

Regular: $197 Founder Offer: $69 — Save $128

Highlights Frame & Caption Export
Analyzing
podcast_episode_42.mp4
Duration: 45:32 · 1080p · Face-tracked
9:16
Clip 00:41 – 01:26 45s
Suggested Clips Groq
92 "The moment everything changed" — a complete, self-contained story with a clear emotional payoff.
87 "Why most people get this backwards" — strong contrarian hook, punchy opening line.
81 "The $10k mistake" — concrete number + curiosity gap, plays well standalone.
—
Local Render Captions On
Export 9:16 Batch Export
Auto Face-Track
Animated Captions
Nothing Uploaded
0AI Providers
0Caption Styles
0Framing Modes
$0Monthly Fees
Built for
YouTube Creators
Podcast Clippers
Streamers
Course Creators
Coaches
Agencies

One AI Video Clipper for Every Long-Form Source

ClipSonic is a desktop AI video clipper that turns long videos into short-form content automatically. Paste a link or drop a file and it handles the full workflow — download, transcribe, find highlights, reframe, caption and export — without uploading your video or charging per clip.

ClipSonic AI video clipper — project dashboard turning long videos into ranked vertical clips
The ClipSonic studio — every long video becomes a set of ranked, ready-to-frame clips.

One App. Every Clipping Workflow.

From a pasted link to an exported short — every step runs on your machine.

Paste a Link. Start Clipping.

Paste any YouTube URL and ClipSonic downloads and transcribes it automatically — no separate downloader, no manual steps.

  • YouTube URLs, auto-downloaded
  • Live progress with cancel support
  • Pick a quality tier (up to 1080p)
  • Auto mono 16kHz audio extraction
https://youtube.com/watch?v=...
Downloading — 64%

Drop Any Video. No Uploads.

Have the file already? Drag it in. ClipSonic reads duration, resolution, and a thumbnail instantly, then transcribes locally — the file never leaves your computer.

  • MP4, MOV, MKV, and more
  • Instant duration/resolution/thumbnail
  • No file ever leaves your machine
  • No size or duration limits
Drop a video file hereMP4, MOV, MKV, AVI, WebM...

AI Picks the Best Moments.

Groq, OpenAI, or a local model on your own machine reads the full transcript and returns ranked clip suggestions — each with a title, a hook reason, and a viral score, so you know why a moment was picked before you even watch it.

  • Hook reason + viral score per clip
  • Cloud or fully-local AI — your choice
  • Handles hour-long videos via chunking
  • Approve, refine, or reject each one
92
"The moment everything changed"Self-contained story, strong emotional payoff
87
"Why most people get this backwards"Contrarian hook, punchy opening line
74
"Three tools I use daily"Listicle format, clear structure

Viral Moments Without an API Key.

Smart mode scores every candidate moment against six on-device signals — keyword weight, hook quality, speech pace, dramatic pauses, audio energy, and clean cut points — then keeps the highest-scoring, non-overlapping windows. No LLM, no key, no per-clip cost.

  • Six signals scored from speech + audio
  • No API key and nothing billed per clip
  • Fully offline with local Whisper
  • Cuts land on sentence ends, not mid-thought
Keyword weight
Hook quality
Speech pace
Dramatic pauses
Audio energy
Clean cut points
Scored on your machine — no model call

Or Skip AI. Auto-Split by Length.

Fixed-Length mode slices the whole video into back-to-back clips — no AI, no API key, no transcript unless you want captions. Cuts snap to the nearest natural silence so you never land mid-word.

  • 30 / 45 / 60 / 90 / 120s presets
  • Cuts snapped to silence
  • Full-coverage, non-overlapping clips
  • Captions still optional, on request
30s45s60s90s120s
Cuts snapped to silence — no mid-word splits

Three Ways to Cut a Video. Pick One Per Project.

Every project starts by choosing how the clips get found. Two of them never touch an LLM — so an API key is a choice, not a requirement.

02

AI Viral Clips

A language model reads the whole transcript and explains every pick.

How it works

After transcription, the full transcript goes to a language model — Groq, OpenAI, or a local model via Ollama — which returns ranked clip candidates. Long videos are chunked and de-duplicated automatically, so an hour-long recording still gets coverage end to end.

What you get back

  • A written title for every clip
  • A plain-English hook reason — why this moment works
  • A viral score out of 100 to sort by
  • Self-contained moments — complete thoughts, not fragments
  • Approve, retrim, or reject each suggestion before rendering
Your own provider key Needs a transcript Or run it locally via Ollama

Best for: when you want reasoning behind each pick — and titles you can paste straight into the upload form.

03

Fixed-Length Clips

You pick the length. It splits the whole video. No AI at all.

How it works

Choose a clip length and ClipSonic slices the entire video into back-to-back clips of about that duration — full coverage, nothing overlapping, nothing skipped. Each cut is nudged to the nearest natural silence so it never lands mid-word.

The details that matter

  • Presets from 30s to 120s, or type your own
  • Cuts snap to silence within a window around each target
  • A short leftover tail is merged into the clip before it — no 4-second stubs
  • Transcription only runs if you want captions
  • Works with zero AI configured — fresh install, straight to clips
No API key No transcript needed Fully offline Fastest mode

Best for: gameplay, vlogs, sermons and streams — anything where you want the whole thing chopped and posted, not curated.

Side by Side

Smart Viral AI Viral Fixed-Length
How clips are found On-device scoring, six signals Language model reads the transcript Even intervals, snapped to silence
AI provider key needed No Yes — or a local model No
Transcription required Yes Yes Only for captions
Can run 100% offline Yes, with Local Whisper Yes, with Ollama Yes, always
Explains each pick Score + key phrases Title + hook reason + score —
Video coverage Best moments only Best moments only Every second of it
Typical speed Fast Depends on your provider Fastest
Cost per video Free Your provider's rate (often free tier) Free

Whichever mode you pick, everything after it is identical: face-tracked 9:16 reframing, animated captions, and local rendering — and the mode is chosen per project, so you can switch on the next video.

Every Clip Framed Like Someone Watched It

A vertical crop usually means guessing where to point the camera for the whole clip. ClipSonic reads every shot instead — who's on screen, where they are, and when your video cuts — then frames each one on its own terms.

Auto Frame: one, two or three people, decided per shot

Choose Auto Frame once and every shot gets the layout that fits how many people are actually in it. Each pane zooms on its own, so a closeup and a wide shot are both framed properly instead of sharing one compromise.

A
One person Fills the frame
A
B
Two people One to each half
A
B
C
Three people Lead takes the big pane

It reframes on the cut, not on a timer

Podcasts and interviews cut between a wide two-shot and a closeup of whoever is talking. ClipSonic finds every one of those edits, changes the layout with them, and cuts the crop instead of gliding across the join.

0:00 – 0:04 Wide shot · two panes
0:04 – 0:28 Closeup · full frame

A 28-second podcast clip: both hosts across the opening wide shot, full frame for the closeup that follows.

Crops sized from the face, not the canvas

A crop wide enough to fill a vertical frame is wide enough to show two people and centre neither. ClipSonic sizes every crop from the face it's following — close enough to frame them properly, and never wide enough to reach the next person along.

Plain center crop Lands between them
Face-aware Frames a person

The bright rectangle is the part of your video that survives the crop.

Six framing modes, switchable per clip

Auto Frame handles most footage. The rest are there for when you'd rather decide yourself.

Auto Frame

Recommended

One, two or three panes, decided per shot and switched on every cut.

Face tracking

Follows one speaker full-frame, with headroom above the head and a camera that eases instead of snapping.

Split duo

Two people stacked, each tracked separately — and full frame when only one is on screen.

Split + gameplay

Face-tracked facecam on top, your own looping background video underneath.

Center crop

A straight centred crop to the output shape. Fast, predictable, no detection.

Original

No reframing at all — keeps your source's own aspect ratio and resolution.

One control, in the editor

Framing is a per-clip choice you can change any time — pick a mode and the clip reframes itself. Every output shape is supported, from vertical to square to landscape.

ClipSonic — Editor
Framing
Auto Frame Face tracking Center crop Split duo Split + gameplay Original
Output shape
9:16 1:1 4:5 3:4 16:9

Runs on your machine. Face detection and reframing happen locally — no upload, no per-clip cost, no queue.

Everything the framing engine does

  • Shot detection — finds every cut in your source and reframes each shot separately.
  • Adaptive layout — one, two or three panes, chosen by who is on screen.
  • Face-sized crops — framed on the subject, never wide enough to catch a neighbour.
  • Per-pane zoom — a closeup and a wide shot each get their own framing.
  • Multi-person tracking — keeps people apart even when they sit close together.
  • Re-identification — someone who turns away is still the same person when they turn back.
  • Smooth virtual camera — eases in and out with a dead zone, so small head movements don't move the frame.
  • Cuts on edits — the camera never pans across a cut in your source.
  • Headroom framing — subjects sit slightly above centre, the way a camera operator would place them.
  • Every output shape — 9:16, 4:5, 3:4, 1:1 and 16:9, up to 4K.
  • Per-clip control — change any clip's framing at any time and re-render just that one.
  • Fully local — detection runs on your machine, so nothing is uploaded and nothing is metered.

A Closer Look at What You Get

The tools that turn a raw upload into a finished, watchable short.

AI Highlight Detection

Know Exactly Which Moments Will Hit.

ClipSonic doesn't just guess timestamps — it reads the transcript and scores each candidate moment for hooks, punchlines, and complete thoughts, returning a viral score and a plain-English reason for every suggestion. Run the model in the cloud (Groq or OpenAI) or fully on-device with a local model via Ollama. Long videos are chunked and de-duplicated automatically, so a 60-minute recording still gets sensible coverage from start to finish.

  • Title, hook reason & viral score per clip
  • 5–15 ranked suggestions per video
  • Cloud or local model — no lock-in
  • Approve, refine, or reject before rendering
95
"I almost quit right here"Vulnerable moment, high rewatch value
83
"The framework in 40 seconds"Dense value, clean structure
78
"That's when the crowd lost it"Reaction spike, strong payoff
Animated Captions

Captions That Actually Get Watched.

Six built-in styles — Clean, Karaoke, Bold Pop, Neon Glow, Bounce, Boxed — each with word-by-word highlight animation as the audio plays. Preview live over the clip, drag the caption to any position on the frame, and the chosen style is saved per-clip, not app-wide.

  • 6 distinct animated presets
  • Word-by-word highlight as it's spoken
  • Drag-to-position, saved per clip
  • Live preview before you render
Clean
Karaoke
Bold Pop
Neon Glow
Bounce
Boxed
One-Pass Export

From Approved Clip to Finished Short — One Pass.

Cutting, reframing, and captions render together in a single ffmpeg pass instead of three separate re-encodes, keeping quality high and export times low. Export one clip or hit "Export all" to batch-render every approved clip with live per-clip progress.

  • Cut + reframe + captions in one render
  • Batch-export every approved clip
  • Quality tiers up to 1080p
  • Custom output folder, non-colliding filenames
clip-01-the-moment.mp4 Done
clip-02-why-most-people.mp4 Done
clip-03-the-framework.mp4 Rendering
clip-04-three-tools.mp4 Queued
Cloud or 100% Local AI

Your Choice: Fast Cloud AI, or Fully Offline.

Every AI step can run in the cloud or right on your own machine. Transcribe with local Whisper (GPU-accelerated whisper.cpp, with the word-level timing the captions need) and detect highlights with a local model via Ollama — no API key, no per-minute fees, no internet required. Nothing, not even the transcript text, ever leaves your computer. Prefer the fastest setup? Point transcription and highlights at Groq or OpenAI instead. Mix and match per feature — it's a setting, not a lock-in.

  • Local Whisper transcription — download a model once, run offline
  • Local highlight detection via Ollama — no key, on-device
  • Or use Groq / OpenAI — pick a provider per feature
  • GPU-accelerated where available, CPU fallback everywhere
Transcription Local Whisper
Highlights Local (Ollama)
Internet Not needed
Data leaving your PC None

Why ClipSonic Over Cloud Clip Tools?

Cloud Clip Tools
$20–$100+/month subscriptions
Your videos uploaded to their servers
Capped clips or minutes per month
Generic auto-crop, no real face tracking
Watermark on lower tiers
Rendering queued on their servers
Cloud AI only — your data leaves your machine
VS
ClipSonic
One-time payment, forever
100% local rendering — nothing uploaded
Unlimited clips, no monthly caps
Real face tracking, incl. duo & gameplay split
No watermark, ever
Renders instantly on your own machine
Optional 100% local AI — runs fully offline

Not Just Cutting. The Whole Pipeline, Automated.

Everything between "long video" and "posted short," handled.

AI Highlight Detection

Finds hooks, punchlines, and complete thoughts — with a viral score and plain-English reason for every clip.

Auto Frame

Detects who is on screen in every shot and reframes on each cut — full frame, two-up, or three. See the layouts →

Animated Captions

Six word-by-word animated styles, draggable to any position on the frame.

One-Pass Export

Cut, reframe, and caption in a single render — plus batch export for every approved clip.

Three Clipping Modes

Smart Viral (on-device scoring, no key), AI Viral (a model ranks and explains), or Fixed-Length (no AI at all). Compare them →

Cloud or Local AI

Transcribe and detect highlights with Groq, OpenAI, or fully on-device — local Whisper plus a local model via Ollama, no API key needed.

Long Video to Short in 3 Steps

01

Paste a URL or Drop a File

Add any YouTube link or local video — ClipSonic downloads and transcribes it automatically.

02

Pick One of Three Modes

Smart Viral scores every moment on your machine with no API key, AI Viral has a model rank them with a hook reason, and Fixed-Length auto-splits the whole video with no AI at all.

03

Review, Frame & Export

Trim if needed, pick a caption style and framing mode, then export ready-to-post 9:16 shorts.

What Early Users Say

"Replaced a $49/month clipping subscription. The hook reasoning alone saves me from scrubbing hour-long uploads."

Marcus R.YouTube Creator

"Nothing about my VODs gets uploaded anywhere to get clipped — that alone was the deciding factor for our stream team."

James L.Streamer

Common Questions

What are the three clipping modes?

Every project starts by picking one. Smart Viral Clips scores every moment on your own machine against six signals — keyword weight, hook quality, speech pace, dramatic pauses, audio energy, and clean cut points — with no LLM and no API key. AI Viral Clips sends the transcript to a language model (Groq, OpenAI, or a local model via Ollama) which returns ranked clips with a title, a hook reason, and a viral score. Fixed-Length Clips uses no AI at all: pick a length and it splits the whole video into back-to-back clips, snapped to natural silence. You choose per project, and everything after — 9:16 reframing, captions, export — is identical.

Which mode should I use?

Use Smart Viral Clips if you want ranked highlights with no API key and no per-clip cost — ideal for podcasts and interviews. Use AI Viral Clips when you want a written reason and a ready-to-paste title for each pick. Use Fixed-Length when you want the entire video chopped and posted rather than curated — gameplay, streams, sermons, vlogs.

How does ClipSonic find the best clips?

Two ways, depending on the mode. Smart Viral Clips scores every candidate window locally against six speech and audio signals — no model call. AI Viral Clips transcribes, then an AI model reads the transcript and scores moments for hooks, punchlines, and complete thoughts — each suggestion comes with a viral score and a plain-English reason. Run that model in the cloud (Groq or OpenAI) or entirely on your own machine (a local model via Ollama) — your choice.

Can I run ClipSonic completely offline?

Yes. Choose local Whisper for transcription and a local model via Ollama for highlight detection, and the entire pipeline — transcribe, find clips, frame, caption, export — runs on your own machine with no API key and no internet at all. Nothing, not even the transcript text, ever leaves your computer.

Does ClipSonic upload my video anywhere?

No. Cutting, face-tracking, captioning, and exporting always happen locally on your machine. If you pick a cloud AI provider, only the transcript text is sent to find highlights — never the video file. Pick the local models instead and even that stays on your device.

Can I make clips without AI?

Yes — two of the three modes need no API key at all. Fixed-Length auto-splits any video into back-to-back clips (30/45/60/90/120s), snapped to natural silence so cuts don't land mid-word, with no transcript required unless you want captions. Smart Viral Clips still picks the best moments for you, scoring them on your own machine with no LLM involved.

Does it track faces automatically?

Yes. ClipSonic detects and follows faces frame-by-frame to keep the speaker centered in a 9:16 crop, including a two-person split-screen and a gameplay-plus-facecam layout.

What do the animated captions look like?

Six presets — Clean, Karaoke, Bold Pop, Neon Glow, Bounce, Boxed — each with word-by-word highlight animation, and you can drag the caption to any position on the frame.

Is this a one-time payment?

Yes. Pay once and own it forever — no monthly subscription, no per-clip fees, no watermark.

Which AI providers are supported?

For transcription: Groq, OpenAI, or local Whisper running on your own machine (GPU-accelerated, no key). For highlight detection: Groq, OpenAI, or a local model via Ollama. Pick a provider per feature in Settings — mix cloud and local however you like.

What is local Whisper transcription?

A speech-to-text engine (whisper.cpp) that runs directly on your computer — GPU-accelerated on most machines — with word-level timing for the animated captions. Download a model once and transcribe unlimited audio offline, with no API key and no per-minute cloud fees. Great for long podcasts and anything you want kept private.

What video sources can I use?

Paste any YouTube URL or drop a local video file (MP4, MOV, MKV, and more). YouTube downloads run automatically in the background.

Do I need a powerful computer?

Any modern Windows or Mac machine. Rendering runs locally via ffmpeg — no cloud render queue, no waiting in line. Local AI transcription is fastest with a GPU; on a CPU-only machine it still works, just slower — or use a cloud provider instead.

Stop Paying Monthly to Clip Your Videos.

AI highlight detection, face tracking, animated captions, batch export — one payment, forever.

Regular: $197 Founder Offer: $69 Save $128

7-Day Money-Back 100% Local Rendering Priority Support