Fragua
A skill for Claude Code and OpenCode that turns a raw recording into a video ready to publish on TikTok, Reels or YouTube Shorts.
You drop the video into the chat and the agent edits it: cuts the silences, burns in karaoke captions, adds shot changes and graphic cards, grades the colour, normalises the audio, and then writes the title, description, tags and captions for each network.
Everything runs on your machine. No API keys, no paid services, and none of your files leave the computer.
Install
Claude Code
Install it like any other plugin, from inside Claude Code:
/plugin marketplace add pedrovelasquez9/fragua
/plugin install fragua@fragua
Then set up the dependencies once — ffmpeg, Python, the transcription model and the fonts. Just ask the agent:
/fragua:setup
Or run it yourself from the installed folder:
powershell -ExecutionPolicy Bypass -File scripts\setup.ps1 # Windows
./scripts/setup.sh # macOS and Linux
The setup downloads about 1.6 GB, installs anything missing through winget,
brew or your package manager, and finishes with a self-check. When you see
todo verde, you are ready.
OpenCode
OpenCode has no marketplace, so clone the repository and run the installer:
git clone https://github.com/pedrovelasquez9/fragua
cd fragua
./install.sh --target opencode # macOS and Linux
powershell -ExecutionPolicy Bypass -File install.ps1 -Target opencode # Windows
The installer copies the skill where OpenCode looks for it, resolves every dependency and runs the verification suite.
Commands
Inside Claude Code, Fragua adds two commands:
| Command | What it does |
|---|---|
/fragua:setup |
Installs and verifies everything it needs. Run once after installing |
/fragua:assets <folder> |
Points Fragua at your music, sound effects, stickers and fonts, and indexes them |
Editing needs no command: just ask.
Installer options
| Option | What it does |
|---|---|
--target claude / --target opencode |
Install for one agent only (default: both) |
--project |
Install into the current folder instead of system-wide |
--model ggml-medium |
Smaller model: 1.4 GB less to download, faster, slightly less accurate |
--version v1.12.0 |
Install that exact version instead of the latest |
On Windows the same options use PowerShell syntax: -Target, -Project, -Model.
Updating
/plugin marketplace update fragua # Claude Code
git pull && ./install.sh # OpenCode
Do not copy the
vendor/folder between machines. Those 1.5 GB are regenerated by the setup on each computer.
Going back to an earlier version
Every release is tagged, so any of them can be installed exactly as it shipped. Useful when an update changes something you liked, or when you want to reproduce an edit you made months ago.
On OpenCode, the installer takes the tag:
./install.sh --version v1.12.0
powershell -ExecutionPolicy Bypass -File install.ps1 -Version v1.12.0
On Claude Code, a marketplace can also be a folder on your disk, so you clone the version you want and point the plugin manager at it:
git clone --branch v1.12.0 https://github.com/pedrovelasquez9/fragua fragua-1.12.0
Then, inside Claude Code:
/plugin marketplace add ./fragua-1.12.0
/plugin install fragua@fragua
To come back to the current version, point the marketplace at the repository
again — /plugin marketplace add pedrovelasquez9/fragua — and reinstall.
The version numbers are in CHANGELOG.md, and git tag lists
them all if you already have the repository cloned.
Which version do I have?
Ask the agent — what version of fragua is this? — and it will tell you, along
with what changed since yours.
CHANGELOG.md lists every version and what it added, newest first. Worth a look after updating: new features usually need nothing from you, so it is easy to miss that they arrived.
How to use it
There are no commands to memorise. Open Claude Code or OpenCode and ask in plain language, pointing at the video:
edit this video for tiktok: C:\Users\me\Videos\recording.mp4
The agent picks up the skill and gets to work. It takes five to fifteen minutes depending on the length, and tells you what it finds along the way.
What happens meanwhile
- Measures the video: duration, brightness, noise floor
- Detects the silences and trims them
- Transcribes what you say so it can be read
- Decides the edit — this is where it warns you if you shot a take twice, if you stumbled, or if the first seconds are dead air
- Designs the cards that summarise your ideas on screen
- Renders and hands you the file
- Writes the copy for YouTube, Instagram and TikTok
What you can ask for
All of this works as typed:
edit this video and give me your opinion on the content
make it for reels instead of tiktok
the captions are too big, shrink them and render again
drop the second card, it covers what I'm showing
just write the youtube title and description
use the take at the end, not the first one
If something is off, say so and it redoes it. You do not need to know which parameter to touch: describe what you see («the shirt looks pixelated», «the audio cuts off at the end») and the agent finds the cause.
What you get back
- The edited video, next to the original with
-EDITin the name - A project folder, so you can tweak something later without redoing everything
- The publishing copy: title, description, 15 tags and the captions for Instagram and TikTok
Your own assets
Fragua can pull music, sound effects, stickers, images and fonts from a folder you own. Point it there once:
/fragua:assets D:/my-assets
On OpenCode there are no slash commands, so just say it:
use D:/my-assets as my assets folder.
Organise it however you like — the classification reads the files, not the folder names — but this is the layout it expects:
my-assets/
music/ long background tracks
sfx/ short hits: whoosh, pop, riser
stickers/ PNGs with transparency
images/ screenshots, logos, backgrounds
fonts/ your own .ttf or .otf
You never have to reindex. Every edit refreshes the catalogue first, so
anything you dropped in the folder this morning is already available this
afternoon. /fragua:assets is still there if you want to see the inventory.
From then on the agent knows what you have and reaches for it on its own: a
whoosh over a whip pan, a pop when a card appears, your track under the voice —
ducked automatically so it never fights the narration. Ask for it explicitly if
you prefer: put the lofi track underneath and a pop on each card.
The catalogue is stored in ~/.fragua/assets.json, outside the plugin, so it
survives updates.
If you never configure anything, Fragua uses its own empty assets/ folder and
everything else works the same.
Music from CapCut or TikTok is licensed for those apps only, and triggers copyright claims on YouTube. Use CC0 sources such as Pixabay, Mixkit or Freesound instead.
Images that appear when you mention something
Drop an image named after a word into your assets folder — youtube.png,
claude-code.png, github.png — and whenever the video says that word, the
image pops up in a corner for a couple of seconds.
Nothing to configure and nothing to run: the filename is the trigger word, and every edit reindexes the folder before starting.
It never covers your face and never fills the screen — corner only, capped at 40% of the width, one every eight seconds at most. The entry sound, if you add one, ducks under your voice so it never competes with what you are saying.
Your own footage as cutaways
Drop video clips into your assets folder and the edit can cut away to them — full screen, while your voice carries on underneath. That is what makes it read as a second camera rather than an interruption.
They fade in and out, and land in the pauses between phrases rather than mid-word — a cutaway that appears while a syllable is being spoken is noticed even when it is the right shot.
Just say so: use the clips in my assets folder at a few points. The agent
picks the moments where your face adds nothing — a long explanation, a list, an
abstract idea — and never over your hook or your punchline. It also matches each
clip's colour to yours, because footage from another camera rarely lands on the
same brightness, and without that it reads as a different video rather than a
different angle.
Long videos get chapters
Anything over about five minutes comes back with YouTube chapters, ready to paste into the description. The timings are the edit's, not the recording's — which is the whole point, because silence removal moves everything and a chapter list written against the raw footage drifts by minutes.
They are checked against what YouTube actually requires before you get them: first one at 0:00, at least three, none shorter than ten seconds. Miss any of those and YouTube quietly shows no chapters at all.
Long videos also come with a subtitle file rather than captions burnt into the picture — on a wide screen those cover the code, and viewers cannot turn them off. Upload it to YouTube and the text becomes searchable.
Long videos come back with thumbnails
The spec to build them, three of them, one per title: the background prompt ready to paste into your generator, what photo of you it needs — expression, framing, hands, the direction of the light — and where it goes in pixels on a 1280×720 frame, and the headline with its font, its colour and its place.
It does not hand you a frame from the video: it tells you what the thumbnail needs, and you decide whether to pull it from the footage, shoot it on purpose or use one you already have.
Do not ask the generator to use your face. Hand it your photo and it will not paste it: it does identity transfer and returns someone similar, which on a personal channel is noticeable. That is why the background prompt asks for no people and leaves that side empty — you composite your own photo on top, so the face is genuinely yours.
The font it names is the one already in the video, so the thumbnail and the captions read as the same channel.
Shot changes and transitions
Where the subject changes, the video shrinks onto black and a headline animates into the space that opens above it, then eases back to full size. A fade to black only covers the join; this one uses it — it puts the phrase you just said on screen, so it can happen three or four times in a video without tiring.
Shot changes are movements, not jumps: close to a second of travel, leaving and arriving at a standstill. When a line deserves a real jump instead, the framing changes between one frame and the next, the way an editor cuts to a second camera.
You never choose any of this by name. The agent reads your transcript, finds the lines that carry weight, and puts the movement on the first word of the phrase — never spaced evenly by the clock, which is what makes an edit feel random.
Animated cards
Cards can be drawn as motion instead of a still: the bullets arrive one at a
time, the rule under a heading draws itself, and a figure counts up to its
value. Just ask — edit this with animated cards — or say it once and it
sticks for that video.
They are React components rendered with Remotion, which
means two things worth knowing. It adds about half a minute to an edit. And if
you write React, a new card type is a component in remotion/src/Card.tsx — you
do not have to touch any Python.
Remotion is optional and is not installed by default, because it pulls in
Node and its own headless Chrome. If you already have Node, /fragua:setup
prepares it for you. Without it, everything works exactly as before with still
cards.
It remembers what it told you
After each edit Fragua records which content recommendations you applied and which you did not. Next time it checks that history before giving an opinion.
So if something has been pending for three videos, it stops mentioning it in passing and makes it the main point. And when you fix something, it says so instead of repeating the same note forever.
Ask for the report any time: how am I improving?
Recording tips
The skill fixes a lot, but some things are cheaper to get right while shooting. These three come up the most:
Watch the glare on your glasses. If your monitor reflects in the lenses it covers your eyes, and no filter fully recovers that. Lower the screen brightness, move it off the camera axis, or tilt the temples slightly.
Repeating a take is fine. The agent spots repetitions and keeps the good one. If you stumble, stop and repeat the whole sentence.
Put the hook first. The strongest line of the video belongs in the first five seconds, not at second forty. If it lands at the end the agent will tell you, but it cannot reorder your argument without it showing.
Troubleshooting
| What you see | What to do |
|---|---|
| The agent says ffmpeg is missing | Run the setup again, and open a new terminal so the PATH refreshes |
| «falta whisper.cpp» or similar | Run the setup again; it is safe to repeat as often as you like |
| Captions do not match what you say | Ask it to review the transcript — it is a text file it can fix |
| It takes forever | Normal on long recordings. --model ggml-medium is noticeably faster |
| Installed but the agent ignores it | Check the folder is named exactly fragua |
For anything else, just ask the agent: it carries the full technical documentation and can diagnose its own failures.
For developers
To understand the pipeline, run the scripts by hand or change the colour, effect and typography parameters, see TECHNICAL_README.md.
Licence
Fragua is released under the MIT licence — use it, modify it and distribute it as you like. See LICENSE.
What the setup downloads carries its own permissive licences: whisper.cpp and the Whisper model under MIT, and the Roboto, Anton and Poppins fonts under the SIL Open Font License.
If you add your own music or stickers under assets/, licensing is on you.
Do not use CapCut or TikTok material: it is licensed for those apps only and
their music triggers copyright claims on YouTube.
No comments yet
Be the first to share your take.