Workflow Ai Generation: Install, Source and Security | FunnelSlayer

Workflow Ai Generation

Published by damionrashford in media-os

No known issues9 installs

What this skill does

Generate media from scratch with 2026 open-source AI — TTS voiceover (Kokoro / OpenVoice / Piper), image gen (FLUX-schnell / Kolors / Sana / ComfyUI), video gen (LTX-Video / CogVideoX / Mochi / Wan), music (Riffusion / YuE), lipsync talking heads (LivePortrait / LatentSync), OCR (PaddleOCR / Tesseract 5 / TrOCR), zero-shot tagging (CLIP / SigLIP / BLIP-2 / LLaVA). Strict commercial-safe license filter. Use when the user says "generate a video", "TTS voiceover", "AI explainer video", "clone my vo

Add Workflow Ai Generation to your agent

Review the source and files first. When you are ready, copy the prompt instruction or use the CLI command supported by your environment.

Install with a prompt

Paste this into a compatible coding agent:

add this skill "workflow-ai-generation" from https://github.com/damionrashford/media-os

Install with the CLI

Run this command in a controlled environment after reviewing the repository:

npx skills add https://github.com/damionrashford/media-os --skill workflow-ai-generation

Skill instructions

Workflow — AI Generation

What: Synthesize new media with open-source, commercial-safe AI models. Strict license filter: Apache-2 / MIT / BSD / GPL. NC / research-only models are always-dropped.

Skills used

media-tts-ai, media-whisper, media-sd, media-svd, media-musicgen, media-lipsync, media-depth, media-ocr-ai, media-tag, media-demucs, media-denoise-ai.

Tool matrix

TTS (media-tts-ai)

ModelLicenseBest for
KokoroApache-2general default
OpenVoiceMITvoice cloning
CosyVoiceApache-2Chinese
ChatterboxMITexpressive
PiperMITembedded / offline
StyleTTS2MITsingle-voice quality
ParlerApache-2style prompting
BarkMITcreative / SFX
OrpheusApache-2modern expressive

DROPPED: XTTS-v2 (CPML NC), F5-TTS (research).

Image (media-sd)

ModelLicenseBest for
FLUX-schnellApache-2default (4-step distilled)
KolorsApache-2bilingual EN/ZH
SanaApache-2fast 4K
ComfyUIGPL-3node-graph workflows

DROPPED: FLUX-dev (NC), SDXL / SD3 base (restrictive).

Video (media-svd)

ModelLicenseBest for
LTX-VideoApache-2fastest
CogVideoXApache-2high quality (slower)
MochiApache-2cinematic
WanApache-2versatile

DROPPED: Stable Video Diffusion (NC research).

Music / SFX (media-musicgen)

ModelLicenseBest for
RiffusionApache-2spectrogram-diffusion
YuEApache-2long-form structured

DROPPED: Meta MusicGen (CC-BY-NC).

Lipsync (media-lipsync)

ModelLicense
LivePortraitMIT
LatentSyncApache-2

DROPPED: Wav2Lip (research), SadTalker (NC).

OCR (media-ocr-ai)

ModelLicenseBest for
PaddleOCRApache-2Latin + Chinese default
Tesseract 5Apache-2fastest classical
TrOCRMIThandwriting
EasyOCRApache-2multilingual

DROPPED: Surya (commercial restriction).

Tagging / captioning (media-tag)

CLIP (MIT), SigLIP (Apache-2), BLIP-2 (BSD), LLaVA (Apache-2 but needs Llama-2 backbone — check license).

Stem separation (media-demucs)

htdemucs 4-stem, htdemucs_6s 6-stem (adds guitar + piano).

Speech-to-text (media-whisper)

whisper.cpp (MIT fastest CPU), faster-whisper (MIT CUDA-accelerated).

Example composite workflows

Explainer video from script

script → TTS voiceover (media-tts-ai Kokoro) → slide images (media-sd FLUX-schnell) → B-roll clips (media-svd LTX-Video) → host animation (media-lipsync LivePortrait) → background music (media-musicgen Riffusion) → assemble (ffmpeg-cut-concat) → mix (ffmpeg-audio-filter + sidechain ducking) → loudness normalize (media-ffmpeg-normalize) → auto-burn subtitles (media-whisper + ffmpeg-subtitles) → transcode H.264 → YouTube upload (media-cloud-upload).

Multilingual voice clone

30 s clean reference → DeepFilterNet denoise → OpenVoice clone to EN/ES/FR/JP → per-track denoise → mux with localized subs.

Book cover + trailer

FLUX-schnell cover variations → LLaVA auto-describe for alt-text → Mochi cinematic trailer → StyleTTS2 narrator → YuE cinematic score → assemble.

Automated podcast chapter thumbnails

scenedetect chapter boundaries → extract frame per chapter → BLIP-2 caption → Kolors generate thumbnail from caption → embed chapter metadata.

Digital human with cloned voice

OpenVoice clone → LivePortrait drives portrait → FLUX-schnell branded background → chromakey + RVM matte composite → Riffusion audio bed → transcode + upload.

Gotchas

  • License discipline. Always check the skill's references/LICENSES.md BEFORE adopting a new model. HuggingFace license field is NOT authoritative. Pin model weights to specific commit hashes.
  • GPL-3 in ComfyUI / RVM — requires source distribution if modified/redistributed. Use dynamically / in pipeline, not bundled into proprietary product.
  • All Layer 9 skills need GPU for reasonable throughput (10–50× slower on CPU).
  • TTS sample rates vary: Kokoro 24 kHz, Piper 22.05 kHz, StyleTTS2 24 kHz, OpenVoice 24 kHz. Resample ALL to 48 kHz before mixing.
  • Voice cloning needs CLEAN reference. Noisy sample → noisy clone. DeepFilterNet the reference first.
  • FLUX-schnell is 4-step distilled. Do NOT push --steps above 4–8 — wastes compute, no quality gain.
  • ComfyUI workflow JSONs pin specific node versions. Use ComfyUI Manager to install matching nodes.
  • LTX-Video prompt adherence is phrasing-sensitive. "cat running" ≠ "running cat".
  • CogVideoX-5b needs ~20 GB VRAM. 2b variant runs on 8 GB at lower quality.
  • Riffusion produces 5.11-second clips natively. Chain with crossfade for longer, or use YuE for structured long-form.
  • LivePortrait expects clean frontal portrait. Angled faces, glasses, occluded mouths degrade output.
  • LivePortrait outputs 512×512 by default. Upscale with media-upscale for larger.
  • Whisper large-v3 is 3 GB. Test with base.en (140 MB) first — quality gap to medium (1.5 GB) is small for clean audio.
  • Whisper hallucinates on silence. Trim leading/trailing with silenceremove.
  • Whisper word-level timestamps require --word_timestamps True (faster-whisper) or --max-len 1 --split-on-word (whisper.cpp).
  • Whisper --language auto is fragile on accented speech. Specify language explicitly.
  • Diarization is external to Whisper. Use pyannote.audio (MIT) or simple-diarizer.
  • OCR selection: PaddleOCR for Latin+Chinese, TrOCR for handwriting, Tesseract for speed on clean docs.
  • CLIP / SigLIP / BLIP-2 need fixed-resolution inputs (224/336/384/448). The tagctl.py script resizes internally.
  • LLaVA needs an LLM backbone (Vicuna / Llama-2 7B+). Llama-2 community license has revenue caps.

Related

  • workflow-ai-enhancement — for enhancing EXISTING footage (not generating new).
  • workflow-podcast-pipeline — for podcast-specific AI workflows.
  • workflow-analysis-quality — VMAF + QC on AI-generated output.

Files included

  • SKILL.md

More skills