feat(skills): ai-presenter-video optional skill (port of lanshu, 955★ MIT)

Ports cclank/lanshu-create-ai-presenter-video (MIT, 955 stars in 7 days)
into optional-skills/creative/ai-presenter-video. Provider-neutral
presenter-video production: locked narration as master clock, avatar
generation with pilot-first cost discipline, lip-sync/identity QA,
captions, deterministic ffmpeg finalization with loudness normalization
and contact-sheet verification.

Hermes adaptations in the hub SKILL.md: SKILL_DIR resolution (upstream
hardcoded ~/.codex/skills), capability mapping to text_to_speech / FAL
video families / vision_analyze / hyperframes, consent-flag JSON paths
(input.* vs root), preflight error-vs-remote-blocker semantics.
References kept substantively verbatim (all-English upstream). Scripts
unmodified. LICENSE carried.

Validated hands-on: init_job -> preflight gating (blocked until manual
review booleans + input.remote_upload_approved) -> finalize_delivery on
a synthetic 1080x1920 render (master+share decode-verified, delivery
report, 9-frame contact sheet). Cold-subagent live test: SHIP; 3
friction fixes folded in (resolution guard, boolean-flip example,
preflight-writes-job note).
This commit is contained in:
Teknium
2026-08-27 22:14:12 -07:00
parent f94a7a1a05
commit a79ff58d65
12 changed files with 1211 additions and 0 deletions

View File

@@ -0,0 +1,21 @@
MIT License
Copyright (c) 2026 lanshu
Permission is hereby granted, free of charge, to any person obtaining a copy
of this software and associated documentation files (the "Software"), to deal
in the Software without restriction, including without limitation the rights
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
copies of the Software, and to permit persons to whom the Software is
furnished to do so, subject to the following conditions:
The above copyright notice and this permission notice shall be included in all
copies or substantial portions of the Software.
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
SOFTWARE.

View File

@@ -0,0 +1,196 @@
---
name: ai-presenter-video
description: "Make a verified AI presenter video from script + image."
version: 1.0.0
author: cclank (https://github.com/cclank/lanshu-create-ai-presenter-video), ported by Hermes Agent
license: MIT
platforms: [linux, macos]
required_commands: [ffmpeg, ffprobe, python3]
metadata:
hermes:
tags: [video, presenter, avatar, lipsync, tts, captions, creative]
category: creative
homepage: https://github.com/cclank/lanshu-create-ai-presenter-video
related_skills: [hyperframes, kanban-video-orchestrator, comfyui]
---
# AI Presenter Video
Turn a topic (or finished script) plus ONE authorized adult presenter image
into a complete, publish-ready presenter-led video: locked narration, avatar
generation with lip-sync QA, captions, deterministic editing, loudness-normalized
master/share encodes, and machine + visual acceptance reports.
Use this skill for new presenter videos AND for continuing, revising,
captioning, lip-sync-repairing, or re-exporting an existing presenter-video
job. The workflow is provider-neutral: pick generation capabilities from what
is actually available in the session (FAL video/image models via
`image_generate` and the video-gen plugin, TTS via `text_to_speech`, ASR via
the whisper/STT tooling, ffmpeg for everything deterministic).
> Ported from cclank/lanshu-create-ai-presenter-video (MIT). Upstream body
> kept substantively verbatim in `references/`; Hermes adaptations live in
> this hub file. Scripts are deterministic (no network, no credentials).
## Hermes adaptations (read first)
- **Skill dir resolution** — upstream hardcoded `~/.codex/skills/...`. In
Hermes resolve it once per session:
```bash
SKILL_DIR="$(dirname "$(find ~/.hermes/skills ~/.hermes/hermes-agent/optional-skills -path '*/ai-presenter-video/SKILL.md' 2>/dev/null | head -1)")"
[ -f "$SKILL_DIR/SKILL.md" ] || echo "skill dir not found — locate ai-presenter-video/SKILL.md manually and set SKILL_DIR to its directory"
```
Shell variables do not persist between tool calls — re-paste the resolution
line (or the expanded path) in each terminal call that uses it.
- **Capability mapping** — where the references say "a voice generation
capability", use `text_to_speech` (OpenAI/Edge/ElevenLabs per user config);
"presenter/avatar generation" → FAL image-to-video families (Kling, Wan,
MiniMax H3 etc.) through the configured video tooling, or an avatar/lipsync
endpoint the user has access to; "word-timestamp ASR" → whisper via the STT
tooling or `faster-whisper` in a venv; "deterministic compositor" → ffmpeg
filtergraphs, or the `hyperframes` skill when installed (the editing
reference has a HyperFrames section that maps directly onto it).
- **Visual QA** — do the "normal-speed visual review" steps with
`vision_analyze` on the generated contact sheet plus sampled frames
(identity, mouth timing, hands, blinking, continuity). Numeric checks come
from the scripts' ffprobe output.
- **Paid-generation consent** — remote avatar/TTS generation is billable.
Follow the upstream operating rules: before the first paid call state the
uploaded assets, requested seconds, known cost, pilot size, and retry
ceiling, and get the user's explicit go-ahead. Never upload the presenter
image to a remote provider before `remote_upload_approved` is true in
`job.json`.
- **Consent flags live under `input`** — `rights_confirmed`,
`adult_presenter_confirmed`, `remote_upload_approved`, and
`voice_clone_approved` sit inside the `input` object of `job.json` (init
flags set them; hand-editing must target `input.*`, not the job root).
`manual_input_review.*` sits at the root. `preflight.py` distinguishes
`errors` (block everything) from `remote_blockers` (block only remote
generation) — local script/audio work may proceed while remote is blocked.
## Workflow
1. **Start or resume a job.** New job:
```bash
python3 "$SKILL_DIR/scripts/init_job.py" \
--job-dir ~/Videos/my-presenter-video \
--presenter-image /path/to/presenter.png \
--topic "explain context engineering in one minute" \
--duration 60 --aspect 9:16 \
--rights-confirmed --adult-presenter-confirmed
```
Use `--script` for an existing script file; other flags: `--voice-sample`,
`--supporting-media`, `--width`, `--height`, `--fps`, `--watermark`,
`--cta`. For an existing job, read `job.json` + QA reports and resume from
the earliest unfinished state — never regenerate accepted work.
2. **Manual input review.** Actually look at the presenter image
(`vision_analyze`) and listen to any voice sample; record findings by
setting the `manual_input_review` booleans in `job.json`, e.g.:
```bash
python3 - <<'PY'
import json
p = "~/Videos/my-presenter-video/job.json" # expand ~ or use an absolute path
import os; p = os.path.expanduser(p)
j = json.load(open(p))
j["manual_input_review"].update(image_viewed=True, single_clear_face=True,
image_has_no_unwanted_text=True)
json.dump(j, open(p, "w"), indent=2)
PY
```
Then gate:
```bash
python3 "$SKILL_DIR/scripts/preflight.py" ~/Videos/my-presenter-video/job.json
```
Proceed only when `ok: true`; do remote generation only when
`remote_ready: true`. Note: preflight also updates `job.json` in place
(records the report path) — re-read it after running rather than editing
a stale copy.
3. **Lock content and audio** — read `references/generation.md`. Script →
full narration via `text_to_speech` → ASR-verify the narration against the
script → record real durations. The locked audio is the master clock for
everything downstream.
4. **Plan and generate the presenter** — read `references/generation.md`.
Short low-cost pilot first; full run only after the pilot passes identity
and mouth-timing review.
5. **Edit** — read `references/editing.md`. Deterministic timeline driven by
the locked audio; captions and keyword callouts only after audio and media
are final.
6. **Verify and deliver** — read `references/qa-recovery.md`, render, then:
```bash
bash "$SKILL_DIR/scripts/finalize_delivery.sh" \
~/Videos/my-presenter-video/renders/rendered.mp4 \
~/Videos/my-presenter-video/outputs my-video
```
The finalizer preserves aspect ratio, runs two-pass loudness normalization
(program ≈ −16 LUFS), produces master + share encodes, decode-verifies
both, writes a delivery report JSON, and emits a nine-frame contact sheet.
Inspect the contact sheet with `vision_analyze` before claiming completion.
## Operating rules (non-negotiable)
- Confirm image rights, adult status, remote-upload approval, and
voice-cloning authorization before the relevant remote action.
- Never infer or clone a real person's voice from an image; use an authorized
sample or a stock TTS voice.
- Lock the complete narration before presenter generation, caption timing, or
final scene boundaries.
- Mute video sources in the final composition; only the approved narration
and intentional mix tracks carry audio.
- Preserve provider request bodies and task IDs (minus credentials/expiring
URLs). Poll interrupted work before resubmitting — avoid double billing.
- Stop after three rejected paid candidates and summarize the failure mode.
- Do not claim completion until the final files fully decode and the contact
sheet or full playback has been reviewed.
## Defaults for minimal input
9:16, 1080×1920, 30fps; topic-derived videos target 45–75s; stock voice when
no authorized sample; presenter-led layout with hook → 2–4 beats → close;
no music/CTA unless requested; language inferred from the request.
## Reference routing
- `references/generation.md` — intake, content, voice, capability selection,
presenter prompts, paid generation, provider changes.
- `references/editing.md` — timeline contract, openings/closes, captions,
keyword-callout presets, HyperFrames composition, exports.
- `references/qa-recovery.md` — technical acceptance, visual acceptance, and
recovery for lip-sync/identity/hands/exposure/freeze/caption/audio faults.
## Pitfalls
- `preflight.py` requires ffprobe; on a bare box install ffmpeg first.
- The consent booleans set by init flags land under `input.*`; editing them
at the job-json root silently does nothing (preflight keeps blocking).
- `finalize_delivery.sh` needs bash + jq + awk and a fully decodable input —
a truncated render fails the decode check by design, not by accident.
- Long avatar clips drift: prefer one continuous presenter source sliced on
the audio timeline over many regenerated chapter clips (identity drift
across regenerations is the #1 visual-QA failure).
- FAL i2v endpoints cap duration (typically 5–15s); plan chapter-level
presenter segments accordingly and reuse the pilot's seed/params for
consistency where the endpoint supports it.
## Verification
Validated hands-on (Aug 2026): `init_job.py` → `job.json` with correct state
machine; `preflight.py` correctly blocked on unreviewed inputs, flipped to
`ok: true` after review booleans, and kept `remote_ready: false` until
`input.remote_upload_approved`; `finalize_delivery.sh` on a synthetic 5s
1080×1920 render produced decode-verified master (631kbit/s) + share encodes,
delivery-report JSON, and a 9-frame contact sheet, exit 0.

View File

@@ -0,0 +1,83 @@
{
"schema_version": 1,
"job_id": "replace-me",
"state": "intake",
"mode": "automation",
"input": {
"topic": "",
"script_path": "",
"presenter_image": "",
"voice_sample": "",
"supporting_media": [],
"rights_confirmed": false,
"adult_presenter_confirmed": false,
"remote_upload_approved": false,
"voice_clone_approved": false
},
"creative": {
"language": "auto",
"audience": "general",
"goal": "explain clearly",
"duration_target_s": 60,
"aspect": "9:16",
"width": 1080,
"height": 1920,
"fps": 30,
"style": "credible contemporary presenter",
"watermark": "",
"cta": ""
},
"voice": {
"strategy": "auto",
"voice_id": "",
"rate": 1.06,
"segment_lufs": -17,
"program_lufs": -16
},
"plan": {
"status": "draft",
"opening_target_s": 4,
"closing_target_s": 5,
"chapters": [],
"requested_generation_seconds": 0,
"estimated_cost": null,
"price_evidence_date": "",
"pilot_approved": false,
"paid_generation_approved": false,
"retry_ceiling": 3
},
"capabilities": {
"voice_generation": {},
"main_presenter": {},
"short_motion": {},
"lipsync_repair": {},
"word_timestamp_asr": {},
"timeline_compositor": {},
"encoder_qa": {}
},
"manual_input_review": {
"image_viewed": false,
"single_clear_face": false,
"image_has_no_unwanted_text": false,
"voice_sample_listened": false,
"single_clear_speaker": false
},
"artifacts": {
"script": "",
"beat_sheet": "",
"timeline": "",
"storyboard": "",
"final_audio": "",
"caption_json": "",
"rendered": "",
"master": "",
"share": ""
},
"qa": {
"preflight_report": "",
"asr_report": "",
"composition_report": "",
"delivery_report": "",
"manual_visual_review": ""
}
}

View File

@@ -0,0 +1,65 @@
# Editing
Read this file for the visual plan, deterministic timeline, openings, closes, screen demos, captions, keyword callouts, covers, previews, and exports.
## Visual routes
- **Presenter-led:** the presenter carries the explanation; add one concise visual idea per beat.
- **Screen-demo:** preserve real UI readability; introduce the presenter large at chapter starts, then shrink the same continuous video into a verified safe region.
- **Mixed explainer:** assign supporting media only where it proves or clarifies narration.
Do not fabricate product evidence. Label conceptual visuals clearly.
## Timeline contract
The locked narration defines duration and semantic boundaries. Every clip keeps three independent values:
- authored start in the final video;
- authored duration;
- source-media start.
When an opening changes length, shift authored starts while preserving presenter source offsets. Keep video sources muted and use approved external narration as the program clock. Prefer one continuous presenter source and source-time slices over regenerated chapter clips.
Use hard cuts at settled semantic pauses. Add transitions only when they clarify structure. Preserve a quiet tail so the last word, gesture, and mouth position finish naturally.
## Opening, body, and close
- Establish the topic and payoff within the first few seconds.
- Design the time-zero frame as a deliberate cover frame.
- Keep opening copy away from eyes, mouth, and the hand-motion path.
- Use the same visual system across the cover, chapters, captions, keywords, progress, and close.
- End with a conclusion or one requested next step. Do not add promotional copy without a brief requirement.
## Captions
Generate word timestamps from the final audio. Split into short semantic phrases, normally one readable line. A practical baseline is about `40ms` visual lead and `120ms` tail hold when neighboring phrases allow it.
Keep captions clear of the face, UI, watermark, progress line, and platform controls. Highlight one meaningful term per phrase. Make entrances and exits quick enough to preserve reading stability.
## Presenter-side keyword presets
Use keyword callouts to amplify selected spoken beats. Avoid repeating one identical card throughout the video.
Rotate a small family of presets such as:
- radial burst;
- tilted ribbon or sticker;
- hand-drawn circle or marker stroke;
- large/small type contrast;
- staggered word-chip cluster;
- double-layer outline lockup.
Bind each callout to a real spoken stress or ASR word anchor. Separate the entry timing of the kicker, main word, secondary word, marker, rays, and outline by a few frames. Let the callout settle briefly, then collapse before the next idea. Keep it inside a consistent presenter-side safe region and verify the complete motion path does not cover the face or hands.
## Composition and preview
Use a deterministic compositor available in the environment. For HyperFrames:
1. load the current HyperFrames workflow and domain instructions;
2. give all timed media explicit start, duration, track, and source offset;
3. run the required composition check with representative emphasis and transition samples;
4. inspect time zero, chapter cuts, layout changes, keyword beats, captions, close, and final frame;
5. open the final Studio preview and obtain approval before rendering;
6. render at delivery quality only after approval.
Measure the assembled stereo program after rendering. Mono narration duplicated into stereo can measure about 3 LU louder, so perform final program normalization from the rendered file.

View File

@@ -0,0 +1,83 @@
# Generation
Read this file for intake, content, narration, capability selection, paid generation, presenter prompts, or provider changes.
## Inputs and authorization
Required:
- a topic or script;
- one local presenter image with one clear adult face;
- confirmed image rights before remote upload;
- confirmed adult status from the user.
Optional:
- authorized voice sample;
- screen recordings, B-roll, images, charts, brand assets, or source links;
- audience, duration, language, aspect, style, watermark, music, CTA, provider preference, or previous episode.
Inspect the real image. Record visible anchors only: framing, hair, glasses, clothing, accessories, hands, furniture, background geometry, camera height, crop, light, exposure, white balance, embedded text, logos, occlusion, multiple faces, and resolution. Do not infer sensitive traits.
When a voice sample exists, listen for one speaker, audible speech duration, language, noise, echo, clipping, music, silence, and codec damage. Record whether cloning is authorized. Keep the original unchanged.
## Content and narration
For a topic, use one clear spine:
```text
hook → promise → 2–4 useful beats → synthesis → close
```
For a supplied script, preserve factual meaning while improving breath, pronunciation, sentence length, and transitions. Flag material factual edits for approval.
Generate the complete approved narration before visual work. Keep one voice identity and one synthesis configuration. Split only for operational limits or clean semantic sections. Preserve raw and normalized copies.
Run ASR against final audio. Review names, numbers, English tokens, omitted words, repeated words, and tail speech. Use actual audio durations for every later cut.
## Capability routing
Resolve capabilities at execution time:
| Capability | Acceptance priority |
|---|---|
| Voice generation | authorization, identity, pronunciation, rate control, clean audio |
| Main presenter | exact audio support, semantic lip sync, identity stability, required duration |
| Short motion | stable face, one controllable gesture, fixed camera, economical retry |
| Lip-sync repair | preserves accepted motion while replacing mouth timing with exact audio |
| Word-timestamp ASR | accurate word start/end times and reviewable output |
| Timeline compositor | explicit source offsets, frame-accurate seek, deterministic render |
| Encoder and QA | local probe, loudness, decode, black-frame and contact-sheet support |
Inspect current installed tools and official documentation when availability, duration, pricing, or request fields may have changed. Prefer a user-selected provider when it passes the gates. Otherwise choose the simplest available capability that does.
Record the actual provider, model, version, region, parameters, price evidence date, task ID, and sanitized request body in `job.json`. Ask before a provider change that affects cost, privacy, voice, appearance, or quality.
## Billing and pilot gate
Before the first remote or paid call, state:
- files or data being uploaded;
- capability and selected tool;
- requested seconds or units;
- known price or that the cost is unknown;
- pilot duration and retry ceiling;
- expected output and main risks.
Generate the smallest useful pilot. Continue to a full run only after technical checks and visual review pass. Do not generate a long clip when the edit uses only a short range.
## Presenter prompt structure
Write prompts in this order:
1. Bind the user image as the sole presenter and scene reference.
2. List only visible identity, clothing, accessory, camera, framing, background, light, exposure, and white-balance anchors.
3. Request realistic skin, hair, eye moisture, fabric, breathing, bilateral blinking, and restrained head motion.
4. Supply exact dialogue through the tool's supported audio or dialogue field.
5. Tie one physically plausible gesture to one phrase and approximate time.
6. Reserve a settled, closed-mouth tail.
7. Exclude identity, face, glasses, hair, wardrobe, skin tone, background, camera, lighting, finger, hand, dialogue, text, logo, watermark, and extra-person drift.
For the main presenter, prioritize mouth timing and identity. Keep hands low or outside frame. For an opening or close, one small gesture may carry the hook or CTA. Keep hands below the collarbone, away from the face and lens, and end the movement before the tail hold.
If the body performance is accepted but mouth timing is visibly late, preserve the motion plate and apply lip-sync repair with the exact locked audio and no duration extension.

View File

@@ -0,0 +1,73 @@
# QA and recovery
Read this file before accepting generated media or delivery, and whenever a production fault appears.
## Acceptance gates
### Inputs
- Topic or script exists.
- Presenter image decodes and contains one reviewed adult face.
- Image rights, adult confirmation, and remote-upload permission are recorded.
- Any voice sample contains one authorized speaker.
- Supporting media is probed and its provenance is recorded.
### Audio
- Final ASR matches the intended script once, without material omissions or additions.
- Sections have consistent loudness, commonly near `-17 LUFS`.
- Assembled program loudness is commonly `-16 ± 0.5 LUFS` unless the destination specifies another target.
- No clipped words, doubled tracks, echo, clicks, unexpected silence, or truncated tail.
- Measure the final stereo file; section measurements alone are insufficient.
### Presenter
- Expected resolution, frame rate, duration, and decodable streams.
- No black frames, long exact-frame freezes, progressive darkening, or duration extension.
- Identity, face geometry, hair, glasses, clothing, accessories, background, crop, exposure, and white balance remain coherent.
- Mouth follows names, numbers, English tokens, plosives, and phrase endings.
- Blinks are sparse and bilateral; gestures occur once and settle; hands remain plausible and away from the face.
- Tail ends with a settled face and resting mouth.
### Composition and delivery
- No flash, duplicate presenter, source-time reset, missing overlay, overflow, face obstruction, caption collision, or unreadable UI.
- Opening, callouts, captions, watermark, progress, and platform safe zones remain compatible.
- Master and share files have the expected duration, dimensions, frame rate, codec, pixel format, and audio rate.
- Both files fully decode.
- Contact sheet covers opening, chapters, emphasis graphics, close, and final frame.
- Watch the complete video at normal speed before delivery.
## Recovery rules
### Interrupted remote task
Poll the saved task ID and download a succeeded result. Resubmit only after a confirmed failure or cancellation.
### Voice identity or loudness changes
Confirm one synthesis identity and configuration. Regenerate the mismatching section, normalize again, and rerun ASR. If the final render is louder than the source, confirm video is muted and apply whole-program two-pass loudness correction.
### Presenter changes between chapters
Use one continuous audio-driven source with source-time slices. When splitting is unavoidable, lock identity, framing, parameters, light, and quiet edge holds.
### Body motion works but mouth is late
Keep the accepted motion plate. Apply lip-sync repair with the exact locked audio and no duration extension. Check speech anchors at numbers, plosives, English tokens, and final syllables.
### Face, glasses, hands, or light drift
Retry from the original image with reduced motion and explicit camera, exposure, and white-balance locks. Simplify the gesture and keep hands low. Structural face or finger faults require regeneration; grading cannot correct them.
### Frozen picture-in-picture
Confirm the layout uses moving video, source offsets advance continuously, and no poster frame or identical-frame padding replaced the presenter.
### Captions or keyword graphics feel wrong
Regenerate timings from the final audio. Split by meaning, shorten visible phrases, and adjust small lead/hold margins. Bind keyword motion to actual spoken anchors and verify its complete path against the face, hands, UI, and captions.
### Three rejected paid candidates
Stop. Preserve candidates, prompts, task IDs, cost, and rejection notes. Summarize the recurring failure, remaining options, expected additional cost, and the single variable proposed for the next attempt.

View File

@@ -0,0 +1,166 @@
#!/usr/bin/env bash
set -euo pipefail
usage() {
printf 'Usage: %s rendered.mp4 output-dir output-stem\n' "$(basename "$0")" >&2
}
die() {
printf 'ERROR: %s\n' "$*" >&2
exit 1
}
[[ "$#" -eq 3 ]] || { usage; exit 64; }
INPUT="$1"
OUTPUT_DIR="$2"
STEM="$3"
[[ -f "$INPUT" && -s "$INPUT" ]] || die "input render is missing or empty: $INPUT"
[[ "$STEM" =~ ^[A-Za-z0-9._-]+$ ]] || die "output stem contains unsupported characters: $STEM"
for command in ffmpeg ffprobe jq mktemp awk sed; do
command -v "$command" >/dev/null 2>&1 || die "required command unavailable: $command"
done
mkdir -p "$OUTPUT_DIR"
OUTPUT_DIR="$(cd "$OUTPUT_DIR" && pwd -P)"
INPUT="$(cd "$(dirname "$INPUT")" && pwd -P)/$(basename "$INPUT")"
MASTER="$OUTPUT_DIR/${STEM}-master.mp4"
SHARE="$OUTPUT_DIR/${STEM}-share.mp4"
REPORT="$OUTPUT_DIR/${STEM}-delivery-report.json"
CONTACT="$OUTPUT_DIR/${STEM}-contact-sheet.png"
# Keep shared reports portable and avoid exposing the developer machine path.
INPUT_REPORT="$(basename "$INPUT")"
MASTER_REPORT="$(basename "$MASTER")"
SHARE_REPORT="$(basename "$SHARE")"
CONTACT_REPORT="$(basename "$CONTACT")"
for output in "$MASTER" "$SHARE" "$REPORT" "$CONTACT"; do
[[ ! -e "$output" ]] || die "refusing to overwrite existing output: $output"
done
TMP_ROOT="${TMPDIR:-/tmp}"
TMP_DIR="$(mktemp -d "${TMP_ROOT%/}/presenter-finalize.XXXXXX")"
trap 'rm -rf -- "$TMP_DIR"' EXIT INT TERM
STREAM_JSON="$TMP_DIR/source-probe.json"
ffprobe -v error -show_streams -show_format -of json "$INPUT" >"$STREAM_JSON"
jq -e '.streams | any(.codec_type == "video") and any(.codec_type == "audio")' "$STREAM_JSON" >/dev/null \
|| die "input must contain decodable video and audio streams"
WIDTH="$(jq -r '[.streams[] | select(.codec_type == "video")][0].width' "$STREAM_JSON")"
HEIGHT="$(jq -r '[.streams[] | select(.codec_type == "video")][0].height' "$STREAM_JSON")"
FPS="$(jq -r '[.streams[] | select(.codec_type == "video")][0].avg_frame_rate' "$STREAM_JSON")"
[[ "$FPS" != "0/0" && -n "$FPS" ]] || FPS="30"
MEASURE_LOG="$TMP_DIR/loudnorm-pass1.log"
ffmpeg -hide_banner -nostdin -nostats -i "$INPUT" \
-map 0:a:0 -vn \
-af 'loudnorm=I=-16:TP=-1.5:LRA=9:print_format=json' \
-f null - >/dev/null 2>"$MEASURE_LOG"
MEASURE_JSON="$TMP_DIR/measure.json"
sed -n '/^{/,/^}/p' "$MEASURE_LOG" >"$MEASURE_JSON"
jq -e . "$MEASURE_JSON" >/dev/null || die "could not parse loudness measurement"
MEASURED_I="$(jq -r .input_i "$MEASURE_JSON")"
MEASURED_TP="$(jq -r .input_tp "$MEASURE_JSON")"
MEASURED_LRA="$(jq -r .input_lra "$MEASURE_JSON")"
MEASURED_THRESH="$(jq -r .input_thresh "$MEASURE_JSON")"
OFFSET="$(jq -r .target_offset "$MEASURE_JSON")"
LOUDNORM="loudnorm=I=-16:TP=-1.5:LRA=9:measured_I=${MEASURED_I}:measured_TP=${MEASURED_TP}:measured_LRA=${MEASURED_LRA}:measured_thresh=${MEASURED_THRESH}:offset=${OFFSET}:linear=true:print_format=summary"
VIDEO_FILTER="scale=trunc(iw/2)*2:trunc(ih/2)*2:flags=lanczos,setsar=1,fps=${FPS},format=yuv420p"
encode() {
local crf="$1"
local preset="$2"
local audio_bitrate="$3"
local destination="$4"
local temporary="$5"
ffmpeg -hide_banner -nostdin -loglevel warning -stats -i "$INPUT" \
-map 0:v:0 -map 0:a:0 -sn -dn \
-vf "$VIDEO_FILTER" -af "$LOUDNORM" \
-c:v libx264 -preset "$preset" -crf "$crf" -profile:v high \
-pix_fmt yuv420p -tag:v avc1 -fps_mode cfr \
-color_primaries bt709 -color_trc bt709 -colorspace bt709 \
-c:a aac -b:a "$audio_bitrate" -ar 48000 -ac 2 \
-map_metadata -1 -map_chapters -1 -movflags +faststart \
"$temporary"
[[ -s "$temporary" ]] || die "encoder produced an empty output"
mv -- "$temporary" "$destination"
}
encode 16 slow 256k "$MASTER" "$TMP_DIR/master.mp4"
encode 24 medium 160k "$SHARE" "$TMP_DIR/share.mp4"
for file in "$MASTER" "$SHARE"; do
ffmpeg -hide_banner -nostdin -v error -xerror -i "$file" \
-map 0:v:0 -map 0:a:0 -f null - >/dev/null
done
ffprobe -v error -show_streams -show_format -of json "$MASTER" >"$TMP_DIR/master-probe.json"
ffprobe -v error -show_streams -show_format -of json "$SHARE" >"$TMP_DIR/share-probe.json"
for probe_file in "$STREAM_JSON" "$TMP_DIR/master-probe.json" "$TMP_DIR/share-probe.json"; do
jq 'if .format.filename? then .format.filename = (.format.filename | split("/") | last) else . end' \
"$probe_file" >"${probe_file}.portable"
mv -- "${probe_file}.portable" "$probe_file"
done
ffmpeg -hide_banner -nostdin -i "$MASTER" \
-vf 'blackdetect=d=0.10:pix_th=0.02' -an -f null - \
>/dev/null 2>"$TMP_DIR/blackdetect.log"
BLACK_EVENTS="$(grep -c 'black_start:' "$TMP_DIR/blackdetect.log" || true)"
DURATION="$(ffprobe -v error -show_entries format=duration -of default=nw=1:nk=1 "$MASTER")"
awk -v duration="$DURATION" 'BEGIN {
p[1]=0.2; p[2]=duration*0.125; p[3]=duration*0.25; p[4]=duration*0.375;
p[5]=duration*0.5; p[6]=duration*0.625; p[7]=duration*0.75;
p[8]=duration*0.875; p[9]=duration-0.2;
for(i=1;i<=9;i++) printf "%.6f\n", p[i]
}' >"$TMP_DIR/timestamps.txt"
if (( WIDTH > HEIGHT )); then
CONTACT_FRAME_FILTER='scale=480:270:force_original_aspect_ratio=decrease:flags=lanczos,pad=480:270:(ow-iw)/2:(oh-ih)/2:color=black'
elif (( HEIGHT > WIDTH )); then
CONTACT_FRAME_FILTER='scale=270:480:force_original_aspect_ratio=decrease:flags=lanczos,pad=270:480:(ow-iw)/2:(oh-ih)/2:color=black'
else
CONTACT_FRAME_FILTER='scale=360:360:force_original_aspect_ratio=decrease:flags=lanczos,pad=360:360:(ow-iw)/2:(oh-ih)/2:color=black'
fi
index=0
while IFS= read -r timestamp; do
index=$((index + 1))
printf -v frame '%s/frame-%02d.png' "$TMP_DIR" "$index"
ffmpeg -hide_banner -nostdin -loglevel error -ss "$timestamp" -i "$MASTER" \
-frames:v 1 -vf "$CONTACT_FRAME_FILTER" -update 1 "$frame"
done <"$TMP_DIR/timestamps.txt"
ffmpeg -hide_banner -nostdin -loglevel error -framerate 1 -start_number 1 \
-i "$TMP_DIR/frame-%02d.png" -frames:v 1 \
-vf 'tile=3x3:padding=12:margin=12:color=0x101218' -update 1 "$TMP_DIR/contact.png"
mv -- "$TMP_DIR/contact.png" "$CONTACT"
jq -n \
--arg status verified \
--arg input "$INPUT_REPORT" \
--arg master "$MASTER_REPORT" \
--arg share "$SHARE_REPORT" \
--arg contact "$CONTACT_REPORT" \
--arg duration "$DURATION" \
--argjson black_events "$BLACK_EVENTS" \
--slurpfile measured "$MEASURE_JSON" \
--slurpfile source_probe "$STREAM_JSON" \
--slurpfile master_probe "$TMP_DIR/master-probe.json" \
--slurpfile share_probe "$TMP_DIR/share-probe.json" \
'{status:$status,input:$input,master:$master,share:$share,contact_sheet:$contact,duration_s:($duration|tonumber),source_loudness:$measured[0],source_probe:$source_probe[0],master_probe:$master_probe[0],share_probe:$share_probe[0],full_decode_passed:true,black_frame_events:$black_events}' \
>"$TMP_DIR/report.json"
mv -- "$TMP_DIR/report.json" "$REPORT"
printf 'Master: %s\nShare: %s\nReport: %s\nContact: %s\n' "$MASTER" "$SHARE" "$REPORT" "$CONTACT"

View File

@@ -0,0 +1,151 @@
#!/usr/bin/env python3
"""Create a non-overwriting presenter-video job from minimal inputs."""
from __future__ import annotations
import argparse
import json
import re
import sys
from pathlib import Path
SKILL_DIR = Path(__file__).resolve().parent.parent
TEMPLATE = SKILL_DIR / "assets" / "job.template.json"
ASPECT_DEFAULTS = {
"9:16": (1080, 1920),
"16:9": (1920, 1080),
"1:1": (1080, 1080),
"4:5": (1080, 1350),
}
def absolute_existing(path_text: str, label: str) -> str:
path = Path(path_text).expanduser().resolve()
if not path.is_file():
raise ValueError(f"{label} does not exist or is not a file: {path}")
return str(path)
def slugify(value: str) -> str:
slug = re.sub(r"[^a-z0-9]+", "-", value.lower()).strip("-")
return slug or "presenter-video"
def parse_args() -> argparse.Namespace:
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument("--job-dir", required=True)
parser.add_argument("--presenter-image", required=True)
source = parser.add_mutually_exclusive_group(required=True)
source.add_argument("--topic")
source.add_argument("--script", help="Existing local script file")
parser.add_argument("--voice-sample")
parser.add_argument("--supporting-media", action="append", default=[])
parser.add_argument("--language", default="auto")
parser.add_argument("--audience", default="general")
parser.add_argument("--duration", type=float, default=60.0)
parser.add_argument("--aspect", choices=sorted(ASPECT_DEFAULTS), default="9:16")
parser.add_argument("--width", type=int)
parser.add_argument("--height", type=int)
parser.add_argument("--fps", type=int, choices=(24, 25, 30, 50, 60), default=30)
parser.add_argument("--style", default="credible contemporary presenter")
parser.add_argument("--watermark", default="")
parser.add_argument("--cta", default="")
parser.add_argument("--rights-confirmed", action="store_true")
parser.add_argument("--adult-presenter-confirmed", action="store_true")
parser.add_argument("--remote-upload-approved", action="store_true")
parser.add_argument("--voice-clone-approved", action="store_true")
return parser.parse_args()
def main() -> int:
args = parse_args()
if not 5 <= args.duration <= 1800:
raise ValueError("--duration must be between 5 and 1800 seconds")
if (args.width is None) != (args.height is None):
raise ValueError("--width and --height must be supplied together")
width, height = ASPECT_DEFAULTS[args.aspect]
if args.width is not None and args.height is not None:
width, height = args.width, args.height
if min(width, height) < 256 or max(width, height) > 7680:
raise ValueError("custom dimensions must be between 256 and 7680 pixels")
job_dir = Path(args.job_dir).expanduser().resolve()
if job_dir.exists() and any(job_dir.iterdir()):
raise ValueError(f"job directory must be absent or empty: {job_dir}")
job_dir.mkdir(parents=True, exist_ok=True)
manifest = json.loads(TEMPLATE.read_text(encoding="utf-8"))
manifest["job_id"] = slugify(job_dir.name)
manifest["input"]["topic"] = (args.topic or "").strip()
manifest["input"]["script_path"] = (
absolute_existing(args.script, "script") if args.script else ""
)
manifest["input"]["presenter_image"] = absolute_existing(
args.presenter_image, "presenter image"
)
manifest["input"]["voice_sample"] = (
absolute_existing(args.voice_sample, "voice sample")
if args.voice_sample
else ""
)
manifest["input"]["supporting_media"] = [
absolute_existing(item, "supporting media") for item in args.supporting_media
]
manifest["input"]["rights_confirmed"] = args.rights_confirmed
manifest["input"]["adult_presenter_confirmed"] = args.adult_presenter_confirmed
manifest["input"]["remote_upload_approved"] = args.remote_upload_approved
manifest["input"]["voice_clone_approved"] = args.voice_clone_approved
creative = manifest["creative"]
creative.update(
{
"language": args.language,
"audience": args.audience,
"duration_target_s": args.duration,
"aspect": args.aspect,
"width": width,
"height": height,
"fps": args.fps,
"style": args.style,
"watermark": args.watermark,
"cta": args.cta,
}
)
directories = [
"docs",
"assets/source",
"assets/audio/reference",
"assets/audio/raw",
"assets/audio/final",
"assets/video/candidates",
"assets/video/selected",
"assets/video/render",
"assets/captions",
"qa/requests",
"qa/asr",
"qa/contacts",
"qa/reports",
"renders",
"outputs",
]
for relative in directories:
(job_dir / relative).mkdir(parents=True, exist_ok=True)
job_path = job_dir / "job.json"
job_path.write_text(
json.dumps(manifest, ensure_ascii=False, indent=2) + "\n",
encoding="utf-8",
)
print(json.dumps({"job": str(job_path), "state": "intake"}, ensure_ascii=False))
return 0
if __name__ == "__main__":
try:
raise SystemExit(main())
except ValueError as exc:
print(f"ERROR: {exc}", file=sys.stderr)
raise SystemExit(2)

View File

@@ -0,0 +1,177 @@
#!/usr/bin/env python3
"""Validate a presenter-video job and write a machine-readable preflight report."""
from __future__ import annotations
import json
import shutil
import subprocess
import sys
from pathlib import Path
from typing import Any
def probe(path: Path) -> dict[str, Any]:
result = subprocess.run(
[
"ffprobe",
"-v",
"error",
"-show_entries",
"stream=index,codec_type,codec_name,width,height,pix_fmt,sample_rate,channels,r_frame_rate:format=duration,size,bit_rate",
"-of",
"json",
str(path),
],
check=True,
capture_output=True,
text=True,
)
return json.loads(result.stdout)
def require_file(value: str, label: str, errors: list[str]) -> Path | None:
if not value:
errors.append(f"missing {label}")
return None
path = Path(value).expanduser().resolve()
if not path.is_file():
errors.append(f"{label} is not a file: {path}")
return None
return path
def write_atomic(path: Path, text: str) -> None:
temporary = path.with_suffix(path.suffix + ".tmp")
temporary.write_text(text, encoding="utf-8")
temporary.replace(path)
def main() -> int:
if len(sys.argv) != 2:
print("Usage: preflight.py ~/Videos/my-presenter-video/job.json", file=sys.stderr)
return 64
if not shutil.which("ffprobe"):
print("ERROR: ffprobe is required", file=sys.stderr)
return 2
job_path = Path(sys.argv[1]).expanduser().resolve()
if not job_path.is_file():
print(f"ERROR: job manifest does not exist: {job_path}", file=sys.stderr)
return 2
job = json.loads(job_path.read_text(encoding="utf-8"))
job_dir = job_path.parent
errors: list[str] = []
remote_blockers: list[str] = []
warnings: list[str] = []
media: dict[str, Any] = {}
input_data = job.get("input", {})
topic = str(input_data.get("topic", "")).strip()
script_text = str(input_data.get("script_path", "")).strip()
if not topic and not script_text:
errors.append("topic or script_path is required")
if script_text:
require_file(script_text, "script", errors)
image = require_file(str(input_data.get("presenter_image", "")), "presenter image", errors)
if image:
try:
media["presenter_image"] = probe(image)
streams = [
stream
for stream in media["presenter_image"].get("streams", [])
if stream.get("codec_type") == "video"
]
if not streams:
errors.append("presenter image has no decodable image/video stream")
else:
width = int(streams[0].get("width") or 0)
height = int(streams[0].get("height") or 0)
if min(width, height) < 512:
warnings.append(f"presenter image is low resolution: {width}x{height}")
except (subprocess.CalledProcessError, json.JSONDecodeError) as exc:
errors.append(f"could not decode presenter image: {exc}")
voice_value = str(input_data.get("voice_sample", "")).strip()
if voice_value:
voice = require_file(voice_value, "voice sample", errors)
if voice:
try:
media["voice_sample"] = probe(voice)
streams = [
stream
for stream in media["voice_sample"].get("streams", [])
if stream.get("codec_type") == "audio"
]
if not streams:
errors.append("voice sample has no audio stream")
duration = float(media["voice_sample"].get("format", {}).get("duration") or 0)
if duration < 4:
warnings.append(f"voice sample is short: {duration:.3f}s")
if duration > 60:
warnings.append(f"voice sample is unusually long: {duration:.3f}s")
except (subprocess.CalledProcessError, json.JSONDecodeError, ValueError) as exc:
errors.append(f"could not decode voice sample: {exc}")
if not input_data.get("voice_clone_approved"):
remote_blockers.append("voice_clone_approved must be true before voice cloning")
else:
warnings.append("no voice sample supplied; use and record a stock voice")
supporting_reports = []
for value in input_data.get("supporting_media", []):
path = require_file(str(value), "supporting media", errors)
if path:
try:
supporting_reports.append({"file": path.name, "probe": probe(path)})
except (subprocess.CalledProcessError, json.JSONDecodeError) as exc:
errors.append(f"could not decode supporting media {path}: {exc}")
media["supporting_media"] = supporting_reports
if not input_data.get("rights_confirmed"):
remote_blockers.append("rights_confirmed must be true before presenter synthesis")
if not input_data.get("adult_presenter_confirmed"):
remote_blockers.append("adult_presenter_confirmed must be true before presenter synthesis")
if not input_data.get("remote_upload_approved"):
remote_blockers.append("remote_upload_approved must be true before remote generation")
manual = job.get("manual_input_review", {})
for key in ("image_viewed", "single_clear_face", "image_has_no_unwanted_text"):
if not manual.get(key):
errors.append(f"manual_input_review.{key} must be true")
if voice_value:
for key in ("voice_sample_listened", "single_clear_speaker"):
if not manual.get(key):
errors.append(f"manual_input_review.{key} must be true")
creative = job.get("creative", {})
duration = float(creative.get("duration_target_s") or 0)
if not 5 <= duration <= 1800:
errors.append("creative.duration_target_s must be between 5 and 1800")
width = int(creative.get("width") or 0)
height = int(creative.get("height") or 0)
if min(width, height) < 256 or max(width, height) > 7680:
errors.append("creative width/height must be between 256 and 7680")
if int(creative.get("fps") or 0) not in (24, 25, 30, 50, 60):
errors.append("creative.fps must be one of 24, 25, 30, 50, or 60")
report_path = job_dir / "qa" / "reports" / "preflight.json"
report_path.parent.mkdir(parents=True, exist_ok=True)
report = {
"ok": not errors,
"remote_ready": not errors and not remote_blockers,
"job": job_path.name,
"errors": errors,
"remote_blockers": remote_blockers,
"warnings": warnings,
"media": media,
}
write_atomic(report_path, json.dumps(report, ensure_ascii=False, indent=2) + "\n")
job.setdefault("qa", {})["preflight_report"] = "qa/reports/preflight.json"
write_atomic(job_path, json.dumps(job, ensure_ascii=False, indent=2) + "\n")
print(json.dumps(report, ensure_ascii=False, indent=2))
return 0 if not errors else 1
if __name__ == "__main__":
raise SystemExit(main())

View File

@@ -60,6 +60,7 @@ hermes skills uninstall <skill-name>
| Skill | Description |
|-------|-------------|
| [**ai-presenter-video**](/docs/user-guide/skills/optional/creative/creative-ai-presenter-video) | Make a verified AI presenter video from script + image. |
| [**archify**](/docs/user-guide/skills/optional/creative/creative-archify) | Validated interactive HTML diagrams, upstream-maintained. |
| [**ascii-art**](/docs/user-guide/skills/optional/creative/creative-ascii-art) | ASCII art: pyfiglet, cowsay, boxes, image-to-ascii. |
| [**audiocraft-audio-generation**](/docs/user-guide/skills/optional/creative/creative-audiocraft-audio-generation) | AudioCraft: MusicGen text-to-music, AudioGen text-to-sound. |

View File

@@ -0,0 +1,194 @@
---
title: "Ai Presenter Video — Make a verified AI presenter video from script + image"
sidebar_label: "Ai Presenter Video"
description: "Make a verified AI presenter video from script + image"
---
{/* This page is auto-generated from the skill's SKILL.md by website/scripts/generate-skill-docs.py. Edit the source SKILL.md, not this page. */}
# Ai Presenter Video
Make a verified AI presenter video from script + image.
## Skill metadata
| | |
|---|---|
| Source | Optional — install with `hermes skills install official/creative/ai-presenter-video` |
| Path | `optional-skills/creative/ai-presenter-video` |
| Version | `1.0.0` |
| Author | cclank (https://github.com/cclank/lanshu-create-ai-presenter-video), ported by Hermes Agent |
| License | MIT |
| Platforms | linux, macos |
| Tags | `video`, `presenter`, `avatar`, `lipsync`, `tts`, `captions`, `creative` |
| Related skills | [`hyperframes`](/docs/user-guide/skills/optional/creative/creative-hyperframes), [`kanban-video-orchestrator`](/docs/user-guide/skills/optional/creative/creative-kanban-video-orchestrator), [`comfyui`](/docs/user-guide/skills/bundled/creative/creative-comfyui) |
## Reference: full SKILL.md
:::info
The following is the complete skill definition that Hermes loads when this skill is triggered. This is what the agent sees as instructions when the skill is active.
:::
# AI Presenter Video
Turn a topic (or finished script) plus ONE authorized adult presenter image
into a complete, publish-ready presenter-led video: locked narration, avatar
generation with lip-sync QA, captions, deterministic editing, loudness-normalized
master/share encodes, and machine + visual acceptance reports.
Use this skill for new presenter videos AND for continuing, revising,
captioning, lip-sync-repairing, or re-exporting an existing presenter-video
job. The workflow is provider-neutral: pick generation capabilities from what
is actually available in the session (FAL video/image models via
`image_generate` and the video-gen plugin, TTS via `text_to_speech`, ASR via
the whisper/STT tooling, ffmpeg for everything deterministic).
> Ported from cclank/lanshu-create-ai-presenter-video (MIT). Upstream body
> kept substantively verbatim in `references/`; Hermes adaptations live in
> this hub file. Scripts are deterministic (no network, no credentials).
## Hermes adaptations (read first)
- **Skill dir resolution** — upstream hardcoded `~/.codex/skills/...`. In
Hermes resolve it once per session:
```bash
SKILL_DIR="$(dirname "$(find ~/.hermes/skills ~/.hermes/hermes-agent/optional-skills -path '*/ai-presenter-video/SKILL.md' 2>/dev/null | head -1)")"
```
Shell variables do not persist between tool calls — re-paste the resolution
line (or the expanded path) in each terminal call that uses it.
- **Capability mapping** — where the references say "a voice generation
capability", use `text_to_speech` (OpenAI/Edge/ElevenLabs per user config);
"presenter/avatar generation" → FAL image-to-video families (Kling, Wan,
MiniMax H3 etc.) through the configured video tooling, or an avatar/lipsync
endpoint the user has access to; "word-timestamp ASR" → whisper via the STT
tooling or `faster-whisper` in a venv; "deterministic compositor" → ffmpeg
filtergraphs, or the `hyperframes` skill when installed (the editing
reference has a HyperFrames section that maps directly onto it).
- **Visual QA** — do the "normal-speed visual review" steps with
`vision_analyze` on the generated contact sheet plus sampled frames
(identity, mouth timing, hands, blinking, continuity). Numeric checks come
from the scripts' ffprobe output.
- **Paid-generation consent** — remote avatar/TTS generation is billable.
Follow the upstream operating rules: before the first paid call state the
uploaded assets, requested seconds, known cost, pilot size, and retry
ceiling, and get the user's explicit go-ahead. Never upload the presenter
image to a remote provider before `remote_upload_approved` is true in
`job.json`.
- **Consent flags live under `input`** — `rights_confirmed`,
`adult_presenter_confirmed`, `remote_upload_approved`, and
`voice_clone_approved` sit inside the `input` object of `job.json` (init
flags set them; hand-editing must target `input.*`, not the job root).
`manual_input_review.*` sits at the root. `preflight.py` distinguishes
`errors` (block everything) from `remote_blockers` (block only remote
generation) — local script/audio work may proceed while remote is blocked.
## Workflow
1. **Start or resume a job.** New job:
```bash
python3 "$SKILL_DIR/scripts/init_job.py" \
--job-dir ~/Videos/my-presenter-video \
--presenter-image /path/to/presenter.png \
--topic "explain context engineering in one minute" \
--duration 60 --aspect 9:16 \
--rights-confirmed --adult-presenter-confirmed
```
Use `--script` for an existing script file; other flags: `--voice-sample`,
`--supporting-media`, `--width`, `--height`, `--fps`, `--watermark`,
`--cta`. For an existing job, read `job.json` + QA reports and resume from
the earliest unfinished state — never regenerate accepted work.
2. **Manual input review.** Actually look at the presenter image
(`vision_analyze`) and listen to any voice sample; record findings by
setting the `manual_input_review` booleans in `job.json`. Then gate:
```bash
python3 "$SKILL_DIR/scripts/preflight.py" ~/Videos/my-presenter-video/job.json
```
Proceed only when `ok: true`; do remote generation only when
`remote_ready: true`.
3. **Lock content and audio** — read `references/generation.md`. Script →
full narration via `text_to_speech` → ASR-verify the narration against the
script → record real durations. The locked audio is the master clock for
everything downstream.
4. **Plan and generate the presenter** — read `references/generation.md`.
Short low-cost pilot first; full run only after the pilot passes identity
and mouth-timing review.
5. **Edit** — read `references/editing.md`. Deterministic timeline driven by
the locked audio; captions and keyword callouts only after audio and media
are final.
6. **Verify and deliver** — read `references/qa-recovery.md`, render, then:
```bash
bash "$SKILL_DIR/scripts/finalize_delivery.sh" \
~/Videos/my-presenter-video/renders/rendered.mp4 \
~/Videos/my-presenter-video/outputs my-video
```
The finalizer preserves aspect ratio, runs two-pass loudness normalization
(program ≈ −16 LUFS), produces master + share encodes, decode-verifies
both, writes a delivery report JSON, and emits a nine-frame contact sheet.
Inspect the contact sheet with `vision_analyze` before claiming completion.
## Operating rules (non-negotiable)
- Confirm image rights, adult status, remote-upload approval, and
voice-cloning authorization before the relevant remote action.
- Never infer or clone a real person's voice from an image; use an authorized
sample or a stock TTS voice.
- Lock the complete narration before presenter generation, caption timing, or
final scene boundaries.
- Mute video sources in the final composition; only the approved narration
and intentional mix tracks carry audio.
- Preserve provider request bodies and task IDs (minus credentials/expiring
URLs). Poll interrupted work before resubmitting — avoid double billing.
- Stop after three rejected paid candidates and summarize the failure mode.
- Do not claim completion until the final files fully decode and the contact
sheet or full playback has been reviewed.
## Defaults for minimal input
9:16, 1080×1920, 30fps; topic-derived videos target 45–75s; stock voice when
no authorized sample; presenter-led layout with hook → 2–4 beats → close;
no music/CTA unless requested; language inferred from the request.
## Reference routing
- `references/generation.md` — intake, content, voice, capability selection,
presenter prompts, paid generation, provider changes.
- `references/editing.md` — timeline contract, openings/closes, captions,
keyword-callout presets, HyperFrames composition, exports.
- `references/qa-recovery.md` — technical acceptance, visual acceptance, and
recovery for lip-sync/identity/hands/exposure/freeze/caption/audio faults.
## Pitfalls
- `preflight.py` requires ffprobe; on a bare box install ffmpeg first.
- The consent booleans set by init flags land under `input.*`; editing them
at the job-json root silently does nothing (preflight keeps blocking).
- `finalize_delivery.sh` needs bash + jq + awk and a fully decodable input —
a truncated render fails the decode check by design, not by accident.
- Long avatar clips drift: prefer one continuous presenter source sliced on
the audio timeline over many regenerated chapter clips (identity drift
across regenerations is the #1 visual-QA failure).
- FAL i2v endpoints cap duration (typically 5–15s); plan chapter-level
presenter segments accordingly and reuse the pilot's seed/params for
consistency where the endpoint supports it.
## Verification
Validated hands-on (Aug 2026): `init_job.py` → `job.json` with correct state
machine; `preflight.py` correctly blocked on unreviewed inputs, flipped to
`ok: true` after review booleans, and kept `remote_ready: false` until
`input.remote_upload_approved`; `finalize_delivery.sh` on a synthetic 5s
1080×1920 render produced decode-verified master (631kbit/s) + share encodes,
delivery-report JSON, and a 9-frame contact sheet, exit 0.

View File

@@ -355,6 +355,7 @@ const sidebars: SidebarsConfig = {
key: 'skills-optional-creative',
collapsed: true,
items: [
'user-guide/skills/optional/creative/creative-ai-presenter-video',
'user-guide/skills/optional/creative/creative-ascii-art',
'user-guide/skills/optional/creative/creative-archify',
'user-guide/skills/optional/creative/creative-audiocraft-audio-generation',